# How are LLM cost and latency controlled during indexing and query?

> Graph RAG — a good answer covers: Caching; batching; model choice per stage; token budgets; small-model or non-LLM shortcuts.

Canonical page: https://llms-technical-reviews.com/graph-rag/q/cost/

## Verdict

The cheapest options avoid the LLM where they can: [graphify](/p/graphify/) parses code with no model, and [Semantica](/p/semantica/) extracts with spaCy by default. Among full LLM pipelines, [LightRAG](/p/lightrag/), [HippoRAG](/p/hipporag/) and [Vector Graph RAG](/p/vector-graph-rag/) keep each query to a few calls. [GraphRAG](/p/graphrag/) spends most of its budget at index time.

**Heavy indexing.** GraphRAG runs extraction plus gleaning on every chunk, a summarization call for merged descriptions, and one report per community. Each stage has its own model, and its `fast` method replaces LLM extraction with NLP. Only indexing calls are cached. [nano-graphrag](/p/nano-graphrag/) splits work between a "best" model and a "cheap" model and caches every call, including queries, but each insert regenerates all community reports. [LLM Graph Builder](/p/llm-graph-builder/) makes one call per `chunks_to_combine` chunks and has no extraction cache. Optional per-user daily and monthly token limits can stop a job before it starts. [AutoFlow](/p/autoflow/) makes two DSPy calls per chunk plus LLM merge calls for entities, and uses a separate `fast_llm` to rewrite questions.

**Lean indexing, cheap queries.** LightRAG writes no reports, caches both extraction and answers, and routes extraction, keyword and answer calls to different models. Its answer cache ignores the retrieved context, so a cached answer can survive changes to the data. HippoRAG caches every request in SQLite. A query costs one filter call plus a PageRank solve. Vector Graph RAG caches every LLM call on disk, uses `gpt-4o-mini` for every step with no per-stage choice, and makes three calls per query. [TrustGraph](/p/trustgraph/) makes two LLM calls per query and filters edges with a MiniLM cross-encoder. It does not cache extraction.

**Skipping the LLM.** graphify packs files for its LLM pass into chunks of up to 60,000 tokens. When a response is truncated, it splits the chunk in half and retries, up to three levels deep. It caches results per file hash. Semantica's LLM stages are opt-in, it caches extraction with a TTL of 3,600 seconds, and it can write extractive community reports with no model.

Pick: graphify or Semantica's defaults when the budget is close to zero.
Pick: LightRAG or HippoRAG for low running cost with LLM-quality extraction.
Pick: GraphRAG when global answers justify paying for indexing up front.

## Per-project answers

### Graphify-Labs/graphify (answered)

**Token-budget chunk packing.** The semantic extraction pipeline (`extract_corpus_parallel`, llm.py:2567) packs files into chunks using a configurable `token_budget` (default 60,000 tokens). Each file's token cost is estimated (characters/4 for code, 1,600 fixed per image), and chunks are closed when adding the next file would exceed the budget (llm.py:2160–2203). This prevents single over-large LLM requests that waste money on truncated retries.

**Adaptive retry with bisection.** When an LLM response hits `finish_reason="length"` (truncation), `_extract_with_adaptive_retry()` (llm.py:2319) bisects the chunk and re-extracts each half recursively, up to a configurable `max_retry_depth` (default 3, controlled by `GRAPHIFY_MAX_RETRY_DEPTH`). This bounds worst-case per-chunk cost at `2^depth` calls. Setting depth to 0 disables all retries — one call per chunk, full stop.

**Hollow-response backoff, not bisection.** A hollow response (HTTP 200 with empty/unparseable content) is retried with backoff (2s, 8s) rather than bisected, because bisecting a backend issue cannot converge and would cost far more (llm.py:2406–2421). After 3 attempts the chunk fails loudly.

**Model choice per stage.** The BACKENDS dict (llm.py:104–224) defines model defaults per provider (claude, openai, gemini, kimi, ollama, deepseek, azure, bedrock, claude-cli) with configurable `pricing` per megatoken. Users can shift cost by selecting cheaper models (e.g., Ollama/local models cost $0). Model overrides are settable per provider via env vars (`ANTHROPIC_MODEL`, `GRAPHIFY_OPENAI_MODEL`, etc.).

**Caching avoids re-extraction.** Both the AST cache (versioned content-addressed) and semantic cache (content-hash + prompt-fingerprinted) skip re-extraction of unchanged files, eliminating LLM calls entirely for files that haven't changed since the last run.

**Concurrency control.** The thread pool defaults to 4 concurrent chunks (`max_concurrency`, llm.py:2594). Ollama defaults to serial (1) because concurrent requests cause VRAM pressure and hollow responses (llm.py:2671–2672). The `GRAPHIFY_MAX_RETRIES` env var controls SDK-level retries on rate limits, defaulting to 6 (llm.py:430).

**Timeout guard.** `GRAPHIFY_API_TIMEOUT` (default 600s) caps total request wall-clock time per chunk (llm.py:410–420). A timed-out chunk triggers adaptive bisection rather than silent failure.


Citations: [graphify/llm.py:2160-2203](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L2160-L2203) · [graphify/llm.py:2567-2610](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L2567-L2610) · [graphify/llm.py:2319-2370](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L2319-L2370) · [graphify/llm.py:104-225](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L104-L225) · [graphify/llm.py:410-464](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L410-L464) · [graphify/cache.py:1274-1350](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/cache.py#L1274-L1350)

### HKUDS/LightRAG (answered)

**LLM response caching (primary cost control).** Both entity extraction and query results are cached. `enable_llm_cache` (default True) caches query answers; `enable_llm_cache_for_entity_extract` (default True) caches extraction results (`lightrag/lightrag.py:1093-1097`). The cache is implemented in `handle_cache` / `save_to_cache` (`lightrag/utils.py:4448-4520`), storing flattened cache keys of format `{mode}:{cache_type}:{hash}` in `llm_response_cache` (a KV storage). On cache hit, the LLM call is entirely skipped — response is returned from the stored value. The extraction cache also has per-chunk write-ahead semantics to avoid orphaned cache rows.

**Concurrency batching.** The number of concurrent LLM extraction calls per document is controlled by `llm_model_max_async` (default 4) via `asyncio.Semaphore(chunk_max_async)` at `lightrag/operate.py:4547-4548`. Document-level parallelism is capped by `max_parallel_insert` (default 3) at `lightrag/pipeline.py:2653`. Embedding requests are batched by `embedding_batch_num` at `lightrag/kg/nano_vector_db_impl.py:133` and various backend implementations.

**Role-based model routing.** LightRAG supports separate LLM bindings per role (`extract`, `keyword`, `query`, `vlm`) via `RoleLLMConfig` (`lightrag/llm_roles.py:57-63`). Each role can use a different model (e.g., a cheaper/faster model for extraction, a more capable one for query), configured via env vars like `EXTRACT_LLM_BINDING`, `QUERY_LLM_BINDING`. This lets operators trade cost vs. quality per stage.

**Token budgets.** Multiple token budgets limit LLM context size at key stages:
- `MAX_EXTRACT_INPUT_TOKENS` limits the gleaning step input (default 128K, `lightrag/operate.py:3992-3996`).
- At query time, `max_entity_tokens`, `max_relation_tokens`, `max_total_tokens` limit what goes into the final prompt (`_apply_token_truncation` at `lightrag/operate.py:5501`).
- Entity/relation descriptions are chunked and recursively summarized via `_handle_entity_relation_summary` map-reduce (`lightrag/operate.py:372`) when they exceed token limits, avoiding giant prompts.
- Document chunks that exceed the embedding model's token limit are truncated by `_truncate_vdb_content` (`lightrag/operate.py:298`).

**Non-LLM shortcuts.** The entire relation weight contract (`docs/ProgramingWithCore.md`) avoids LLM calls by simply counting distinct source IDs for the weight floor. The `kg_chunk_pick_method` (`DEFAULT_KG_CHUNK_PICK_METHOD`) supports weighted polling (based on occurrence count in entity/relation source_id lists) as a non-LLM alternative to vector-similarity chunk selection. The `NaiveQuery` mode bypasses the KG entirely and does a vector-only search. `bypass` mode skips all retrieval.

> **Editor's note.** Correction: the default `MAX_EXTRACT_INPUT_TOKENS` is 20,480 (`constants.py`), not 128K. Note also that the query-answer cache key does not include the retrieved context, so a cached answer can outlive changes to the data.

Citations: [lightrag/lightrag.py:1093-1097](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L1093-L1097) · [lightrag/utils.py:4448-4477](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/utils.py#L4448-L4477) · [lightrag/operate.py:4546-4558](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4546-L4558) · [lightrag/llm_roles.py:52-57](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/llm_roles.py#L52-L57) · [lightrag/operate.py:5501-5525](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5501-L5525) · [lightrag/operate.py:372-420](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L372-L420)

### microsoft/graphrag (answered)

**LLM response caching.** The `with_cache` middleware wraps every LLM completion and embedding call. Before calling the model, it computes a cache key from full request args, checks `cache.get(key)`, and on a hit returns the cached response. Applies to both indexing and query. Streaming and mocked responses are excluded.

**Batching.** `embed_text` buffers rows and dispatches batches at configurable `batch_size` and `batch_max_tokens`, with flush size sized to saturate `num_threads × batch_size`. For LLM completions, `derive_from_rows` parallelizes across `num_threads`.

**Model choice per stage.** Separate model IDs per stage: `extract_graph.completion_model_id`, `summarize_descriptions.completion_model_id`, `community_reports.completion_model_id`, `embed_text.embedding_model_id`. Each can be a cheap/fast model via litellm model strings.

**Token budgets.** All context builders enforce `max_context_tokens=8000`. `build_entity_context` truncates when cumulative tokens exceed budget. Community reports cap `max_input_length`/`max_report_length`. Global search enforces `max_data_tokens=8000` and per-stage max-length params. `max_gleanings` (default configurable) controls extra LLM rounds during extraction — setting it to 0 skips gleaning.

**Rate limiting and retries.** `with_rate_limiting` uses a sliding-window rate limiter per model, counting tokens. `with_retries` adds exponential backoff. Both compose via middleware pipeline.

**Metrics.** `MetricsProcessor` and `MetricsStore` track token usage, response times, and cache hit rates. Query results return per-category `llm_calls`, `prompt_tokens`, and `output_tokens`.

> **Editor's note.** Correction: only indexing calls are cached. The query engines are built with `create_completion(model_settings)` and no cache, so query calls are never cached. The configured context budget is 12,000 tokens, not 8,000.

Citations: [packages/graphrag-llm/graphrag_llm/middleware/with_cache.py:107-152](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-llm/graphrag_llm/middleware/with_cache.py#L107-L152) · [packages/graphrag/graphrag/index/operations/embed_text/embed_text.py:23-153](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/operations/embed_text/embed_text.py#L23-L153) · [packages/graphrag/graphrag/index/workflows/extract_graph.py:32-75](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/workflows/extract_graph.py#L32-L75) · [packages/graphrag-llm/graphrag_llm/middleware/with_rate_limiting.py:17-79](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-llm/graphrag_llm/middleware/with_rate_limiting.py#L17-L79) · [packages/graphrag/graphrag/query/context_builder/local_context.py:30-91](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/query/context_builder/local_context.py#L30-L91) · [packages/graphrag-llm/graphrag_llm/metrics/metrics_processor.py:21-59](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-llm/graphrag_llm/metrics/metrics_processor.py#L21-L59)

### semantica-agi/semantica (answered)

Semantica controls LLM cost and latency through several mechanisms. **Extraction-level caching** (`semantica/semantic_extract/cache.py`) stores entity, relation, and triplet extraction results keyed by text hash and method, with configurable TTL (default 3600s) and dual backends: in-memory LRU or persistent SQLite. The cache sits behind all extractor calls so repeated texts never reach the LLM. **Batching** is configurable in the default optimization config: `batch_size: 10`, `max_tokens_per_batch: 2000`, `enable_batching: True` (`semantica/semantic_extract/config.py:58-71`). The `max_workers` setting (default 8, auto-clamped to CPU count and capped at 32) controls parallelism in pipelines. **Model choice per stage** is fully configurable: the GraphBuilder allows separate selection of `ner_method`, `relation_method`, and `triplet_method` — each can be `"ml"` (spaCy, free), `"pattern"` (regex, free), or `"llm"` (paid). The default stack uses free local methods (`"ml"` for NER, `"pattern"` for relations/triplets), requiring no API keys (graph_builder.py:519-521). LLM enhancement is opt-in. The **LLM provider layer** (`semantica/llms/` and `semantica/semantic_extract/providers.py`) supports 10+ providers including local/cheap options: Ollama for local open-source models, Groq for fast cheap inference, and LiteLLM for routing to the cheapest available model. Each provider exposes `generate_typed()` for structured output via Pydantic schemas, reducing the need for multiple retries. **Community summarization** has a `max_tokens` budget (default 4000) that limits prompt size, and centrality-based token budgeting (`_pack_context`) prioritizes the most central entities and highest-impact child reports within that budget (`community_summarizer.py:1169-1728`). The `_extractive_fallback` method produces deterministic reports with no LLM call when no model is configured. **Global query** uses similar token budgeting (`max_context_tokens: 4000`, `response_token_budget: 600`) with dynamic level promotion — if level-0 community reports exceed budget, it promotes to coarser levels where fewer but denser reports exist (`global_retriever.py:428-534`). The Map phase runs concurrently across community reports with a configurable `max_workers` (default 4).


Citations: [semantica/semantic_extract/cache.py:1-45](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/semantic_extract/cache.py#L1-L45) · [semantica/semantic_extract/config.py:50-72](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/semantic_extract/config.py#L50-L72) · [semantica/kg/graph_builder.py:505-525](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/graph_builder.py#L505-L525) · [semantica/kg/community_summarizer.py:1169-1210](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/community_summarizer.py#L1169-L1210) · [semantica/context/global_retriever.py:428-470](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/context/global_retriever.py#L428-L470) · [semantica/context/global_retriever.py:844-870](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/context/global_retriever.py#L844-L870)

### neo4j-labs/llm-graph-builder (answered)

**Caching.** There is no disk or database caching of LLM extraction outputs; every `get_graph_from_llm()` call sends chunk text to the LLM. At the file level, the `GCS_FILE_CACHE` option (`main.py:56-60`) stores uploaded files in GCS instead of a local temp directory, which is a storage cache, not an LLM cache.

**Batching.** Chunks are processed in configurable batches controlled by `UPDATE_GRAPH_CHUNKS_PROCESSED` environment variable (default 20, `main.py:515`). Within each batch, multiple chunks are combined via `get_combined_chunks()` (`llm.py:158-182`) into a single document with configurable `chunks_to_combine` — meaning up to `UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine` chunks are sent to the LLM in one extraction call, reducing the number of LLM invocations proportionally.

**Model choice per stage.** Different models can be assigned to different pipeline stages. The default community-creation model is a smaller/cheaper model (`openai_gpt_5.4_mini`, `communities.py:17`). The `GRAPH_CLEANUP_MODEL` env var (`post_processing.py:157`) controls which LLM merges node labels. The chat response also uses a model-per-deployment while the extraction model is configured per-source. For embeddings, the default is the free/small `sentence-transformers/all-MiniLM-L6-v2` (384-dim) (`common_fn.py:719`).

**Token budgets.** When `TRACK_USER_USAGE` is enabled (`main.py:504`), the system enforces preconfigured daily (`DAILY_TOKENS_LIMIT`, default 250k) and monthly (`MONTHLY_TOKENS_LIMIT`, default 1M) token limits per user (`common_fn.py:491-492`). Before processing begins, `track_token_usage()` with `operation_type="precheck"` (`common_fn.py:460-477`) checks if the user has exceeded their limits and raises `LLMGraphBuilderException` if so. After each batch, actual token usage is persisted to a `User` node in a separate token-tracker Neo4j database.

**Non-LLM shortcuts.** The KNN `SIMILAR` relationship between chunks (`graphDB_dataAccess.py:184-195`) is computed purely via vector cosine similarity in Cypher (no LLM involved). Entity deduplication (`graphDB_dataAccess.py:470-538`) uses embedding similarity, substring matching, and Levenshtein distance — not an LLM. The local/global search modes that return raw chunk text or entity subgraphs without summarization avoid per-query LLM overhead for retrieval.

**Embedding filter at query.** The `EmbeddingsFilter` in the query pipeline (`QA_integration.py:316-320`) uses a `similarity_threshold` (default 0.10, `constants.py:247`) and `TokenTextSplitter` to pre-filter retrieved documents before the LLM sees them, reducing LLM context length and cost.

> **Editor's note.** Correction: batching does not put `UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine` chunks into one LLM call. Each extraction call covers `chunks_to_combine` concatenated chunks, so a batch of 20 chunks costs 20 / `chunks_to_combine` LLM calls; the batch size only controls how often progress and embeddings are written.

Citations: [backend/src/main.py:515-515](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L515-L515) · [backend/src/communities.py:15-18](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L15-L18) · [backend/src/shared/common_fn.py:443-477](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/common_fn.py#L443-L477) · [backend/src/QA_integration.py:302-325](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/QA_integration.py#L302-L325) · [backend/src/graphDB_dataAccess.py:182-196](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L182-L196) · [backend/src/post_processing.py:149-158](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/post_processing.py#L149-L158)

### OSU-NLP-Group/HippoRAG (answered)

**LLM response caching.** The `CacheOpenAI` class (src/hipporag/llm/openai_gpt.py:93–220) wraps every `infer()` call with `@cache_response` (33–91), which SHA-256-hashes the serialized request (messages, model, endpoint, generation params) and checks a SQLite cache. On hit it returns `(message, metadata, True)` with zero API cost. The cache file is stored per-LLM name under `save_dir/llm_cache/`. Cache-aware token accounting (prompt/completion/billable vs logical) is displayed during batch OpenIE (src/hipporag/information_extraction/openie_openai.py:222–238), separating hit vs miss costs.

**Batching.** Embeddings are batched through `embedding_batch_size` (default 16, config_utils.py:140–141). OpenIE runs via `ThreadPoolExecutor` with `openie_max_workers` (default 8, config_utils.py:311), allowing concurrent NER and triple-extraction calls. KNN for synonymy edges is batched with configurable `synonymy_edge_query_batch_size` (default 1000) and `synonymy_edge_key_batch_size` (default 10000, config_utils.py:164–170).

**Model choice per stage.** The extraction/OpenIE LLM (`extraction_llm`) and the QA LLM (`qa_llm`) are independently configurable (HippoRAG.py:97–98, 220–229). By default they share the same model, but users can pass a cheaper model for extraction and a stronger one for QA. The embedding model is also independently configurable (`embedding_model_name`, default `nvidia/NV-Embed-v2`, config_utils.py:136).

**Non-LLM shortcuts.** The DSPy reranking filter (rerank.py) is the only LLM call at query time beyond QA — it uses the extraction LLM with `max_new_tokens=512` to filter the candidate fact list. PPR is pure NumPy + iGraph C-extension (no LLM). Fact scoring is a simple `np.dot` (HippoRAG.py:1926). When no facts survive reranking, the system falls back to dense passage retrieval which is entirely embedding-based (1930–1967), avoiding any LLM call for that query.

**Token budgets.** `openie_ner_max_tokens=2048` and `openie_triple_max_tokens=4096` are configurable (config_utils.py:315–322). The `DSPyFilter` uses `max_new_tokens=512` (rerank.py:36). QA inference uses `max_new_tokens=2048` (config_utils.py:38). The model is configured with `temperature=0` by default for deterministic extraction.


Citations: [src/hipporag/llm/openai_gpt.py:33-91](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/llm/openai_gpt.py#L33-L91) · [src/hipporag/information_extraction/openie_openai.py:210-286](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/information_extraction/openie_openai.py#L210-L286) · [src/hipporag/HippoRAG.py:226-229](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L226-L229) · [src/hipporag/utils/config_utils.py:136-170](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/utils/config_utils.py#L136-L170) · [src/hipporag/rerank.py:36-36](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/rerank.py#L36-L36) · [src/hipporag/HippoRAG.py:760-764](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L760-L764)

### gusye1234/nano-graphrag (answered)

**Two-tier model routing.** The system separates a `best_model_func` (default gpt-4o) for expensive core tasks — entity extraction, query response generation, community report generation — from a `cheap_model_func` (default gpt-4o-mini) used only for entity/relation description summarization when descriptions exceed `entity_summary_to_max_tokens`. Both functions are wrapped by `limit_async_func_call` which caps concurrent requests (`best_model_max_async` and `cheap_model_max_async`, both default 16) via a spin-loop semaphore (`_utils.py:276-295`) rather than `asyncio.Semaphore` for compatibility with nested event loops.

**LLM response caching.** All LLM calls pass through `openai_complete_if_cache` (`_llm.py:50-75`). Before making an API call, it computes an MD5 hash of `(model, messages)` via `compute_args_hash` and checks the `hashing_kv` (a `JsonKVStorage` persisted as `kv_store_llm_response_cache.json`). On a cache hit, the stored response is returned with zero latency and zero cost. On a miss, the API is called and the result is stored. The cache is checked on every LLM call — extraction, summarization, community report generation, and query — and persists across sessions. The `enable_llm_cache` flag (default True) controls whether the cache storage is created.

**Retry with backoff.** Network calls to OpenAI, Azure OpenAI, and Amazon Bedrock are decorated with `tenacity.retry` using exponential backoff (`wait_exponential` multiplier=1s, min=4s, max=10s) and up to 5 attempts, retrying only on `RateLimitError` and `APIConnectionError` (`_llm.py:45-49`). This avoids wasting tokens on failed calls and handles API throttling gracefully.

**Token budget truncation.** Throughout retrieval and context assembly, `truncate_list_by_token_size` (`_utils.py:169-183`) is called repeatedly with configurable `max_token_size` parameters: `local_max_token_for_text_unit` (4000), `local_max_token_for_local_context` (4800), `local_max_token_for_community_report` (3200), `global_max_token_for_community_report` (16384), `naive_max_token_for_text_unit` (12000). These per-section caps bound the prompt size sent to the LLM, preventing context overflow and limiting per-call cost.

**Embedding batching.** Embedding requests are batched by `embedding_batch_num` (default 32) in `NanoVectorDBStorage.upsert` (`_storage/vdb_nanovectordb.py:40-47`). Batches are sent concurrently via `asyncio.gather`, with an overall concurrency limit of `embedding_func_max_async` (default 16) enforced by the wrapper on the embedding function. This optimizes throughput while respecting API rate limits.


Citations: [nano_graphrag/_llm.py:45-75](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_llm.py#L45-L75) · [nano_graphrag/_utils.py:169-183](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_utils.py#L169-L183) · [nano_graphrag/_utils.py:276-295](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_utils.py#L276-L295) · [nano_graphrag/_storage/vdb_nanovectordb.py:28-52](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_storage/vdb_nanovectordb.py#L28-L52) · [nano_graphrag/graphrag.py:108-123](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/graphrag.py#L108-L123) · [nano_graphrag/graphrag.py:181-186](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/graphrag.py#L181-L186)

### pingcap/autoflow (answered)

AutoFlow employs several strategies to control LLM cost and latency.

**Two-tier LLM model:** The `ChatEngineConfig` distinguishes between a primary `llm` (used for the final answer synthesis) and a `fast_llm` (used for cheaper operations like question refinement, clarification, and intent generation) (`backend/app/rag/chat/config.py:133-158`). The fast LLM is typically a smaller/cheaper model, while the primary LLM handles the expensive QA generation.

**Embedding model for similarity, not LLM for most retrieval:** The core graph retrieval is driven by **vector similarity** (cosine distance on TiDB HNSW indexes), not by LLM calls. The `WeightedGraphRetriever` (`core/autoflow/knowledge_graph/retrievers/weighted.py`) and `TiDBGraphStore.search_relationships_weight()` (`backend/.../tidb_graph_store.py:632-753`) use embedding distance as the primary signal, only ranking results with the weighted formula. The LLM is only invoked during extraction (indexing time) and during the final response generation (query time).

**Entity deduplication with conditional LLM merge:** The `get_or_create_entity()` method first checks cosine similarity of description embeddings. Only when an entity with the same name is found *and* its description/metadata are different does it call the `MergeEntities` DSPy program (an LLM call) — otherwise it simply reuses the existing row (`tidb_graph_store.py:397-442`). This avoids an LLM call for duplicate entities.

**Singleflight cache for model factories:** The `@singleflight_cache` decorator on `get_dynamic_entity_model`, `get_dynamic_relationship_model`, and `get_dynamic_chunk_model` ensures that per-namespace model creation (which involves schema reflection) runs only once across concurrent threads (`backend/app/utils/singleflight_cache.py:5-45`).

**Semantic cache for external engines:** The `SemanticCacheManager` (`backend/app/rag/semantic_cache/base.py:91-199`) stores question-answer pairs in TiDB with vector embeddings. Before invoking the external chat engine, the system checks if a semantically similar question has been answered before, using cosine-distance search (threshold 0.5) followed by a DSPy `QASemanticSearchModule` to classify exact/similar/no-match. This avoids redundant LLM calls for repeated queries.

**Indexing batching via Celery:** Indexing tasks are dispatched to Celery workers (`build_index_for_document.delay(...)`, `build_kg_index_for_chunk.delay(...)`), enabling parallel processing of documents/chunks across workers (`backend/app/tasks/build_index.py`).

**No per-chunk LLM call deduplication:** Every chunk that has not been indexed (no existing relationships for its `chunk_id`) gets a fresh extraction call, which is the main cost driver at indexing time.

**Configurable model choice:** The system uses OpenAI by default (GPT-4o for indexing extraction, text-embedding-3-small for embeddings), but the `LLMConfig` and `EmbeddingModelConfig` abstractions support pluggable providers (`core/autoflow/configs/models/llms/base.py:6-13`, `embeddings/base.py`), so operators can choose cheaper models.


Citations: [backend/app/rag/chat/config.py:147-158](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/config.py#L147-L158) · [backend/app/rag/semantic_cache/base.py:91-199](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/semantic_cache/base.py#L91-L199) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:367-477](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L367-L477) · [core/autoflow/knowledge_graph/retrievers/weighted.py:45-68](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/retrievers/weighted.py#L45-L68) · [backend/app/utils/singleflight_cache.py:5-45](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/utils/singleflight_cache.py#L5-L45)

### trustgraph-ai/trustgraph (answered)

**Caching.** Graph RAG uses a per-request `LRUCacheWithTTL` (`retrieval/graph_rag/graph_rag.py:94-134`) for entity label lookups (max_size=5000, TTL=300s). OntoRAG has `CacheManager` (`query/ontology/cache.py:463-610`) with in-memory/file backends, TTL expiry, LRU eviction, and `@cache_result` decorator. Ontology loading cached per flow component (`extract/kg/ontology/extract.py:237-243`). IAM auth cached on gateway. SPARQL EXISTS cached (`query/sparql/algebra.py:410-417`).

**Batching.** Extraction processors batch output (`triples_batch_size=50`, `entity_batch_size=5`). `OntologyEmbedder` (`ontology_embedder.py:149-170`) batches at 50. Embeddings client accepts lists. Graph RAG does concurrent batch traversal (`graph_rag.py:269-316`).

**Model choice per stage.** `LlmService` (`base/llm_service.py:81-91`) reads a `model` flow parameter per request, configurable per pipeline stage. Supports API (OpenAI, Claude, Azure, Bedrock, Vertex AI) and local (vLLM, TGI, Ollama, LM Studio, llamafile) models.

**Token budgets.** Graph traversal caps at `max_reranker_input=350` edges per hop, `max_reranker_text_length=240` chars (`retrieval/graph_rag/rag.py:40-43`). Only `edge_limit=25` edges reach synthesis. Document RAG limits `doc_limit` (`retrieval/document_rag/rag.py:31-40`). Response objects carry `in_token`/`out_token`/`model`.

**Non-LLM shortcuts.** Ontology selector bypassed when <5 elements (`extract/kg/ontology/extract.py:161`). FTS5 BM25 provides zero-LLM-cost sparse retrieval. OntoRAG answers SPARQL/Cypher queries structurally. Cross-encoder reranker is a small encoder model.

> **Editor's note.** Correction: the 25-edge cap is per hop, not per query; with the default max_path_length of 2, synthesis can receive up to 50 edges plus source-document metadata edges.

Citations: [trustgraph-flow/trustgraph/retrieval/graph_rag/graph_rag.py:94-134](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/retrieval/graph_rag/graph_rag.py#L94-L134) · [trustgraph-flow/trustgraph/query/ontology/cache.py:94-135](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/query/ontology/cache.py#L94-L135) · [trustgraph-flow/trustgraph/extract/kg/ontology/extract.py:237-243](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/extract.py#L237-L243) · [trustgraph-base/trustgraph/base/llm_service.py:80-92](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-base/trustgraph/base/llm_service.py#L80-L92) · [trustgraph-flow/trustgraph/retrieval/graph_rag/rag.py:35-43](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/retrieval/graph_rag/rag.py#L35-L43) · [trustgraph-flow/trustgraph/extract/kg/ontology/extract.py:158-162](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/extract.py#L158-L162)

### zilliztech/vector-graph-rag (answered)

Cost and latency are managed through several mechanisms:

**LLM Response Caching** (`llm/cache.py:14-165`): Every LLM call (triplet extraction, entity extraction, reranking, answer generation) is cached to disk. The cache uses MD5 hashing of the prompt with temperature as a cache key, so identical prompts hit the cache. Extraction and NER each have their own caching paths (`extractor.py:147-168` and `extractor.py:369-395`).

**Model choice per stage**: The default `llm_model` is `gpt-4o-mini` (`config.py:28-31`), a small, cheap model. The reranker and extractor reuse the same model unless overridden — there is no separate model parameter for different pipeline stages. The `AnswerGenerator` uses the same model. Temperature is 0 for deterministic outputs (`config.py:118`).

**Batching**: Embedding generation uses batch calls with a configurable `batch_size` (default 32, `config.py:133`). Database inserts use the same batching. The eviction strategy (`retriever.py:264-336`) limits the reranker input to at most `relation_number_threshold` (default 1000) relations, preventing unbounded prompt sizes.

**Optional Jev reranker**: Instead of the LLM-based `LLMReranker`, the system supports `JevReranker` (`llm/jev.py:27-33`), a specialized high-throughput scoring model that may be cheaper per-relation than an LLM call. Configured via `reranker_provider: "jev"` (`config.py:123-128`).

**Single-pass design**: The project's key algorithmic bet is replacing iterative multi-step LLM agent loops (as in IRCoT or GraphRAG iterative reflection) with a single-pass LLM reranking step. This caps the number of LLM calls per query to: NER (1), reranking (1), answer generation (1) — versus potentially dozens in iterative approaches. The CLAUDE.md explicitly states this philosophy.

**NER Cache TSV**: Query-time entity extraction supports a pre-computed NER cache in HippoRAG TSV format (`extractor.py:296-332`), eliminating LLM calls for known questions in evaluation settings.


Citations: [src/vector_graph_rag/llm/extractor.py:138-168](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/llm/extractor.py#L138-L168) · [src/vector_graph_rag/graph/retriever.py:264-336](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/graph/retriever.py#L264-L336) · [src/vector_graph_rag/config.py:28-31](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/config.py#L28-L31) · [src/vector_graph_rag/llm/jev.py:27-33](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/llm/jev.py#L27-L33)
