LLMs Technical Reviews

How are LLM cost and latency controlled during indexing and query?

Caching; batching; model choice per stage; token budgets; small-model or non-LLM shortcuts.

Verdict

The cheapest options avoid the LLM where they can: graphify parses code with no model, and Semantica extracts with spaCy by default. Among full LLM pipelines, LightRAG, HippoRAG and Vector Graph RAG keep each query to a few calls. GraphRAG spends most of its budget at index time.

Heavy indexing. GraphRAG runs extraction plus gleaning on every chunk, a summarization call for merged descriptions, and one report per community. Each stage has its own model, and its fast method replaces LLM extraction with NLP. Only indexing calls are cached. nano-graphrag splits work between a “best” model and a “cheap” model and caches every call, including queries, but each insert regenerates all community reports. LLM Graph Builder makes one call per chunks_to_combine chunks and has no extraction cache. Optional per-user daily and monthly token limits can stop a job before it starts. AutoFlow makes two DSPy calls per chunk plus LLM merge calls for entities, and uses a separate fast_llm to rewrite questions.

Lean indexing, cheap queries. LightRAG writes no reports, caches both extraction and answers, and routes extraction, keyword and answer calls to different models. Its answer cache ignores the retrieved context, so a cached answer can survive changes to the data. HippoRAG caches every request in SQLite. A query costs one filter call plus a PageRank solve. Vector Graph RAG caches every LLM call on disk, uses gpt-4o-mini for every step with no per-stage choice, and makes three calls per query. TrustGraph makes two LLM calls per query and filters edges with a MiniLM cross-encoder. It does not cache extraction.

Skipping the LLM. graphify packs files for its LLM pass into chunks of up to 60,000 tokens. When a response is truncated, it splits the chunk in half and retries, up to three levels deep. It caches results per file hash. Semantica’s LLM stages are opt-in, it caches extraction with a TTL of 3,600 seconds, and it can write extractive community reports with no model.

Pick: graphify or Semantica’s defaults when the budget is close to zero. Pick: LightRAG or HippoRAG for low running cost with LLM-quality extraction. Pick: GraphRAG when global answers justify paying for indexing up front.

Per-project answers

Graphify-Labs/graphify

answered

Token-budget chunk packing. The semantic extraction pipeline (extract_corpus_parallel, llm.py:2567) packs files into chunks using a configurable token_budget (default 60,000 tokens). Each file's token cost is estimated (characters/4 for code, 1,600 fixed per image), and chunks are closed when adding the next file would exceed the budget (llm.py:2160–2203). This prevents single over-large LLM requests that waste money on truncated retries.

Adaptive retry with bisection. When an LLM response hits finish_reason="length" (truncation), _extract_with_adaptive_retry() (llm.py:2319) bisects the chunk and re-extracts each half recursively, up to a configurable max_retry_depth (default 3, controlled by GRAPHIFY_MAX_RETRY_DEPTH). This bounds worst-case per-chunk cost at 2^depth calls. Setting depth to 0 disables all retries — one call per chunk, full stop.

Hollow-response backoff, not bisection. A hollow response (HTTP 200 with empty/unparseable content) is retried with backoff (2s, 8s) rather than bisected, because bisecting a backend issue cannot converge and would cost far more (llm.py:2406–2421). After 3 attempts the chunk fails loudly.

Model choice per stage. The BACKENDS dict (llm.py:104–224) defines model defaults per provider (claude, openai, gemini, kimi, ollama, deepseek, azure, bedrock, claude-cli) with configurable pricing per megatoken. Users can shift cost by selecting cheaper models (e.g., Ollama/local models cost $0). Model overrides are settable per provider via env vars (ANTHROPIC_MODEL, GRAPHIFY_OPENAI_MODEL, etc.).

Caching avoids re-extraction. Both the AST cache (versioned content-addressed) and semantic cache (content-hash + prompt-fingerprinted) skip re-extraction of unchanged files, eliminating LLM calls entirely for files that haven't changed since the last run.

Concurrency control. The thread pool defaults to 4 concurrent chunks (max_concurrency, llm.py:2594). Ollama defaults to serial (1) because concurrent requests cause VRAM pressure and hollow responses (llm.py:2671–2672). The GRAPHIFY_MAX_RETRIES env var controls SDK-level retries on rate limits, defaulting to 6 (llm.py:430).

Timeout guard. GRAPHIFY_API_TIMEOUT (default 600s) caps total request wall-clock time per chunk (llm.py:410–420). A timed-out chunk triggers adaptive bisection rather than silent failure.

HKUDS/LightRAG

answered

LLM response caching (primary cost control). Both entity extraction and query results are cached. enable_llm_cache (default True) caches query answers; enable_llm_cache_for_entity_extract (default True) caches extraction results (lightrag/lightrag.py:1093-1097). The cache is implemented in handle_cache / save_to_cache (lightrag/utils.py:4448-4520), storing flattened cache keys of format {mode}:{cache_type}:{hash} in llm_response_cache (a KV storage). On cache hit, the LLM call is entirely skipped — response is returned from the stored value. The extraction cache also has per-chunk write-ahead semantics to avoid orphaned cache rows.

Concurrency batching. The number of concurrent LLM extraction calls per document is controlled by llm_model_max_async (default 4) via asyncio.Semaphore(chunk_max_async) at lightrag/operate.py:4547-4548. Document-level parallelism is capped by max_parallel_insert (default 3) at lightrag/pipeline.py:2653. Embedding requests are batched by embedding_batch_num at lightrag/kg/nano_vector_db_impl.py:133 and various backend implementations.

Role-based model routing. LightRAG supports separate LLM bindings per role (extract, keyword, query, vlm) via RoleLLMConfig (lightrag/llm_roles.py:57-63). Each role can use a different model (e.g., a cheaper/faster model for extraction, a more capable one for query), configured via env vars like EXTRACT_LLM_BINDING, QUERY_LLM_BINDING. This lets operators trade cost vs. quality per stage.

Token budgets. Multiple token budgets limit LLM context size at key stages:

  • MAX_EXTRACT_INPUT_TOKENS limits the gleaning step input (default 128K, lightrag/operate.py:3992-3996).
  • At query time, max_entity_tokens, max_relation_tokens, max_total_tokens limit what goes into the final prompt (_apply_token_truncation at lightrag/operate.py:5501).
  • Entity/relation descriptions are chunked and recursively summarized via _handle_entity_relation_summary map-reduce (lightrag/operate.py:372) when they exceed token limits, avoiding giant prompts.
  • Document chunks that exceed the embedding model's token limit are truncated by _truncate_vdb_content (lightrag/operate.py:298).

Non-LLM shortcuts. The entire relation weight contract (docs/ProgramingWithCore.md) avoids LLM calls by simply counting distinct source IDs for the weight floor. The kg_chunk_pick_method (DEFAULT_KG_CHUNK_PICK_METHOD) supports weighted polling (based on occurrence count in entity/relation source_id lists) as a non-LLM alternative to vector-similarity chunk selection. The NaiveQuery mode bypasses the KG entirely and does a vector-only search. bypass mode skips all retrieval.

Editor's note. Correction: the default MAX_EXTRACT_INPUT_TOKENS is 20,480 (constants.py), not 128K. Note also that the query-answer cache key does not include the retrieved context, so a cached answer can outlive changes to the data.

microsoft/graphrag

answered

LLM response caching. The with_cache middleware wraps every LLM completion and embedding call. Before calling the model, it computes a cache key from full request args, checks cache.get(key), and on a hit returns the cached response. Applies to both indexing and query. Streaming and mocked responses are excluded.

Batching. embed_text buffers rows and dispatches batches at configurable batch_size and batch_max_tokens, with flush size sized to saturate num_threads × batch_size. For LLM completions, derive_from_rows parallelizes across num_threads.

Model choice per stage. Separate model IDs per stage: extract_graph.completion_model_id, summarize_descriptions.completion_model_id, community_reports.completion_model_id, embed_text.embedding_model_id. Each can be a cheap/fast model via litellm model strings.

Token budgets. All context builders enforce max_context_tokens=8000. build_entity_context truncates when cumulative tokens exceed budget. Community reports cap max_input_length/max_report_length. Global search enforces max_data_tokens=8000 and per-stage max-length params. max_gleanings (default configurable) controls extra LLM rounds during extraction — setting it to 0 skips gleaning.

Rate limiting and retries. with_rate_limiting uses a sliding-window rate limiter per model, counting tokens. with_retries adds exponential backoff. Both compose via middleware pipeline.

Metrics. MetricsProcessor and MetricsStore track token usage, response times, and cache hit rates. Query results return per-category llm_calls, prompt_tokens, and output_tokens.

Editor's note. Correction: only indexing calls are cached. The query engines are built with create_completion(model_settings) and no cache, so query calls are never cached. The configured context budget is 12,000 tokens, not 8,000.

semantica-agi/semantica

answered

Semantica controls LLM cost and latency through several mechanisms. Extraction-level caching (semantica/semantic_extract/cache.py) stores entity, relation, and triplet extraction results keyed by text hash and method, with configurable TTL (default 3600s) and dual backends: in-memory LRU or persistent SQLite. The cache sits behind all extractor calls so repeated texts never reach the LLM. Batching is configurable in the default optimization config: batch_size: 10, max_tokens_per_batch: 2000, enable_batching: True (semantica/semantic_extract/config.py:58-71). The max_workers setting (default 8, auto-clamped to CPU count and capped at 32) controls parallelism in pipelines. Model choice per stage is fully configurable: the GraphBuilder allows separate selection of ner_method, relation_method, and triplet_method — each can be "ml" (spaCy, free), "pattern" (regex, free), or "llm" (paid). The default stack uses free local methods ("ml" for NER, "pattern" for relations/triplets), requiring no API keys (graph_builder.py:519-521). LLM enhancement is opt-in. The LLM provider layer (semantica/llms/ and semantica/semantic_extract/providers.py) supports 10+ providers including local/cheap options: Ollama for local open-source models, Groq for fast cheap inference, and LiteLLM for routing to the cheapest available model. Each provider exposes generate_typed() for structured output via Pydantic schemas, reducing the need for multiple retries. Community summarization has a max_tokens budget (default 4000) that limits prompt size, and centrality-based token budgeting (_pack_context) prioritizes the most central entities and highest-impact child reports within that budget (community_summarizer.py:1169-1728). The _extractive_fallback method produces deterministic reports with no LLM call when no model is configured. Global query uses similar token budgeting (max_context_tokens: 4000, response_token_budget: 600) with dynamic level promotion — if level-0 community reports exceed budget, it promotes to coarser levels where fewer but denser reports exist (global_retriever.py:428-534). The Map phase runs concurrently across community reports with a configurable max_workers (default 4).

neo4j-labs/llm-graph-builder

answered

Caching. There is no disk or database caching of LLM extraction outputs; every get_graph_from_llm() call sends chunk text to the LLM. At the file level, the GCS_FILE_CACHE option (main.py:56-60) stores uploaded files in GCS instead of a local temp directory, which is a storage cache, not an LLM cache.

Batching. Chunks are processed in configurable batches controlled by UPDATE_GRAPH_CHUNKS_PROCESSED environment variable (default 20, main.py:515). Within each batch, multiple chunks are combined via get_combined_chunks() (llm.py:158-182) into a single document with configurable chunks_to_combine — meaning up to UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine chunks are sent to the LLM in one extraction call, reducing the number of LLM invocations proportionally.

Model choice per stage. Different models can be assigned to different pipeline stages. The default community-creation model is a smaller/cheaper model (openai_gpt_5.4_mini, communities.py:17). The GRAPH_CLEANUP_MODEL env var (post_processing.py:157) controls which LLM merges node labels. The chat response also uses a model-per-deployment while the extraction model is configured per-source. For embeddings, the default is the free/small sentence-transformers/all-MiniLM-L6-v2 (384-dim) (common_fn.py:719).

Token budgets. When TRACK_USER_USAGE is enabled (main.py:504), the system enforces preconfigured daily (DAILY_TOKENS_LIMIT, default 250k) and monthly (MONTHLY_TOKENS_LIMIT, default 1M) token limits per user (common_fn.py:491-492). Before processing begins, track_token_usage() with operation_type="precheck" (common_fn.py:460-477) checks if the user has exceeded their limits and raises LLMGraphBuilderException if so. After each batch, actual token usage is persisted to a User node in a separate token-tracker Neo4j database.

Non-LLM shortcuts. The KNN SIMILAR relationship between chunks (graphDB_dataAccess.py:184-195) is computed purely via vector cosine similarity in Cypher (no LLM involved). Entity deduplication (graphDB_dataAccess.py:470-538) uses embedding similarity, substring matching, and Levenshtein distance — not an LLM. The local/global search modes that return raw chunk text or entity subgraphs without summarization avoid per-query LLM overhead for retrieval.

Embedding filter at query. The EmbeddingsFilter in the query pipeline (QA_integration.py:316-320) uses a similarity_threshold (default 0.10, constants.py:247) and TokenTextSplitter to pre-filter retrieved documents before the LLM sees them, reducing LLM context length and cost.

Editor's note. Correction: batching does not put UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine chunks into one LLM call. Each extraction call covers chunks_to_combine concatenated chunks, so a batch of 20 chunks costs 20 / chunks_to_combine LLM calls; the batch size only controls how often progress and embeddings are written.

OSU-NLP-Group/HippoRAG

answered

LLM response caching. The CacheOpenAI class (src/hipporag/llm/openai_gpt.py:93–220) wraps every infer() call with @cache_response (33–91), which SHA-256-hashes the serialized request (messages, model, endpoint, generation params) and checks a SQLite cache. On hit it returns (message, metadata, True) with zero API cost. The cache file is stored per-LLM name under save_dir/llm_cache/. Cache-aware token accounting (prompt/completion/billable vs logical) is displayed during batch OpenIE (src/hipporag/information_extraction/openie_openai.py:222–238), separating hit vs miss costs.

Batching. Embeddings are batched through embedding_batch_size (default 16, config_utils.py:140–141). OpenIE runs via ThreadPoolExecutor with openie_max_workers (default 8, config_utils.py:311), allowing concurrent NER and triple-extraction calls. KNN for synonymy edges is batched with configurable synonymy_edge_query_batch_size (default 1000) and synonymy_edge_key_batch_size (default 10000, config_utils.py:164–170).

Model choice per stage. The extraction/OpenIE LLM (extraction_llm) and the QA LLM (qa_llm) are independently configurable (HippoRAG.py:97–98, 220–229). By default they share the same model, but users can pass a cheaper model for extraction and a stronger one for QA. The embedding model is also independently configurable (embedding_model_name, default nvidia/NV-Embed-v2, config_utils.py:136).

Non-LLM shortcuts. The DSPy reranking filter (rerank.py) is the only LLM call at query time beyond QA — it uses the extraction LLM with max_new_tokens=512 to filter the candidate fact list. PPR is pure NumPy + iGraph C-extension (no LLM). Fact scoring is a simple np.dot (HippoRAG.py:1926). When no facts survive reranking, the system falls back to dense passage retrieval which is entirely embedding-based (1930–1967), avoiding any LLM call for that query.

Token budgets. openie_ner_max_tokens=2048 and openie_triple_max_tokens=4096 are configurable (config_utils.py:315–322). The DSPyFilter uses max_new_tokens=512 (rerank.py:36). QA inference uses max_new_tokens=2048 (config_utils.py:38). The model is configured with temperature=0 by default for deterministic extraction.

gusye1234/nano-graphrag

answered

Two-tier model routing. The system separates a best_model_func (default gpt-4o) for expensive core tasks — entity extraction, query response generation, community report generation — from a cheap_model_func (default gpt-4o-mini) used only for entity/relation description summarization when descriptions exceed entity_summary_to_max_tokens. Both functions are wrapped by limit_async_func_call which caps concurrent requests (best_model_max_async and cheap_model_max_async, both default 16) via a spin-loop semaphore (_utils.py:276-295) rather than asyncio.Semaphore for compatibility with nested event loops.

LLM response caching. All LLM calls pass through openai_complete_if_cache (_llm.py:50-75). Before making an API call, it computes an MD5 hash of (model, messages) via compute_args_hash and checks the hashing_kv (a JsonKVStorage persisted as kv_store_llm_response_cache.json). On a cache hit, the stored response is returned with zero latency and zero cost. On a miss, the API is called and the result is stored. The cache is checked on every LLM call — extraction, summarization, community report generation, and query — and persists across sessions. The enable_llm_cache flag (default True) controls whether the cache storage is created.

Retry with backoff. Network calls to OpenAI, Azure OpenAI, and Amazon Bedrock are decorated with tenacity.retry using exponential backoff (wait_exponential multiplier=1s, min=4s, max=10s) and up to 5 attempts, retrying only on RateLimitError and APIConnectionError (_llm.py:45-49). This avoids wasting tokens on failed calls and handles API throttling gracefully.

Token budget truncation. Throughout retrieval and context assembly, truncate_list_by_token_size (_utils.py:169-183) is called repeatedly with configurable max_token_size parameters: local_max_token_for_text_unit (4000), local_max_token_for_local_context (4800), local_max_token_for_community_report (3200), global_max_token_for_community_report (16384), naive_max_token_for_text_unit (12000). These per-section caps bound the prompt size sent to the LLM, preventing context overflow and limiting per-call cost.

Embedding batching. Embedding requests are batched by embedding_batch_num (default 32) in NanoVectorDBStorage.upsert (_storage/vdb_nanovectordb.py:40-47). Batches are sent concurrently via asyncio.gather, with an overall concurrency limit of embedding_func_max_async (default 16) enforced by the wrapper on the embedding function. This optimizes throughput while respecting API rate limits.

pingcap/autoflow

answered

AutoFlow employs several strategies to control LLM cost and latency.

Two-tier LLM model: The ChatEngineConfig distinguishes between a primary llm (used for the final answer synthesis) and a fast_llm (used for cheaper operations like question refinement, clarification, and intent generation) (backend/app/rag/chat/config.py:133-158). The fast LLM is typically a smaller/cheaper model, while the primary LLM handles the expensive QA generation.

Embedding model for similarity, not LLM for most retrieval: The core graph retrieval is driven by vector similarity (cosine distance on TiDB HNSW indexes), not by LLM calls. The WeightedGraphRetriever (core/autoflow/knowledge_graph/retrievers/weighted.py) and TiDBGraphStore.search_relationships_weight() (backend/.../tidb_graph_store.py:632-753) use embedding distance as the primary signal, only ranking results with the weighted formula. The LLM is only invoked during extraction (indexing time) and during the final response generation (query time).

Entity deduplication with conditional LLM merge: The get_or_create_entity() method first checks cosine similarity of description embeddings. Only when an entity with the same name is found and its description/metadata are different does it call the MergeEntities DSPy program (an LLM call) — otherwise it simply reuses the existing row (tidb_graph_store.py:397-442). This avoids an LLM call for duplicate entities.

Singleflight cache for model factories: The @singleflight_cache decorator on get_dynamic_entity_model, get_dynamic_relationship_model, and get_dynamic_chunk_model ensures that per-namespace model creation (which involves schema reflection) runs only once across concurrent threads (backend/app/utils/singleflight_cache.py:5-45).

Semantic cache for external engines: The SemanticCacheManager (backend/app/rag/semantic_cache/base.py:91-199) stores question-answer pairs in TiDB with vector embeddings. Before invoking the external chat engine, the system checks if a semantically similar question has been answered before, using cosine-distance search (threshold 0.5) followed by a DSPy QASemanticSearchModule to classify exact/similar/no-match. This avoids redundant LLM calls for repeated queries.

Indexing batching via Celery: Indexing tasks are dispatched to Celery workers (build_index_for_document.delay(...), build_kg_index_for_chunk.delay(...)), enabling parallel processing of documents/chunks across workers (backend/app/tasks/build_index.py).

No per-chunk LLM call deduplication: Every chunk that has not been indexed (no existing relationships for its chunk_id) gets a fresh extraction call, which is the main cost driver at indexing time.

Configurable model choice: The system uses OpenAI by default (GPT-4o for indexing extraction, text-embedding-3-small for embeddings), but the LLMConfig and EmbeddingModelConfig abstractions support pluggable providers (core/autoflow/configs/models/llms/base.py:6-13, embeddings/base.py), so operators can choose cheaper models.

trustgraph-ai/trustgraph

answered

Caching. Graph RAG uses a per-request LRUCacheWithTTL (retrieval/graph_rag/graph_rag.py:94-134) for entity label lookups (max_size=5000, TTL=300s). OntoRAG has CacheManager (query/ontology/cache.py:463-610) with in-memory/file backends, TTL expiry, LRU eviction, and @cache_result decorator. Ontology loading cached per flow component (extract/kg/ontology/extract.py:237-243). IAM auth cached on gateway. SPARQL EXISTS cached (query/sparql/algebra.py:410-417).

Batching. Extraction processors batch output (triples_batch_size=50, entity_batch_size=5). OntologyEmbedder (ontology_embedder.py:149-170) batches at 50. Embeddings client accepts lists. Graph RAG does concurrent batch traversal (graph_rag.py:269-316).

Model choice per stage. LlmService (base/llm_service.py:81-91) reads a model flow parameter per request, configurable per pipeline stage. Supports API (OpenAI, Claude, Azure, Bedrock, Vertex AI) and local (vLLM, TGI, Ollama, LM Studio, llamafile) models.

Token budgets. Graph traversal caps at max_reranker_input=350 edges per hop, max_reranker_text_length=240 chars (retrieval/graph_rag/rag.py:40-43). Only edge_limit=25 edges reach synthesis. Document RAG limits doc_limit (retrieval/document_rag/rag.py:31-40). Response objects carry in_token/out_token/model.

Non-LLM shortcuts. Ontology selector bypassed when <5 elements (extract/kg/ontology/extract.py:161). FTS5 BM25 provides zero-LLM-cost sparse retrieval. OntoRAG answers SPARQL/Cypher queries structurally. Cross-encoder reranker is a small encoder model.

Editor's note. Correction: the 25-edge cap is per hop, not per query; with the default max_path_length of 2, synthesis can receive up to 50 edges plus source-document metadata edges.

zilliztech/vector-graph-rag

answered

Cost and latency are managed through several mechanisms:

LLM Response Caching (llm/cache.py:14-165): Every LLM call (triplet extraction, entity extraction, reranking, answer generation) is cached to disk. The cache uses MD5 hashing of the prompt with temperature as a cache key, so identical prompts hit the cache. Extraction and NER each have their own caching paths (extractor.py:147-168 and extractor.py:369-395).

Model choice per stage: The default llm_model is gpt-4o-mini (config.py:28-31), a small, cheap model. The reranker and extractor reuse the same model unless overridden — there is no separate model parameter for different pipeline stages. The AnswerGenerator uses the same model. Temperature is 0 for deterministic outputs (config.py:118).

Batching: Embedding generation uses batch calls with a configurable batch_size (default 32, config.py:133). Database inserts use the same batching. The eviction strategy (retriever.py:264-336) limits the reranker input to at most relation_number_threshold (default 1000) relations, preventing unbounded prompt sizes.

Optional Jev reranker: Instead of the LLM-based LLMReranker, the system supports JevReranker (llm/jev.py:27-33), a specialized high-throughput scoring model that may be cheaper per-relation than an LLM call. Configured via reranker_provider: "jev" (config.py:123-128).

Single-pass design: The project's key algorithmic bet is replacing iterative multi-step LLM agent loops (as in IRCoT or GraphRAG iterative reflection) with a single-pass LLM reranking step. This caps the number of LLM calls per query to: NER (1), reranking (1), answer generation (1) — versus potentially dozens in iterative approaches. The CLAUDE.md explicitly states this philosophy.

NER Cache TSV: Query-time entity extraction supports a pre-computed NER cache in HippoRAG TSV format (extractor.py:296-332), eliminating LLM calls for known questions in evaluation settings.

← How are updates and incremental indexing handled?