LLMs Technical Reviews

How does query-time retrieval use the graph?

Local / global / hybrid modes; traversal; combining graph and vector hits; how context is assembled for the LLM.

Verdict

For broad questions about a whole corpus, GraphRAG’s global and DRIFT search are the most complete. For multi-hop questions over passages, HippoRAG’s Personalized PageRank and Vector Graph RAG’s single rerank call are the cheapest paths.

Community reports plus local search. GraphRAG has local (12,000-token context in fixed shares), global map-reduce, DRIFT and basic engines. DRIFT picks a random community report as its template, so results vary between runs. nano-graphrag builds local context from the top 20 entities as four CSV tables, each with its own token budget. Its default global mode maps over 16,384-token groups of reports, and still makes those map calls with only_need_context. Semantica offers local, global, DRIFT and hybrid modes. In local mode its vector, graph and memory retrievers run one after another. LLM Graph Builder defaults to graph_vector_fulltext: chunk hits, then one or two hops from each chunk’s entities depending on their similarity to the query. Other modes search community summaries or generate Cypher.

Vector-seeded neighbourhoods. LightRAG makes one LLM call to extract keywords. Low-level keywords search entity vectors and high-level keywords search relation vectors. The default mix mode adds chunk hits, with top_k 40. AutoFlow splits the question into sub-questions by default. For each one it walks relationships by weight, two levels deep. It then retrieves chunks with a rewritten question. TrustGraph grounds the question’s concepts in entity vectors and walks two hops. A cross-encoder keeps 25 edges per hop, so up to 50 edges reach synthesis. It has no global mode.

Graph as a passage ranker. HippoRAG scores facts, lets an LLM filter the top 5, then runs PPR with damping 0.5. If no fact survives the filter, it quietly falls back to dense retrieval. Vector Graph RAG extracts entities from the question, keeps entity hits above 0.9 similarity, expands one degree, asks the LLM for five relations, and answers from three passages.

Lexical traversal. graphify scores node labels by token match behind a trigram prefilter, then runs BFS or DFS to depth 3. It returns 2,000 tokens of text to the calling agent and makes no LLM or vector call.

Pick: GraphRAG or Semantica for synthesis across the corpus. Pick: LightRAG mix or TrustGraph for fast answers grounded in entities. Pick: HippoRAG or Vector Graph RAG for multi-hop QA over passages.

Per-project answers

Graphify-Labs/graphify

answered

MCP server tools. The query surface is a set of MCP tools defined in serve.py (lines 1959–2114): query_graph (the primary search+traverse tool), get_node, get_neighbors, get_community, shortest_path, god_nodes, and graph_stats. There is no distinct local/global/hybrid mode — the single query_graph tool supports BFS and DFS traversal modes.

Scoring and seed selection. A query in _query_graph_text() (serve.py:1361–1420) first splits the question into search terms via _query_terms() (serve.py:293), which handles Chinese segmentation via jieba when available. These terms are scored against every node in _score_query() (serve.py:554–680), which computes a combined TF-IDF–style score using per-token frequency (via _compute_idf) with a full-query exact-match bonus. Seeds are selected by _pick_seeds() via a gap-based heuristic over the ranked list, guaranteeing at least one seed per distinct query term. Relational-intent verbs ("calls", "uses", etc.) are demoted from the per-term guarantee to prevent an incidental match on a verb string from seating a decoy root (serve.py:1386–1391).

Trigram prefilter for speed. Before scoring each node, a trigram index (_get_trigram_index, serve.py:451–472) prunes the candidate set to nodes whose normalised text trigrams overlap with query trigrams, reducing passes over the full graph. On sparse queries the fallback is a full scan.

Traversal. Starting from selected seeds, the graph is traversed via BFS (breadth search — the default) or DFS (depth-first trace). Depth is capped (default 3, max 6). Edge relations can be filtered via explicit context_filter or auto-inferred from query intent words (_infer_context_filters, serve.py:940). The traversal runs on a filtered subgraph view.

Context assembly. The visited subgraph is rendered to plain text by _subgraph_to_text() — nodes as labelled entries with their source file, edges as directional relation statements — and capped at a token_budget (default 2000 tokens) so the output fits in an LLM context window. Seeds are rendered first so they survive truncation.

No vector search. There is no vector/hybrid search layer. All retrieval is text-based (trigram prefilter → token/scored exact & prefix matching → graph traversal). The output is returned as text directly to the calling agent (the MCP client), not fed into any secondary LLM query pipeline within graphify.

HKUDS/LightRAG

answered

Five query modes. Defined in QueryParam.mode at lightrag/base.py:90-100 ("local", "global", "hybrid", "naive", "mix", "bypass"). aquery_llm (lightrag/lightrag.py:5114-5212) dispatches to kg_query (modes local/global/hybrid/mix) or naive_query (mode naive) or a direct LLM call (mode bypass).

Keyword extraction. Every KG-based query first extracts high-level and low-level keywords from the query text via get_keywords_from_query (lightrag/operate.py:4975), which calls the LLM with the keywords_extraction prompt template (lightrag/prompt.py:484-498). Low-level keywords are specific named entities; high-level keywords capture conceptual themes.

Local mode (lightrag/operate.py:5366). Low-level keywords are embedded and used to query entities_vdb (vector DB of entity descriptions). The top-K entities are retrieved from the graph via _get_node_data (lightrag/operate.py:6200), which fetches node properties and degrees, then _find_most_related_edges_from_entities traverses the graph to find connected edges, sorting by rank × weight.

Global mode (lightrag/operate.py:5377). High-level keywords are embedded and used to query relationships_vdb (vector DB of relation descriptions). The top-K relations are retrieved via _get_edge_data (lightrag/operate.py:6529), which fetches edge properties from the graph, then _find_most_related_entities_from_relationships collects the endpoint entities.

Hybrid mode (lightrag/operate.py:5388). Runs both local and global retrieval in parallel, then performs a round-robin merge of entities and relations, deduplicating by name/edge-key (lightrag/operate.py:5432-5486).

Mix mode (lightrag/operate.py:5411). Same as hybrid but additionally queries chunks_vdb (document chunk vector store) directly with the query embedding. The result includes vector-retrieved text chunks alongside graph-retrieved entities and relations.

Naive mode. Direct vector search over document chunks, no graph involvement.

Context assembly (_build_query_context at lightrag/operate.py:6075-6197): A 4-stage pipeline — (1) _perform_kg_search retrieves entities, relations, and vector chunks; (2) _apply_token_truncation reduces results to fit LLM token budgets (max_entity_tokens, max_relation_tokens, max_total_tokens); (3) _merge_all_chunks selects the most relevant text chunks connected to the surviving entities/relations; (4) _build_context_str formats everything into the final prompt with the rag_response or naive_rag_response template (lightrag/prompt.py:334-440), which includes a Knowledge Graph Data section and a Document Chunks section with citation references. The prompt is then sent to the query-role LLM.

Editor's note. Correction: in hybrid and mix modes the entity-side and relation-side searches are awaited one after the other, not in parallel. Only the embeddings are computed in one batch. The default mode is mix and the default top_k is 40.

microsoft/graphrag

answered

GraphRAG provides three query modes.

Local Search (LocalSearch at search.py:31). The LocalContextBuilder selects candidate entities by vector similarity. build_entity_context formats entity data (id, title, description, rank) into a delimited table. Relationships are filtered via _filter_relationships (at local_context.py:232): in-network edges first, then out-network edges prioritized by shared-link count. Source text units are included via build_text_unit_context. Covariates (claims) are optionally added. All builders enforce max_context_tokens=8000. The context is injected into LOCAL_SEARCH_SYSTEM_PROMPT and the LLM generates the answer.

Global Search (GlobalSearch at search.py:55) operates on community reports only. A map stage (lines 172-180) sends batches of reports to the LLM in parallel via asyncio.gather, each returning JSON {description, score} key points. A reduce stage (line 191) collects, scores, sorts, and concatenates points within max_data_tokens, then feeds them to the LLM for synthesis. DynamicCommunitySelection (at dynamic_community_selection.py:26) can optionally rate each community's relevance by LLM call, descending the hierarchy for relevant ones.

DRIFT Search extends local search with iterative query expansion — follow-up queries generated from the initial answer, then re-searched.

Hybrid approach. Vector embeddings select candidates; the graph structure (neighbor relationships, community membership, text units) expands the context.

Editor's note. Correction: 8,000 tokens is only the function default. The configured default max_context_tokens for local, global and basic search is 12,000 (config/defaults.py). DRIFT picks its query-expansion template from a random community report, so its results are not deterministic.

semantica-agi/semantica

answered

Query-time retrieval supports three modes — local, global, and hybrid — selected via the mode parameter of ContextRetriever.retrieve() (context_retriever.py:210-295). Local mode runs three retrievers in parallel: (1) vector search via _retrieve_from_vector() which queries the vector store (FAISS/Qdrant/etc.) for semantically similar content; (2) graph retrieval via _retrieve_from_graph() which first extracts query intent (relationship types, entity types, keywords), then computes cosine similarity between the query embedding and every entity/relationship embedding — boosting scores when entity types or relation types match the intent — and finally performs BFS-based graph expansion (max_expansion_hops, default 2) to pull in related entities. All results are merged via _rank_and_merge() which normalizes scores per source, applies contextual boosting (graph results with more neighbors get up to 20% boost; results found in both vector and graph get 20% cross-source boost), and re-ranks using 70% original + 30% query-content similarity. Global mode delegates to GlobalGraphRetriever.search() (global_retriever.py:329+), which implements a Map-Reduce pattern over community reports: it dynamically selects a hierarchy level (promoting to coarser levels if token budget is exceeded), runs Map LLM calls concurrently via ThreadPoolExecutor (each extracting key points from a community report into a structured MapPointSchema with relevance scores), then synthesizes the Reduce response from top-scored points within a token budget, with automatic evidence/citation validation. Hybrid mode runs both local and global retrievers, concatenates results, and re-ranks. DRIFT mode (drift_search.py) combines global thematic framing with directed facet reasoning and local entity-hop graph traversal. The assembled context is returned as RetrievedContext objects with content, score, source, related_entities, and related_relationships — ready for LLM consumption.

Editor's note. Correction: in local mode the vector, graph and memory retrievers run sequentially, not in parallel; only the global map phase uses a thread pool.

neo4j-labs/llm-graph-builder

answered

Query-time retrieval is handled by QA_RAG() (QA_integration.py:710-748), which supports multiple modes configured in CHAT_MODE_CONFIG_MAP (constants.py:718-780).

Local vector mode (vector, fulltext). Uses a Neo4jVector retriever on the vector index (Chunk.embedding). VECTOR_SEARCH_QUERY (constants.py:302-326) retrieves top-k matching chunks, their documents, and associated chunk details. For fulltext, a keyword full-text index is added for hybrid vector+keyword search.

Graph-augmented modes (graph_vector, graph_vector_fulltext). VECTOR_GRAPH_SEARCH_QUERY (constants.py:329-507) first retrieves top-k chunks, then for each chunk's entities, performs a 2-hop neighborhood traversal. Entity embedding similarity to the query vector determines traversal depth: entities with mid-range scores (0.3-0.9) get 1-hop neighbors (limit 20), entities with high scores (>0.9) get 2-hop neighbors (limit 40). This produces a combined context of chunk text plus entity/relationship subgraph, assembled into a structured text block with Text Content:, Entities: and Relationships: sections (constants.py:482-498).

Entity vector mode (entity_vector). LOCAL_COMMUNITY_SEARCH_QUERY (constants.py:515-558) retrieves top-k __Entity__ nodes by vector similarity, then fetches their associated chunks (up to 3), communities (up to 3), internal relationships, and outside entity connections (up to 10). This gives a local neighborhood view around matching entities.

Global community mode (global_vector). GLOBAL_VECTOR_SEARCH_QUERY (constants.py:681-695) searches the community_vector index (Community.summary), combining full-text (via community_keyword index) and vector search. It returns community summaries as the context, providing a document-set overview rather than chunk-level detail.

Graph (Cypher) mode (graph). Uses GraphCypherQAChain (QA_integration.py:562-583) which generates a Cypher query from the user's natural language question, executes it against Neo4j, and then answers with the results via a QA LLM.

Context assembly. In all RAG modes, retrieved documents pass through format_documents() (QA_integration.py:181-228) which sorts by similarity score, truncates to a model-specific token cutoff (CHAT_TOKEN_CUT_OFF, constants.py:249-253), formats each as "Document start / source / content / Document end", and collects source metadata, entity IDs, and community IDs. The final context is passed to a ChatPromptTemplate with the CHAT_SYSTEM_TEMPLATE (constants.py:256-297). Default mode is graph_vector_fulltext (CHAT_DEFAULT_MODE, constants.py:716).

OSU-NLP-Group/HippoRAG

answered

No local/global/hybrid modes. HippoRAG has a single retrieval mode: it scores facts against the query, reranks, then runs PPR over the graph seeded by entity weights. The user-visible entry points are retrieve() (HippoRAG.py:704–794), retrieve_ircot() (804–857) and retrieve_dpr() (968–1039), plus QA helpers rag_qa() and answer_with_ircot().

Step 1 — Fact retrieval. Queries are encoded through the embedding model with a query_to_fact instruction (get_query_instruction, prompts/linking.py:5). Dot-product similarity scores are computed against all fact embeddings and min-max normalized (HippoRAG.py:1892–1928).

Step 2 — Recognition-memory reranking. The DSPyFilter (rerank.py:14–126) uses the extraction LLM to filter the top linking_top_k facts (default 5), passing each candidate fact list to the LLM with a prompt asking it to select only relevant facts. The LLM's selection is matched back by text similarity (difflib.get_close_matches).

Step 3 — Graph search with PPR. graph_search_with_fact_entities (HippoRAG.py:2008–2121) assigns each entity a weight from the average score of the facts containing it. Simultaneously, dense_passage_retrieval (1930–1967) computes passage–query dot-product scores, scaled by passage_node_weight (default 0.05). Entity and passage weights are combined into one node-weight vector, then run_ppr (2177–2218) calls igraph.Graph.personalized_pagerank with damping=0.5 (default), using the weight attribute on edges and an undirected projection (directed=False). Passage nodes are extracted from the PPR scores and sorted.

Step 4 — Context for the LLM. The top num_to_retrieve passages are collected via _build_retrieval_result (796–802). The QA step (qa, 1117–1177) passes the top qa_top_k (default 5) passages to the LLM with a prompt formatted as "Wikipedia Title: {passage}" + "Question: {query}", using dataset-specific prompt templates (e.g., rag_qa_musique).

Fallback. If no facts survive reranking, dense_passage_retrieval is used as a fallback (HippoRAG.py:762–764).

gusye1234/nano-graphrag

answered

Three query modes are selected via QueryParam.mode: "local", "global", or "naive" (graphrag.py:236-275).

Local mode (_op.py:935-967, builder at 844-933). Step 1: embed the query string and search entities_vdb (cosine similarity, top_k=20) to get the most relevant entity vectors. Step 2: look up those entities in the graph and expand to one-hop neighbors via get_nodes_edges_batch. Step 3: from those entity nodes, gather community reports via _find_most_related_community_from_entities — filters to communities at or below query_param.level (default 2), sorts by occurrence count and rating, and truncates to local_max_token_for_community_report tokens (3200). Step 4: gather source text chunks from the entities and their neighbors via _find_most_related_text_unit_from_entities, merging them with relation counts for deduplication and sorting, truncated to local_max_token_for_text_unit (4000). Step 5: gather and rank edges via _find_most_related_edges_from_entities, truncated to local_max_token_for_local_context (4800). All four sections are assembled into CSV tables (Reports, Entities, Relationships, Sources) and fed as context to the LLM with the local_rag_response prompt. Optional local_community_single_one flag limits to the top community.

Global mode (_op.py:1017-1104). Retrieves all community schemas from the graph, filters to level <= query_param.level, sorts by occurrence, caps at global_max_consider_community (512), and filters by global_min_community_rating. Remaining communities are grouped into token-budget-limited batches (global_max_token_for_community_report, 16384). For each group, an LLM (acting as "analyst") extracts key points with importance scores using global_map_rag_points prompt. Results are flattened, filtered by score>0, sorted, truncated again, then a final LLM call with global_reduce_rag_response synthesizes the analysts' perspectives into a unified answer.

Naive mode (_op.py:1107-1140). Simple vector search on chunks_vdb, retrieves top_k chunks, truncates to naive_max_token_for_text_unit (12000), and feeds them to the LLM.

All three modes support only_need_context=True to return the raw retrieved context without an LLM generation call.

Editor's note. Correction: only_need_context=True avoids LLM calls only in local and naive modes; global mode still runs the map-stage LLM call per community group and returns the scored analyst points, skipping only the reduce call.

pingcap/autoflow

answered

Query-time retrieval is a hybrid process that combines graph traversal, vector similarity, and chunk retrieval, assembled into context for the LLM.

How it's invoked (chat flow): In ChatFlow._builtin_chat() (backend/app/rag/chat/chat_flow.py:200-258), the pipeline is: (1) search the knowledge graph, (2) refine the user question using graph context, (3) optionally clarify, (4) search for relevant chunks, (5) generate the answer with both graph context and chunks.

Graph retrieval modes: The KnowledgeGraphOption config has two modes controlled by using_intent_search (backend/app/rag/chat/config.py:53-55):

  • Normal mode (using_intent_search=False): The KnowledgeGraphSimpleRetriever wraps TiDBGraphStore.retrieve_with_weight() which performs a multi-depth traversal. At depth 0, it vector-search-ranks relationships against the query embedding using a weighted score: alpha * (1/embedding_distance) + weight_score + degree_score. Then for subsequent depths, it progressively explores neighbors, breaking the search into distance-range bands ([(0,0.25), (0.25,0.35), (0.35,0.45), (0.45,0.55)]) with decreasing search ratios, ensuring broad coverage. Synopsis entities are appended via fetch_similar_entities(entity_type=synopsis). Results are rendered into a prompt template (normal_graph_knowledge) listing entities and relationships.
  • Intent mode (using_intent_search=True): The KnowledgeGraphFusionRetriever first decomposes the question into sub-queries (via DecomposedFactors schema, found in schema.py:113-115), retrieves per sub-query, then fuses the results — merging entities by set union and relationships by rag_description key (accumulating weight). Each sub-graph is rendered via intent_graph_knowledge prompt template.

Combining graph + vector hits: After graph retrieval, the refined question (enriched with graph context) is used by ChunkFusionRetriever (retrieve_flow.py:134-142) to search the vector index for relevant chunks. Both graph context (as a formatted string knowledge_graph_context) and chunks (as NodeWithScore[]) are passed to the final LLM prompt (_generate_answer at chat_flow.py:460-524) via the text_qa_template prompt, which includes {graph_knowledges} as a partial format variable.

Fusion across multiple knowledge bases: The KnowledgeGraphFusionRetriever._knowledge_graph_fusion() (backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py:97-137) merges entities/relationships from multiple KB retrievers into a single result with child subgraph metadata.

No true local/global modes: Unlike Microsoft's GraphRAG, AutoFlow does not have a separate "global" search over community summaries. It always retrieves graph context by traversing from query-similar relationships.

trustgraph-ai/trustgraph

answered

TrustGraph has three query modes:

Graph RAG (retrieval/graph_rag/): A five-stage pipeline:

  1. Concept extraction (graph_rag.py:156-178): LLM prompt (extract-concepts) decomposes the query into concept strings. Falls back to raw query if empty.

  2. Entity grounding (graph_rag.py:192-241): Each concept is embedded and used to query the graph embeddings vector store. Results are deduplicated into seed entity URIs.

  3. Iterative hop-and-filter (graph_rag.py:323-496): For up to max_path_length (2) hops, retrieves triples where frontier entities appear as subject, predicate, or object. Schema predicates (RDF/RDFS/OWL) filtered out. Labels resolved in batch with a per-request LRU cache. Each edge is rendered as text and scored against concepts by a cross-encoder reranker. Top edge_limit (25) edges survive, capped at max_reranker_input (350) candidates and max_reranker_text_length (240) chars.

  4. Source tracing (graph_rag.py:498-618): Edges traced to source documents via provenance chain (tg:contains -> prov:wasDerivedFrom).

  5. Synthesis (graph_rag.py:791-828): Edges and source metadata assembled into a kg-synthesis LLM prompt. Supports streaming. Explainability triples at every stage.

Document RAG (retrieval/document_rag/document_rag.py:216-530): Queries document embeddings. Supports vector, keyword (BM25), and hybrid modes (RRF fusion). Optional cross-encoder reranker with MMR diversity.

OntoRAG Query (query/ontology/query_service.py:144-218): Analyzes the question, matches ontology, generates SPARQL/Cypher, executes, and answers.

No local/global modes -- always local traversal.

Editor's note. Correction: edge_limit (25) is applied per hop, so the default two-hop traversal can pass up to 50 selected edges, plus traced source-document metadata, to the kg-synthesis prompt.

zilliztech/vector-graph-rag

answered

Query-time retrieval is a multi-step pipeline orchestrated by GraphRetriever.retrieve() (graph/retriever.py:388-487):

  1. Entity extraction: EntityExtractor (llm/extractor.py:252-403) uses an OpenAI LLM call to extract named entities from the question, with a HippoRAG-compatible NER cache fallback from TSV files.

  2. Bi-directional vector search: Extracted entities are embedded and searched against the entity collection (_search_entities), and the raw question is embedded and searched against the relation collection (_search_relations). Both apply configurable similarity thresholds (entity default 0.9, relation default -1.0 i.e. keep all) (retriever.py:114-222).

  3. Subgraph expansion: Seed entities and relations are loaded into a SubGraph (graph/knowledge_graph.py:152-591) which lazily fetches neighbor relations and entities from Milvus. Default expansion degree is 1. The expansion walks from seed entities to their relations, merges with seed relations, then for each degree expands relations-to-entities-to-relations.

  4. Eviction strategy: If expanded relations exceed relation_number_threshold (default 1000), a vector search re-ranks them down to the threshold. Otherwise, relations are sorted by ID for deterministic behavior (HippoRAG compat) (retriever.py:264-336).

  5. LLM reranking (optional): LLMReranker (llm/reranker.py:96-294) takes the candidate relations and uses GPT with three few-shot multi-hop examples to select the most useful ones via chain-of-thought JSON output. This is the single-pass replacement for iterative agent loops.

  6. Passage retrieval: Selected relation IDs are used to fetch associated passages, calling MilvusStore.get_passages_by_ids() after resolving passage IDs from relation metadata (rag.py:735-782). If fewer than final_top_k passages are found, graph passages are supplemented with naive vector search results (rag.py:1465-1475).

  7. Answer generation: AnswerGenerator (llm/reranker.py:296-392) feeds the final passages as context to the LLM.

There are no separate local/global modes — only graph retrieval, with naive RAG available as a comparison baseline.

← Are communities, summaries or hierarchies built over the graph? · How are updates and incremental indexing handled? →