LLMs Technical Reviews
Home / Graph RAG / graphrag

microsoft/graphrag

Python pipeline that turns text into an LLM-extracted entity graph with Leiden community reports, queried by local, global and DRIFT search.

GitHub ↗★ 36kPythonMITcommit 769542f · 2026-09-23homepage ↗

Overview

GraphRAG is Microsoft’s reference implementation of graph-based retrieval-augmented generation. It is a Python library and a graphrag CLI. You point it at a folder of documents, and an indexing pipeline uses an LLM to extract entities and relationships from every chunk. It then clusters the resulting graph into a hierarchy of communities with Leiden and asks the LLM to write a report for each community. The output is a set of tables (Parquet by default), plus embeddings in a vector store. Four query engines read those tables: local, global, DRIFT and basic.

The design suits corpus-level questions such as “what are the main themes across these documents?”. Plain chunk retrieval cannot answer those well, because no single chunk holds the answer. GraphRAG answers them by map-reducing over the community reports. The price is indexing cost: every chunk gets at least one extraction call, every entity with several descriptions gets a summarization call, and every community gets a report call.

At this SHA the repository is a monorepo of eight packages (graphrag, graphrag-llm, graphrag-storage, graphrag-vectors, graphrag-cache, graphrag-chunking, graphrag-input, graphrag-common). The README says the project is largely in maintenance mode: bug fixes and dependency updates, no new features.

Architecture

flowchart LR
  D["Input documents"] --> C["Chunker (text_units)"]
  C --> X["extract_graph (LLM)"]
  X --> S["summarize_descriptions (LLM)"]
  S --> F["finalize_graph (degrees)"]
  F --> L["create_communities (Leiden)"]
  L --> R["community_reports (LLM)"]
  R --> E["generate_text_embeddings"]
  E --> V["Vector store (LanceDB)"]
  F --> T["Tables (Parquet / CSV / Cosmos)"]
  R --> T
  T --> Q["Query engines"]
  V --> Q
  Q --> LS["local"]
  Q --> GS["global (map-reduce)"]
  Q --> DS["DRIFT"]
  Q --> BS["basic"]
Component Path Role
CLI packages/graphrag/graphrag/cli/main.py init, index, update, prompt-tune, query commands (Typer)
Indexing API packages/graphrag/graphrag/api/index.py build_index picks a pipeline and runs it
Pipeline registry packages/graphrag/graphrag/index/workflows/factory.py Named workflow lists: standard, fast, and their -update variants
Workflows packages/graphrag/graphrag/index/workflows/ One module per step; each reads and writes tables
Graph extraction packages/graphrag/graphrag/index/operations/extract_graph/ Per-chunk LLM extraction, gleaning, merge by name
Clustering packages/graphrag/graphrag/index/operations/cluster_graph.py Hierarchical Leiden via graspologic-native
Query engines packages/graphrag/graphrag/query/structured_search/ local_search, global_search, drift_search, basic_search
LLM layer packages/graphrag-llm/graphrag_llm/ LiteLLM wrapper with cache, retry, rate-limit and metrics middleware
Storage packages/graphrag-storage/graphrag_storage/ File, Azure Blob, Cosmos DB and memory storage; Parquet, CSV and Cosmos table providers
Vectors packages/graphrag-vectors/graphrag_vectors/ LanceDB (default), Azure AI Search, Cosmos DB

How a request flows

Indexing (graphrag index)

  1. build_index resolves the method name (for example standard or standard-update) and gets a workflow list from PipelineFactory.create_pipeline (api/index.py, factory.py).
  2. run_pipeline creates the input storage, output storage, table provider and cache. Then it runs each workflow in order and writes stats.json and context.json after each step (run_pipeline.py).
  3. create_base_text_units streams documents through the chunker. The default is 1,200 tokens with 100 overlap. Each chunk’s ID is a SHA-512 hash of its text (create_base_text_units.py).
  4. The extract_graph workflow builds two models, one for extraction and one for summarization, each with its own cache namespace (extract_graph.py). GraphExtractor sends one prompt per chunk and parses ("entity"<|>NAME<|>TYPE<|>description) and ("relationship"<|>...) records (graph_extractor.py).
  5. Results are merged across chunks by exact (title, type) for entities and (source, target) for relationships. Relationship weights are summed (extract_graph.py). Then the description lists are summarized by the LLM.
  6. finalize_graph computes node degrees. create_communities runs hierarchical Leiden. create_community_reports writes a report per community, and generate_text_embeddings embeds entity descriptions, chunk text and report content.

Query (graphrag query --method local|global|drift|basic)

  1. The API loads the tables into dataclasses. For local search, map_query_to_entities embeds the query and finds the nearest entity descriptions. LocalSearchMixedContext.build_context then fills the token budget in fixed shares: community reports, the entity and relationship tables, and source text units (mixed_context.py).
  2. For global search, report batches go to the LLM in parallel (map), each returning scored key points. A reduce call merges the top points into the answer (global_search/search.py).

Key components

Extraction and gleaning

The extraction prompt asks for entities of configured types (default organization, person, geo, event) and relationships with a numeric strength. It also says “Return output in English”. Names and types are upper-cased before merging, so Acme and ACME become one entity, but ACME CORP stays separate. There is no further entity resolution. With max_gleanings > 0 (default 1), the extractor asks “continue” and then a yes/no “loop” question until the limit is reached (graph_extractor.py). The fast method replaces LLM extraction with noun-phrase extraction and co-occurrence edges (extract_graph_nlp, build_noun_graph), then prunes the graph.

Communities

cluster_graph makes the edge list undirected, optionally keeps only the largest connected component (use_lcc, default true), and calls graspologic_native.hierarchical_leiden with max_cluster_size 10 and a fixed seed (cluster_graph.py, hierarchical_leiden.py). Each community row keeps its level, parent and children. Reports are generated level by level, deepest level first.

Query engines

  • Local: entity vector search, then the graph neighbourhood, community reports and text units in one prompt (default budget 12,000 tokens).
  • Global: map-reduce over the reports at one community level. Optional dynamic community selection asks the LLM to rate reports and descends into the children of relevant ones.
  • DRIFT: writes a hypothetical answer using a randomly chosen community report as the template, embeds it to rank reports, and asks a primer prompt for intermediate answers and follow-up questions. It then runs up to n_depth (3) rounds of local search on the top follow-ups and reduces the results (drift_search/search.py).
  • Basic: vector search over text units only.

LLM layer

create_completion wraps a LiteLLM model in a middleware pipeline: request counts, cache, retries, rate limiting, metrics (completion_factory.py). The cache key is the full request. Streaming and mocked calls are not cached (with_cache.py). The indexing workflows pass a cache. The query factory does not, so query calls are never cached.

Extending it

  • Custom workflows: register a function with PipelineFactory.register and a new pipeline with register_pipeline. Or set workflows: in settings.yaml to run a custom list.
  • Prompts: every LLM step takes a prompt file. graphrag prompt-tune generates domain-adapted extraction and report prompts from a sample of your data.
  • Backends: storage, table providers, vector stores and caches each have a factory in their package. Cosmos DB, Azure Blob and Azure AI Search ship in the box.
  • Models: any LiteLLM model string. Each stage has its own completion_model_id, so extraction can use a cheaper model than report writing.

Running it

pip install graphrag, then graphrag init --root ./proj writes settings.yaml and default prompts. The default models are gpt-4.1 and text-embedding-3-large from provider openai. Put files in input/, run graphrag index --root ./proj, then graphrag query --root ./proj --method global "...". graphrag update adds new documents to an existing index. A separate Streamlit app in unified-search-app/ compares the query methods side by side.

Strengths and caveats

  • Strength: the reference design for hierarchical community summaries (Leiden levels, each with LLM-written reports), which give global search a whole-corpus answer path. nano-graphrag, Semantica and LLM Graph Builder build similar hierarchies.
  • Strength: a clean table-based output. You can inspect, version and load the index into any dataframe tool.
  • Strength: per-stage model choice, a request-level cache and token metrics on every query result.
  • Caveat: the oversized-community fallback does not work as designed. summarize_communities builds the context for every level before it generates any report, so build_level_context always gets an empty report table. Communities that exceed max_input_length are trimmed instead of being filled with sub-community reports (summarize_communities.py, context_builder.py).
  • Caveat: incremental update is append-only. New documents are found by title, and deleted_inputs is computed but never read, so removed or edited documents stay in the index (incremental_index.py). New communities are clustered from the delta graph alone and appended with shifted IDs. Old communities are not re-clustered.
  • Caveat: no graph database. Query engines load all tables into memory, and traversal is one hop around the matched entities.
  • Caveat: indexing cost grows with chunks, entities and communities. DRIFT is not deterministic, because it picks its template report at random.

Sources: code at 769542f, deepwiki-open wiki (11 pages), verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit 769542f. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Chunking. Documents are split by TokenChunker with configurable token size and overlap, or by SentenceChunker. Chunks are written to the text_units table via create_base_text_units, with a SHA-512 hash as ID and a token count. Optional document metadata can be prepended to each chunk.

Extraction prompt. The prompt at extract_graph.py:6-126 asks the LLM to identify entities of specified types (ORGANIZATION, PERSON, GEO, etc.) with name, type, and description, then identify pairwise relationships with description and numeric strength. Output is a delimited format with ## separators and <|COMPLETE|> termination.

Multi-round gleaning. GraphExtractor (at graph_extractor.py:38) extracts entities per chunk. A gleaning loop (lines 99-120) re-prompts the model up to max_gleanings times with CONTINUE_PROMPT to catch missed items, using LOOP_PROMPT as a stopping gate.

Merging. Results from all chunks are grouped by exact-match on (title, type) for entities and (source, target) for relationships (see _merge_entities and _merge_relationships at extract_graph.py:104-129). The finalize_graph workflow computes undirected node degrees from the deduplicated edge set.

Summarization. Raw descriptions are re-summarized by a second LLM via the prompt at summarize_descriptions.py:6-20, limited to max_summary_length words.

Schema. Final columns defined in schemas.py:70-159: entities = (id, title, type, description, text_unit_ids, frequency, degree); relationships = (id, source, target, description, weight, combined_degree, text_unit_ids).

Where and how is the graph stored?

answered

The graph is stored as tabular files (parquet/CSV) or in Azure CosmosDB — not as a graph database. The base Storage abstraction provides key-value file access with implementations for local disk, Azure Blob, Azure Cosmos DB, and in-memory. On top sits TableProvider for table-level operations: read_dataframe, write_dataframe, and open for streaming row operations.

Table schema. Six tables are produced with fixed column schemas in schemas.py:70-159. Entities: id, title, type, description, text_unit_ids, frequency, degree. Relationships: id, source, target, description, weight, combined_degree, text_unit_ids. Communities: id, community, level, parent, children, entity_ids, relationship_ids, text_unit_ids. TextUnits: id, text, n_tokens, document_id, entity_ids, relationship_ids.

Embeddings. The generate_text_embeddings workflow (at generate_text_embeddings.py:52) creates embeddings for three configured fields: text_unit text, entity title-description, and community report full content. The embed_text operation reads rows, batches them (configurable batch_size/batch_max_tokens), calls the embedding model, and loads vectors into a pluggable VectorStore. Optional snapshot tables store vectors alongside base tables.

No graph DB at query time. Query sessions load entities/relationships/communities into Python dataclass objects from serialized tables at startup.

Are communities, summaries or hierarchies built over the graph?

answered

Yes — GraphRAG builds a hierarchical community structure using the Leiden algorithm.

Detection. The create_communities workflow calls cluster_graph, which normalizes edges to undirected, optionally restricts to the largest connected component (use_lcc), then calls hierarchical_leiden via graspologic-native at hierarchical_leiden.py:11. Parameters: max_cluster_size (default 10), resolution 1.0, randomness 0.001, configurable seed.

Hierarchy. _compute_leiden_communities (at cluster_graph.py:51) builds level-to-community and cluster-to-parent mappings. Each community stores level (0=root), community ID, parent, children, plus aggregated entity_ids, relationship_ids, and text_unit_ids (see create_communities.py:86-192).

Report generation. The create_community_reports workflow builds a local context (nodes, edges, claims) per community and calls an LLM to generate structured reports. Each report has summary, full_content (narrative), findings (structured JSON), rank, and rating_explanation. Reports can be weighted by text-unit count of their entities (at community_context.py:189).

When computed. Communities and reports are built during indexing, not at query time.

Editor's note. Addition: all level contexts are built before any report is generated (summarize_communities.py L61-L73), so build_level_context always gets an empty report table. Oversized communities are trimmed instead of being filled with sub-community reports.

How does query-time retrieval use the graph?

answered

GraphRAG provides three query modes.

Local Search (LocalSearch at search.py:31). The LocalContextBuilder selects candidate entities by vector similarity. build_entity_context formats entity data (id, title, description, rank) into a delimited table. Relationships are filtered via _filter_relationships (at local_context.py:232): in-network edges first, then out-network edges prioritized by shared-link count. Source text units are included via build_text_unit_context. Covariates (claims) are optionally added. All builders enforce max_context_tokens=8000. The context is injected into LOCAL_SEARCH_SYSTEM_PROMPT and the LLM generates the answer.

Global Search (GlobalSearch at search.py:55) operates on community reports only. A map stage (lines 172-180) sends batches of reports to the LLM in parallel via asyncio.gather, each returning JSON {description, score} key points. A reduce stage (line 191) collects, scores, sorts, and concatenates points within max_data_tokens, then feeds them to the LLM for synthesis. DynamicCommunitySelection (at dynamic_community_selection.py:26) can optionally rate each community's relevance by LLM call, descending the hierarchy for relevant ones.

DRIFT Search extends local search with iterative query expansion — follow-up queries generated from the initial answer, then re-searched.

Hybrid approach. Vector embeddings select candidates; the graph structure (neighbor relationships, community membership, text units) expands the context.

Editor's note. Correction: 8,000 tokens is only the function default. The configured default max_context_tokens for local, global and basic search is 12,000 (config/defaults.py). DRIFT picks its query-expansion template from a random community report, so its results are not deterministic.

How are updates and incremental indexing handled?

answered

Incremental indexing has partial support with title-based delta detection.

Delta detection. get_delta_docs at incremental_index.py:29 compares the input dataset against previously-indexed documents by title (lines 49-50). InputDelta captures new_inputs (titles not in previous docs) and deleted_inputs (previous titles missing from input).

Re-indexing. update/entities.py, update/relationships.py, and update/communities.py run extraction and clustering on only the delta documents. concat_dataframes (line 61) merges old and delta tables, assigning new human_readable_id values starting at max(old) + 1.

Deletion. The deleted_inputs DataFrame is populated, but the system does not fully re-index downstream artifacts — orphaned entities/edges/communities remain unless a full rebuild is run.

Caching. The LLM cache (with_cache at cache_middleware.py:21) caches by full request content, so re-extracting the same text on unchanged documents hits cache. But this is content-addressable, not incrementally aware.

Limitations. The approach is title-based so two documents with the same title are treated as the same. Editing documents in place requires manual deletion or full rebuild.

Editor's note. Correction: deleted_inputs is computed but never read anywhere, so deleted or edited documents stay in the index. The update/*.py modules only merge the old and delta tables (extraction runs earlier, on the delta documents). Delta communities are clustered on the delta graph alone and appended with shifted IDs; the old hierarchy is not re-clustered.

How are LLM cost and latency controlled during indexing and query?

answered

LLM response caching. The with_cache middleware wraps every LLM completion and embedding call. Before calling the model, it computes a cache key from full request args, checks cache.get(key), and on a hit returns the cached response. Applies to both indexing and query. Streaming and mocked responses are excluded.

Batching. embed_text buffers rows and dispatches batches at configurable batch_size and batch_max_tokens, with flush size sized to saturate num_threads × batch_size. For LLM completions, derive_from_rows parallelizes across num_threads.

Model choice per stage. Separate model IDs per stage: extract_graph.completion_model_id, summarize_descriptions.completion_model_id, community_reports.completion_model_id, embed_text.embedding_model_id. Each can be a cheap/fast model via litellm model strings.

Token budgets. All context builders enforce max_context_tokens=8000. build_entity_context truncates when cumulative tokens exceed budget. Community reports cap max_input_length/max_report_length. Global search enforces max_data_tokens=8000 and per-stage max-length params. max_gleanings (default configurable) controls extra LLM rounds during extraction — setting it to 0 skips gleaning.

Rate limiting and retries. with_rate_limiting uses a sliding-window rate limiter per model, counting tokens. with_retries adds exponential backoff. Both compose via middleware pipeline.

Metrics. MetricsProcessor and MetricsStore track token usage, response times, and cache hit rates. Query results return per-category llm_calls, prompt_tokens, and output_tokens.

Editor's note. Correction: only indexing calls are cached. The query engines are built with create_completion(model_settings) and no cache, so query calls are never cached. The configured context budget is 12,000 tokens, not 8,000.