LLMs Technical Reviews
Home / Graph RAG / nano-graphrag

gusye1234/nano-graphrag

A small, readable Python reimplementation of Microsoft GraphRAG with local, global and naive query modes over a NetworkX graph.

GitHub ↗★ 4.0kPythonMITcommit acb35c0 · 2026-01-27

Overview

nano-graphrag is a compact reimplementation of the Microsoft GraphRAG paper. It keeps the core recipe and drops almost everything else. Text is chunked, an LLM extracts entities and relationships from each chunk, the results are merged into a NetworkX graph, hierarchical Leiden finds communities, and an LLM writes a JSON report for each community. At query time you pick local (entity vectors plus graph neighbourhood), global (map-reduce over community reports) or naive (plain chunk RAG).

The whole engine is two files. graphrag.py holds the GraphRAG dataclass, which is both the configuration object and the orchestrator. _op.py holds every pipeline step as a plain async function. Storage, LLM and embedding are injected as classes or callables, so most customisation is a constructor argument, not a subclass. The README’s claim of about 1,100 lines of core code is roughly right: _op.py is 1,140 lines and the rest is small.

It is a library for people who want to read, fork and change a GraphRAG pipeline. It has no server, no UI, no deletion API and no multi-tenant story. Persistence is a working directory of JSON and GraphML files. Think of it as a reference implementation you build on, not a service you deploy.

Architecture

flowchart LR
  IN["insert(text)"] --> GR["GraphRAG.ainsert"]
  GR --> CH["get_chunks (token windows)"]
  CH --> EX["extract_entities (best LLM + gleaning)"]
  EX --> MG["merge nodes / edges"]
  MG --> G["NetworkXStorage (GraphML)"]
  MG --> EV["entities_vdb (nano-vectordb)"]
  G --> LD["hierarchical_leiden on LCC"]
  LD --> CR["generate_community_report"]
  CR --> KV["community_reports (JSON KV)"]
  Q["query(q, mode)"] --> LQ["local_query"]
  Q --> GQ["global_query (map-reduce)"]
  Q --> NQ["naive_query"]
  LQ --> EV
  LQ --> G
  LQ --> KV
  GQ --> KV
Component Path Role
Orchestrator nano_graphrag/graphrag.py GraphRAG dataclass: all defaults, storage wiring, ainsert, aquery
Pipeline ops nano_graphrag/_op.py Chunking, extraction, merging, community reports, the three query modes
Interfaces nano_graphrag/base.py QueryParam, BaseKVStorage, BaseVectorStorage, BaseGraphStorage
Graph stores nano_graphrag/_storage/gdb_networkx.py, gdb_neo4j.py NetworkX + GraphML (default) and Neo4j + GDS Leiden
Vector stores nano_graphrag/_storage/vdb_nanovectordb.py, vdb_hnswlib.py nano-vectordb JSON file (default) and hnswlib
KV store nano_graphrag/_storage/kv_json.py One JSON file per namespace
LLM adapters nano_graphrag/_llm.py OpenAI, Azure OpenAI and Bedrock completions and embeddings, cache and retry
Prompts nano_graphrag/prompt.py PROMPTS dict: extraction, community report, local/global/naive answers
DSPy extractor nano_graphrag/entity_extraction/ Alternative typed extractor with self-refine, plus dataset tools

How a request flows

Insert (graph.insert(text)):

  1. Dedup documents. ainsert hashes each document with MD5 into a doc- id and drops ids already in full_docs (graphrag.py).
  2. Chunk. get_chunks tokenizes and calls chunk_func. The default chunking_by_token_size cuts 1,200-token windows with 100 tokens of overlap. Chunk ids are MD5 hashes of chunk text, so repeated chunks are skipped (_op.py, L94-L108).
  3. Drop old reports. community_reports.drop() runs before extraction, because communities are always rebuilt from scratch.
  4. Extract. extract_entities sends every chunk through the “best” model with the entity_extraction prompt, then runs up to entity_extract_max_gleaning continuation passes. The model emits ("entity"<|>NAME<|>TYPE<|>DESC) and ("relationship"<|>SRC<|>TGT<|>DESC<|>WEIGHT) records, parsed with regex and delimiter splits (_op.py).
  5. Merge. Names are upper-cased. _merge_nodes_then_upsert takes the majority type and joins distinct descriptions with <SEP>. _merge_edges_then_upsert sums weights and creates missing endpoints as UNKNOWN nodes. When a merged description passes 500 tokens, the “cheap” model summarises it (_op.py, L182-L279). Each entity’s name plus description is embedded into entities_vdb.
  6. Cluster. NetworkXStorage._leiden_clustering runs graspologic’s hierarchical_leiden (max cluster size 10, fixed seed) on the stable largest connected component and writes a JSON clusters list onto each node (gdb_networkx.py).
  7. Report. generate_community_report walks levels from the finest to the root. For each community it packs entities and edges as CSV, ordered by degree, and reuses sub-community reports when a community has over 100 nodes or edges. The best model returns JSON with title, summary, rating and findings (_op.py, L625-L697).
  8. Persist. _insert_done calls index_done_callback on every store, which writes the JSON, GraphML and vector files.

Query (graph.query(q, QueryParam(mode=...)), default mode global): aquery dispatches to one of three functions (graphrag.py). Local mode is described below. Global mode filters communities to level <= 2, sorts by occurrence, keeps up to 512, then runs map and reduce LLM calls. Naive mode is top-k chunk retrieval.

Key components

Local query

_build_local_query_context is the most interesting retrieval code (_op.py). It takes the top 20 entities by vector similarity and builds four CSV tables, each with its own token budget:

  • Reports. Communities the entities belong to, ranked by how many hits they got and then by rating (3,200 tokens).
  • Entities. The hits, with graph degree as rank.
  • Relationships. Every edge touching a hit, ranked by edge degree and weight (4,800 tokens).
  • Sources. The source chunks of the hits, preferring chunks that one-hop neighbours also mention (4,000 tokens).

Traversal is exactly one hop. There is no multi-hop path search.

Global query

global_query packs community reports into groups of 16,384 tokens. It asks the model, once per group, for scored “points” in JSON, keeps points with score above 0, and sends the best ones to a final reduce prompt (_op.py). With only_need_context=True it still makes the map calls. It skips only the reduce call.

Storage contracts

base.py defines three async interfaces. The graph interface is the largest: node and edge getters, batch variants, degrees, upsert_*, clustering and community_schema. NetworkXStorage.community_schema rebuilds the community table from the node clusters attributes. A community’s sub-communities are the communities on the next level whose nodes are a subset of its own (gdb_networkx.py). The Neo4j backend projects the graph into GDS and calls gds.leiden.write. It passes max_graph_cluster_size as maxLevels, so that knob means something different on the two backends (gdb_neo4j.py).

LLM layer, cache and concurrency

Every built-in completion function checks an MD5 hash of (model, messages) against the llm_response_cache KV store before it calls the API. It retries rate-limit and connection errors up to five times (_llm.py). On a miss, it writes the whole cache file back to disk after each call. Concurrency is capped by limit_async_func_call, a busy-wait counter instead of an asyncio.Semaphore (16 by default per model and for embeddings) (_utils.py).

Extending it

  • Models. Pass any async fn(prompt, system_prompt=None, history_messages=[], **kwargs) as best_model_func or cheap_model_func. Pop hashing_kv from kwargs if you want the cache. The examples folder covers DeepSeek, Ollama, Bedrock and fully local setups.
  • Embeddings. Wrap a function with wrap_embedding_func_with_attrs(embedding_dim=..., max_token_size=...) and pass it as embedding_func.
  • Storage. Pass graph_storage_cls, vector_db_storage_cls or key_string_value_json_storage_cls. The built-ins are Neo4j and hnswlib. FAISS, Milvus and Qdrant ship as examples.
  • Extraction. entity_extraction_func is swappable. entity_extraction/extract.py provides a DSPy-based extractor with typed entities and optional self-refine. chunk_func can be chunking_by_seperators or your own.
  • Prompts. Edit PROMPTS in place. Entity types default to organization, person, geo and event.
  • Malformed JSON. convert_response_to_json_func lets you plug in a repair function for models that return broken JSON.

Running it

  • pip install nano-graphrag (or pip install -e .). Python 3.9+. requirements.txt installs openai, graspologic, nano-vectordb, hnswlib, dspy-ai, neo4j and aioboto3 unconditionally.
  • Set OPENAI_API_KEY, or pass using_azure_openai=True or using_amazon_bedrock=True with model ids. The defaults are gpt-4o, gpt-4o-mini and text-embedding-3-small.
  • GraphRAG(working_dir="./x") then insert(...) and query(...). Reopening the same directory reloads everything. Naive mode needs enable_naive_rag=True at construction time. Neo4j needs addon_params={"neo4j_url": ..., "neo4j_auth": ...} and the GDS plugin.

Strengths and caveats

  • Strength: readable. The pipeline is a few hundred lines of plain async functions with GraphRAG’s prompts. It is the easiest GraphRAG codebase in this category to read end to end.
  • Strength: easy to swap parts. Model, embedding, chunker, extractor and all three stores are constructor arguments with small interfaces.
  • Strength: idempotent re-inserts. Content hashing for documents and chunks, plus the persistent LLM cache, make repeated runs cheap.
  • Caveat: every insert redoes all communities. Each ainsert drops all reports, re-clusters and re-summarises every community. Report cost grows with the whole graph, not with the new document. If extraction finds no entities, the reports are dropped and not rebuilt.
  • Caveat: only the largest connected component is clustered. Entities outside it get no community and never reach global search.
  • Caveat: no deletion or update. There is no API to remove a document. Rebuilding means deleting the working directory.
  • Caveat: single-process files. JSON and GraphML stores are rewritten in full on each callback, and the LLM cache is rewritten after every uncached call. That is fine for a notebook and poor for large corpora or concurrent writers.
  • Caveat: some dead code. node2vec_params and embed_nodes exist, but no pipeline step calls them.

Sources: code at acb35c0, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (31 pages), verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit acb35c0. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Chunking. Documents are split via chunking_by_token_size (_op.py:31-58), the default strategy: each doc is tokenized (tiktoken or HuggingFace), then a sliding window of chunk_token_size (default 1200) tokens with chunk_overlap_token_size (default 100) overlap produces chunks. An alternative chunking_by_seperators method is also implemented, splitting on sentence/paragraph boundaries. Each chunk gets an md5-hash ID prefixed "chunk-".

Entity and relation extraction. The primary extraction function extract_entities (_op.py:282-414) sends each chunk through the LLM designated as best_model_func (default gpt-4o). The prompt entity_extraction (prompt.py:195-294) asks the model to output ("entity"<|>NAME<|>TYPE<|>DESC) and ("relationship"<|>SRC<|>TGT<|>DESC<|>STRENGTH) tuples separated by ## delimiters. Default entity types are organization, person, geo, event. Gleaning runs entity_extract_max_gleaning (default 1) additional passes: the model receives its prior output and a continuation prompt, then decides YES/NO whether more entities remain via entiti_if_loop_extraction (prompt.py:320-322). An alternative DSPy-based pipeline exists in entity_extraction/module.py using dspy.ChainOfThought with Pydantic schemas, and can optionally run critique-then-refine cycles.

Entity resolution / deduplication. _merge_nodes_then_upsert (_op.py:182-227) is called per unique entity name across all chunks. It loads any existing node from the graph, merges descriptions (sorted, deduplicated, joined by <SEP>), picks the majority entity type, aggregates source IDs, and calls _handle_entity_relation_summary which conditionally uses the cheap_model_func to condense descriptions that exceed entity_summary_to_max_tokens (500). Edges are merged similarly in _merge_edges_then_upsert (_op.py:230-279), summing weights, taking the minimum order, and merging descriptions. If an edge mentions nodes not yet in the graph, they are auto-created as type "UNKNOWN". The undirected graph is enforced by sorting edge tuples before storage.

Schema. There is no formal ontology beyond the four default entity types; the prompt passes entity types as a comma-separated string. Node attributes: entity_type, description, source_id. Edge attributes: weight, description, source_id, order.

Where and how is the graph stored?

answered

Graph storage. The default graph backend is NetworkXStorage (_storage/gdb_networkx.py), an in-memory networkx.Graph that is serialized to a GraphML XML file on disk (graph_{namespace}.graphml) at index_done_callback time (gdb_networkx.py:80-98). On startup, if the GraphML file exists it is deserialized back. Nodes are stored as NetworkX node attributes: each upsert_node call adds a node with attributes dict containing entity_type, description, source_id, and after clustering a serialized JSON clusters field. Edges are NetworkX edge attributes with weight, description, source_id, order. An optional Neo4j backend (gdb_neo4j.py) mirrors the same operations via Cypher queries, storing all nodes under a single label derived from the working directory path with community IDs stored in a communityIds array property.

Key-value storage. JsonKVStorage (_storage/kv_json.py) stores Python dicts as flat JSON files under the working directory: kv_store_full_docs.json, kv_store_text_chunks.json, kv_store_llm_response_cache.json, kv_store_community_reports.json. These are loaded on init and flushed to disk at callbacks.

Vector storage. The default vector DB is NanoVectorDBStorage (_storage/vdb_nanovectordb.py), backed by the nano-vectordb library which persists to a JSON file (vdb_entities.json or vdb_chunks.json). Two vector DBs are created per GraphRAG instance: entities_vdb stores entity embeddings (content = entity_name + description, with entity_name as a meta field) for local-mode queries, and optionally chunks_vdb stores chunk embeddings for naive RAG mode. The embedding function (default OpenAI text-embedding-3-small, 1536-dim) is called in batches of embedding_batch_num (32), and concurrency is limited by embedding_func_max_async (16). Embeddings are concatenated into numpy arrays and upserted into NanoVectorDB, which supports cosine-similarity query with a better_than_threshold filter (default 0.2).

Are communities, summaries or hierarchies built over the graph?

answered

Community detection. The clustering method on NetworkXStorage (gdb_networkx.py:165-168,230-252) runs the Leiden algorithm via graspologic.partition.hierarchical_leiden. Before running, the graph is passed through stable_largest_connected_component which discards disconnected subgraphs and stabilizes node ordering. The Leiden parameters include max_cluster_size (default 10) and graph_cluster_seed (0xDEADBEEF). The hierarchical partition produces multiple levels; each node is annotated with {"level": level_key, "cluster": cluster_id} entries stored as a JSON-serialized clusters attribute. The Neo4j backend (gdb_neo4j.py:387-440) instead runs Leiden via GDS (Graph Data Science) library procedures, projecting the graph and writing community IDs back into node properties.

Community schema. After clustering, community_schema() (gdb_networkx.py:170-224) iterates all nodes, groups them by (level, cluster) key, collects member nodes, edges (from node adjacency), source-chunk IDs, and computes occurrence as the ratio of chunk IDs to the max across all communities. It then builds the hierarchy by linking each community to its sub-communities (strict subset of nodes at the next level down).

Community reports. generate_community_report (_op.py:625-697) processes communities level-by-level from highest (most fine-grained) to lowest (root). For each community, _pack_single_community_describe (_op.py:464-600) assembles a CSV of entities, relationships, and optional sub-community reports (used for large communities with >100 nodes or edges), truncating each section by token budget. The LLM (best_model_func) is prompted with community_report (prompt.py:63-192) asking for a JSON-structured report with title, summary, an impact severity rating (0-10), and 5-10 detailed findings. Results are stored as both report_string (markdown) and report_json in the community_reports KV store. Reports are recomputed on every document insertion — the code explicitly comments "TODO: don't support incremental update for communities now, so we have to drop all" (graphrag.py:318).

How does query-time retrieval use the graph?

answered

Three query modes are selected via QueryParam.mode: "local", "global", or "naive" (graphrag.py:236-275).

Local mode (_op.py:935-967, builder at 844-933). Step 1: embed the query string and search entities_vdb (cosine similarity, top_k=20) to get the most relevant entity vectors. Step 2: look up those entities in the graph and expand to one-hop neighbors via get_nodes_edges_batch. Step 3: from those entity nodes, gather community reports via _find_most_related_community_from_entities — filters to communities at or below query_param.level (default 2), sorts by occurrence count and rating, and truncates to local_max_token_for_community_report tokens (3200). Step 4: gather source text chunks from the entities and their neighbors via _find_most_related_text_unit_from_entities, merging them with relation counts for deduplication and sorting, truncated to local_max_token_for_text_unit (4000). Step 5: gather and rank edges via _find_most_related_edges_from_entities, truncated to local_max_token_for_local_context (4800). All four sections are assembled into CSV tables (Reports, Entities, Relationships, Sources) and fed as context to the LLM with the local_rag_response prompt. Optional local_community_single_one flag limits to the top community.

Global mode (_op.py:1017-1104). Retrieves all community schemas from the graph, filters to level <= query_param.level, sorts by occurrence, caps at global_max_consider_community (512), and filters by global_min_community_rating. Remaining communities are grouped into token-budget-limited batches (global_max_token_for_community_report, 16384). For each group, an LLM (acting as "analyst") extracts key points with importance scores using global_map_rag_points prompt. Results are flattened, filtered by score>0, sorted, truncated again, then a final LLM call with global_reduce_rag_response synthesizes the analysts' perspectives into a unified answer.

Naive mode (_op.py:1107-1140). Simple vector search on chunks_vdb, retrieves top_k chunks, truncates to naive_max_token_for_text_unit (12000), and feeds them to the LLM.

All three modes support only_need_context=True to return the raw retrieved context without an LLM generation call.

Editor's note. Correction: only_need_context=True avoids LLM calls only in local and naive modes; global mode still runs the map-stage LLM call per community group and returns the scored analyst points, skipping only the reduce call.

How are updates and incremental indexing handled?

answered

Incremental indexing has limited support with known gaps. On each call to ainsert (graphrag.py:277-348), new documents are checked against full_docs KV storage by their md5-hash ID: if the hash already exists, the document is skipped (filter_keys on line 287). Similarly, chunks are checked against text_chunks storage (line 304). So re-inserting the same document is idempotent — a no-op if all hashes match.

Graph merging. When entities or relations from new chunks overlap with existing graph nodes, the merge functions _merge_nodes_then_upsert and _merge_edges_then_upsert (_op.py:182-279) read existing node/edge data from the graph, merge descriptions (deduplicated, sorted, joined by <SEP>), and update. Edge weights are summed across extraction runs. This means the graph grows incrementally without data loss.

No incremental communities. The code explicitly states: # TODO: don't support incremental update for communities now, so we have to drop all (graphrag.py:318), followed by await self.community_reports.drop(). Every call to ainsert wipes the community reports and recomputes from scratch (clustering → community report generation). This is the biggest gap in incremental support.

Embedding re-upsert. Entity embeddings are recomputed for all extracted entities and upserted into the vector DB; the vector DB storage (NanoVectorDBStorage.upsert) uses NanoVectorDB.upsert which updates existing entries by matching on __id__.

Deletion. There is no API for deleting documents, entities, or edges. The BaseKVStorage.drop() method exists and wipes the in-memory dict (used only for community reports), but the full_docs and text_chunks stores are append-only within a session. A deletion would require manual file removal and re-indexing.

LLM cache persistence. The llm_response_cache KV store persists across sessions (it is a JSON file loaded at init), so repeated extraction or query calls benefit from cached LLM responses even after a restart — but cache entries are keyed by (model, messages) hash, not by document ID, so a re-insert of the same document with the same chunk content would reuse cached extractions.

How are LLM cost and latency controlled during indexing and query?

answered

Two-tier model routing. The system separates a best_model_func (default gpt-4o) for expensive core tasks — entity extraction, query response generation, community report generation — from a cheap_model_func (default gpt-4o-mini) used only for entity/relation description summarization when descriptions exceed entity_summary_to_max_tokens. Both functions are wrapped by limit_async_func_call which caps concurrent requests (best_model_max_async and cheap_model_max_async, both default 16) via a spin-loop semaphore (_utils.py:276-295) rather than asyncio.Semaphore for compatibility with nested event loops.

LLM response caching. All LLM calls pass through openai_complete_if_cache (_llm.py:50-75). Before making an API call, it computes an MD5 hash of (model, messages) via compute_args_hash and checks the hashing_kv (a JsonKVStorage persisted as kv_store_llm_response_cache.json). On a cache hit, the stored response is returned with zero latency and zero cost. On a miss, the API is called and the result is stored. The cache is checked on every LLM call — extraction, summarization, community report generation, and query — and persists across sessions. The enable_llm_cache flag (default True) controls whether the cache storage is created.

Retry with backoff. Network calls to OpenAI, Azure OpenAI, and Amazon Bedrock are decorated with tenacity.retry using exponential backoff (wait_exponential multiplier=1s, min=4s, max=10s) and up to 5 attempts, retrying only on RateLimitError and APIConnectionError (_llm.py:45-49). This avoids wasting tokens on failed calls and handles API throttling gracefully.

Token budget truncation. Throughout retrieval and context assembly, truncate_list_by_token_size (_utils.py:169-183) is called repeatedly with configurable max_token_size parameters: local_max_token_for_text_unit (4000), local_max_token_for_local_context (4800), local_max_token_for_community_report (3200), global_max_token_for_community_report (16384), naive_max_token_for_text_unit (12000). These per-section caps bound the prompt size sent to the LLM, preventing context overflow and limiting per-call cost.

Embedding batching. Embedding requests are batched by embedding_batch_num (default 32) in NanoVectorDBStorage.upsert (_storage/vdb_nanovectordb.py:40-47). Batches are sent concurrently via asyncio.gather, with an overall concurrency limit of embedding_func_max_async (default 16) enforced by the wrapper on the embedding function. This optimizes throughput while respecting API rate limits.