gusye1234/nano-graphrag
A small, readable Python reimplementation of Microsoft GraphRAG with local, global and naive query modes over a NetworkX graph.
Overview
nano-graphrag is a compact reimplementation of the Microsoft GraphRAG paper. It keeps the core recipe and drops almost everything else. Text is chunked, an LLM extracts entities and relationships from each chunk, the results are merged into a NetworkX graph, hierarchical Leiden finds communities, and an LLM writes a JSON report for each community. At query time you pick local (entity vectors plus graph neighbourhood), global (map-reduce over community reports) or naive (plain chunk RAG).
The whole engine is two files. graphrag.py holds the GraphRAG dataclass, which is both the configuration object and the orchestrator. _op.py holds every pipeline step as a plain async function. Storage, LLM and embedding are injected as classes or callables, so most customisation is a constructor argument, not a subclass. The README’s claim of about 1,100 lines of core code is roughly right: _op.py is 1,140 lines and the rest is small.
It is a library for people who want to read, fork and change a GraphRAG pipeline. It has no server, no UI, no deletion API and no multi-tenant story. Persistence is a working directory of JSON and GraphML files. Think of it as a reference implementation you build on, not a service you deploy.
Architecture
flowchart LR
IN["insert(text)"] --> GR["GraphRAG.ainsert"]
GR --> CH["get_chunks (token windows)"]
CH --> EX["extract_entities (best LLM + gleaning)"]
EX --> MG["merge nodes / edges"]
MG --> G["NetworkXStorage (GraphML)"]
MG --> EV["entities_vdb (nano-vectordb)"]
G --> LD["hierarchical_leiden on LCC"]
LD --> CR["generate_community_report"]
CR --> KV["community_reports (JSON KV)"]
Q["query(q, mode)"] --> LQ["local_query"]
Q --> GQ["global_query (map-reduce)"]
Q --> NQ["naive_query"]
LQ --> EV
LQ --> G
LQ --> KV
GQ --> KV
| Component | Path | Role |
|---|---|---|
| Orchestrator | nano_graphrag/graphrag.py |
GraphRAG dataclass: all defaults, storage wiring, ainsert, aquery |
| Pipeline ops | nano_graphrag/_op.py |
Chunking, extraction, merging, community reports, the three query modes |
| Interfaces | nano_graphrag/base.py |
QueryParam, BaseKVStorage, BaseVectorStorage, BaseGraphStorage |
| Graph stores | nano_graphrag/_storage/gdb_networkx.py, gdb_neo4j.py |
NetworkX + GraphML (default) and Neo4j + GDS Leiden |
| Vector stores | nano_graphrag/_storage/vdb_nanovectordb.py, vdb_hnswlib.py |
nano-vectordb JSON file (default) and hnswlib |
| KV store | nano_graphrag/_storage/kv_json.py |
One JSON file per namespace |
| LLM adapters | nano_graphrag/_llm.py |
OpenAI, Azure OpenAI and Bedrock completions and embeddings, cache and retry |
| Prompts | nano_graphrag/prompt.py |
PROMPTS dict: extraction, community report, local/global/naive answers |
| DSPy extractor | nano_graphrag/entity_extraction/ |
Alternative typed extractor with self-refine, plus dataset tools |
How a request flows
Insert (graph.insert(text)):
- Dedup documents.
ainserthashes each document with MD5 into adoc-id and drops ids already infull_docs(graphrag.py). - Chunk.
get_chunkstokenizes and callschunk_func. The defaultchunking_by_token_sizecuts 1,200-token windows with 100 tokens of overlap. Chunk ids are MD5 hashes of chunk text, so repeated chunks are skipped (_op.py, L94-L108). - Drop old reports.
community_reports.drop()runs before extraction, because communities are always rebuilt from scratch. - Extract.
extract_entitiessends every chunk through the “best” model with theentity_extractionprompt, then runs up toentity_extract_max_gleaningcontinuation passes. The model emits("entity"<|>NAME<|>TYPE<|>DESC)and("relationship"<|>SRC<|>TGT<|>DESC<|>WEIGHT)records, parsed with regex and delimiter splits (_op.py). - Merge. Names are upper-cased.
_merge_nodes_then_upserttakes the majority type and joins distinct descriptions with<SEP>._merge_edges_then_upsertsums weights and creates missing endpoints asUNKNOWNnodes. When a merged description passes 500 tokens, the “cheap” model summarises it (_op.py, L182-L279). Each entity’s name plus description is embedded intoentities_vdb. - Cluster.
NetworkXStorage._leiden_clusteringruns graspologic’shierarchical_leiden(max cluster size 10, fixed seed) on the stable largest connected component and writes a JSONclusterslist onto each node (gdb_networkx.py). - Report.
generate_community_reportwalks levels from the finest to the root. For each community it packs entities and edges as CSV, ordered by degree, and reuses sub-community reports when a community has over 100 nodes or edges. The best model returns JSON with title, summary, rating and findings (_op.py, L625-L697). - Persist.
_insert_donecallsindex_done_callbackon every store, which writes the JSON, GraphML and vector files.
Query (graph.query(q, QueryParam(mode=...)), default mode global): aquery dispatches to one of three functions (graphrag.py). Local mode is described below. Global mode filters communities to level <= 2, sorts by occurrence, keeps up to 512, then runs map and reduce LLM calls. Naive mode is top-k chunk retrieval.
Key components
Local query
_build_local_query_context is the most interesting retrieval code (_op.py). It takes the top 20 entities by vector similarity and builds four CSV tables, each with its own token budget:
- Reports. Communities the entities belong to, ranked by how many hits they got and then by rating (3,200 tokens).
- Entities. The hits, with graph degree as
rank. - Relationships. Every edge touching a hit, ranked by edge degree and weight (4,800 tokens).
- Sources. The source chunks of the hits, preferring chunks that one-hop neighbours also mention (4,000 tokens).
Traversal is exactly one hop. There is no multi-hop path search.
Global query
global_query packs community reports into groups of 16,384 tokens. It asks the model, once per group, for scored “points” in JSON, keeps points with score above 0, and sends the best ones to a final reduce prompt (_op.py). With only_need_context=True it still makes the map calls. It skips only the reduce call.
Storage contracts
base.py defines three async interfaces. The graph interface is the largest: node and edge getters, batch variants, degrees, upsert_*, clustering and community_schema. NetworkXStorage.community_schema rebuilds the community table from the node clusters attributes. A community’s sub-communities are the communities on the next level whose nodes are a subset of its own (gdb_networkx.py). The Neo4j backend projects the graph into GDS and calls gds.leiden.write. It passes max_graph_cluster_size as maxLevels, so that knob means something different on the two backends (gdb_neo4j.py).
LLM layer, cache and concurrency
Every built-in completion function checks an MD5 hash of (model, messages) against the llm_response_cache KV store before it calls the API. It retries rate-limit and connection errors up to five times (_llm.py). On a miss, it writes the whole cache file back to disk after each call. Concurrency is capped by limit_async_func_call, a busy-wait counter instead of an asyncio.Semaphore (16 by default per model and for embeddings) (_utils.py).
Extending it
- Models. Pass any async
fn(prompt, system_prompt=None, history_messages=[], **kwargs)asbest_model_funcorcheap_model_func. Pophashing_kvfrom kwargs if you want the cache. The examples folder covers DeepSeek, Ollama, Bedrock and fully local setups. - Embeddings. Wrap a function with
wrap_embedding_func_with_attrs(embedding_dim=..., max_token_size=...)and pass it asembedding_func. - Storage. Pass
graph_storage_cls,vector_db_storage_clsorkey_string_value_json_storage_cls. The built-ins are Neo4j and hnswlib. FAISS, Milvus and Qdrant ship as examples. - Extraction.
entity_extraction_funcis swappable.entity_extraction/extract.pyprovides a DSPy-based extractor with typed entities and optional self-refine.chunk_funccan bechunking_by_seperatorsor your own. - Prompts. Edit
PROMPTSin place. Entity types default to organization, person, geo and event. - Malformed JSON.
convert_response_to_json_funclets you plug in a repair function for models that return broken JSON.
Running it
pip install nano-graphrag(orpip install -e .). Python 3.9+.requirements.txtinstalls openai, graspologic, nano-vectordb, hnswlib, dspy-ai, neo4j and aioboto3 unconditionally.- Set
OPENAI_API_KEY, or passusing_azure_openai=Trueorusing_amazon_bedrock=Truewith model ids. The defaults aregpt-4o,gpt-4o-miniandtext-embedding-3-small. GraphRAG(working_dir="./x")theninsert(...)andquery(...). Reopening the same directory reloads everything. Naive mode needsenable_naive_rag=Trueat construction time. Neo4j needsaddon_params={"neo4j_url": ..., "neo4j_auth": ...}and the GDS plugin.
Strengths and caveats
- Strength: readable. The pipeline is a few hundred lines of plain async functions with GraphRAG’s prompts. It is the easiest GraphRAG codebase in this category to read end to end.
- Strength: easy to swap parts. Model, embedding, chunker, extractor and all three stores are constructor arguments with small interfaces.
- Strength: idempotent re-inserts. Content hashing for documents and chunks, plus the persistent LLM cache, make repeated runs cheap.
- Caveat: every insert redoes all communities. Each
ainsertdrops all reports, re-clusters and re-summarises every community. Report cost grows with the whole graph, not with the new document. If extraction finds no entities, the reports are dropped and not rebuilt. - Caveat: only the largest connected component is clustered. Entities outside it get no community and never reach global search.
- Caveat: no deletion or update. There is no API to remove a document. Rebuilding means deleting the working directory.
- Caveat: single-process files. JSON and GraphML stores are rewritten in full on each callback, and the LLM cache is rewritten after every uncached call. That is fine for a notebook and poor for large corpora or concurrent writers.
- Caveat: some dead code.
node2vec_paramsandembed_nodesexist, but no pipeline step calls them.
Sources: code at acb35c0, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (31 pages), verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit acb35c0. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredChunking. Documents are split via chunking_by_token_size (_op.py:31-58), the default strategy: each doc is tokenized (tiktoken or HuggingFace), then a sliding window of chunk_token_size (default 1200) tokens with chunk_overlap_token_size (default 100) overlap produces chunks. An alternative chunking_by_seperators method is also implemented, splitting on sentence/paragraph boundaries. Each chunk gets an md5-hash ID prefixed "chunk-".
Entity and relation extraction. The primary extraction function extract_entities (_op.py:282-414) sends each chunk through the LLM designated as best_model_func (default gpt-4o). The prompt entity_extraction (prompt.py:195-294) asks the model to output ("entity"<|>NAME<|>TYPE<|>DESC) and ("relationship"<|>SRC<|>TGT<|>DESC<|>STRENGTH) tuples separated by ## delimiters. Default entity types are organization, person, geo, event. Gleaning runs entity_extract_max_gleaning (default 1) additional passes: the model receives its prior output and a continuation prompt, then decides YES/NO whether more entities remain via entiti_if_loop_extraction (prompt.py:320-322). An alternative DSPy-based pipeline exists in entity_extraction/module.py using dspy.ChainOfThought with Pydantic schemas, and can optionally run critique-then-refine cycles.
Entity resolution / deduplication. _merge_nodes_then_upsert (_op.py:182-227) is called per unique entity name across all chunks. It loads any existing node from the graph, merges descriptions (sorted, deduplicated, joined by <SEP>), picks the majority entity type, aggregates source IDs, and calls _handle_entity_relation_summary which conditionally uses the cheap_model_func to condense descriptions that exceed entity_summary_to_max_tokens (500). Edges are merged similarly in _merge_edges_then_upsert (_op.py:230-279), summing weights, taking the minimum order, and merging descriptions. If an edge mentions nodes not yet in the graph, they are auto-created as type "UNKNOWN". The undirected graph is enforced by sorting edge tuples before storage.
Schema. There is no formal ontology beyond the four default entity types; the prompt passes entity types as a comma-separated string. Node attributes: entity_type, description, source_id. Edge attributes: weight, description, source_id, order.
Where and how is the graph stored?
answeredGraph storage. The default graph backend is NetworkXStorage (_storage/gdb_networkx.py), an in-memory networkx.Graph that is serialized to a GraphML XML file on disk (graph_{namespace}.graphml) at index_done_callback time (gdb_networkx.py:80-98). On startup, if the GraphML file exists it is deserialized back. Nodes are stored as NetworkX node attributes: each upsert_node call adds a node with attributes dict containing entity_type, description, source_id, and after clustering a serialized JSON clusters field. Edges are NetworkX edge attributes with weight, description, source_id, order. An optional Neo4j backend (gdb_neo4j.py) mirrors the same operations via Cypher queries, storing all nodes under a single label derived from the working directory path with community IDs stored in a communityIds array property.
Key-value storage. JsonKVStorage (_storage/kv_json.py) stores Python dicts as flat JSON files under the working directory: kv_store_full_docs.json, kv_store_text_chunks.json, kv_store_llm_response_cache.json, kv_store_community_reports.json. These are loaded on init and flushed to disk at callbacks.
Vector storage. The default vector DB is NanoVectorDBStorage (_storage/vdb_nanovectordb.py), backed by the nano-vectordb library which persists to a JSON file (vdb_entities.json or vdb_chunks.json). Two vector DBs are created per GraphRAG instance: entities_vdb stores entity embeddings (content = entity_name + description, with entity_name as a meta field) for local-mode queries, and optionally chunks_vdb stores chunk embeddings for naive RAG mode. The embedding function (default OpenAI text-embedding-3-small, 1536-dim) is called in batches of embedding_batch_num (32), and concurrency is limited by embedding_func_max_async (16). Embeddings are concatenated into numpy arrays and upserted into NanoVectorDB, which supports cosine-similarity query with a better_than_threshold filter (default 0.2).
Are communities, summaries or hierarchies built over the graph?
answeredCommunity detection. The clustering method on NetworkXStorage (gdb_networkx.py:165-168,230-252) runs the Leiden algorithm via graspologic.partition.hierarchical_leiden. Before running, the graph is passed through stable_largest_connected_component which discards disconnected subgraphs and stabilizes node ordering. The Leiden parameters include max_cluster_size (default 10) and graph_cluster_seed (0xDEADBEEF). The hierarchical partition produces multiple levels; each node is annotated with {"level": level_key, "cluster": cluster_id} entries stored as a JSON-serialized clusters attribute. The Neo4j backend (gdb_neo4j.py:387-440) instead runs Leiden via GDS (Graph Data Science) library procedures, projecting the graph and writing community IDs back into node properties.
Community schema. After clustering, community_schema() (gdb_networkx.py:170-224) iterates all nodes, groups them by (level, cluster) key, collects member nodes, edges (from node adjacency), source-chunk IDs, and computes occurrence as the ratio of chunk IDs to the max across all communities. It then builds the hierarchy by linking each community to its sub-communities (strict subset of nodes at the next level down).
Community reports. generate_community_report (_op.py:625-697) processes communities level-by-level from highest (most fine-grained) to lowest (root). For each community, _pack_single_community_describe (_op.py:464-600) assembles a CSV of entities, relationships, and optional sub-community reports (used for large communities with >100 nodes or edges), truncating each section by token budget. The LLM (best_model_func) is prompted with community_report (prompt.py:63-192) asking for a JSON-structured report with title, summary, an impact severity rating (0-10), and 5-10 detailed findings. Results are stored as both report_string (markdown) and report_json in the community_reports KV store. Reports are recomputed on every document insertion — the code explicitly comments "TODO: don't support incremental update for communities now, so we have to drop all" (graphrag.py:318).
How does query-time retrieval use the graph?
answeredThree query modes are selected via QueryParam.mode: "local", "global", or "naive" (graphrag.py:236-275).
Local mode (_op.py:935-967, builder at 844-933). Step 1: embed the query string and search entities_vdb (cosine similarity, top_k=20) to get the most relevant entity vectors. Step 2: look up those entities in the graph and expand to one-hop neighbors via get_nodes_edges_batch. Step 3: from those entity nodes, gather community reports via _find_most_related_community_from_entities — filters to communities at or below query_param.level (default 2), sorts by occurrence count and rating, and truncates to local_max_token_for_community_report tokens (3200). Step 4: gather source text chunks from the entities and their neighbors via _find_most_related_text_unit_from_entities, merging them with relation counts for deduplication and sorting, truncated to local_max_token_for_text_unit (4000). Step 5: gather and rank edges via _find_most_related_edges_from_entities, truncated to local_max_token_for_local_context (4800). All four sections are assembled into CSV tables (Reports, Entities, Relationships, Sources) and fed as context to the LLM with the local_rag_response prompt. Optional local_community_single_one flag limits to the top community.
Global mode (_op.py:1017-1104). Retrieves all community schemas from the graph, filters to level <= query_param.level, sorts by occurrence, caps at global_max_consider_community (512), and filters by global_min_community_rating. Remaining communities are grouped into token-budget-limited batches (global_max_token_for_community_report, 16384). For each group, an LLM (acting as "analyst") extracts key points with importance scores using global_map_rag_points prompt. Results are flattened, filtered by score>0, sorted, truncated again, then a final LLM call with global_reduce_rag_response synthesizes the analysts' perspectives into a unified answer.
Naive mode (_op.py:1107-1140). Simple vector search on chunks_vdb, retrieves top_k chunks, truncates to naive_max_token_for_text_unit (12000), and feeds them to the LLM.
All three modes support only_need_context=True to return the raw retrieved context without an LLM generation call.
only_need_context=True avoids LLM calls only in local and naive modes; global mode still runs the map-stage LLM call per community group and returns the scored analyst points, skipping only the reduce call.How are updates and incremental indexing handled?
answeredIncremental indexing has limited support with known gaps. On each call to ainsert (graphrag.py:277-348), new documents are checked against full_docs KV storage by their md5-hash ID: if the hash already exists, the document is skipped (filter_keys on line 287). Similarly, chunks are checked against text_chunks storage (line 304). So re-inserting the same document is idempotent — a no-op if all hashes match.
Graph merging. When entities or relations from new chunks overlap with existing graph nodes, the merge functions _merge_nodes_then_upsert and _merge_edges_then_upsert (_op.py:182-279) read existing node/edge data from the graph, merge descriptions (deduplicated, sorted, joined by <SEP>), and update. Edge weights are summed across extraction runs. This means the graph grows incrementally without data loss.
No incremental communities. The code explicitly states: # TODO: don't support incremental update for communities now, so we have to drop all (graphrag.py:318), followed by await self.community_reports.drop(). Every call to ainsert wipes the community reports and recomputes from scratch (clustering → community report generation). This is the biggest gap in incremental support.
Embedding re-upsert. Entity embeddings are recomputed for all extracted entities and upserted into the vector DB; the vector DB storage (NanoVectorDBStorage.upsert) uses NanoVectorDB.upsert which updates existing entries by matching on __id__.
Deletion. There is no API for deleting documents, entities, or edges. The BaseKVStorage.drop() method exists and wipes the in-memory dict (used only for community reports), but the full_docs and text_chunks stores are append-only within a session. A deletion would require manual file removal and re-indexing.
LLM cache persistence. The llm_response_cache KV store persists across sessions (it is a JSON file loaded at init), so repeated extraction or query calls benefit from cached LLM responses even after a restart — but cache entries are keyed by (model, messages) hash, not by document ID, so a re-insert of the same document with the same chunk content would reuse cached extractions.
How are LLM cost and latency controlled during indexing and query?
answeredTwo-tier model routing. The system separates a best_model_func (default gpt-4o) for expensive core tasks — entity extraction, query response generation, community report generation — from a cheap_model_func (default gpt-4o-mini) used only for entity/relation description summarization when descriptions exceed entity_summary_to_max_tokens. Both functions are wrapped by limit_async_func_call which caps concurrent requests (best_model_max_async and cheap_model_max_async, both default 16) via a spin-loop semaphore (_utils.py:276-295) rather than asyncio.Semaphore for compatibility with nested event loops.
LLM response caching. All LLM calls pass through openai_complete_if_cache (_llm.py:50-75). Before making an API call, it computes an MD5 hash of (model, messages) via compute_args_hash and checks the hashing_kv (a JsonKVStorage persisted as kv_store_llm_response_cache.json). On a cache hit, the stored response is returned with zero latency and zero cost. On a miss, the API is called and the result is stored. The cache is checked on every LLM call — extraction, summarization, community report generation, and query — and persists across sessions. The enable_llm_cache flag (default True) controls whether the cache storage is created.
Retry with backoff. Network calls to OpenAI, Azure OpenAI, and Amazon Bedrock are decorated with tenacity.retry using exponential backoff (wait_exponential multiplier=1s, min=4s, max=10s) and up to 5 attempts, retrying only on RateLimitError and APIConnectionError (_llm.py:45-49). This avoids wasting tokens on failed calls and handles API throttling gracefully.
Token budget truncation. Throughout retrieval and context assembly, truncate_list_by_token_size (_utils.py:169-183) is called repeatedly with configurable max_token_size parameters: local_max_token_for_text_unit (4000), local_max_token_for_local_context (4800), local_max_token_for_community_report (3200), global_max_token_for_community_report (16384), naive_max_token_for_text_unit (12000). These per-section caps bound the prompt size sent to the LLM, preventing context overflow and limiting per-call cost.
Embedding batching. Embedding requests are batched by embedding_batch_num (default 32) in NanoVectorDBStorage.upsert (_storage/vdb_nanovectordb.py:40-47). Batches are sent concurrently via asyncio.gather, with an overall concurrency limit of embedding_func_max_async (default 16) enforced by the wrapper on the embedding function. This optimizes throughput while respecting API rate limits.