HKUDS/LightRAG
Python graph RAG engine that merges LLM-extracted entities into one graph and retrieves by keyword-matched entities, relations and chunks.
Overview
LightRAG is a graph RAG engine from HKU’s data science lab. It ships as a Python library (lightrag-hku), a FastAPI server (lightrag-server) and a React web UI. Like GraphRAG, it has an LLM read every chunk and extract entities and relationships. The difference is what happens next. LightRAG merges the results straight into one growing graph keyed by entity name, and it embeds every entity and every relationship into their own vector stores. It never clusters the graph or writes community summaries. All “global” reasoning happens at query time.
A query starts with one LLM call that splits the question into low-level keywords (specific names) and high-level keywords (themes). Low-level keywords search the entity vectors. High-level keywords search the relationship vectors. Both expand one hop through the graph. The default mix mode also adds plain vector hits over chunks. Because documents merge into the graph one at a time, adding a document needs no rebuild, and deleting one is a supported operation.
The project has grown well beyond the paper. At this SHA it includes a multi-stage ingestion pipeline with crash-recovery journaling, document parsers (native, MinerU, Docling), a vision-model analysis step for images and tables, four chunking strategies, and about 25 storage implementations. The top-level modules of the package alone are about 46,000 lines of Python.
Architecture
flowchart LR
U["Client / Web UI"] --> API["FastAPI server"]
API --> LR["LightRAG class"]
LR --> P["Pipeline: parse, analyze, process"]
P --> CH["Chunker"]
CH --> EX["extract_entities (LLM)"]
EX --> M["merge_nodes_and_edges"]
M --> G["Graph storage"]
M --> EV["Entity vectors"]
M --> RV["Relation vectors"]
CH --> CV["Chunk vectors + KV"]
LR --> KQ["kg_query"]
KQ --> KW["Keyword extraction (LLM)"]
KW --> EV
KW --> RV
KQ --> G
KQ --> CV
KQ --> ANS["Answer (LLM)"]
| Component | Path | Role |
|---|---|---|
| Core class | lightrag/lightrag.py |
LightRAG dataclass: config, storage wiring, insert, query, delete and graph edit APIs |
| Pipeline | lightrag/pipeline.py |
Enqueue, parse, analyze and process workers; document status; crash recovery |
| Operators | lightrag/operate.py |
Extraction, merging, summarization, kg_query, naive_query, context building |
| Prompts | lightrag/prompt.py |
Extraction, keyword and answer templates; default entity types |
| Chunkers | lightrag/chunker/ |
Fixed-token (default), recursive-character, semantic-vector, paragraph-semantic |
| Storage registry | lightrag/kg/__init__.py |
KV, vector, graph and doc-status implementations |
| LLM roles | lightrag/llm_roles.py |
Separate extract, keyword, query and vlm model bindings |
| Provider bindings | lightrag/llm/ |
OpenAI, Azure, Ollama, Anthropic, Gemini, Bedrock, Hugging Face and others |
| API server | lightrag/api/ |
Document, query and graph routes, plus an Ollama-compatible chat API |
| Web UI | lightrag_webui/ |
Document manager, graph viewer and query panel |
How a request flows
Insert
ainsertturns its arguments into per-document chunk options, then callsapipeline_enqueue_documentsandapipeline_process_enqueue_documents(lightrag.py). The API server calls those two methods directly, so it can choose a chunking strategy per document.- Enqueue gives each document an ID from an MD5 hash of its content (or file path) and skips duplicates. The process worker chunks the text (default 1,200 tokens with 100 overlap).
- Chunks are written to the chunk vector store and the chunk KV store first. Then
_process_extract_entitiesruns extraction, unless the document opted out of the graph (pipeline.py). extract_entitiessends each chunk to theextractrole model, with up tollm_model_max_asynccalls at once. Every call goes through the LLM cache (operate.py).merge_nodes_and_edgesmerges the per-chunk results into the graph. It writes the document’s entity and relation lists first as recovery anchors, then upserts nodes, edges and their vectors (pipeline.py).
Query
aquery_llmsendslocal,global,hybridandmixtokg_query,naivetonaive_query, andbypassstraight to the query model (lightrag.py).kg_queryextracts keywords with thekeywordmodel, then_perform_kg_searchembeds the query and both keyword strings in one batch. It runs the entity-side search and the relation-side search, merges them round-robin, and inmixmode adds chunk vector hits (operate.py)._build_query_contextcuts entities and relations to their token budgets, picks the chunks linked to what survived, and formats one context with knowledge-graph and document-chunk sections. Thequerymodel answers. The answer is cached under a hash of the query, mode, retrieval settings, keywords and model identity (operate.py).
Key components
Extraction and merging
The default prompt asks for entities with a type from an 11-class list (Person, Organization, Location, Event, Concept, Method and others, plus Other) and for relationships with keywords and a description. Records are separated by <|#|> (prompt.py). A JSON output mode is also available. Gleaning runs at most once. It is skipped when the replayed conversation would exceed MAX_EXTRACT_INPUT_TOKENS (default 20,480), and a gleaned description replaces the first one only if it is longer (operate.py).
Entities merge across chunks and documents by normalized name. The entity type is decided by majority vote. Descriptions are joined without an LLM call until there are 8 fragments or they pass the token limit. After that, _handle_entity_relation_summary summarizes them, recursively if needed (operate.py).
Retrieval
_get_node_data takes the top-k entities from the entity vector store (top_k default 40), reads their degrees, and collects all their edges, ranked by edge degree and weight (operate.py). The relation-side search does the reverse: top-k relations, then their endpoint entities. Traversal is one hop. “Global” in LightRAG means relation-centric search, not corpus-wide summaries. Token budgets default to 6,000 for entities, 8,000 for relations and 30,000 in total. Reranking of chunks is on by default when a rerank model is configured (base.py).
Storage
Four storage roles each take a class name: KV, vector, graph and document status (kg/init.py). The defaults are JSON files, NanoVectorDB, NetworkX and a JSON doc-status store, all held in memory and flushed to the working directory (lightrag.py). NetworkXStorage rewrites one GraphML file and allows only one writer per workspace (networkx_impl.py). PostgreSQL, MongoDB and OpenSearch can each serve all four roles. Neo4j and Memgraph can serve as the graph store, and Milvus, Qdrant and Faiss as vector stores.
Extending it
- Storage: implement the
BaseKVStorage,BaseVectorStorage,BaseGraphStorageorDocStatusStorageinterface, then register the class inSTORAGES(kg/__init__.py), whichkg/factory.pyloads by name. - Models: pass any async
llm_model_funcandembedding_func, or set per-role overrides withRoleLLMConfigorEXTRACT_LLM_BINDING,QUERY_LLM_BINDINGand similar env vars (llm_roles.py). - Prompts and schema: override
addon_params['entity_types_guidance']or replace entries inPROMPTS. Akg_extraction_validatorhook can check each chunk’s extraction. - Graph edits:
acreate_entity,aedit_entity,amerge_entities,acreate_relationandadelete_by_entitychange the graph by hand and keep the vectors in sync.
Running it
Library: pip install lightrag-hku, create LightRAG(working_dir=..., llm_model_func=..., embedding_func=...), await rag.initialize_storages(), then await rag.ainsert(text) and await rag.aquery(q, param=QueryParam(mode="mix")). Server: install the api extra, copy env.example to .env, and run lightrag-server (or lightrag-gunicorn, or the Docker Compose files). The server serves the web UI, a REST API and an Ollama-style /api/chat endpoint, so chat clients that speak Ollama can use it as a model.
Strengths and caveats
- Strength: cheap, true incremental indexing. A new document only adds and merges. Deletion removes nodes and edges that only that document produced, and rebuilds shared ones from the cached extraction results instead of re-reading the chunks.
- Strength: wide backend choice, per-role model routing, and caches for both extraction and query answers.
- Strength: five retrieval modes plus
bypassbehind one parameter, withonly_need_contextto get the context without an answer. - Caveat: no community detection or hierarchical summaries, so whole-corpus “what are the themes” questions depend on how well the high-level keywords match relationship descriptions.
- Caveat: entity resolution is exact name matching after normalization. Spelling variants become separate nodes unless you merge them with
amerge_entities. - Caveat: the default stores keep everything in process memory, and NetworkX allows a single writer. Multi-worker deployments need a database backend.
- Caveat: the answer-cache key does not include the retrieved context. With the default
enable_llm_cache=True, a repeated question can return an answer cached before new documents were added. Thelightrag-clean-llmqctool clears these entries by hand. - Caveat: the code base is large and moves fast. The pipeline has many journaling and recovery paths, which makes it harder to read and to change safely.
Sources: code at 453dce8, deepwiki-open wiki (14 pages), verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit 453dce8. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredChunking. Documents are split into overlapping token-sized windows by the default fixed-token chunker (chunking_by_token_size / chunking_by_fixed_token in lightrag/chunker/token_size.py), which segments on a character delimiter if provided then windows at a configurable chunk_token_size (default 1200) + chunk_overlap_token_size (default 100). Alternative strategies — chunking_by_recursive_character (R), chunking_by_semantic_vector (V), and chunking_by_paragraph_semantic (P) — are selectable per document via process_options in the file pipeline (lightrag/chunker/__init__.py).
Entity and relation extraction. Each chunk is sent to the LLM (role "extract", configured by role_llm_funcs at lightrag/llm_roles.py:52-57) through extract_entities in lightrag/operate.py:3942. Two prompt formats exist: the default text-delimiter format (lightrag/prompt.py:56-207) using tuple_delimiter (<|#|>) and completion_delimiter (<|COMPLETE|>) separators, or JSON structured output (lightrag/prompt.py:175-295) when entity_extraction_use_json is set. Both extract entity_name, entity_type (from a guided type taxonomy with 11 classes like Person, Organization, Concept, etc.), and entity_description, plus binary (source, target, keywords, description) relations. A gleaning step (entity_extract_max_gleaning > 0) re-queries the LLM with the original prompt + prior response to find missed or corrected entities, merging by preferring longer descriptions (lightrag/operate.py:4290-4363).
Entity resolution / deduplication. Entities extracted from different chunks for the same document are merged in merge_nodes_and_edges (lightrag/operate.py:3514). Edges are deduplicated by sorting their endpoint tuple and using a seen set (lightrag/operate.py:3652-3654). On upsert to the graph, the graph store's upsert_node and upsert_edge in NetworkX (lightrag/kg/networkx_impl.py:801-845) are idempotent — calling graph.add_node(id, **data) on an existing node overwrites its attributes. There is no cross-document entity resolution or explicit coreference resolution beyond the LLM's per-prompt consistency instructions ("Ensure consistent naming across the entire extraction process"). Entity names are normalized with normalize_entity_name from lightrag/utils.py.
Schema / ontology. Not user-extensible in the code itself; the 11 entity types (Person, Creature, Organization, Location, Event, Concept, Method, Content, Data, Artifact, NaturalObject) are defined in PROMPTS["default_entity_types_guidance"] at lightrag/prompt.py:20-34 and can be overridden via addon_params['entity_types_guidance']. Relations are unlabeled binary tuples with keyword tags and free-text descriptions.
_merge_nodes_then_upsert reads the existing node by its normalized name and merges into it, with the entity type picked by majority vote. What is missing is fuzzy or alias resolution. Gleaning runs at most once, whatever entity_extract_max_gleaning is set to.Where and how is the graph stored?
answeredGraph storage backends. The graph is stored via a pluggable BaseGraphStorage (lightrag/base.py:671). Implementations registered in lightrag/kg/__init__.py:12-23 include NetworkXStorage (default, file-based), Neo4JStorage (Neo4j graph DB), PGGraphStorage / PGTableGraphStorage (PostgreSQL), MongoGraphStorage, MemgraphStorage, and OpenSearchGraphStorage.
Default: NetworkX (in-memory + file). NetworkXStorage (lightrag/kg/networkx_impl.py:38-85) keeps the entire graph in process memory using networkx.Graph. It persists by rewriting one GraphML file (graph_<namespace>.graphml) per workspace via index_done_callback. It declares requires_single_writer = True, meaning only one process may mutate it at a time, enforced by the pipeline's busy reservation or LightRAG._admin_write_gate. Concurrent readers detect peer writes through a two-channel fence: file (st_mtime_ns, st_size) fingerprint plus a storage_updated flag. A commit publishes the whole namespace, so partial mutations from other in-flight writers can be published unintentionally — accepted as a documented residue.
Node/edge schema. Nodes and edges are NetworkX graph.add_node(id, **data) / graph.add_edge(src, tgt, **data) calls with arbitrary string-keyed dictionaries (lightrag/kg/networkx_impl.py:820-821, 844-845). Node data typically includes entity_name, entity_type, description, source_id, file_path, created_at. Edge data includes src_id, tgt_id, weight, source_id, keywords, description, file_path, created_at. XML attribute validation (validate_xml_attributes) is applied before every write since GraphML serialization can't handle nested structures.
Embeddings: separate vector stores. Graph nodes (entities) and edges (relations) are dually stored: once in the graph for structural traversal, and once in entity/relation vector databases (entities_vdb / relationships_vdb). The default NanoVectorDBStorage (lightrag/kg/nano_vector_db_impl.py:62-97) is an in-memory store serialized to a JSON file per workspace. Each entity vector record holds {content, entity_name, source_id, description, entity_type, file_path} (lightrag/operate.py:1805-1813). Each relation vector record holds {src_id, tgt_id, source_id, content, keywords, description, weight, file_path} (lightrag/operate.py:2324-2335). Other supported vector DBs: Milvus, PGVector, Faiss, Qdrant, Mongo, OpenSearch. NoopVectorDBStorage disables vector storage.
Are communities, summaries or hierarchies built over the graph?
answeredLightRAG does not perform community detection, community summarization, or hierarchical graph summarization. A grep for "leiden", "louvain", "community detection", "community summar", and "hierarchical summary" across the entire lightrag/ Python source tree returns zero matches in the core logic. The only hit for "community" in non-test Python code is inside lightrag/chunker/paragraph_semantic.py, which is about paragraph-level document chunking (a semantic text-splitting technique) and has nothing to do with graph community detection.
Instead of building communities, LightRAG takes a retrieval-time approach to knowledge synthesis. The longest path in the knowledge graph was present in earlier GraphRAG papers (like Microsoft's GraphRAG) which pioneered the Leiden-based community summary approach, but LightRAG opted for a simpler design where the query process itself bridges local and global context:
- Local mode (
lightrag/operate.py:5366-5376): starts from the most vector-similar entities and retrieves their neighbor edges via_get_node_data(entity vector DB → graph node degree → related edges). - Global mode (
lightrag/operate.py:5377-5386): starts from the most vector-similar relationships and collects their endpoint entities via_get_edge_data(relation vector DB → related entities). - Hybrid mode (
lightrag/operate.py:5388-5408): runs both and round-robin merges results. - Mix mode (
lightrag/operate.py:5411-5430): adds direct vector search over document chunks on top.
The context assembly stages (_build_query_context → _perform_kg_search → _apply_token_truncation → _merge_all_chunks → _build_context_str at lightrag/operate.py:6075-6197) feed all retrieved entities, relations, and document chunks into the LLM prompt for final answer synthesis. There is no pre-computed community-level summary at any point.
How does query-time retrieval use the graph?
answeredFive query modes. Defined in QueryParam.mode at lightrag/base.py:90-100 ("local", "global", "hybrid", "naive", "mix", "bypass"). aquery_llm (lightrag/lightrag.py:5114-5212) dispatches to kg_query (modes local/global/hybrid/mix) or naive_query (mode naive) or a direct LLM call (mode bypass).
Keyword extraction. Every KG-based query first extracts high-level and low-level keywords from the query text via get_keywords_from_query (lightrag/operate.py:4975), which calls the LLM with the keywords_extraction prompt template (lightrag/prompt.py:484-498). Low-level keywords are specific named entities; high-level keywords capture conceptual themes.
Local mode (lightrag/operate.py:5366). Low-level keywords are embedded and used to query entities_vdb (vector DB of entity descriptions). The top-K entities are retrieved from the graph via _get_node_data (lightrag/operate.py:6200), which fetches node properties and degrees, then _find_most_related_edges_from_entities traverses the graph to find connected edges, sorting by rank × weight.
Global mode (lightrag/operate.py:5377). High-level keywords are embedded and used to query relationships_vdb (vector DB of relation descriptions). The top-K relations are retrieved via _get_edge_data (lightrag/operate.py:6529), which fetches edge properties from the graph, then _find_most_related_entities_from_relationships collects the endpoint entities.
Hybrid mode (lightrag/operate.py:5388). Runs both local and global retrieval in parallel, then performs a round-robin merge of entities and relations, deduplicating by name/edge-key (lightrag/operate.py:5432-5486).
Mix mode (lightrag/operate.py:5411). Same as hybrid but additionally queries chunks_vdb (document chunk vector store) directly with the query embedding. The result includes vector-retrieved text chunks alongside graph-retrieved entities and relations.
Naive mode. Direct vector search over document chunks, no graph involvement.
Context assembly (_build_query_context at lightrag/operate.py:6075-6197): A 4-stage pipeline — (1) _perform_kg_search retrieves entities, relations, and vector chunks; (2) _apply_token_truncation reduces results to fit LLM token budgets (max_entity_tokens, max_relation_tokens, max_total_tokens); (3) _merge_all_chunks selects the most relevant text chunks connected to the surviving entities/relations; (4) _build_context_str formats everything into the final prompt with the rag_response or naive_rag_response template (lightrag/prompt.py:334-440), which includes a Knowledge Graph Data section and a Document Chunks section with citation references. The prompt is then sent to the query-role LLM.
mix and the default top_k is 40.How are updates and incremental indexing handled?
answeredDocument addition (incremental). New documents are inserted via ainsert (lightrag/lightrag.py:2310-2377) or the pipeline apipeline_enqueue_documents. Each document is chunked independently, extracted against the LLM, and the resulting entities/relations are upserted into the existing graph. Entity and relation upsert is inherently additive — upsert_node / upsert_edge in NetworkX merge properties into the existing node/edge. There is no global re-clustering or full rebuild. The pipeline process (manage_pipeline_status in lightrag/pipeline.py) coordinates concurrent document processing via a busy reservation but allows concurrent enqueue + processing.
Document deletion. adelete_by_doc_id (lightrag/lightrag.py:6727-6800) removes a document and its derived graph contributions. It acquires the pipeline busy slot, then calls _purge_kg_contributions (lightrag/lightrag.py:6015-6090) which: (1) reads the per-document write-ahead anchors (full_entities / full_relations — stored in dedicated KV stores and written before any graph mutation in merge_nodes_and_edges, lightrag/operate.py:3666-3703); (2) for each affected entity/relation, checks whether other documents still reference it via chunk-tracking (entity_chunks / relation_chunks) — if no remaining sources, it deletes the graph node/edge outright; if other documents share it, it rebuilds the entity/relation description from the surviving chunks' cache (LLM re-extraction from cache); (3) deletes the chunks themselves only after all graph contributions are removed; (4) removes the recovery anchors last. The process is journaled in doc_status.metadata so a crash mid-purge is resumable (_resolve_purge_recovery_proof, lightrag/lightrag.py).
Custom chunk patching. ainsert_custom_chunks (lightrag/lightrag.py:2394-2509) supports patching documents without full deletion — it writes new chunks, re-extracts, and merges them into the graph without touching the original chunks. The operation is journaled for crash recovery.
Extraction caching. LLM extraction results are cached in llm_response_cache (default JsonKVStorage) and keyed by prompt hash + model identity. On re-insertion or document deletion, cached extractions are replayed to rebuild entities/relations without re-calling the LLM. The extraction cache is partitioned by llm_cache_identity (model/provider) and each cache row is attached to its owning chunk before being written (use_llm_func_with_cache at lightrag/utils.py:5515-5577). The cache_keys_collector mechanism allows batch cache operations.
How are LLM cost and latency controlled during indexing and query?
answeredLLM response caching (primary cost control). Both entity extraction and query results are cached. enable_llm_cache (default True) caches query answers; enable_llm_cache_for_entity_extract (default True) caches extraction results (lightrag/lightrag.py:1093-1097). The cache is implemented in handle_cache / save_to_cache (lightrag/utils.py:4448-4520), storing flattened cache keys of format {mode}:{cache_type}:{hash} in llm_response_cache (a KV storage). On cache hit, the LLM call is entirely skipped — response is returned from the stored value. The extraction cache also has per-chunk write-ahead semantics to avoid orphaned cache rows.
Concurrency batching. The number of concurrent LLM extraction calls per document is controlled by llm_model_max_async (default 4) via asyncio.Semaphore(chunk_max_async) at lightrag/operate.py:4547-4548. Document-level parallelism is capped by max_parallel_insert (default 3) at lightrag/pipeline.py:2653. Embedding requests are batched by embedding_batch_num at lightrag/kg/nano_vector_db_impl.py:133 and various backend implementations.
Role-based model routing. LightRAG supports separate LLM bindings per role (extract, keyword, query, vlm) via RoleLLMConfig (lightrag/llm_roles.py:57-63). Each role can use a different model (e.g., a cheaper/faster model for extraction, a more capable one for query), configured via env vars like EXTRACT_LLM_BINDING, QUERY_LLM_BINDING. This lets operators trade cost vs. quality per stage.
Token budgets. Multiple token budgets limit LLM context size at key stages:
MAX_EXTRACT_INPUT_TOKENSlimits the gleaning step input (default 128K,lightrag/operate.py:3992-3996).- At query time,
max_entity_tokens,max_relation_tokens,max_total_tokenslimit what goes into the final prompt (_apply_token_truncationatlightrag/operate.py:5501). - Entity/relation descriptions are chunked and recursively summarized via
_handle_entity_relation_summarymap-reduce (lightrag/operate.py:372) when they exceed token limits, avoiding giant prompts. - Document chunks that exceed the embedding model's token limit are truncated by
_truncate_vdb_content(lightrag/operate.py:298).
Non-LLM shortcuts. The entire relation weight contract (docs/ProgramingWithCore.md) avoids LLM calls by simply counting distinct source IDs for the weight floor. The kg_chunk_pick_method (DEFAULT_KG_CHUNK_PICK_METHOD) supports weighted polling (based on occurrence count in entity/relation source_id lists) as a non-LLM alternative to vector-similarity chunk selection. The NaiveQuery mode bypasses the KG entirely and does a vector-only search. bypass mode skips all retrieval.
MAX_EXTRACT_INPUT_TOKENS is 20,480 (constants.py), not 128K. Note also that the query-answer cache key does not include the retrieved context, so a cached answer can outlive changes to the data.