# HKUDS/LightRAG

> Python graph RAG engine that merges LLM-extracted entities into one graph and retrieves by keyword-matched entities, relations and chunks.

- Category: [Graph RAG](https://llms-technical-reviews.com/graph-rag/)
- Repository: https://github.com/HKUDS/LightRAG (reviewed at commit `453dce83d6d0354a06e46c8d4029a0895c4e054b`, 2026-09-26)
- Stars: 39998 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/lightrag/

## Overview

LightRAG is a graph RAG engine from HKU's data science lab. It ships as a Python library (`lightrag-hku`), a FastAPI server (`lightrag-server`) and a React web UI. Like GraphRAG, it has an LLM read every chunk and extract entities and relationships. The difference is what happens next. LightRAG merges the results straight into one growing graph keyed by entity name, and it embeds every entity and every relationship into their own vector stores. It never clusters the graph or writes community summaries. All "global" reasoning happens at query time.

A query starts with one LLM call that splits the question into low-level keywords (specific names) and high-level keywords (themes). Low-level keywords search the entity vectors. High-level keywords search the relationship vectors. Both expand one hop through the graph. The default `mix` mode also adds plain vector hits over chunks. Because documents merge into the graph one at a time, adding a document needs no rebuild, and deleting one is a supported operation.

The project has grown well beyond the paper. At this SHA it includes a multi-stage ingestion pipeline with crash-recovery journaling, document parsers (native, MinerU, Docling), a vision-model analysis step for images and tables, four chunking strategies, and about 25 storage implementations. The top-level modules of the package alone are about 46,000 lines of Python.

## Architecture

```mermaid
flowchart LR
  U["Client / Web UI"] --> API["FastAPI server"]
  API --> LR["LightRAG class"]
  LR --> P["Pipeline: parse, analyze, process"]
  P --> CH["Chunker"]
  CH --> EX["extract_entities (LLM)"]
  EX --> M["merge_nodes_and_edges"]
  M --> G["Graph storage"]
  M --> EV["Entity vectors"]
  M --> RV["Relation vectors"]
  CH --> CV["Chunk vectors + KV"]
  LR --> KQ["kg_query"]
  KQ --> KW["Keyword extraction (LLM)"]
  KW --> EV
  KW --> RV
  KQ --> G
  KQ --> CV
  KQ --> ANS["Answer (LLM)"]
```

| Component | Path | Role |
|---|---|---|
| Core class | `lightrag/lightrag.py` | `LightRAG` dataclass: config, storage wiring, insert, query, delete and graph edit APIs |
| Pipeline | `lightrag/pipeline.py` | Enqueue, parse, analyze and process workers; document status; crash recovery |
| Operators | `lightrag/operate.py` | Extraction, merging, summarization, `kg_query`, `naive_query`, context building |
| Prompts | `lightrag/prompt.py` | Extraction, keyword and answer templates; default entity types |
| Chunkers | `lightrag/chunker/` | Fixed-token (default), recursive-character, semantic-vector, paragraph-semantic |
| Storage registry | `lightrag/kg/__init__.py` | KV, vector, graph and doc-status implementations |
| LLM roles | `lightrag/llm_roles.py` | Separate `extract`, `keyword`, `query` and `vlm` model bindings |
| Provider bindings | `lightrag/llm/` | OpenAI, Azure, Ollama, Anthropic, Gemini, Bedrock, Hugging Face and others |
| API server | `lightrag/api/` | Document, query and graph routes, plus an Ollama-compatible chat API |
| Web UI | `lightrag_webui/` | Document manager, graph viewer and query panel |

## How a request flows

**Insert**

1. `ainsert` turns its arguments into per-document chunk options, then calls `apipeline_enqueue_documents` and `apipeline_process_enqueue_documents` ([lightrag.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L2310-L2377)). The API server calls those two methods directly, so it can choose a chunking strategy per document.
2. Enqueue gives each document an ID from an MD5 hash of its content (or file path) and skips duplicates. The process worker chunks the text (default 1,200 tokens with 100 overlap).
3. Chunks are written to the chunk vector store and the chunk KV store first. Then `_process_extract_entities` runs extraction, unless the document opted out of the graph ([pipeline.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/pipeline.py#L5720-L5750)).
4. `extract_entities` sends each chunk to the `extract` role model, with up to `llm_model_max_async` calls at once. Every call goes through the LLM cache ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3942-L4000)).
5. `merge_nodes_and_edges` merges the per-chunk results into the graph. It writes the document's entity and relation lists first as recovery anchors, then upserts nodes, edges and their vectors ([pipeline.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/pipeline.py#L5804-L5815)).

**Query**

6. `aquery_llm` sends `local`, `global`, `hybrid` and `mix` to `kg_query`, `naive` to `naive_query`, and `bypass` straight to the query model ([lightrag.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L5114-L5215)).
7. `kg_query` extracts keywords with the `keyword` model, then `_perform_kg_search` embeds the query and both keyword strings in one batch. It runs the entity-side search and the relation-side search, merges them round-robin, and in `mix` mode adds chunk vector hits ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5280-L5498)).
8. `_build_query_context` cuts entities and relations to their token budgets, picks the chunks linked to what survived, and formats one context with knowledge-graph and document-chunk sections. The `query` model answers. The answer is cached under a hash of the query, mode, retrieval settings, keywords and model identity ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4842-L4895)).

## Key components

### Extraction and merging

The default prompt asks for entities with a type from an 11-class list (Person, Organization, Location, Event, Concept, Method and others, plus `Other`) and for relationships with keywords and a description. Records are separated by `<|#|>` ([prompt.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L14-L34)). A JSON output mode is also available. Gleaning runs at most once. It is skipped when the replayed conversation would exceed `MAX_EXTRACT_INPUT_TOKENS` (default 20,480), and a gleaned description replaces the first one only if it is longer ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4252-L4364)).

Entities merge across chunks and documents by normalized name. The entity type is decided by majority vote. Descriptions are joined without an LLM call until there are 8 fragments or they pass the token limit. After that, `_handle_entity_relation_summary` summarizes them, recursively if needed ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L372-L440)).

### Retrieval

`_get_node_data` takes the top-k entities from the entity vector store (`top_k` default 40), reads their degrees, and collects all their edges, ranked by edge degree and weight ([operate.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6200-L6313)). The relation-side search does the reverse: top-k relations, then their endpoint entities. Traversal is one hop. "Global" in LightRAG means relation-centric search, not corpus-wide summaries. Token budgets default to 6,000 for entities, 8,000 for relations and 30,000 in total. Reranking of chunks is on by default when a rerank model is configured ([base.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/base.py#L89-L130)).

### Storage

Four storage roles each take a class name: KV, vector, graph and document status ([kg/__init__.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/__init__.py#L1-L48)). The defaults are JSON files, NanoVectorDB, NetworkX and a JSON doc-status store, all held in memory and flushed to the working directory ([lightrag.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L696-L705)). `NetworkXStorage` rewrites one GraphML file and allows only one writer per workspace ([networkx_impl.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/networkx_impl.py#L38-L60)). PostgreSQL, MongoDB and OpenSearch can each serve all four roles. Neo4j and Memgraph can serve as the graph store, and Milvus, Qdrant and Faiss as vector stores.

## Extending it

- **Storage:** implement the `BaseKVStorage`, `BaseVectorStorage`, `BaseGraphStorage` or `DocStatusStorage` interface, then register the class in `STORAGES` (`kg/__init__.py`), which `kg/factory.py` loads by name.
- **Models:** pass any async `llm_model_func` and `embedding_func`, or set per-role overrides with `RoleLLMConfig` or `EXTRACT_LLM_BINDING`, `QUERY_LLM_BINDING` and similar env vars ([llm_roles.py](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/llm_roles.py#L52-L57)).
- **Prompts and schema:** override `addon_params['entity_types_guidance']` or replace entries in `PROMPTS`. A `kg_extraction_validator` hook can check each chunk's extraction.
- **Graph edits:** `acreate_entity`, `aedit_entity`, `amerge_entities`, `acreate_relation` and `adelete_by_entity` change the graph by hand and keep the vectors in sync.

## Running it

Library: `pip install lightrag-hku`, create `LightRAG(working_dir=..., llm_model_func=..., embedding_func=...)`, `await rag.initialize_storages()`, then `await rag.ainsert(text)` and `await rag.aquery(q, param=QueryParam(mode="mix"))`. Server: install the `api` extra, copy `env.example` to `.env`, and run `lightrag-server` (or `lightrag-gunicorn`, or the Docker Compose files). The server serves the web UI, a REST API and an Ollama-style `/api/chat` endpoint, so chat clients that speak Ollama can use it as a model.

## Strengths and caveats

- **Strength:** cheap, true incremental indexing. A new document only adds and merges. Deletion removes nodes and edges that only that document produced, and rebuilds shared ones from the cached extraction results instead of re-reading the chunks.
- **Strength:** wide backend choice, per-role model routing, and caches for both extraction and query answers.
- **Strength:** five retrieval modes plus `bypass` behind one parameter, with `only_need_context` to get the context without an answer.
- **Caveat:** no community detection or hierarchical summaries, so whole-corpus "what are the themes" questions depend on how well the high-level keywords match relationship descriptions.
- **Caveat:** entity resolution is exact name matching after normalization. Spelling variants become separate nodes unless you merge them with `amerge_entities`.
- **Caveat:** the default stores keep everything in process memory, and NetworkX allows a single writer. Multi-worker deployments need a database backend.
- **Caveat:** the answer-cache key does not include the retrieved context. With the default `enable_llm_cache=True`, a repeated question can return an answer cached before new documents were added. The `lightrag-clean-llmqc` tool clears these entries by hand.
- **Caveat:** the code base is large and moves fast. The pipeline has many journaling and recovery paths, which makes it harder to read and to change safely.

*Sources: code at 453dce8, deepwiki-open wiki (14 pages), verified Q&A.*

## How HKUDS/LightRAG answers the Graph RAG questions

### How is the knowledge graph extracted from documents? (answered)

**Chunking.** Documents are split into overlapping token-sized windows by the default fixed-token chunker (`chunking_by_token_size` / `chunking_by_fixed_token` in `lightrag/chunker/token_size.py`), which segments on a character delimiter if provided then windows at a configurable `chunk_token_size` (default 1200) + `chunk_overlap_token_size` (default 100). Alternative strategies — `chunking_by_recursive_character` (R), `chunking_by_semantic_vector` (V), and `chunking_by_paragraph_semantic` (P) — are selectable per document via `process_options` in the file pipeline (`lightrag/chunker/__init__.py`).

**Entity and relation extraction.** Each chunk is sent to the LLM (role `"extract"`, configured by `role_llm_funcs` at `lightrag/llm_roles.py:52-57`) through `extract_entities` in `lightrag/operate.py:3942`. Two prompt formats exist: the default text-delimiter format (`lightrag/prompt.py:56-207`) using `tuple_delimiter` (`<|#|>`) and `completion_delimiter` (`<|COMPLETE|>`) separators, or JSON structured output (`lightrag/prompt.py:175-295`) when `entity_extraction_use_json` is set. Both extract `entity_name`, `entity_type` (from a guided type taxonomy with 11 classes like Person, Organization, Concept, etc.), and `entity_description`, plus binary `(source, target, keywords, description)` relations. A **gleaning** step (`entity_extract_max_gleaning > 0`) re-queries the LLM with the original prompt + prior response to find missed or corrected entities, merging by preferring longer descriptions (`lightrag/operate.py:4290-4363`).

**Entity resolution / deduplication.** Entities extracted from different chunks for the same document are merged in `merge_nodes_and_edges` (`lightrag/operate.py:3514`). Edges are deduplicated by sorting their endpoint tuple and using a `seen` set (`lightrag/operate.py:3652-3654`). On upsert to the graph, the graph store's `upsert_node` and `upsert_edge` in NetworkX (`lightrag/kg/networkx_impl.py:801-845`) are idempotent — calling `graph.add_node(id, **data)` on an existing node overwrites its attributes. There is no cross-document entity resolution or explicit coreference resolution beyond the LLM's per-prompt consistency instructions ("Ensure consistent naming across the entire extraction process"). Entity names are normalized with `normalize_entity_name` from `lightrag/utils.py`.

**Schema / ontology.** Not user-extensible in the code itself; the 11 entity types (Person, Creature, Organization, Location, Event, Concept, Method, Content, Data, Artifact, NaturalObject) are defined in `PROMPTS["default_entity_types_guidance"]` at `lightrag/prompt.py:20-34` and can be overridden via `addon_params['entity_types_guidance']`. Relations are unlabeled binary tuples with keyword tags and free-text descriptions.

> **Editor's note.** Correction: entities do merge across documents. `_merge_nodes_then_upsert` reads the existing node by its normalized name and merges into it, with the entity type picked by majority vote. What is missing is fuzzy or alias resolution. Gleaning runs at most once, whatever `entity_extract_max_gleaning` is set to.

Citations: [lightrag/operate.py:3942-3965](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3942-L3965) · [lightrag/prompt.py:56-117](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L56-L117) · [lightrag/prompt.py:20-34](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L20-L34) · [lightrag/chunker/token_size.py:133-180](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/chunker/token_size.py#L133-L180) · [lightrag/operate.py:3514-3545](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3514-L3545) · [lightrag/operate.py:4290-4365](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4290-L4365)

### Where and how is the graph stored? (answered)

**Graph storage backends.** The graph is stored via a pluggable `BaseGraphStorage` (`lightrag/base.py:671`). Implementations registered in `lightrag/kg/__init__.py:12-23` include NetworkXStorage (default, file-based), Neo4JStorage (Neo4j graph DB), PGGraphStorage / PGTableGraphStorage (PostgreSQL), MongoGraphStorage, MemgraphStorage, and OpenSearchGraphStorage.

**Default: NetworkX (in-memory + file).** `NetworkXStorage` (`lightrag/kg/networkx_impl.py:38-85`) keeps the entire graph in process memory using `networkx.Graph`. It persists by rewriting one GraphML file (`graph_<namespace>.graphml`) per workspace via `index_done_callback`. It declares `requires_single_writer = True`, meaning only one process may mutate it at a time, enforced by the pipeline's `busy` reservation or `LightRAG._admin_write_gate`. Concurrent readers detect peer writes through a two-channel fence: file `(st_mtime_ns, st_size)` fingerprint plus a `storage_updated` flag. A commit publishes the whole namespace, so partial mutations from other in-flight writers can be published unintentionally — accepted as a documented residue.

**Node/edge schema.** Nodes and edges are NetworkX `graph.add_node(id, **data)` / `graph.add_edge(src, tgt, **data)` calls with arbitrary string-keyed dictionaries (`lightrag/kg/networkx_impl.py:820-821, 844-845`). Node data typically includes `entity_name`, `entity_type`, `description`, `source_id`, `file_path`, `created_at`. Edge data includes `src_id`, `tgt_id`, `weight`, `source_id`, `keywords`, `description`, `file_path`, `created_at`. XML attribute validation (`validate_xml_attributes`) is applied before every write since GraphML serialization can't handle nested structures.

**Embeddings: separate vector stores.** Graph nodes (entities) and edges (relations) are **dually stored**: once in the graph for structural traversal, and once in entity/relation **vector databases** (`entities_vdb` / `relationships_vdb`). The default `NanoVectorDBStorage` (`lightrag/kg/nano_vector_db_impl.py:62-97`) is an in-memory store serialized to a JSON file per workspace. Each entity vector record holds `{content, entity_name, source_id, description, entity_type, file_path}` (`lightrag/operate.py:1805-1813`). Each relation vector record holds `{src_id, tgt_id, source_id, content, keywords, description, weight, file_path}` (`lightrag/operate.py:2324-2335`). Other supported vector DBs: Milvus, PGVector, Faiss, Qdrant, Mongo, OpenSearch. `NoopVectorDBStorage` disables vector storage.


Citations: [lightrag/kg/networkx_impl.py:38-85](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/networkx_impl.py#L38-L85) · [lightrag/kg/__init__.py:1-48](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/__init__.py#L1-L48) · [lightrag/base.py:671-714](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/base.py#L671-L714) · [lightrag/kg/nano_vector_db_impl.py:62-97](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/nano_vector_db_impl.py#L62-L97) · [lightrag/operate.py:1800-1835](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L1800-L1835) · [lightrag/operate.py:2310-2340](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L2310-L2340)

### Are communities, summaries or hierarchies built over the graph? (answered)

LightRAG does **not** perform community detection, community summarization, or hierarchical graph summarization. A grep for "leiden", "louvain", "community detection", "community summar", and "hierarchical summary" across the entire `lightrag/` Python source tree returns zero matches in the core logic. The only hit for "community" in non-test Python code is inside `lightrag/chunker/paragraph_semantic.py`, which is about paragraph-level document chunking (a semantic text-splitting technique) and has nothing to do with graph community detection.

Instead of building communities, LightRAG takes a **retrieval-time approach** to knowledge synthesis. The longest path in the knowledge graph was present in earlier GraphRAG papers (like Microsoft's GraphRAG) which pioneered the Leiden-based community summary approach, but LightRAG opted for a simpler design where the query process itself bridges local and global context:

- **Local mode** (`lightrag/operate.py:5366-5376`): starts from the most vector-similar entities and retrieves their neighbor edges via `_get_node_data` (entity vector DB → graph node degree → related edges).
- **Global mode** (`lightrag/operate.py:5377-5386`): starts from the most vector-similar relationships and collects their endpoint entities via `_get_edge_data` (relation vector DB → related entities).
- **Hybrid mode** (`lightrag/operate.py:5388-5408`): runs both and round-robin merges results.
- **Mix mode** (`lightrag/operate.py:5411-5430`): adds direct vector search over document chunks on top.

The context assembly stages (`_build_query_context` → `_perform_kg_search` → `_apply_token_truncation` → `_merge_all_chunks` → `_build_context_str` at `lightrag/operate.py:6075-6197`) feed all retrieved entities, relations, and document chunks into the LLM prompt for final answer synthesis. There is no pre-computed community-level summary at any point.

> **Editor's note.** Correction: the sentence about a 'longest path' in earlier GraphRAG papers is not supported by the code or the cited lines. The rest stands: there is no community detection or hierarchical summary.

Citations: [lightrag/operate.py:5280-5498](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5280-L5498) · [lightrag/operate.py:6075-6100](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6075-L6100) · [lightrag/operate.py:6529-6585](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6529-L6585) · [lightrag/operate.py:5366-5430](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5366-L5430) · [lightrag/operate.py:6200-6257](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6200-L6257)

### How does query-time retrieval use the graph? (answered)

**Five query modes.** Defined in `QueryParam.mode` at `lightrag/base.py:90-100` (`"local"`, `"global"`, `"hybrid"`, `"naive"`, `"mix"`, `"bypass"`). `aquery_llm` (`lightrag/lightrag.py:5114-5212`) dispatches to `kg_query` (modes local/global/hybrid/mix) or `naive_query` (mode naive) or a direct LLM call (mode bypass).

**Keyword extraction.** Every KG-based query first extracts **high-level** and **low-level keywords** from the query text via `get_keywords_from_query` (`lightrag/operate.py:4975`), which calls the LLM with the `keywords_extraction` prompt template (`lightrag/prompt.py:484-498`). Low-level keywords are specific named entities; high-level keywords capture conceptual themes.

**Local mode** (`lightrag/operate.py:5366`). Low-level keywords are embedded and used to query `entities_vdb` (vector DB of entity descriptions). The top-K entities are retrieved from the graph via `_get_node_data` (`lightrag/operate.py:6200`), which fetches node properties and degrees, then `_find_most_related_edges_from_entities` traverses the graph to find connected edges, sorting by rank × weight.

**Global mode** (`lightrag/operate.py:5377`). High-level keywords are embedded and used to query `relationships_vdb` (vector DB of relation descriptions). The top-K relations are retrieved via `_get_edge_data` (`lightrag/operate.py:6529`), which fetches edge properties from the graph, then `_find_most_related_entities_from_relationships` collects the endpoint entities.

**Hybrid mode** (`lightrag/operate.py:5388`). Runs both local and global retrieval in parallel, then performs a round-robin merge of entities and relations, deduplicating by name/edge-key (`lightrag/operate.py:5432-5486`).

**Mix mode** (`lightrag/operate.py:5411`). Same as hybrid but additionally queries `chunks_vdb` (document chunk vector store) directly with the query embedding. The result includes vector-retrieved text chunks alongside graph-retrieved entities and relations.

**Naive mode.** Direct vector search over document chunks, no graph involvement.

**Context assembly** (`_build_query_context` at `lightrag/operate.py:6075-6197`): A 4-stage pipeline — (1) `_perform_kg_search` retrieves entities, relations, and vector chunks; (2) `_apply_token_truncation` reduces results to fit LLM token budgets (`max_entity_tokens`, `max_relation_tokens`, `max_total_tokens`); (3) `_merge_all_chunks` selects the most relevant text chunks connected to the surviving entities/relations; (4) `_build_context_str` formats everything into the final prompt with the `rag_response` or `naive_rag_response` template (`lightrag/prompt.py:334-440`), which includes a `Knowledge Graph Data` section and a `Document Chunks` section with citation references. The prompt is then sent to the query-role LLM.

> **Editor's note.** Correction: in hybrid and mix modes the entity-side and relation-side searches are awaited one after the other, not in parallel. Only the embeddings are computed in one batch. The default mode is `mix` and the default `top_k` is 40.

Citations: [lightrag/lightrag.py:5114-5170](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L5114-L5170) · [lightrag/operate.py:5280-5498](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5280-L5498) · [lightrag/operate.py:6075-6197](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6075-L6197) · [lightrag/operate.py:6200-6257](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6200-L6257) · [lightrag/operate.py:6529-6585](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L6529-L6585) · [lightrag/prompt.py:334-386](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L334-L386)

### How are updates and incremental indexing handled? (answered)

**Document addition (incremental).** New documents are inserted via `ainsert` (`lightrag/lightrag.py:2310-2377`) or the pipeline `apipeline_enqueue_documents`. Each document is chunked independently, extracted against the LLM, and the resulting entities/relations are upserted into the existing graph. Entity and relation upsert is inherently additive — `upsert_node` / `upsert_edge` in NetworkX merge properties into the existing node/edge. There is no global re-clustering or full rebuild. The pipeline process (`manage_pipeline_status` in `lightrag/pipeline.py`) coordinates concurrent document processing via a `busy` reservation but allows concurrent enqueue + processing.

**Document deletion.** `adelete_by_doc_id` (`lightrag/lightrag.py:6727-6800`) removes a document and its derived graph contributions. It acquires the pipeline `busy` slot, then calls `_purge_kg_contributions` (`lightrag/lightrag.py:6015-6090`) which: (1) reads the per-document write-ahead anchors (`full_entities` / `full_relations` — stored in dedicated KV stores and written before any graph mutation in `merge_nodes_and_edges`, `lightrag/operate.py:3666-3703`); (2) for each affected entity/relation, checks whether other documents still reference it via chunk-tracking (`entity_chunks` / `relation_chunks`) — if no remaining sources, it deletes the graph node/edge outright; if other documents share it, it rebuilds the entity/relation description from the surviving chunks' cache (LLM re-extraction from cache); (3) deletes the chunks themselves only after all graph contributions are removed; (4) removes the recovery anchors last. The process is journaled in `doc_status.metadata` so a crash mid-purge is resumable (`_resolve_purge_recovery_proof`, `lightrag/lightrag.py`).

**Custom chunk patching.** `ainsert_custom_chunks` (`lightrag/lightrag.py:2394-2509`) supports patching documents without full deletion — it writes new chunks, re-extracts, and merges them into the graph without touching the original chunks. The operation is journaled for crash recovery.

**Extraction caching.** LLM extraction results are cached in `llm_response_cache` (default `JsonKVStorage`) and keyed by prompt hash + model identity. On re-insertion or document deletion, cached extractions are replayed to rebuild entities/relations without re-calling the LLM. The extraction cache is partitioned by `llm_cache_identity` (model/provider) and each cache row is attached to its owning chunk before being written (`use_llm_func_with_cache` at `lightrag/utils.py:5515-5577`). The `cache_keys_collector` mechanism allows batch cache operations.


Citations: [lightrag/lightrag.py:2310-2377](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L2310-L2377) · [lightrag/lightrag.py:6015-6090](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L6015-L6090) · [lightrag/lightrag.py:6727-6826](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L6727-L6826) · [lightrag/lightrag.py:2394-2466](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L2394-L2466) · [lightrag/operate.py:3514-3591](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3514-L3591) · [lightrag/utils.py:5515-5577](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/utils.py#L5515-L5577)

### How are LLM cost and latency controlled during indexing and query? (answered)

**LLM response caching (primary cost control).** Both entity extraction and query results are cached. `enable_llm_cache` (default True) caches query answers; `enable_llm_cache_for_entity_extract` (default True) caches extraction results (`lightrag/lightrag.py:1093-1097`). The cache is implemented in `handle_cache` / `save_to_cache` (`lightrag/utils.py:4448-4520`), storing flattened cache keys of format `{mode}:{cache_type}:{hash}` in `llm_response_cache` (a KV storage). On cache hit, the LLM call is entirely skipped — response is returned from the stored value. The extraction cache also has per-chunk write-ahead semantics to avoid orphaned cache rows.

**Concurrency batching.** The number of concurrent LLM extraction calls per document is controlled by `llm_model_max_async` (default 4) via `asyncio.Semaphore(chunk_max_async)` at `lightrag/operate.py:4547-4548`. Document-level parallelism is capped by `max_parallel_insert` (default 3) at `lightrag/pipeline.py:2653`. Embedding requests are batched by `embedding_batch_num` at `lightrag/kg/nano_vector_db_impl.py:133` and various backend implementations.

**Role-based model routing.** LightRAG supports separate LLM bindings per role (`extract`, `keyword`, `query`, `vlm`) via `RoleLLMConfig` (`lightrag/llm_roles.py:57-63`). Each role can use a different model (e.g., a cheaper/faster model for extraction, a more capable one for query), configured via env vars like `EXTRACT_LLM_BINDING`, `QUERY_LLM_BINDING`. This lets operators trade cost vs. quality per stage.

**Token budgets.** Multiple token budgets limit LLM context size at key stages:
- `MAX_EXTRACT_INPUT_TOKENS` limits the gleaning step input (default 128K, `lightrag/operate.py:3992-3996`).
- At query time, `max_entity_tokens`, `max_relation_tokens`, `max_total_tokens` limit what goes into the final prompt (`_apply_token_truncation` at `lightrag/operate.py:5501`).
- Entity/relation descriptions are chunked and recursively summarized via `_handle_entity_relation_summary` map-reduce (`lightrag/operate.py:372`) when they exceed token limits, avoiding giant prompts.
- Document chunks that exceed the embedding model's token limit are truncated by `_truncate_vdb_content` (`lightrag/operate.py:298`).

**Non-LLM shortcuts.** The entire relation weight contract (`docs/ProgramingWithCore.md`) avoids LLM calls by simply counting distinct source IDs for the weight floor. The `kg_chunk_pick_method` (`DEFAULT_KG_CHUNK_PICK_METHOD`) supports weighted polling (based on occurrence count in entity/relation source_id lists) as a non-LLM alternative to vector-similarity chunk selection. The `NaiveQuery` mode bypasses the KG entirely and does a vector-only search. `bypass` mode skips all retrieval.

> **Editor's note.** Correction: the default `MAX_EXTRACT_INPUT_TOKENS` is 20,480 (`constants.py`), not 128K. Note also that the query-answer cache key does not include the retrieved context, so a cached answer can outlive changes to the data.

Citations: [lightrag/lightrag.py:1093-1097](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L1093-L1097) · [lightrag/utils.py:4448-4477](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/utils.py#L4448-L4477) · [lightrag/operate.py:4546-4558](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4546-L4558) · [lightrag/llm_roles.py:52-57](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/llm_roles.py#L52-L57) · [lightrag/operate.py:5501-5525](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L5501-L5525) · [lightrag/operate.py:372-420](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L372-L420)
