# How are updates and incremental indexing handled?

> Graph RAG — a good answer covers: Adding or changing documents without a full rebuild; deletion; caching of extraction results.

Canonical page: https://llms-technical-reviews.com/graph-rag/q/incremental/

## Verdict

[LightRAG](/p/lightrag/) and [HippoRAG](/p/hipporag/) handle change best. Both add documents without a rebuild and delete them with reference-counted cleanup of shared entities. [GraphRAG](/p/graphrag/) and [nano-graphrag](/p/nano-graphrag/) are in practice append-only.

**Add and delete per document.** LightRAG merges each document into the live graph. `adelete_by_doc_id` rebuilds shared entities from cached extraction results, and deletion is journaled so a crash can resume. HippoRAG sends only unseen chunk hashes to OpenIE and runs synonymy KNN only for new entities. Its manifest refuses to mix state from two different models. [Vector Graph RAG](/p/vector-graph-rag/) has `upsert_documents_by_source` and a cascading `delete_documents_by_source`, but only in Python: its REST import endpoints always rebuild the whole graph. [LLM Graph Builder](/p/llm-graph-builder/) processes each document on its own, can resume from the last processed chunk, and deletes only the entities no other document references. It has no extraction cache, so a retry pays again. [AutoFlow](/p/autoflow/) deletes a document's relationships and then any orphaned entities, but merged descriptions stay behind. Re-indexing skips chunks that already have relationships.

**Append-only.** GraphRAG's `update` finds new documents by title. It computes deleted documents but never uses that list, and it appends delta communities without re-clustering. nano-graphrag skips known document and chunk hashes, but every insert drops and regenerates all community reports. It has no delete API. [Semantica](/p/semantica/) merges new entities with `incremental_merge` and caches extraction with a TTL. It has no document-level delete, only entity-level purge and erasure.

**Driven by file changes.** [graphify](/p/graphify/) caches both AST and LLM results by file-content hash, and the LLM cache also includes a prompt fingerprint. `graphify update` re-parses changed code with no LLM and removes nodes for deleted files. Communities are recomputed on every build.

**Collection-level only.** [TrustGraph](/p/trustgraph/) has no incremental path: there is no extraction cache, graph data is deleted per collection, and deleting a document in the librarian leaves its triples in place.

Pick: LightRAG or HippoRAG for a corpus that changes daily.
Pick: graphify for repositories that change on every commit.
Pick: GraphRAG or nano-graphrag only when you can rebuild in batches.

## Per-project answers

### Graphify-Labs/graphify (answered)

**Per-file AST cache.** The AST extraction cache (`cache.py`) is content-addressed: each file's extraction result is keyed by a hash of its bytes and namespaced by package version (`cache/ast/v{version}-s{schema}/`). On re-extraction, `load_cached()` returns the cached result for unchanged files; only new or modified files are re-parsed. Stale entries from older versions are cleaned up eagerly (`_cleanup_stale_ast_entries`, cache.py:45–69).

**Per-file semantic cache.** LLM extraction results are cached individually under `cache/semantic/` (or `cache/semantic-{mode}/` for deep mode), keyed by file-content hash. The cache also fingerprints the extraction prompt text (`_PROMPT_FP_LEN = 12`, cache.py:82), so entries invalidate when the prompt changes without re-billing unchanged files on every patch release. The MCP server checkpoints each chunk's results to the semantic cache as it completes, so an interrupted run loses only in-flight chunks (`_checkpoint_chunk`, llm.py:2677–2700). Corrupt entries are detected and reported without silently serving stale data.

**File-system watch mode.** `graphify watch` (`watch.py`) monitors the filesystem for changes and writes changed paths to a `.pending_changes` file. The pending drain mechanism (`_drain_pending`, watch.py:49–78) reads and deduplicates these paths. A co-operative lock (`_rebuild_lock`) prevents concurrent rebuilds. Post-commit hooks can append to the pending file without owning the lock.

**Incremental merge.** The `build_merge()` function (build.py:1953) takes an existing graph plus new extraction results and produces a merged graph without a full rebuild. Node deduplication and edge rewiring handle ID collisions deterministically. The `resolution_context_nodes` and `resolution_context_edges` parameters (extract.py:7417–7457) let incremental re-extraction pass unchanged file metadata as read-only context, so cross-file resolvers can still bind calls to unchanged callees without re-parsing them.

**Deletion.** The `--update` CLI flag deletes graph nodes for files that no longer exist on disk before merging new extractions. There is no explicit per-entity deletion API.

**Global graph.** `global_graph.py` tracks per-repo commits and allows adding/updating a repo's graph in a cross-repo global graph via `global_add()`.


Citations: [graphify/cache.py:22-98](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/cache.py#L22-L98) · [graphify/cache.py:1274-1350](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/cache.py#L1274-L1350) · [graphify/watch.py:1-78](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/watch.py#L1-L78) · [graphify/extract.py:7410-7458](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/extract.py#L7410-L7458) · [graphify/llm.py:2677-2700](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L2677-L2700)

### HKUDS/LightRAG (answered)

**Document addition (incremental).** New documents are inserted via `ainsert` (`lightrag/lightrag.py:2310-2377`) or the pipeline `apipeline_enqueue_documents`. Each document is chunked independently, extracted against the LLM, and the resulting entities/relations are upserted into the existing graph. Entity and relation upsert is inherently additive — `upsert_node` / `upsert_edge` in NetworkX merge properties into the existing node/edge. There is no global re-clustering or full rebuild. The pipeline process (`manage_pipeline_status` in `lightrag/pipeline.py`) coordinates concurrent document processing via a `busy` reservation but allows concurrent enqueue + processing.

**Document deletion.** `adelete_by_doc_id` (`lightrag/lightrag.py:6727-6800`) removes a document and its derived graph contributions. It acquires the pipeline `busy` slot, then calls `_purge_kg_contributions` (`lightrag/lightrag.py:6015-6090`) which: (1) reads the per-document write-ahead anchors (`full_entities` / `full_relations` — stored in dedicated KV stores and written before any graph mutation in `merge_nodes_and_edges`, `lightrag/operate.py:3666-3703`); (2) for each affected entity/relation, checks whether other documents still reference it via chunk-tracking (`entity_chunks` / `relation_chunks`) — if no remaining sources, it deletes the graph node/edge outright; if other documents share it, it rebuilds the entity/relation description from the surviving chunks' cache (LLM re-extraction from cache); (3) deletes the chunks themselves only after all graph contributions are removed; (4) removes the recovery anchors last. The process is journaled in `doc_status.metadata` so a crash mid-purge is resumable (`_resolve_purge_recovery_proof`, `lightrag/lightrag.py`).

**Custom chunk patching.** `ainsert_custom_chunks` (`lightrag/lightrag.py:2394-2509`) supports patching documents without full deletion — it writes new chunks, re-extracts, and merges them into the graph without touching the original chunks. The operation is journaled for crash recovery.

**Extraction caching.** LLM extraction results are cached in `llm_response_cache` (default `JsonKVStorage`) and keyed by prompt hash + model identity. On re-insertion or document deletion, cached extractions are replayed to rebuild entities/relations without re-calling the LLM. The extraction cache is partitioned by `llm_cache_identity` (model/provider) and each cache row is attached to its owning chunk before being written (`use_llm_func_with_cache` at `lightrag/utils.py:5515-5577`). The `cache_keys_collector` mechanism allows batch cache operations.


Citations: [lightrag/lightrag.py:2310-2377](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L2310-L2377) · [lightrag/lightrag.py:6015-6090](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L6015-L6090) · [lightrag/lightrag.py:6727-6826](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L6727-L6826) · [lightrag/lightrag.py:2394-2466](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/lightrag.py#L2394-L2466) · [lightrag/operate.py:3514-3591](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3514-L3591) · [lightrag/utils.py:5515-5577](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/utils.py#L5515-L5577)

### microsoft/graphrag (answered)

Incremental indexing has **partial support** with title-based delta detection.

**Delta detection.** `get_delta_docs` at `incremental_index.py:29` compares the input dataset against previously-indexed documents by `title` (lines 49-50). `InputDelta` captures `new_inputs` (titles not in previous docs) and `deleted_inputs` (previous titles missing from input).

**Re-indexing.** `update/entities.py`, `update/relationships.py`, and `update/communities.py` run extraction and clustering on only the delta documents. `concat_dataframes` (line 61) merges old and delta tables, assigning new `human_readable_id` values starting at `max(old) + 1`.

**Deletion.** The deleted_inputs DataFrame is populated, but the system does not fully re-index downstream artifacts — orphaned entities/edges/communities remain unless a full rebuild is run.

**Caching.** The LLM cache (`with_cache` at `cache_middleware.py:21`) caches by full request content, so re-extracting the same text on unchanged documents hits cache. But this is content-addressable, not incrementally aware.

**Limitations.** The approach is title-based so two documents with the same title are treated as the same. Editing documents in place requires manual deletion or full rebuild.

> **Editor's note.** Correction: `deleted_inputs` is computed but never read anywhere, so deleted or edited documents stay in the index. The `update/*.py` modules only merge the old and delta tables (extraction runs earlier, on the delta documents). Delta communities are clustered on the delta graph alone and appended with shifted IDs; the old hierarchy is not re-clustered.

Citations: [packages/graphrag/graphrag/index/update/incremental_index.py:29-78](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/update/incremental_index.py#L29-L78) · [packages/graphrag-llm/graphrag_llm/middleware/with_cache.py:21-154](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-llm/graphrag_llm/middleware/with_cache.py#L21-L154) · [packages/graphrag/graphrag/index/update/entities.py:1-10](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/update/entities.py#L1-L10) · [packages/graphrag/graphrag/index/update/communities.py:1-10](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/update/communities.py#L1-L10)

### semantica-agi/semantica (answered)

Incremental updates are handled at multiple levels. The `EntityMerger.incremental_merge()` method (`semantica/deduplication/entity_merger.py:388-473`) merges new entities into an existing set by running `DuplicateDetector.incremental_detect()` between new and existing entities, then merging each candidate pair with configurable strategies — avoiding duplicate merges via processed-ID tracking. The `GraphBuilder.build()` method is designed to be re-runnable: it merges sources and can be called repeatedly on new documents, with `merge_entities=True` enabling deduplication against previously built entity sets. The builder's entity resolver (`EntityResolver`) can operate on the full accumulated set each time, and `_remap_relationship_endpoints()` rewrites relation endpoints after entity merges change canonical IDs (graph_builder.py:370-454). **Ingestion** is source-agnostic: `FileIngestor`, `StreamIngestor` (Kafka, RabbitMQ, Pulsar, Kinesis), `WebIngestor`, `DBIngestor`, and dozens of other ingestors in `semantica/ingest/` each produce documents that feed into extraction. The **Pipeline** module (`pipeline_builder.py`, `execution_engine.py`) provides full orchestration with retry, parallelism, and resource scheduling. **Version management** via `TemporalVersionManager` (`change_management/managers.py`) supports snapshot creation, diffs between versions, and persistent storage (SQLite or in-memory). **Deletion** is not explicitly handled as a first-class operation — there is no built-in mechanism to remove a document's extracted entities/relations after they are added. The system is additive: incremental merges consolidate but do not remove stale data. Extraction caching (`semantica/semantic_extract/cache.py`) provides TTL-based caching for entity, relation, and triplet extraction results, with pluggable backends (in-memory LRU or persistent SQLite). The cache is keyed by text hash and method parameters, avoiding re-extraction from unchanged source texts across rebuilds.

> **Editor's note.** Correction: there is no document-level delete, but entity-level deletion exists via `ContextGraph.purge_node` (tombstoned) and `ErasureCoordinator.erase_entity`, which cascades to agent memory and vector stores and returns a receipt.

Citations: [semantica/deduplication/entity_merger.py:388-473](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/deduplication/entity_merger.py#L388-L473) · [semantica/kg/graph_builder.py:370-455](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/graph_builder.py#L370-L455) · [semantica/change_management/managers.py:1-100](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/change_management/managers.py#L1-L100) · [semantica/semantic_extract/cache.py:1-80](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/semantic_extract/cache.py#L1-L80) · [semantica/ingest/file_ingestor.py:1-60](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/ingest/file_ingestor.py#L1-L60) · [semantica/ingest/stream_ingestor.py:1-60](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/ingest/stream_ingestor.py#L1-L60)

### neo4j-labs/llm-graph-builder (answered)

**Retry conditions for existing documents.** The system supports three retry modes per document (`constants.py:823-825`): `start_from_beginning` (reprocesses all chunks), `delete_entities_and_start_from_beginning` (deletes entity nodes unique to this document then reprocesses), and `start_from_last_processed_position` (resumes from the last chunk that lacks an embedding or has no `HAS_ENTITY` relationship). These are handled in `get_chunkId_chunkDoc_list()` (`main.py:685-744`), which either re-splits the document into chunks or reuses existing `Chunk` nodes from Neo4j. For resumption, it queries `QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION` (`constants.py:801-808`) to find the first chunk without an embedding, then processes from there. `set_status_retry()` (`main.py:945-981`) resets document counters and optionally deletes entities via `QUERY_TO_DELETE_EXISTING_ENTITIES` (`constants.py:791-799`) which only deletes entities not referenced by other documents.

**Adding new documents.** Each document is independent — uploading a new file creates a new `:Document` node and processes it end-to-end. There is no overall corpus index that requires rebuilding. `claim_document_for_processing()` (`graphDB_dataAccess.py:334-360`) uses an atomic `WHERE d.status <> 'Processing'` guard so concurrent requests don't duplicate work.

**Deletion.** `delete_file_from_graph()` (`graphDB_dataAccess.py:362-428`) removes a document and optionally its unique entities (those not referenced by other docs) or just the document+chunks. Orphaned `__Community__` nodes are also cleaned up via `query_to_delete_communities`.

**Caching of extraction results.** There is no extraction-result cache at the LLM level; every call to `get_graph_from_llm()` sends chunk text to the LLM. However, chunk embeddings are not re-computed when resuming: `create_chunk_embeddings()` sets `c.embedding` on each `Chunk` node, and the resume logic (`QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION`) skips chunks that already have embeddings. The optional `GCS_FILE_CACHE` (`main.py:56-60`) caches uploaded files in Google Cloud Storage rather than locally, but this is file-level storage, not extraction caching.


Citations: [backend/src/main.py:685-744](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L685-L744) · [backend/src/main.py:945-981](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L945-L981) · [backend/src/shared/constants.py:783-825](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L783-L825) · [backend/src/graphDB_dataAccess.py:334-428](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L334-L428) · [backend/src/graphDB_dataAccess.py:362-428](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L362-L428)

### OSU-NLP-Group/HippoRAG (answered)

**Document addition.** `index()` (HippoRAG.py:499–593) is designed for incremental use. It calls `load_existing_openie` (1354–1404), which loads any prior OpenIE results from `openie_state.json` or the alternate `openie_results_path`. Only chunks whose hash IDs are not already in the persisted state (`chunk_keys_to_save`) are sent to `batch_openie` for NER+triple extraction. This means re-indexing the same documents is a no-op, and adding new documents only processes the new chunks. The graph is then augmented: `add_fact_edges` and `add_passage_edges` only process new chunks (they check `current_graph_nodes` at lines 1204 and 1256). `add_synonymy_edges` is optimized: when `_pending_synonymy_entity_ids` is set (only for new entities), it runs KNN only for the new entities against all existing entities (1329–1350), rather than recomputing all-vs-all (src/hipporag/HippoRAG.py:1301, 1316–1317).

**Document deletion.** `delete()` (595–682) finds the chunks to remove via hash, looks up their triples and entities, checks `proc_triples_to_docs` to see if a triple is referenced by other chunks (via `remove_sources_from_mapping`), and only removes unreferenced facts and entities. It then deletes vertices from the iGraph after cleaning up `_fact_edge_source_counts` and edge weights (684–702). Embedding store deletions are done separately per namespace.

**State consistency.** The `index_manifest.json` (317–344) binds the OpenIE provenance (model, endpoint, parameters) and embedding config to the persisted state. Loading with a different model/endpoint raises `StateConsistencyError`. The `force_index_from_scratch` and `force_openie_from_scratch` flags allow explicit rebuilds.

**Caching of extraction results.** OpenIE results (NER + triples) are saved to `openie_state.json` and optionally to a human-readable `openie_results_path`. On subsequent index calls, only missing chunk hashes are re-processed — the cached NER and triple outputs for unchanged chunks are reused at the graph-building stage (HippoRAG.py:544–553).


Citations: [src/hipporag/HippoRAG.py:499-553](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L499-L553) · [src/hipporag/HippoRAG.py:595-702](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L595-L702) · [src/hipporag/HippoRAG.py:1354-1404](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1354-L1404) · [src/hipporag/HippoRAG.py:1296-1320](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1296-L1320) · [src/hipporag/HippoRAG.py:317-344](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L317-L344)

### gusye1234/nano-graphrag (answered)

Incremental indexing has limited support with known gaps. On each call to `ainsert` (`graphrag.py:277-348`), new documents are checked against `full_docs` KV storage by their md5-hash ID: if the hash already exists, the document is skipped (`filter_keys` on line 287). Similarly, chunks are checked against `text_chunks` storage (line 304). So re-inserting the same document is idempotent — a no-op if all hashes match.

**Graph merging.** When entities or relations from new chunks overlap with existing graph nodes, the merge functions `_merge_nodes_then_upsert` and `_merge_edges_then_upsert` (`_op.py:182-279`) read existing node/edge data from the graph, merge descriptions (deduplicated, sorted, joined by `<SEP>`), and update. Edge weights are summed across extraction runs. This means the graph grows incrementally without data loss.

**No incremental communities.** The code explicitly states: `# TODO: don't support incremental update for communities now, so we have to drop all` (`graphrag.py:318`), followed by `await self.community_reports.drop()`. Every call to `ainsert` wipes the community reports and recomputes from scratch (clustering → community report generation). This is the biggest gap in incremental support.

**Embedding re-upsert.** Entity embeddings are recomputed for all extracted entities and upserted into the vector DB; the vector DB storage (`NanoVectorDBStorage.upsert`) uses `NanoVectorDB.upsert` which updates existing entries by matching on `__id__`.

**Deletion.** There is no API for deleting documents, entities, or edges. The `BaseKVStorage.drop()` method exists and wipes the in-memory dict (used only for community reports), but the full_docs and text_chunks stores are append-only within a session. A deletion would require manual file removal and re-indexing.

**LLM cache persistence.** The `llm_response_cache` KV store persists across sessions (it is a JSON file loaded at init), so repeated extraction or query calls benefit from cached LLM responses even after a restart — but cache entries are keyed by (model, messages) hash, not by document ID, so a re-insert of the same document with the same chunk content would reuse cached extractions.


Citations: [nano_graphrag/graphrag.py:277-348](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/graphrag.py#L277-L348) · [nano_graphrag/graphrag.py:318-319](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/graphrag.py#L318-L319) · [nano_graphrag/_op.py:182-279](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_op.py#L182-L279) · [nano_graphrag/_storage/kv_json.py:39-46](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_storage/kv_json.py#L39-L46)

### pingcap/autoflow (answered)

AutoFlow supports **per-document incremental indexing** and **re-indexing of failed tasks**, but has no incremental graph update for individual entity/relationship changes at the chunk level.

**Adding new documents:** Documents are added per knowledge base. The `import_documents_for_knowledge_base` Celery task feeds documents through `IndexService.build_vector_index_for_document()` and `IndexService.build_kg_index_for_chunk()` (`backend/app/rag/build_index.py:51-159`). For vector indexing, each document's content is chunked by LlamaIndex's `SentenceSplitter`/`MarkdownNodeParser`, embedded, and inserted into the `chunks_{namespace}` table. For the KG index, each chunk becomes a `TextNode`, entities and relationships are extracted via the LLM extraction pipeline, and saved to the `entities`/`relationships` tables.

**Re-indexing (update):** Documents can be reindexed via the admin API (`POST /admin/knowledge_bases/{kb_id}/documents/reindex`, `backend/app/api/admin_routes/knowledge_base/document/routes.py:134-222`). The process sets the document's `index_status` to `PENDING` and enqueues `build_index_for_document` (vector) and `build_kg_index_for_chunk` (KG) Celery tasks. It only reindexes documents whose status is `FAILED` (or when `reindex_completed_task=True`). For chunks that already have `COMPLETED` KG index, the reindex is skipped (`routes.py:201-215`).

**Deletion:** Deleting a document (`DELETE /admin/knowledge_bases/{kb_id}/documents/{document_id}`, `routes.py:91-131`) first calls `graph_repo.delete_document_relationships()` which deletes all relationships whose `chunk_id` belongs to that document's chunks (`backend/app/repositories/graph.py:57-64`). Then `delete_orphaned_entities()` removes any entity left with no relationships. Finally the chunk entries and document itself are deleted.

**Idempotency / duplicate detection:** `KnowledgeGraphIndex.add_chunk()` (`core/autoflow/knowledge_graph/index.py:38-51`) checks `list_relationships(chunk_id=...)` before extraction — if any relationships already exist for that chunk, it skips. The backend `TiDBGraphStore.save()` similarly checks if any relationship with the same `chunk_id` meta already exists (`tidb_graph_store.py:206-213`). This prevents double-extraction.

**No true incremental graph merge:** When a new chunk is added and its entities overlap with existing ones, they are resolved via the embedding-similarity + LLM-merge mechanism in `get_or_create_entity()` — existing entities are updated (merged) rather than duplicated. But there is no mechanism to re-extract or update relationships for previously-indexed chunks when a new document changes the overall understanding of an entity.

**Namespace isolation:** Each knowledge base gets its own set of tables (`chunks_{kb_id}`, `entities_{kb_id}`, `relationships_{kb_id}`), so operations on one KB do not affect others.


Citations: [backend/app/api/admin_routes/knowledge_base/document/routes.py:91-131](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/api/admin_routes/knowledge_base/document/routes.py#L91-L131) · [backend/app/api/admin_routes/knowledge_base/document/routes.py:173-222](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/api/admin_routes/knowledge_base/document/routes.py#L173-L222) · [backend/app/repositories/graph.py:45-64](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/repositories/graph.py#L45-L64) · [core/autoflow/knowledge_graph/index.py:38-51](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/index.py#L38-L51) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:198-217](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L198-L217)

### trustgraph-ai/trustgraph (insufficient evidence)

This repository does not implement incremental indexing. Across the storage backends, the librarian, flow definitions and extraction processors:

- **No incremental document changes**: Documents flow through a linear pipeline (document -> chunker -> extractors -> storage). Any new or changed document must be re-ingested through the entire pipeline. No change detection, fingerprinting, or differential update exists.

- **No per-document deletion**: The only deletion mechanism is collection-level (`delete_collection`, e.g. `storage/triples/neo4j/write.py:298-327`), removing all data for a workspace+collection pair.

- **No extraction result cache**: Extraction processors call the LLM on every chunk. Re-ingesting a document duplicates all previous extractions. The `EntityRegistry` (`entity_normalizer.py:113-165`) is per-request-scoped, not persistent.

- **No index staleness tracking**: No document versioning or tracking which chunks/triples derive from which document version. Provenance triples are append-only.

Updates require full re-ingestion of the collection; deletion is collection-granularity only.


Citations: [trustgraph-flow/trustgraph/storage/triples/neo4j/write.py:298-327](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/storage/triples/neo4j/write.py#L298-L327) · [trustgraph-flow/trustgraph/extract/kg/ontology/entity_normalizer.py:113-165](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/entity_normalizer.py#L113-L165) · [trustgraph-flow/trustgraph/extract/kg/ontology/extract.py:349-467](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/extract.py#L349-L467)

### zilliztech/vector-graph-rag (answered)

**Full rebuild is the default.** `rebuild_documents()` (`rag.py:942-1012`) drops all three Milvus collections and recreates them, then indexes everything fresh. The legacy `add_documents()` (line 888-940) calls `rebuild_documents()` internally.

**Source-level incremental upsert** is implemented via `upsert_documents_by_source()` (`rag.py:1014-1122`). It operates at the granularity of a "source" (a file, URL, or business record identified by a `source` metadata field). The flow: (1) extracts triplets from new documents, (2) builds graph records, (3) calls `delete_documents_by_source()` to remove all passages for that source, (4) calls `_insert_incremental_graph()` which checks for **existing entities and relations by normalized text** (`storage/milvus.py:646-739`). When an entity or relation already exists, its metadata is **merged** (passage/relation IDs are appended using `_merge_unique`), and `upsert` writes updated records. When they don't exist, they are inserted as new records (`rag.py:596-731`).

**Deletion** via `delete_documents_by_source()` (`rag.py:1137-1250`) cascades: it finds passages matching the source, then updates or deletes relations (removing the passage ID from the relation's adjacency list; deleting the relation entirely if it has no remaining passages), then similarly updates or deletes orphaned entities, and finally deletes the passages themselves.

**LLM response caching** (`llm/cache.py:1-165`) persists extraction results to disk, avoiding re-extraction when rebuilding or re-indexing. The NER (named entity recognition) cache also supports a TSV file format for HippoRAG evaluation compatibility (`llm/extractor.py:296-332`).

The cached result `_extraction_result` and the lazy retriever are reset after any mutation via `self._retriever = None` (rag.py:1010, 1121, 1249).


Citations: [src/vector_graph_rag/rag.py:942-1012](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/rag.py#L942-L1012) · [src/vector_graph_rag/rag.py:1014-1122](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/rag.py#L1014-L1122) · [src/vector_graph_rag/rag.py:1137-1250](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/rag.py#L1137-L1250)
