LLMs Technical Reviews

How are updates and incremental indexing handled?

Adding or changing documents without a full rebuild; deletion; caching of extraction results.

Verdict

LightRAG and HippoRAG handle change best. Both add documents without a rebuild and delete them with reference-counted cleanup of shared entities. GraphRAG and nano-graphrag are in practice append-only.

Add and delete per document. LightRAG merges each document into the live graph. adelete_by_doc_id rebuilds shared entities from cached extraction results, and deletion is journaled so a crash can resume. HippoRAG sends only unseen chunk hashes to OpenIE and runs synonymy KNN only for new entities. Its manifest refuses to mix state from two different models. Vector Graph RAG has upsert_documents_by_source and a cascading delete_documents_by_source, but only in Python: its REST import endpoints always rebuild the whole graph. LLM Graph Builder processes each document on its own, can resume from the last processed chunk, and deletes only the entities no other document references. It has no extraction cache, so a retry pays again. AutoFlow deletes a document’s relationships and then any orphaned entities, but merged descriptions stay behind. Re-indexing skips chunks that already have relationships.

Append-only. GraphRAG’s update finds new documents by title. It computes deleted documents but never uses that list, and it appends delta communities without re-clustering. nano-graphrag skips known document and chunk hashes, but every insert drops and regenerates all community reports. It has no delete API. Semantica merges new entities with incremental_merge and caches extraction with a TTL. It has no document-level delete, only entity-level purge and erasure.

Driven by file changes. graphify caches both AST and LLM results by file-content hash, and the LLM cache also includes a prompt fingerprint. graphify update re-parses changed code with no LLM and removes nodes for deleted files. Communities are recomputed on every build.

Collection-level only. TrustGraph has no incremental path: there is no extraction cache, graph data is deleted per collection, and deleting a document in the librarian leaves its triples in place.

Pick: LightRAG or HippoRAG for a corpus that changes daily. Pick: graphify for repositories that change on every commit. Pick: GraphRAG or nano-graphrag only when you can rebuild in batches.

Per-project answers

Graphify-Labs/graphify

answered

Per-file AST cache. The AST extraction cache (cache.py) is content-addressed: each file's extraction result is keyed by a hash of its bytes and namespaced by package version (cache/ast/v{version}-s{schema}/). On re-extraction, load_cached() returns the cached result for unchanged files; only new or modified files are re-parsed. Stale entries from older versions are cleaned up eagerly (_cleanup_stale_ast_entries, cache.py:45–69).

Per-file semantic cache. LLM extraction results are cached individually under cache/semantic/ (or cache/semantic-{mode}/ for deep mode), keyed by file-content hash. The cache also fingerprints the extraction prompt text (_PROMPT_FP_LEN = 12, cache.py:82), so entries invalidate when the prompt changes without re-billing unchanged files on every patch release. The MCP server checkpoints each chunk's results to the semantic cache as it completes, so an interrupted run loses only in-flight chunks (_checkpoint_chunk, llm.py:2677–2700). Corrupt entries are detected and reported without silently serving stale data.

File-system watch mode. graphify watch (watch.py) monitors the filesystem for changes and writes changed paths to a .pending_changes file. The pending drain mechanism (_drain_pending, watch.py:49–78) reads and deduplicates these paths. A co-operative lock (_rebuild_lock) prevents concurrent rebuilds. Post-commit hooks can append to the pending file without owning the lock.

Incremental merge. The build_merge() function (build.py:1953) takes an existing graph plus new extraction results and produces a merged graph without a full rebuild. Node deduplication and edge rewiring handle ID collisions deterministically. The resolution_context_nodes and resolution_context_edges parameters (extract.py:7417–7457) let incremental re-extraction pass unchanged file metadata as read-only context, so cross-file resolvers can still bind calls to unchanged callees without re-parsing them.

Deletion. The --update CLI flag deletes graph nodes for files that no longer exist on disk before merging new extractions. There is no explicit per-entity deletion API.

Global graph. global_graph.py tracks per-repo commits and allows adding/updating a repo's graph in a cross-repo global graph via global_add().

HKUDS/LightRAG

answered

Document addition (incremental). New documents are inserted via ainsert (lightrag/lightrag.py:2310-2377) or the pipeline apipeline_enqueue_documents. Each document is chunked independently, extracted against the LLM, and the resulting entities/relations are upserted into the existing graph. Entity and relation upsert is inherently additive — upsert_node / upsert_edge in NetworkX merge properties into the existing node/edge. There is no global re-clustering or full rebuild. The pipeline process (manage_pipeline_status in lightrag/pipeline.py) coordinates concurrent document processing via a busy reservation but allows concurrent enqueue + processing.

Document deletion. adelete_by_doc_id (lightrag/lightrag.py:6727-6800) removes a document and its derived graph contributions. It acquires the pipeline busy slot, then calls _purge_kg_contributions (lightrag/lightrag.py:6015-6090) which: (1) reads the per-document write-ahead anchors (full_entities / full_relations — stored in dedicated KV stores and written before any graph mutation in merge_nodes_and_edges, lightrag/operate.py:3666-3703); (2) for each affected entity/relation, checks whether other documents still reference it via chunk-tracking (entity_chunks / relation_chunks) — if no remaining sources, it deletes the graph node/edge outright; if other documents share it, it rebuilds the entity/relation description from the surviving chunks' cache (LLM re-extraction from cache); (3) deletes the chunks themselves only after all graph contributions are removed; (4) removes the recovery anchors last. The process is journaled in doc_status.metadata so a crash mid-purge is resumable (_resolve_purge_recovery_proof, lightrag/lightrag.py).

Custom chunk patching. ainsert_custom_chunks (lightrag/lightrag.py:2394-2509) supports patching documents without full deletion — it writes new chunks, re-extracts, and merges them into the graph without touching the original chunks. The operation is journaled for crash recovery.

Extraction caching. LLM extraction results are cached in llm_response_cache (default JsonKVStorage) and keyed by prompt hash + model identity. On re-insertion or document deletion, cached extractions are replayed to rebuild entities/relations without re-calling the LLM. The extraction cache is partitioned by llm_cache_identity (model/provider) and each cache row is attached to its owning chunk before being written (use_llm_func_with_cache at lightrag/utils.py:5515-5577). The cache_keys_collector mechanism allows batch cache operations.

microsoft/graphrag

answered

Incremental indexing has partial support with title-based delta detection.

Delta detection. get_delta_docs at incremental_index.py:29 compares the input dataset against previously-indexed documents by title (lines 49-50). InputDelta captures new_inputs (titles not in previous docs) and deleted_inputs (previous titles missing from input).

Re-indexing. update/entities.py, update/relationships.py, and update/communities.py run extraction and clustering on only the delta documents. concat_dataframes (line 61) merges old and delta tables, assigning new human_readable_id values starting at max(old) + 1.

Deletion. The deleted_inputs DataFrame is populated, but the system does not fully re-index downstream artifacts — orphaned entities/edges/communities remain unless a full rebuild is run.

Caching. The LLM cache (with_cache at cache_middleware.py:21) caches by full request content, so re-extracting the same text on unchanged documents hits cache. But this is content-addressable, not incrementally aware.

Limitations. The approach is title-based so two documents with the same title are treated as the same. Editing documents in place requires manual deletion or full rebuild.

Editor's note. Correction: deleted_inputs is computed but never read anywhere, so deleted or edited documents stay in the index. The update/*.py modules only merge the old and delta tables (extraction runs earlier, on the delta documents). Delta communities are clustered on the delta graph alone and appended with shifted IDs; the old hierarchy is not re-clustered.

semantica-agi/semantica

answered

Incremental updates are handled at multiple levels. The EntityMerger.incremental_merge() method (semantica/deduplication/entity_merger.py:388-473) merges new entities into an existing set by running DuplicateDetector.incremental_detect() between new and existing entities, then merging each candidate pair with configurable strategies — avoiding duplicate merges via processed-ID tracking. The GraphBuilder.build() method is designed to be re-runnable: it merges sources and can be called repeatedly on new documents, with merge_entities=True enabling deduplication against previously built entity sets. The builder's entity resolver (EntityResolver) can operate on the full accumulated set each time, and _remap_relationship_endpoints() rewrites relation endpoints after entity merges change canonical IDs (graph_builder.py:370-454). Ingestion is source-agnostic: FileIngestor, StreamIngestor (Kafka, RabbitMQ, Pulsar, Kinesis), WebIngestor, DBIngestor, and dozens of other ingestors in semantica/ingest/ each produce documents that feed into extraction. The Pipeline module (pipeline_builder.py, execution_engine.py) provides full orchestration with retry, parallelism, and resource scheduling. Version management via TemporalVersionManager (change_management/managers.py) supports snapshot creation, diffs between versions, and persistent storage (SQLite or in-memory). Deletion is not explicitly handled as a first-class operation — there is no built-in mechanism to remove a document's extracted entities/relations after they are added. The system is additive: incremental merges consolidate but do not remove stale data. Extraction caching (semantica/semantic_extract/cache.py) provides TTL-based caching for entity, relation, and triplet extraction results, with pluggable backends (in-memory LRU or persistent SQLite). The cache is keyed by text hash and method parameters, avoiding re-extraction from unchanged source texts across rebuilds.

Editor's note. Correction: there is no document-level delete, but entity-level deletion exists via ContextGraph.purge_node (tombstoned) and ErasureCoordinator.erase_entity, which cascades to agent memory and vector stores and returns a receipt.

neo4j-labs/llm-graph-builder

answered

Retry conditions for existing documents. The system supports three retry modes per document (constants.py:823-825): start_from_beginning (reprocesses all chunks), delete_entities_and_start_from_beginning (deletes entity nodes unique to this document then reprocesses), and start_from_last_processed_position (resumes from the last chunk that lacks an embedding or has no HAS_ENTITY relationship). These are handled in get_chunkId_chunkDoc_list() (main.py:685-744), which either re-splits the document into chunks or reuses existing Chunk nodes from Neo4j. For resumption, it queries QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION (constants.py:801-808) to find the first chunk without an embedding, then processes from there. set_status_retry() (main.py:945-981) resets document counters and optionally deletes entities via QUERY_TO_DELETE_EXISTING_ENTITIES (constants.py:791-799) which only deletes entities not referenced by other documents.

Adding new documents. Each document is independent — uploading a new file creates a new :Document node and processes it end-to-end. There is no overall corpus index that requires rebuilding. claim_document_for_processing() (graphDB_dataAccess.py:334-360) uses an atomic WHERE d.status <> 'Processing' guard so concurrent requests don't duplicate work.

Deletion. delete_file_from_graph() (graphDB_dataAccess.py:362-428) removes a document and optionally its unique entities (those not referenced by other docs) or just the document+chunks. Orphaned __Community__ nodes are also cleaned up via query_to_delete_communities.

Caching of extraction results. There is no extraction-result cache at the LLM level; every call to get_graph_from_llm() sends chunk text to the LLM. However, chunk embeddings are not re-computed when resuming: create_chunk_embeddings() sets c.embedding on each Chunk node, and the resume logic (QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION) skips chunks that already have embeddings. The optional GCS_FILE_CACHE (main.py:56-60) caches uploaded files in Google Cloud Storage rather than locally, but this is file-level storage, not extraction caching.

OSU-NLP-Group/HippoRAG

answered

Document addition. index() (HippoRAG.py:499–593) is designed for incremental use. It calls load_existing_openie (1354–1404), which loads any prior OpenIE results from openie_state.json or the alternate openie_results_path. Only chunks whose hash IDs are not already in the persisted state (chunk_keys_to_save) are sent to batch_openie for NER+triple extraction. This means re-indexing the same documents is a no-op, and adding new documents only processes the new chunks. The graph is then augmented: add_fact_edges and add_passage_edges only process new chunks (they check current_graph_nodes at lines 1204 and 1256). add_synonymy_edges is optimized: when _pending_synonymy_entity_ids is set (only for new entities), it runs KNN only for the new entities against all existing entities (1329–1350), rather than recomputing all-vs-all (src/hipporag/HippoRAG.py:1301, 1316–1317).

Document deletion. delete() (595–682) finds the chunks to remove via hash, looks up their triples and entities, checks proc_triples_to_docs to see if a triple is referenced by other chunks (via remove_sources_from_mapping), and only removes unreferenced facts and entities. It then deletes vertices from the iGraph after cleaning up _fact_edge_source_counts and edge weights (684–702). Embedding store deletions are done separately per namespace.

State consistency. The index_manifest.json (317–344) binds the OpenIE provenance (model, endpoint, parameters) and embedding config to the persisted state. Loading with a different model/endpoint raises StateConsistencyError. The force_index_from_scratch and force_openie_from_scratch flags allow explicit rebuilds.

Caching of extraction results. OpenIE results (NER + triples) are saved to openie_state.json and optionally to a human-readable openie_results_path. On subsequent index calls, only missing chunk hashes are re-processed — the cached NER and triple outputs for unchanged chunks are reused at the graph-building stage (HippoRAG.py:544–553).

gusye1234/nano-graphrag

answered

Incremental indexing has limited support with known gaps. On each call to ainsert (graphrag.py:277-348), new documents are checked against full_docs KV storage by their md5-hash ID: if the hash already exists, the document is skipped (filter_keys on line 287). Similarly, chunks are checked against text_chunks storage (line 304). So re-inserting the same document is idempotent — a no-op if all hashes match.

Graph merging. When entities or relations from new chunks overlap with existing graph nodes, the merge functions _merge_nodes_then_upsert and _merge_edges_then_upsert (_op.py:182-279) read existing node/edge data from the graph, merge descriptions (deduplicated, sorted, joined by <SEP>), and update. Edge weights are summed across extraction runs. This means the graph grows incrementally without data loss.

No incremental communities. The code explicitly states: # TODO: don't support incremental update for communities now, so we have to drop all (graphrag.py:318), followed by await self.community_reports.drop(). Every call to ainsert wipes the community reports and recomputes from scratch (clustering → community report generation). This is the biggest gap in incremental support.

Embedding re-upsert. Entity embeddings are recomputed for all extracted entities and upserted into the vector DB; the vector DB storage (NanoVectorDBStorage.upsert) uses NanoVectorDB.upsert which updates existing entries by matching on __id__.

Deletion. There is no API for deleting documents, entities, or edges. The BaseKVStorage.drop() method exists and wipes the in-memory dict (used only for community reports), but the full_docs and text_chunks stores are append-only within a session. A deletion would require manual file removal and re-indexing.

LLM cache persistence. The llm_response_cache KV store persists across sessions (it is a JSON file loaded at init), so repeated extraction or query calls benefit from cached LLM responses even after a restart — but cache entries are keyed by (model, messages) hash, not by document ID, so a re-insert of the same document with the same chunk content would reuse cached extractions.

pingcap/autoflow

answered

AutoFlow supports per-document incremental indexing and re-indexing of failed tasks, but has no incremental graph update for individual entity/relationship changes at the chunk level.

Adding new documents: Documents are added per knowledge base. The import_documents_for_knowledge_base Celery task feeds documents through IndexService.build_vector_index_for_document() and IndexService.build_kg_index_for_chunk() (backend/app/rag/build_index.py:51-159). For vector indexing, each document's content is chunked by LlamaIndex's SentenceSplitter/MarkdownNodeParser, embedded, and inserted into the chunks_{namespace} table. For the KG index, each chunk becomes a TextNode, entities and relationships are extracted via the LLM extraction pipeline, and saved to the entities/relationships tables.

Re-indexing (update): Documents can be reindexed via the admin API (POST /admin/knowledge_bases/{kb_id}/documents/reindex, backend/app/api/admin_routes/knowledge_base/document/routes.py:134-222). The process sets the document's index_status to PENDING and enqueues build_index_for_document (vector) and build_kg_index_for_chunk (KG) Celery tasks. It only reindexes documents whose status is FAILED (or when reindex_completed_task=True). For chunks that already have COMPLETED KG index, the reindex is skipped (routes.py:201-215).

Deletion: Deleting a document (DELETE /admin/knowledge_bases/{kb_id}/documents/{document_id}, routes.py:91-131) first calls graph_repo.delete_document_relationships() which deletes all relationships whose chunk_id belongs to that document's chunks (backend/app/repositories/graph.py:57-64). Then delete_orphaned_entities() removes any entity left with no relationships. Finally the chunk entries and document itself are deleted.

Idempotency / duplicate detection: KnowledgeGraphIndex.add_chunk() (core/autoflow/knowledge_graph/index.py:38-51) checks list_relationships(chunk_id=...) before extraction — if any relationships already exist for that chunk, it skips. The backend TiDBGraphStore.save() similarly checks if any relationship with the same chunk_id meta already exists (tidb_graph_store.py:206-213). This prevents double-extraction.

No true incremental graph merge: When a new chunk is added and its entities overlap with existing ones, they are resolved via the embedding-similarity + LLM-merge mechanism in get_or_create_entity() — existing entities are updated (merged) rather than duplicated. But there is no mechanism to re-extract or update relationships for previously-indexed chunks when a new document changes the overall understanding of an entity.

Namespace isolation: Each knowledge base gets its own set of tables (chunks_{kb_id}, entities_{kb_id}, relationships_{kb_id}), so operations on one KB do not affect others.

zilliztech/vector-graph-rag

answered

Full rebuild is the default. rebuild_documents() (rag.py:942-1012) drops all three Milvus collections and recreates them, then indexes everything fresh. The legacy add_documents() (line 888-940) calls rebuild_documents() internally.

Source-level incremental upsert is implemented via upsert_documents_by_source() (rag.py:1014-1122). It operates at the granularity of a "source" (a file, URL, or business record identified by a source metadata field). The flow: (1) extracts triplets from new documents, (2) builds graph records, (3) calls delete_documents_by_source() to remove all passages for that source, (4) calls _insert_incremental_graph() which checks for existing entities and relations by normalized text (storage/milvus.py:646-739). When an entity or relation already exists, its metadata is merged (passage/relation IDs are appended using _merge_unique), and upsert writes updated records. When they don't exist, they are inserted as new records (rag.py:596-731).

Deletion via delete_documents_by_source() (rag.py:1137-1250) cascades: it finds passages matching the source, then updates or deletes relations (removing the passage ID from the relation's adjacency list; deleting the relation entirely if it has no remaining passages), then similarly updates or deletes orphaned entities, and finally deletes the passages themselves.

LLM response caching (llm/cache.py:1-165) persists extraction results to disk, avoiding re-extraction when rebuilding or re-indexing. The NER (named entity recognition) cache also supports a TSV file format for HippoRAG evaluation compatibility (llm/extractor.py:296-332).

The cached result _extraction_result and the lazy retriever are reset after any mutation via self._retriever = None (rag.py:1010, 1121, 1249).

trustgraph-ai/trustgraph

insufficient evidence

This repository does not implement incremental indexing. Across the storage backends, the librarian, flow definitions and extraction processors:

  • No incremental document changes: Documents flow through a linear pipeline (document -> chunker -> extractors -> storage). Any new or changed document must be re-ingested through the entire pipeline. No change detection, fingerprinting, or differential update exists.

  • No per-document deletion: The only deletion mechanism is collection-level (delete_collection, e.g. storage/triples/neo4j/write.py:298-327), removing all data for a workspace+collection pair.

  • No extraction result cache: Extraction processors call the LLM on every chunk. Re-ingesting a document duplicates all previous extractions. The EntityRegistry (entity_normalizer.py:113-165) is per-request-scoped, not persistent.

  • No index staleness tracking: No document versioning or tracking which chunks/triples derive from which document version. Provenance triples are append-only.

Updates require full re-ingestion of the collection; deletion is collection-granularity only.

← How does query-time retrieval use the graph? · How are LLM cost and latency controlled during indexing and query? →