zilliztech/vector-graph-rag
Python Graph RAG library that stores entities, relations and passages as three Milvus collections and walks the graph by ID lookups.
Overview
Vector Graph RAG is a Zilliz-maintained Python library that does Graph RAG without a graph database. An LLM extracts subject-predicate-object triplets from each passage. The library then stores entities, relations and passages as three Milvus collections, each row carrying an embedding plus ID lists that point at its neighbours. “Traversal” is a series of Milvus query calls by ID. There is no Cypher, no graph engine and no community layer.
The retrieval design is borrowed from HippoRAG and says so in the code: the phrase normaliser is “same as HippoRAG’s processing_phrases”, the query NER prompt is HippoRAG-style, and candidate relations are sorted by ID “to match HippoRAG’s behavior”. The project’s own idea is to replace iterative multi-hop agent loops with one expansion step and one LLM reranking call that picks the relations worth reading.
It targets developers who already run Milvus (or are happy with Milvus Lite, the default local file) and want multi-hop QA over a modest corpus. It ships as a pip package with an optional FastAPI server and a React graph viewer that animates each retrieval step.
Architecture
flowchart LR
D["Documents / chunks"] --> TE["TripletExtractor (LLM)"]
TE --> GB["GraphBuilder (IDs + adjacency)"]
GB --> EM["EmbeddingModel"]
EM --> MS["MilvusStore: entities / relations / passages"]
Q["Question"] --> EX["EntityExtractor (LLM NER)"]
EX --> GR["GraphRetriever"]
GR --> MS
GR --> SG["SubGraph.expand"]
SG --> MS
GR --> RR["LLMReranker or JevReranker"]
RR --> AG["AnswerGenerator (LLM)"]
API["FastAPI app + React UI"] --> RAG["VectorGraphRAG facade"]
RAG --> TE
RAG --> GR
| Component | Path | Role |
|---|---|---|
| Facade | src/vector_graph_rag/rag.py |
VectorGraphRAG: rebuild, source-level upsert/delete, query, retrieve |
| Triplet / NER extraction | src/vector_graph_rag/llm/extractor.py |
One JSON-mode LLM call per passage; question NER at query time |
| Graph builder | src/vector_graph_rag/graph/builder.py |
Assigns IDs, dedupes entities by normalised name, builds the six adjacency maps |
| Storage | src/vector_graph_rag/storage/milvus.py |
Three collections, VARCHAR id + vector + text + dynamic fields, IP metric |
| Retriever | src/vector_graph_rag/graph/retriever.py |
Entity and relation vector search, expansion, eviction, metadata filters |
| Subgraph | src/vector_graph_rag/graph/knowledge_graph.py |
Lazy relation-entity-relation expansion with a recorded history |
| Reranking / answer | src/vector_graph_rag/llm/reranker.py, llm/jev.py |
Few-shot “pick 5 relations” call, or per-relation Jev scoring; final answer prompt |
| Embeddings | src/vector_graph_rag/storage/embedding_providers/ |
OpenAI, HuggingFace, Google, Voyage, Jina, Mistral, Ollama, local, ONNX |
| Server and loaders | src/vector_graph_rag/api/app.py, loaders/ |
REST endpoints, file/URL import, chunking, SPA hosting |
How a request flows
Indexing with rebuild_documents(docs):
- Every document without an id gets a UUID, then
TripletExtractor.extract_from_documentscalls the LLM once per document with a one-shot prompt injson_objectmode and stores[[s, p, o], ...]inmetadata["triplets"](extractor.py). GraphBuilderturns triplets into IDs. Entities are deduplicated by exact match on a normalised name (lower-cased, non-alphanumerics replaced by spaces); relations by their"s p o"text (builder.py, extractor.py).- Entity, relation and passage texts are embedded in batches. Only then are the three collections dropped and recreated, and the rows inserted with their ID lists as metadata (rag.py). Doing the expensive work before the drop means a failed extraction does not wipe the old graph.
Querying with query(question) (rag.py):
GraphRetriever.retrieveresolves a metadatafilterinto allowed passage IDs, then asks the LLM for the named entities in the question (retriever.py).- Each question entity is embedded and searched against the entity collection; hits must score above
entity_similarity_threshold(0.9) (retriever.py). Then, in a separate step, the whole question is searched against the relation collection, keeping the top 20 with no threshold by default. SubGraph.expandmerges the seed entities’ relations with the seed relations, then for each degree walks relations to entities to relations, fetching rows from Milvus by ID (knowledge_graph.py). Default degree is 1.- If the expanded relation set exceeds
relation_number_threshold(1000), a filtered vector search keeps the 1000 closest to the question; otherwise relations are sorted by ID (retriever.py). LLMRerankersends all candidates as[id] textlines with three few-shot multi-hop examples and asks for exactly five useful relations (reranker.py). There is no fallback if the model returns nothing usable.- The selected relations’
passage_idsare resolved to passages, the firstfinal_top_k(3) go toAnswerGenerator, and aQueryResultreturns the answer plus the subgraph, eviction stats and rerank result for visualisation.
Key components
Storage model
Each collection has a VARCHAR(64) primary key, a float vector indexed with AUTOINDEX and the inner-product metric, a text field, and Milvus dynamic fields for everything else (milvus.py). Entities carry relation_ids and passage_ids; relations carry entity_ids, passage_ids and subject/predicate/object; passages carry entity_ids, relation_ids and user metadata. collection_prefix gives you several independent graphs in one Milvus database, which is how the server’s graph_name works.
Incremental updates
Full rebuilds are not the only option. upsert_documents_by_source builds a fresh graph for one source’s chunks, deletes that source’s old passages, then merges by text: existing entities and relations are upserted with merged ID lists, new ones inserted (rag.py, L533-L733). delete_documents_by_source cascades: relations and entities lose the deleted passage IDs and are removed when nothing references them (rag.py). Both are documented as not atomic but convergent on retry. add_documents and add_texts are deprecated aliases of the full rebuild.
Configuration
Settings is a pydantic-settings class with the VGRAG_ prefix (config.py). Defaults: gpt-4o-mini for every LLM step, text-embedding-3-large at 3072 dimensions, Milvus Lite at ./vector_graph_rag.db, and a disk LLM cache under ./llm_cache that is on by default. An OpenAI-compatible key is required even with a non-OpenAI embedding provider, because extraction, NER, reranking and answering all use the openai client.
Extending it
- Bring your own triplets.
rebuild_documents_with_tripletsandmetadata["triplets"]skip extraction, so a better extractor or a curated KG can be plugged in. - Embedding providers. A small registry maps names to provider classes behind a three-member protocol (
model_name,dimension,encode). - Reranker.
reranker_provider="jev"swaps the list-selection call for a per-relation scoring service with a threshold. - Any OpenAI-compatible LLM via
openai_base_url. There is no other LLM abstraction. - Low-level CRUD.
graph/graph.pyexposes passage-level create/update/delete with cascading entity and relation maintenance.
Running it
pip install vector-graph-rag (extras: api, loaders, hf, google, voyage, jev), set OPENAI_API_KEY, then VectorGraphRAG().rebuild_texts([...]) and .query(...). Milvus Lite needs no server; point milvus_uri at a Milvus or Zilliz Cloud endpoint for anything shared. The Dockerfile builds the React UI and runs uvicorn vector_graph_rag.api.app:app. Spans are emitted through OpenTelemetry when it is installed.
Strengths and caveats
- Strength: one store. Vectors and adjacency live in the same Milvus rows, so there is no second database to keep in sync, and filters on passage metadata apply to graph retrieval too.
- Strength: cheap, bounded queries. Two LLM calls for retrieval (NER and rerank) plus one for the answer, with a hard cap on candidates.
- Strength: inspectable.
QueryResultcarries the subgraph and its expansion history, which the UI replays. - Caveat: the REST API always rebuilds.
/add_documents,/importand/uploadall callrebuild_*, so each import replaces the named graph. Incremental updates are a Python-only feature. - Caveat: shallow entity resolution. Entities merge only on exact normalised text; “IBM” and “International Business Machines” stay separate, and non-Latin names are mostly stripped by the
[^A-Za-z0-9 ]normaliser. - Caveat: traversal by ID filters. Each hop is an
id in [...]query. That is fine at degree 1 or 2 and at HotpotQA scale; it is not a graph engine, and there are no path queries, community summaries or global questions. - Caveat: fixed shape. The reranker always asks for five relations and answers from three passages; tuning means changing prompts and settings, not plugging in strategies.
Sources: code at 07acd79, deepwiki-open wiki (10 pages), verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit 07acd79. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredDocuments are processed through a three-stage pipeline: extraction, deduplication, and embedding.
Chunking is handled by TextChunker (loaders/chunker.py:19-98) which splits documents using semantic separators (paragraphs, sentences) at a default chunk size of 1000 characters with 200-character overlap.
Triplet extraction uses TripletExtractor (llm/extractor.py:88-250), which calls an OpenAI GPT model (default gpt-4o-mini) with a system prompt and one-shot example asking for JSON output of [subject, predicate, object] arrays. The prompt uses response_format: {type: json_object} (llm/extractor.py:155). Each document chunk gets one LLM call; triplets are stored in Document.metadata["triplets"]. Users can supply pre-extracted triplets to skip the LLM call.
Deduplication is done by GraphBuilder (graph/builder.py:25-215). Entity names are normalized via processing_phrases() (llm/extractor.py:22-33), which replaces non-alphanumeric characters with spaces and lowercases. A entity_name_to_id dict maps normalized names to UUIDs so repeated entities across chunks resolve to the same node. Relations are deduplicated by their full normalized text (relation_text_to_id).
Building the graph happens in GraphBuilder._process_documents() (graph/builder.py:136-157): it iterates triplets, calls _add_relation() which calls _add_entity() for subject and object, building six adjacency maps (entity-to-relation, entity-to-passage, relation-to-entity, relation-to-passage, passage-to-entity, passage-to-relation).
No formal schema or ontology is enforced beyond the subject-predicate-object structure. There is no community detection, no hierarchical summary, and no graph database (Neo4j etc).
Where and how is the graph stored?
answeredThe graph is stored in Milvus, managed by MilvusStore (storage/milvus.py:35-1478). Three collections are created per named graph prefix: {prefix}_vgrag_entities, {prefix}_vgrag_relations, and {prefix}_vgrag_passages. There is no graph database.
Each collection has the same schema: a VARCHAR primary key id (string, UUID or user-provided), a FLOAT_VECTOR field named vector storing embeddings, and a VARCHAR field named text. Collections use the IP (inner product) metric type with configurable index type ("AUTOINDEX" default) (storage/milvus.py:209-225). Additional metadata fields are stored via Milvus's dynamic schema.
Entity records store relation_ids and passage_ids in metadata. Relation records store entity_ids (head and tail), passage_ids, plus structured triplet fields subject, predicate, object (storage/milvus.py:394-438). Passage records store entity_ids, relation_ids, and arbitrary user metadata (e.g. filterable fields like source). The graph adjacency is thus encoded entirely as ID lists in the metadata of each node — there are no explicit edge objects beyond what is captured in the relation collection.
Embeddings sit inside each record as the vector field, generated by EmbeddingModel (storage/embeddings.py) which supports multiple providers (OpenAI, HuggingFace, Voyage, Jina, Ollama, Mistral, Google). Entity, relation, and passage embeddings are all stored in the same vector index, albeit in separate collections. The default embedding model is text-embedding-3-large (3072 dimensions) in Settings, while the factory function defaults to text-embedding-3-small (as noted in CLAUDE.md).
Are communities, summaries or hierarchies built over the graph?
answeredCommunities, hierarchical summaries, and community detection (e.g. Leiden) are NOT implemented.
The project intentionally takes a different approach from Microsoft's GraphRAG. Rather than detecting communities and computing hierarchical summaries over them, Vector Graph RAG stores the graph's adjacency as ID lists in metadata and performs subgraph expansion at query time via SubGraph.expand() (graph/knowledge_graph.py:261-361).
There is no offline graph partitioning, no community detection algorithm (Leiden or otherwise), no summary-of-summaries pyramid, and no global search mode. The system has a single retrieval path: entity extraction from the question, vector search for similar entities and relations, subgraph expansion from seed nodes, optional LLM reranking, and passage retrieval.
The SubGraph class (graph/knowledge_graph.py:152-591) implements lazy expansion: it holds entity and relation IDs and fetches neighbor records from Milvus on demand during expand(degree=...). Each expansion step goes from current relations to their connected entities, then from those entities to their connected relations — this is graph traversal at query time, not community detection at index time.
A naive RAG baseline (VectorGraphRAG.query_naive()) is available for comparison, which skips the graph entirely and does direct passage vector search.
How does query-time retrieval use the graph?
answeredQuery-time retrieval is a multi-step pipeline orchestrated by GraphRetriever.retrieve() (graph/retriever.py:388-487):
Entity extraction:
EntityExtractor(llm/extractor.py:252-403) uses an OpenAI LLM call to extract named entities from the question, with a HippoRAG-compatible NER cache fallback from TSV files.Bi-directional vector search: Extracted entities are embedded and searched against the entity collection (
_search_entities), and the raw question is embedded and searched against the relation collection (_search_relations). Both apply configurable similarity thresholds (entity default 0.9, relation default -1.0 i.e. keep all) (retriever.py:114-222).Subgraph expansion: Seed entities and relations are loaded into a
SubGraph(graph/knowledge_graph.py:152-591) which lazily fetches neighbor relations and entities from Milvus. Default expansion degree is 1. The expansion walks from seed entities to their relations, merges with seed relations, then for each degree expands relations-to-entities-to-relations.Eviction strategy: If expanded relations exceed
relation_number_threshold(default 1000), a vector search re-ranks them down to the threshold. Otherwise, relations are sorted by ID for deterministic behavior (HippoRAG compat) (retriever.py:264-336).LLM reranking (optional):
LLMReranker(llm/reranker.py:96-294) takes the candidate relations and uses GPT with three few-shot multi-hop examples to select the most useful ones via chain-of-thought JSON output. This is the single-pass replacement for iterative agent loops.Passage retrieval: Selected relation IDs are used to fetch associated passages, calling
MilvusStore.get_passages_by_ids()after resolving passage IDs from relation metadata (rag.py:735-782). If fewer thanfinal_top_kpassages are found, graph passages are supplemented with naive vector search results (rag.py:1465-1475).Answer generation:
AnswerGenerator(llm/reranker.py:296-392) feeds the final passages as context to the LLM.
There are no separate local/global modes — only graph retrieval, with naive RAG available as a comparison baseline.
How are updates and incremental indexing handled?
answeredFull rebuild is the default. rebuild_documents() (rag.py:942-1012) drops all three Milvus collections and recreates them, then indexes everything fresh. The legacy add_documents() (line 888-940) calls rebuild_documents() internally.
Source-level incremental upsert is implemented via upsert_documents_by_source() (rag.py:1014-1122). It operates at the granularity of a "source" (a file, URL, or business record identified by a source metadata field). The flow: (1) extracts triplets from new documents, (2) builds graph records, (3) calls delete_documents_by_source() to remove all passages for that source, (4) calls _insert_incremental_graph() which checks for existing entities and relations by normalized text (storage/milvus.py:646-739). When an entity or relation already exists, its metadata is merged (passage/relation IDs are appended using _merge_unique), and upsert writes updated records. When they don't exist, they are inserted as new records (rag.py:596-731).
Deletion via delete_documents_by_source() (rag.py:1137-1250) cascades: it finds passages matching the source, then updates or deletes relations (removing the passage ID from the relation's adjacency list; deleting the relation entirely if it has no remaining passages), then similarly updates or deletes orphaned entities, and finally deletes the passages themselves.
LLM response caching (llm/cache.py:1-165) persists extraction results to disk, avoiding re-extraction when rebuilding or re-indexing. The NER (named entity recognition) cache also supports a TSV file format for HippoRAG evaluation compatibility (llm/extractor.py:296-332).
The cached result _extraction_result and the lazy retriever are reset after any mutation via self._retriever = None (rag.py:1010, 1121, 1249).
How are LLM cost and latency controlled during indexing and query?
answeredCost and latency are managed through several mechanisms:
LLM Response Caching (llm/cache.py:14-165): Every LLM call (triplet extraction, entity extraction, reranking, answer generation) is cached to disk. The cache uses MD5 hashing of the prompt with temperature as a cache key, so identical prompts hit the cache. Extraction and NER each have their own caching paths (extractor.py:147-168 and extractor.py:369-395).
Model choice per stage: The default llm_model is gpt-4o-mini (config.py:28-31), a small, cheap model. The reranker and extractor reuse the same model unless overridden — there is no separate model parameter for different pipeline stages. The AnswerGenerator uses the same model. Temperature is 0 for deterministic outputs (config.py:118).
Batching: Embedding generation uses batch calls with a configurable batch_size (default 32, config.py:133). Database inserts use the same batching. The eviction strategy (retriever.py:264-336) limits the reranker input to at most relation_number_threshold (default 1000) relations, preventing unbounded prompt sizes.
Optional Jev reranker: Instead of the LLM-based LLMReranker, the system supports JevReranker (llm/jev.py:27-33), a specialized high-throughput scoring model that may be cheaper per-relation than an LLM call. Configured via reranker_provider: "jev" (config.py:123-128).
Single-pass design: The project's key algorithmic bet is replacing iterative multi-step LLM agent loops (as in IRCoT or GraphRAG iterative reflection) with a single-pass LLM reranking step. This caps the number of LLM calls per query to: NER (1), reranking (1), answer generation (1) — versus potentially dozens in iterative approaches. The CLAUDE.md explicitly states this philosophy.
NER Cache TSV: Query-time entity extraction supports a pre-computed NER cache in HippoRAG TSV format (extractor.py:296-332), eliminating LLM calls for known questions in evaluation settings.