Graph RAG: graphify vs LightRAG vs graphrag vs semantica vs llm-graph-builder vs HippoRAG vs nano-graphrag vs autoflow vs trustgraph vs vector-graph-rag
Retrieval-augmented generation over knowledge graphs — extracting entities and relations from documents and querying the graph alongside vectors. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.
At a glance
● answered from code · — not applicable (the project does not do this) · ? insufficient evidence
How is the knowledge graph extracted from documents?
Most of these tools have an LLM read each chunk and then merge entities only when their normalized names match exactly. The two exceptions are AutoFlow, which compares description embeddings and asks an LLM before it merges, and graphify, which merges near-identical names with fuzzy matching.
GraphRAG-style delimited prompts. GraphRAG merges entities by (title, type), so one name with two types becomes two nodes. It has four default entity types, and its fast method swaps the LLM for noun-phrase extraction. nano-graphrag uses the same prompt family and upper-cases names. Missing relationship endpoints become UNKNOWN nodes. LightRAG merges by name, picks the type by majority vote, runs gleaning at most once, and calls the LLM to summarize descriptions only after 8 fragments pile up.
Triples and OpenIE. HippoRAG runs NER on every chunk first, then triple extraction based on that output. By default it does not chunk: one document becomes one passage. It links aliases with synonymy edges (similarity of 0.8 or more) and does not merge them. Vector Graph RAG makes one JSON-mode call per 1,000-character chunk. Its name normalizer drops every character outside [A-Za-z0-9 ]. TrustGraph runs several extractors on each chunk, including ontology-guided OntoRAG, which checks domain and range. Its IRIs are built from the name, so the same name always maps to the same node.
Schema-driven and framework pipelines. LLM Graph Builder wraps LangChain's LLMGraphTransformer and accepts optional allowed node and relationship lists. Fuzzy merging happens only when a user approves it in the UI. Unless the user's email ends in @neo4j.com, chunks beyond MAX_TOKEN_CHUNK_SIZE are silently dropped. AutoFlow makes two DSPy calls per chunk. Semantica defaults to spaCy NER plus pattern triplets. Relation extraction is off unless extract_relations=True, and its conflict step only counts conflicts and changes nothing.
Code first. graphify parses code with tree-sitter and no LLM. It uses an LLM only for documents, PDFs and images, and merges names with MinHash plus Jaro-Winkler (threshold 92).
Pick: GraphRAG or LightRAG for general prose with tunable prompts. Pick: TrustGraph or LLM Graph Builder when a schema or ontology must constrain extraction. Pick: graphify for codebases, and AutoFlow when duplicate entities are the main worry.
Where and how is the graph stored?
For a production database that holds both the graph and its vectors, LLM Graph Builder (Neo4j) and AutoFlow (TiDB) are the most integrated. For a portable file you can open and version, use GraphRAG's tables or graphify's graph.json.
Files and in-memory graphs. GraphRAG has no graph store. It writes entities, relationships and communities as Parquet or CSV tables and puts the embeddings in LanceDB. nano-graphrag keeps a NetworkX graph in a GraphML file, with JSON key-value files and nano-vectordb beside it. Neo4j is an option. graphify serialises NetworkX to node-link JSON and stores no embeddings at all. Its Neo4j and FalkorDB pushes are one-way exports. HippoRAG pickles an igraph graph of entity and passage nodes, written atomically under a file lock. Facts are not graph nodes. They live, with chunks and entities, in three Parquet embedding stores.
Pluggable backends. LightRAG has four storage roles. The default graph is a single-writer GraphML file, and Neo4j, Memgraph, PostgreSQL, MongoDB or OpenSearch can replace it. Entities and relations are also stored as vectors. In Semantica the graph is a plain Python dict unless you configure a GraphStore (Neo4j, FalkorDB, Apache AGE or Neptune). Vectors go to a separate store. TrustGraph stores RDF triples in Cassandra, Neo4j, Memgraph or FalkorDB, and entity vectors in Qdrant, Milvus or Pinecone. Its Neo4j writer drops the named-graph field, so the provenance features need Cassandra.
One store for graph and vectors. LLM Graph Builder keeps Document, Chunk, entity and community nodes in Neo4j, each with an embedding property and a vector index. AutoFlow creates entity and relationship tables per knowledge base, with HNSW vector columns in the same rows. TiDB is its only backend. Vector Graph RAG has no graph database at all. It uses three Milvus collections and stores adjacency as ID lists in dynamic fields. Each hop is an id in [...] query.
Pick: GraphRAG or HippoRAG for a batch artifact you rebuild and inspect offline. Pick: LightRAG, TrustGraph or LLM Graph Builder for a live graph in a real database. Pick: Vector Graph RAG or AutoFlow if you already run Milvus or TiDB and want no second system.
Are communities, summaries or hierarchies built over the graph?
GraphRAG and its compact clone nano-graphrag are the reference here: they build hierarchical Leiden communities and an LLM report for each one at index time. Semantica and LLM Graph Builder can also build summarized hierarchies, but you have to trigger them, and in LLM Graph Builder the summary step is broken at the reviewed commit.
Hierarchy plus reports at index time. GraphRAG runs hierarchical Leiden (max_cluster_size 10) on the largest connected component and writes reports deepest level first. Its fallback for oversized communities never receives sub-community reports, so those communities are trimmed instead. nano-graphrag uses the same graspologic call, also on the largest component only. It drops every report and regenerates all of them on each insert.
Hierarchies you trigger. LLM Graph Builder runs GDS Leiden with up to 3 levels and summarizes parent communities from their children. This happens only when the /post_processing job runs with enable_communities. That route passes the embedding provider name in the position of the summary-model argument, so summaries fail unless an LLM config exists under that name. Semantica's CommunityHierarchyBuilder runs real Louvain or Leiden, and CommunitySummarizer writes reports bottom-up, falling back to an extractive report when no LLM is set. Its simpler CommunityDetector labels greedy modularity as both "Louvain" and "Leiden".
Flat clusters without summaries. graphify runs Leiden, falling back to Louvain. It splits any community larger than 25% of the graph, and re-splits communities of 50 or more nodes whose cohesion is below 0.05. The LLM or the host agent writes only short names for them.
None. LightRAG has no clustering step, and its "global" mode searches relationship vectors. HippoRAG uses Personalized PageRank instead. TrustGraph does not apply here: it relies on reranked two-hop traversal. AutoFlow has no detection step. Admins can create "synopsis" entities by hand. Vector Graph RAG expands a subgraph at query time only.
Pick: GraphRAG for corpus-wide thematic questions, if you can afford the reports. Pick: nano-graphrag to study or fork the same design in about 1,100 lines. Pick: LightRAG or HippoRAG when questions are about specific entities and the corpus changes often.
How does query-time retrieval use the graph?
For broad questions about a whole corpus, GraphRAG's global and DRIFT search are the most complete. For multi-hop questions over passages, HippoRAG's Personalized PageRank and Vector Graph RAG's single rerank call are the cheapest paths.
Community reports plus local search. GraphRAG has local (12,000-token context in fixed shares), global map-reduce, DRIFT and basic engines. DRIFT picks a random community report as its template, so results vary between runs. nano-graphrag builds local context from the top 20 entities as four CSV tables, each with its own token budget. Its default global mode maps over 16,384-token groups of reports, and still makes those map calls with only_need_context. Semantica offers local, global, DRIFT and hybrid modes. In local mode its vector, graph and memory retrievers run one after another. LLM Graph Builder defaults to graph_vector_fulltext: chunk hits, then one or two hops from each chunk's entities depending on their similarity to the query. Other modes search community summaries or generate Cypher.
Vector-seeded neighbourhoods. LightRAG makes one LLM call to extract keywords. Low-level keywords search entity vectors and high-level keywords search relation vectors. The default mix mode adds chunk hits, with top_k 40. AutoFlow splits the question into sub-questions by default. For each one it walks relationships by weight, two levels deep. It then retrieves chunks with a rewritten question. TrustGraph grounds the question's concepts in entity vectors and walks two hops. A cross-encoder keeps 25 edges per hop, so up to 50 edges reach synthesis. It has no global mode.
Graph as a passage ranker. HippoRAG scores facts, lets an LLM filter the top 5, then runs PPR with damping 0.5. If no fact survives the filter, it quietly falls back to dense retrieval. Vector Graph RAG extracts entities from the question, keeps entity hits above 0.9 similarity, expands one degree, asks the LLM for five relations, and answers from three passages.
Lexical traversal. graphify scores node labels by token match behind a trigram prefilter, then runs BFS or DFS to depth 3. It returns 2,000 tokens of text to the calling agent and makes no LLM or vector call.
Pick: GraphRAG or Semantica for synthesis across the corpus.
Pick: LightRAG mix or TrustGraph for fast answers grounded in entities.
Pick: HippoRAG or Vector Graph RAG for multi-hop QA over passages.
How are updates and incremental indexing handled?
LightRAG and HippoRAG handle change best. Both add documents without a rebuild and delete them with reference-counted cleanup of shared entities. GraphRAG and nano-graphrag are in practice append-only.
Add and delete per document. LightRAG merges each document into the live graph. adelete_by_doc_id rebuilds shared entities from cached extraction results, and deletion is journaled so a crash can resume. HippoRAG sends only unseen chunk hashes to OpenIE and runs synonymy KNN only for new entities. Its manifest refuses to mix state from two different models. Vector Graph RAG has upsert_documents_by_source and a cascading delete_documents_by_source, but only in Python: its REST import endpoints always rebuild the whole graph. LLM Graph Builder processes each document on its own, can resume from the last processed chunk, and deletes only the entities no other document references. It has no extraction cache, so a retry pays again. AutoFlow deletes a document's relationships and then any orphaned entities, but merged descriptions stay behind. Re-indexing skips chunks that already have relationships.
Append-only. GraphRAG's update finds new documents by title. It computes deleted documents but never uses that list, and it appends delta communities without re-clustering. nano-graphrag skips known document and chunk hashes, but every insert drops and regenerates all community reports. It has no delete API. Semantica merges new entities with incremental_merge and caches extraction with a TTL. It has no document-level delete, only entity-level purge and erasure.
Driven by file changes. graphify caches both AST and LLM results by file-content hash, and the LLM cache also includes a prompt fingerprint. graphify update re-parses changed code with no LLM and removes nodes for deleted files. Communities are recomputed on every build.
Collection-level only. TrustGraph has no incremental path: there is no extraction cache, graph data is deleted per collection, and deleting a document in the librarian leaves its triples in place.
Pick: LightRAG or HippoRAG for a corpus that changes daily. Pick: graphify for repositories that change on every commit. Pick: GraphRAG or nano-graphrag only when you can rebuild in batches.
How are LLM cost and latency controlled during indexing and query?
The cheapest options avoid the LLM where they can: graphify parses code with no model, and Semantica extracts with spaCy by default. Among full LLM pipelines, LightRAG, HippoRAG and Vector Graph RAG keep each query to a few calls. GraphRAG spends most of its budget at index time.
Heavy indexing. GraphRAG runs extraction plus gleaning on every chunk, a summarization call for merged descriptions, and one report per community. Each stage has its own model, and its fast method replaces LLM extraction with NLP. Only indexing calls are cached. nano-graphrag splits work between a "best" model and a "cheap" model and caches every call, including queries, but each insert regenerates all community reports. LLM Graph Builder makes one call per chunks_to_combine chunks and has no extraction cache. Optional per-user daily and monthly token limits can stop a job before it starts. AutoFlow makes two DSPy calls per chunk plus LLM merge calls for entities, and uses a separate fast_llm to rewrite questions.
Lean indexing, cheap queries. LightRAG writes no reports, caches both extraction and answers, and routes extraction, keyword and answer calls to different models. Its answer cache ignores the retrieved context, so a cached answer can survive changes to the data. HippoRAG caches every request in SQLite. A query costs one filter call plus a PageRank solve. Vector Graph RAG caches every LLM call on disk, uses gpt-4o-mini for every step with no per-stage choice, and makes three calls per query. TrustGraph makes two LLM calls per query and filters edges with a MiniLM cross-encoder. It does not cache extraction.
Skipping the LLM. graphify packs files for its LLM pass into chunks of up to 60,000 tokens. When a response is truncated, it splits the chunk in half and retries, up to three levels deep. It caches results per file hash. Semantica's LLM stages are opt-in, it caches extraction with a TTL of 3,600 seconds, and it can write extractive community reports with no model.
Pick: graphify or Semantica's defaults when the budget is close to zero. Pick: LightRAG or HippoRAG for low running cost with LLM-quality extraction. Pick: GraphRAG when global answers justify paying for indexing up front.
The projects
- Graphify-Labs/graphify: Coding-assistant skill and CLI that turns a repo into a tree-sitter code graph plus LLM-extracted doc nodes, queried over MCP by traversal.
- HKUDS/LightRAG: Python graph RAG engine that merges LLM-extracted entities into one graph and retrieves by keyword-matched entities, relations and chunks.
- microsoft/graphrag: Python pipeline that turns text into an LLM-extracted entity graph with Leiden community reports, queried by local, global and DRIFT search.
- semantica-agi/semantica: Large Python knowledge-graph toolkit with spaCy/LLM extraction, pluggable graph stores, and GraphRAG-style community and global search.
- neo4j-labs/llm-graph-builder: FastAPI and React app that turns files, URLs and transcripts into a Neo4j knowledge graph via LLMGraphTransformer, plus GraphRAG chat.
- OSU-NLP-Group/HippoRAG: Research RAG library that links OpenIE triples, entities and passages in one igraph and ranks passages with Personalized PageRank.
- gusye1234/nano-graphrag: A small, readable Python reimplementation of Microsoft GraphRAG with local, global and naive query modes over a NetworkX graph.
- pingcap/autoflow: Self-hosted Graph RAG chat app on TiDB that extracts a DSPy knowledge graph per chunk and fuses it with vector search.
- trustgraph-ai/trustgraph: Microservice Graph RAG platform that extracts RDF triples onto a message bus and answers by cross-encoder-filtered graph hops.
- zilliztech/vector-graph-rag: Python Graph RAG library that stores entities, relations and passages as three Milvus collections and walks the graph by ID lookups.