# Graph RAG: comparison

> Retrieval-augmented generation over knowledge graphs — extracting entities and relations from documents and querying the graph alongside vectors.

Canonical page: https://llms-technical-reviews.com/compare/graph-rag/

## How is the knowledge graph extracted from documents?

Most of these tools have an LLM read each chunk and then merge entities only when their normalized names match exactly. The two exceptions are [AutoFlow](/p/autoflow/), which compares description embeddings and asks an LLM before it merges, and [graphify](/p/graphify/), which merges near-identical names with fuzzy matching.

**GraphRAG-style delimited prompts.** [GraphRAG](/p/graphrag/) merges entities by `(title, type)`, so one name with two types becomes two nodes. It has four default entity types, and its `fast` method swaps the LLM for noun-phrase extraction. [nano-graphrag](/p/nano-graphrag/) uses the same prompt family and upper-cases names. Missing relationship endpoints become `UNKNOWN` nodes. [LightRAG](/p/lightrag/) merges by name, picks the type by majority vote, runs gleaning at most once, and calls the LLM to summarize descriptions only after 8 fragments pile up.

**Triples and OpenIE.** [HippoRAG](/p/hipporag/) runs NER on every chunk first, then triple extraction based on that output. By default it does not chunk: one document becomes one passage. It links aliases with synonymy edges (similarity of 0.8 or more) and does not merge them. [Vector Graph RAG](/p/vector-graph-rag/) makes one JSON-mode call per 1,000-character chunk. Its name normalizer drops every character outside `[A-Za-z0-9 ]`. [TrustGraph](/p/trustgraph/) runs several extractors on each chunk, including ontology-guided OntoRAG, which checks domain and range. Its IRIs are built from the name, so the same name always maps to the same node.

**Schema-driven and framework pipelines.** [LLM Graph Builder](/p/llm-graph-builder/) wraps LangChain's `LLMGraphTransformer` and accepts optional allowed node and relationship lists. Fuzzy merging happens only when a user approves it in the UI. Unless the user's email ends in `@neo4j.com`, chunks beyond `MAX_TOKEN_CHUNK_SIZE` are silently dropped. AutoFlow makes two DSPy calls per chunk. [Semantica](/p/semantica/) defaults to spaCy NER plus pattern triplets. Relation extraction is off unless `extract_relations=True`, and its conflict step only counts conflicts and changes nothing.

**Code first.** graphify parses code with tree-sitter and no LLM. It uses an LLM only for documents, PDFs and images, and merges names with MinHash plus Jaro-Winkler (threshold 92).

Pick: GraphRAG or LightRAG for general prose with tunable prompts.
Pick: TrustGraph or LLM Graph Builder when a schema or ontology must constrain extraction.
Pick: graphify for codebases, and AutoFlow when duplicate entities are the main worry.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/graph-construction/index.md

## Where and how is the graph stored?

For a production database that holds both the graph and its vectors, [LLM Graph Builder](/p/llm-graph-builder/) (Neo4j) and [AutoFlow](/p/autoflow/) (TiDB) are the most integrated. For a portable file you can open and version, use [GraphRAG](/p/graphrag/)'s tables or [graphify](/p/graphify/)'s `graph.json`.

**Files and in-memory graphs.** GraphRAG has no graph store. It writes entities, relationships and communities as Parquet or CSV tables and puts the embeddings in LanceDB. [nano-graphrag](/p/nano-graphrag/) keeps a NetworkX graph in a GraphML file, with JSON key-value files and nano-vectordb beside it. Neo4j is an option. graphify serialises NetworkX to node-link JSON and stores no embeddings at all. Its Neo4j and FalkorDB pushes are one-way exports. [HippoRAG](/p/hipporag/) pickles an igraph graph of entity and passage nodes, written atomically under a file lock. Facts are not graph nodes. They live, with chunks and entities, in three Parquet embedding stores.

**Pluggable backends.** [LightRAG](/p/lightrag/) has four storage roles. The default graph is a single-writer GraphML file, and Neo4j, Memgraph, PostgreSQL, MongoDB or OpenSearch can replace it. Entities and relations are also stored as vectors. In [Semantica](/p/semantica/) the graph is a plain Python dict unless you configure a `GraphStore` (Neo4j, FalkorDB, Apache AGE or Neptune). Vectors go to a separate store. [TrustGraph](/p/trustgraph/) stores RDF triples in Cassandra, Neo4j, Memgraph or FalkorDB, and entity vectors in Qdrant, Milvus or Pinecone. Its Neo4j writer drops the named-graph field, so the provenance features need Cassandra.

**One store for graph and vectors.** LLM Graph Builder keeps Document, Chunk, entity and community nodes in Neo4j, each with an embedding property and a vector index. AutoFlow creates entity and relationship tables per knowledge base, with HNSW vector columns in the same rows. TiDB is its only backend. [Vector Graph RAG](/p/vector-graph-rag/) has no graph database at all. It uses three Milvus collections and stores adjacency as ID lists in dynamic fields. Each hop is an `id in [...]` query.

Pick: GraphRAG or HippoRAG for a batch artifact you rebuild and inspect offline.
Pick: LightRAG, TrustGraph or LLM Graph Builder for a live graph in a real database.
Pick: Vector Graph RAG or AutoFlow if you already run Milvus or TiDB and want no second system.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/graph-storage/index.md

## Are communities, summaries or hierarchies built over the graph?

[GraphRAG](/p/graphrag/) and its compact clone [nano-graphrag](/p/nano-graphrag/) are the reference here: they build hierarchical Leiden communities and an LLM report for each one at index time. [Semantica](/p/semantica/) and [LLM Graph Builder](/p/llm-graph-builder/) can also build summarized hierarchies, but you have to trigger them, and in LLM Graph Builder the summary step is broken at the reviewed commit.

**Hierarchy plus reports at index time.** GraphRAG runs hierarchical Leiden (`max_cluster_size` 10) on the largest connected component and writes reports deepest level first. Its fallback for oversized communities never receives sub-community reports, so those communities are trimmed instead. nano-graphrag uses the same graspologic call, also on the largest component only. It drops every report and regenerates all of them on each insert.

**Hierarchies you trigger.** LLM Graph Builder runs GDS Leiden with up to 3 levels and summarizes parent communities from their children. This happens only when the `/post_processing` job runs with `enable_communities`. That route passes the embedding provider name in the position of the summary-model argument, so summaries fail unless an LLM config exists under that name. Semantica's `CommunityHierarchyBuilder` runs real Louvain or Leiden, and `CommunitySummarizer` writes reports bottom-up, falling back to an extractive report when no LLM is set. Its simpler `CommunityDetector` labels greedy modularity as both "Louvain" and "Leiden".

**Flat clusters without summaries.** [graphify](/p/graphify/) runs Leiden, falling back to Louvain. It splits any community larger than 25% of the graph, and re-splits communities of 50 or more nodes whose cohesion is below 0.05. The LLM or the host agent writes only short names for them.

**None.** [LightRAG](/p/lightrag/) has no clustering step, and its "global" mode searches relationship vectors. [HippoRAG](/p/hipporag/) uses Personalized PageRank instead. [TrustGraph](/p/trustgraph/) does not apply here: it relies on reranked two-hop traversal. [AutoFlow](/p/autoflow/) has no detection step. Admins can create "synopsis" entities by hand. [Vector Graph RAG](/p/vector-graph-rag/) expands a subgraph at query time only.

Pick: GraphRAG for corpus-wide thematic questions, if you can afford the reports.
Pick: nano-graphrag to study or fork the same design in about 1,100 lines.
Pick: LightRAG or HippoRAG when questions are about specific entities and the corpus changes often.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/communities/index.md

## How does query-time retrieval use the graph?

For broad questions about a whole corpus, [GraphRAG](/p/graphrag/)'s global and DRIFT search are the most complete. For multi-hop questions over passages, [HippoRAG](/p/hipporag/)'s Personalized PageRank and [Vector Graph RAG](/p/vector-graph-rag/)'s single rerank call are the cheapest paths.

**Community reports plus local search.** GraphRAG has local (12,000-token context in fixed shares), global map-reduce, DRIFT and basic engines. DRIFT picks a random community report as its template, so results vary between runs. [nano-graphrag](/p/nano-graphrag/) builds local context from the top 20 entities as four CSV tables, each with its own token budget. Its default global mode maps over 16,384-token groups of reports, and still makes those map calls with `only_need_context`. [Semantica](/p/semantica/) offers local, global, DRIFT and hybrid modes. In local mode its vector, graph and memory retrievers run one after another. [LLM Graph Builder](/p/llm-graph-builder/) defaults to `graph_vector_fulltext`: chunk hits, then one or two hops from each chunk's entities depending on their similarity to the query. Other modes search community summaries or generate Cypher.

**Vector-seeded neighbourhoods.** [LightRAG](/p/lightrag/) makes one LLM call to extract keywords. Low-level keywords search entity vectors and high-level keywords search relation vectors. The default `mix` mode adds chunk hits, with `top_k` 40. [AutoFlow](/p/autoflow/) splits the question into sub-questions by default. For each one it walks relationships by weight, two levels deep. It then retrieves chunks with a rewritten question. [TrustGraph](/p/trustgraph/) grounds the question's concepts in entity vectors and walks two hops. A cross-encoder keeps 25 edges per hop, so up to 50 edges reach synthesis. It has no global mode.

**Graph as a passage ranker.** HippoRAG scores facts, lets an LLM filter the top 5, then runs PPR with damping 0.5. If no fact survives the filter, it quietly falls back to dense retrieval. Vector Graph RAG extracts entities from the question, keeps entity hits above 0.9 similarity, expands one degree, asks the LLM for five relations, and answers from three passages.

**Lexical traversal.** [graphify](/p/graphify/) scores node labels by token match behind a trigram prefilter, then runs BFS or DFS to depth 3. It returns 2,000 tokens of text to the calling agent and makes no LLM or vector call.

Pick: GraphRAG or Semantica for synthesis across the corpus.
Pick: LightRAG `mix` or TrustGraph for fast answers grounded in entities.
Pick: HippoRAG or Vector Graph RAG for multi-hop QA over passages.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/query/index.md

## How are updates and incremental indexing handled?

[LightRAG](/p/lightrag/) and [HippoRAG](/p/hipporag/) handle change best. Both add documents without a rebuild and delete them with reference-counted cleanup of shared entities. [GraphRAG](/p/graphrag/) and [nano-graphrag](/p/nano-graphrag/) are in practice append-only.

**Add and delete per document.** LightRAG merges each document into the live graph. `adelete_by_doc_id` rebuilds shared entities from cached extraction results, and deletion is journaled so a crash can resume. HippoRAG sends only unseen chunk hashes to OpenIE and runs synonymy KNN only for new entities. Its manifest refuses to mix state from two different models. [Vector Graph RAG](/p/vector-graph-rag/) has `upsert_documents_by_source` and a cascading `delete_documents_by_source`, but only in Python: its REST import endpoints always rebuild the whole graph. [LLM Graph Builder](/p/llm-graph-builder/) processes each document on its own, can resume from the last processed chunk, and deletes only the entities no other document references. It has no extraction cache, so a retry pays again. [AutoFlow](/p/autoflow/) deletes a document's relationships and then any orphaned entities, but merged descriptions stay behind. Re-indexing skips chunks that already have relationships.

**Append-only.** GraphRAG's `update` finds new documents by title. It computes deleted documents but never uses that list, and it appends delta communities without re-clustering. nano-graphrag skips known document and chunk hashes, but every insert drops and regenerates all community reports. It has no delete API. [Semantica](/p/semantica/) merges new entities with `incremental_merge` and caches extraction with a TTL. It has no document-level delete, only entity-level purge and erasure.

**Driven by file changes.** [graphify](/p/graphify/) caches both AST and LLM results by file-content hash, and the LLM cache also includes a prompt fingerprint. `graphify update` re-parses changed code with no LLM and removes nodes for deleted files. Communities are recomputed on every build.

**Collection-level only.** [TrustGraph](/p/trustgraph/) has no incremental path: there is no extraction cache, graph data is deleted per collection, and deleting a document in the librarian leaves its triples in place.

Pick: LightRAG or HippoRAG for a corpus that changes daily.
Pick: graphify for repositories that change on every commit.
Pick: GraphRAG or nano-graphrag only when you can rebuild in batches.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/incremental/index.md

## How are LLM cost and latency controlled during indexing and query?

The cheapest options avoid the LLM where they can: [graphify](/p/graphify/) parses code with no model, and [Semantica](/p/semantica/) extracts with spaCy by default. Among full LLM pipelines, [LightRAG](/p/lightrag/), [HippoRAG](/p/hipporag/) and [Vector Graph RAG](/p/vector-graph-rag/) keep each query to a few calls. [GraphRAG](/p/graphrag/) spends most of its budget at index time.

**Heavy indexing.** GraphRAG runs extraction plus gleaning on every chunk, a summarization call for merged descriptions, and one report per community. Each stage has its own model, and its `fast` method replaces LLM extraction with NLP. Only indexing calls are cached. [nano-graphrag](/p/nano-graphrag/) splits work between a "best" model and a "cheap" model and caches every call, including queries, but each insert regenerates all community reports. [LLM Graph Builder](/p/llm-graph-builder/) makes one call per `chunks_to_combine` chunks and has no extraction cache. Optional per-user daily and monthly token limits can stop a job before it starts. [AutoFlow](/p/autoflow/) makes two DSPy calls per chunk plus LLM merge calls for entities, and uses a separate `fast_llm` to rewrite questions.

**Lean indexing, cheap queries.** LightRAG writes no reports, caches both extraction and answers, and routes extraction, keyword and answer calls to different models. Its answer cache ignores the retrieved context, so a cached answer can survive changes to the data. HippoRAG caches every request in SQLite. A query costs one filter call plus a PageRank solve. Vector Graph RAG caches every LLM call on disk, uses `gpt-4o-mini` for every step with no per-stage choice, and makes three calls per query. [TrustGraph](/p/trustgraph/) makes two LLM calls per query and filters edges with a MiniLM cross-encoder. It does not cache extraction.

**Skipping the LLM.** graphify packs files for its LLM pass into chunks of up to 60,000 tokens. When a response is truncated, it splits the chunk in half and retries, up to three levels deep. It caches results per file hash. Semantica's LLM stages are opt-in, it caches extraction with a TTL of 3,600 seconds, and it can write extractive community reports with no model.

Pick: graphify or Semantica's defaults when the budget is close to zero.
Pick: LightRAG or HippoRAG for low running cost with LLM-quality extraction.
Pick: GraphRAG when global answers justify paying for indexing up front.

Per-project answers: https://llms-technical-reviews.com/graph-rag/q/cost/index.md

## Projects

- [Graphify-Labs/graphify](https://llms-technical-reviews.com/p/graphify/index.md) — Coding-assistant skill and CLI that turns a repo into a tree-sitter code graph plus LLM-extracted doc nodes, queried over MCP by traversal.
- [HKUDS/LightRAG](https://llms-technical-reviews.com/p/lightrag/index.md) — Python graph RAG engine that merges LLM-extracted entities into one graph and retrieves by keyword-matched entities, relations and chunks.
- [microsoft/graphrag](https://llms-technical-reviews.com/p/graphrag/index.md) — Python pipeline that turns text into an LLM-extracted entity graph with Leiden community reports, queried by local, global and DRIFT search.
- [semantica-agi/semantica](https://llms-technical-reviews.com/p/semantica/index.md) — Large Python knowledge-graph toolkit with spaCy/LLM extraction, pluggable graph stores, and GraphRAG-style community and global search.
- [neo4j-labs/llm-graph-builder](https://llms-technical-reviews.com/p/llm-graph-builder/index.md) — FastAPI and React app that turns files, URLs and transcripts into a Neo4j knowledge graph via LLMGraphTransformer, plus GraphRAG chat.
- [OSU-NLP-Group/HippoRAG](https://llms-technical-reviews.com/p/hipporag/index.md) — Research RAG library that links OpenIE triples, entities and passages in one igraph and ranks passages with Personalized PageRank.
- [gusye1234/nano-graphrag](https://llms-technical-reviews.com/p/nano-graphrag/index.md) — A small, readable Python reimplementation of Microsoft GraphRAG with local, global and naive query modes over a NetworkX graph.
- [pingcap/autoflow](https://llms-technical-reviews.com/p/autoflow/index.md) — Self-hosted Graph RAG chat app on TiDB that extracts a DSPy knowledge graph per chunk and fuses it with vector search.
- [trustgraph-ai/trustgraph](https://llms-technical-reviews.com/p/trustgraph/index.md) — Microservice Graph RAG platform that extracts RDF triples onto a message bus and answers by cross-encoder-filtered graph hops.
- [zilliztech/vector-graph-rag](https://llms-technical-reviews.com/p/vector-graph-rag/index.md) — Python Graph RAG library that stores entities, relations and passages as three Milvus collections and walks the graph by ID lookups.