LLMs Technical Reviews
Home / Graph RAG

Graph RAG

Retrieval-augmented generation over knowledge graphs — extracting entities and relations from documents and querying the graph alongside vectors.

Graph RAG tools build a knowledge graph from your documents and use it during retrieval. An LLM reads each chunk and extracts entities and relationships. These are merged into a graph, usually with embeddings stored beside it. At query time the tool matches the question to graph elements, follows their links, and gives the LLM that structured context along with the source text. Some tools also cluster the graph and write summaries, so they can answer questions about a whole corpus.

When choosing, check five things. How much does indexing cost in LLM calls? Can you add and delete documents without a rebuild? Are communities or summaries built, or does all synthesis happen at query time? Is the graph a file, an in-memory object or a real database? How are near-duplicate entity names resolved?

Projects (2)

ProjectStarsLanguageLicense
HKUDS/LightRAGPython graph RAG engine that merges LLM-extracted entities into one graph and retrieves by keyword-matched entities, relations and chunks.★ 40kPythonMIT
microsoft/graphragPython pipeline that turns text into an LLM-extracted entity graph with Leiden community reports, queried by local, global and DRIFT search.★ 36kPythonMIT

In the research queue: Graphify-Labs/graphify, semantica-agi/semantica, gusye1234/nano-graphrag, neo4j-labs/llm-graph-builder, OSU-NLP-Group/HippoRAG, trustgraph-ai/trustgraph, pingcap/autoflow, zilliztech/vector-graph-rag.

Comparison questions

Each question is answered separately for every project in this category, from that project's source code.

All verdicts on one page →

  1. How is the knowledge graph extracted from documents?Chunking; entity and relation extraction prompts or models; entity resolution / deduplication; schema or ontology.
  2. Where and how is the graph stored?Graph database vs files vs in-memory; node/edge schema; how embeddings sit next to the graph.
  3. Are communities, summaries or hierarchies built over the graph?Community detection (e.g. Leiden); hierarchical summaries; when they are computed; if not done, say so.
  4. How does query-time retrieval use the graph?Local / global / hybrid modes; traversal; combining graph and vector hits; how context is assembled for the LLM.
  5. How are updates and incremental indexing handled?Adding or changing documents without a full rebuild; deletion; caching of extraction results.
  6. How are LLM cost and latency controlled during indexing and query?Caching; batching; model choice per stage; token budgets; small-model or non-LLM shortcuts.