LLMs Technical Reviews
Home / Graph RAG / HippoRAG

OSU-NLP-Group/HippoRAG

Research RAG library that links OpenIE triples, entities and passages in one igraph and ranks passages with Personalized PageRank.

GitHub ↗★ 4.0kPythonMITcommit 2bfd831 · 2026-10-01homepage ↗

Overview

HippoRAG is the reference implementation of two OSU NLP papers. The current main branch is HippoRAG 2 (“From RAG to Memory”); version 1 lives on a legacy branch. It is a Python library, not a service. You give a HippoRAG object a list of passages and it builds a small knowledge graph from them. At query time it does not answer from the graph directly. It uses the graph to re-rank passages, then hands the top passages to an LLM.

The core idea is easy to state. An LLM extracts (subject, predicate, object) triples from each passage. Entities become graph nodes, passages become graph nodes, and edges link entities that share a fact, passages to the entities they mention, and entities whose embeddings are close (“synonymy”). A query is matched against fact embeddings. An LLM filter keeps the relevant facts, and their entities, plus a lightly weighted dense-retrieval score for every passage, seed a Personalized PageRank (PPR). Passages are ranked by their PPR score. There are no communities, no summaries and no local/global modes. Multi-hop recall comes from PPR spreading relevance along entity paths.

Almost everything lives in one 2,200-line class, HippoRAG.py. Recent commits add a lot of engineering on top of the research code: provenance-aware edges, an index manifest that rejects mixed model state, incremental insert and delete, more LLM gateways, and an MCP server. It is a good choice when you want a well-understood, benchmarked multi-hop retriever over a corpus of short passages. It is a weaker fit for millions of documents or for “summarise the whole corpus” questions.

Architecture

flowchart LR
  D["Documents"] --> P["TextPreprocessor"]
  P --> CS["chunk EmbeddingStore"]
  P --> OIE["OpenIE: NER then triples"]
  OIE --> ST["openie_state.json"]
  OIE --> ES["entity + fact EmbeddingStores"]
  ES --> SYN["KNN synonymy edges"]
  OIE --> G["igraph: entity + passage nodes"]
  SYN --> G
  G --> PK["graph.pickle"]
  Q["Query"] --> FS["Fact scoring (dot product)"]
  FS --> RR["DSPyFilter (LLM)"]
  RR --> PPR["Personalized PageRank"]
  CS --> DPR["Dense passage scores"]
  DPR --> PPR
  G --> PPR
  PPR --> QA["QA LLM over top passages"]
Component Path Role
Facade src/hipporag/HippoRAG.py index, delete, retrieve, rag_qa, graph building, PPR
OpenIE src/hipporag/information_extraction/ NER and triple extraction (online API, vLLM offline, Transformers offline)
Prompts src/hipporag/prompts/templates/ One-shot NER and triple prompts, dataset-specific QA prompts
Recognition memory src/hipporag/rerank.py DSPyFilter: LLM keeps only query-relevant facts
Embedding stores src/hipporag/embedding_store.py, vector_stores/ Parquet by default; Qdrant, Chroma, Milvus optional
Embedding models src/hipporag/embedding_model/ NV-Embed-v2 (default), GritLM, Contriever, OpenAI, Cohere, vLLM, Transformers
LLM clients src/hipporag/llm/ OpenAI with SQLite cache, Bedrock, gateways, Transformers
Config src/hipporag/utils/config_utils.py BaseConfig dataclass with every tunable
Baseline src/hipporag/StandardRAG.py Dense-retrieval RAG used for comparison
MCP server src/hipporag/mcp_server.py hipporag-mcp with retrieve, QA, index, delete tools

How a request flows

Indexing, hipporag.index(docs):

  1. Guard state. index() refuses to run on a graph without source-aware edge metadata, or when embedding stores exist but the graph pickle is missing (HippoRAG.py).
  2. Chunk and embed passages. The default TextPreprocessor keeps one chunk per document (preprocessing.py). Chunks go into the chunk store, keyed by chunk-<md5>.
  3. OpenIE on new chunks only. load_existing_openie returns the chunk ids that have no saved extraction. batch_openie runs NER for all of them in a thread pool, then triple extraction conditioned on the NER output in a second pool. Any failed chunk aborts the batch (openie_openai.py).
  4. Encode entities and facts. Triples are normalised with text_processing (casefold, punctuation to spaces). Unique subjects/objects go into the entity store and stringified triples into the fact store (HippoRAG.py).
  5. Collect edges. add_fact_edges counts entity-entity pairs per source chunk. add_passage_edges links each new chunk to its entities. add_synonymy_edges runs a batched KNN of new entities against all entities and keeps neighbours with similarity at or above 0.8 (HippoRAG.py).
  6. Merge and save. add_new_nodes adds entity and passage vertices. add_new_edges merges contributions by logical edge key, so incremental and one-shot builds give the same weights. save_igraph writes the pickle atomically under a file lock (HippoRAG.py).

Querying, hipporag.retrieve(queries):

  1. Load everything. prepare_retrieval_objects loads all ids and embeddings into NumPy arrays and checks that graph vertices, stores and OpenIE state agree exactly (HippoRAG.py).
  2. Score facts. get_fact_scores takes a dot product of the query embedding with every fact embedding and min-max normalises it (HippoRAG.py).
  3. Filter facts. rerank_facts sends the top linking_top_k (5) facts to DSPyFilter, which asks the extraction LLM to keep only the relevant ones (rerank.py).
  4. Seed and run PPR. If no facts survive, the query falls back to plain dense passage retrieval. Otherwise graph_search_with_fact_entities gives each entity the mean score of its facts, divided by how many chunks mention it. It also gives every passage its dense score times passage_node_weight (0.05). run_ppr calls igraph’s personalized_pagerank with damping 0.5 over the undirected graph (HippoRAG.py).
  5. Answer. rag_qa passes the top qa_top_k (5) passages to the QA LLM as Wikipedia Title: ... blocks and splits the reply on Answer: (HippoRAG.py).

Key components

The graph

The igraph graph has two vertex types: entity nodes (entity-<md5> of the normalised phrase) and passage nodes (chunk-<md5>). Facts are not vertices. They live only in the fact embedding store and act as the entry point at query time. Each edge stores weight, edge_kind (fact, passage, synonym or a + combination), per-chunk fact_source_counts, synonym_score and passage_source. The weight is the largest of the fact count, 1.0 for a passage link, and the synonym score (HippoRAG.py). Entity resolution is exact string match after normalisation. Fuzzy equivalence exists only as synonymy edges.

Recognition memory

DSPyFilter is a fixed few-shot prompt saved from a DSPy optimisation run (prompts/filter_default_prompt.py, or a JSON file via rerank_dspy_file_path; rerank.py). The LLM returns a fact list. Each returned fact is snapped to the closest candidate with difflib.get_close_matches(..., cutoff=0.0) (rerank.py). This is the only LLM call in retrieval.

State and provenance

Every index directory holds graph.pickle, three embedding stores, openie_state.json, chunk_metadata.json and an index_manifest.json that pins the LLM, endpoint, embedding model and OpenIE token limits. A mismatch raises StateConsistencyError instead of mixing vectors from two models. delete() removes chunks and then removes only the facts and entities no other chunk still references. It also subtracts the deleted chunks from fact_source_counts on shared edges (HippoRAG.py).

Model plumbing

_get_llm_class picks a gateway client, Bedrock Mantle, Bedrock, Transformers or the default CacheOpenAI from the model-name prefix (llm/init.py). The OpenAI client caches responses in SQLite keyed by a hash of the request. Embedding classes are picked from the provider or name and imported lazily (embedding_model/init.py).

Extending it

  • Configuration. All knobs are fields on BaseConfig: linking_top_k, passage_node_weight, damping, synonymy_edge_sim_threshold, retrieval_top_k, qa_top_k, vector_store_type and others.
  • Injected components. HippoRAG(...) accepts your own extraction_llm, qa_llm, embedding_model and text_preprocessor. With any of these you must pass index_identity, so the manifest can tell incompatible state apart (HippoRAG.py). A real chunker belongs in a BaseTextPreprocessor subclass.
  • Vector backends. get_embedding_store switches between Parquet, Qdrant, Chroma and Milvus (embedding_store.py).
  • Prompts. NER, triple, QA and IRCoT templates are plain Python modules under prompts/templates/. QA templates are picked by rag_qa_<dataset> and fall back to the MuSiQue one.
  • MCP. hipporag-mcp exposes retrieve, rag_qa, index_stats and, unless --read_only is set, index and delete over stdio or streamable HTTP (mcp_server.py).

Running it

  • Install. pip install hipporag into Python 3.10. Extras such as [vllm] and [gritlm] add local-model support.
  • Hosted models. Set OPENAI_API_KEY (or point llm_base_url / embedding_base_url at any OpenAI-compatible endpoint) and call HippoRAG(save_dir=..., llm_model_name=..., embedding_model_name=...), then index() and rag_qa().
  • Local models. The default embedder is nvidia/NV-Embed-v2, which needs a GPU. openie_mode="offline" runs vLLM batch extraction first and then stops on purpose, asking you to re-run indexing in online mode.
  • Benchmarks. main.py and reproduce/ replay the paper datasets (MuSiQue, 2Wiki, HotpotQA and others).
  • Services. None required. A Qdrant, Chroma or Milvus server is optional.

Strengths and caveats

  • Strength: a clear, cheap query path. One embedding pass, one small LLM filter call and one PPR solve per query. No community summaries to build or refresh.
  • Strength: honest incremental updates. Re-indexing skips known chunks. New entities are KNN-matched only against existing ones. Deletion is reference-counted down to individual edge sources.
  • Strength: defensive state handling. The manifest, atomic writes and node-count checks make silent index corruption unlikely.
  • Caveat: everything is in memory. The graph is one pickle, and retrieval loads all entity, passage and fact embeddings into NumPy. Fact scoring is a brute-force dot product even when Qdrant or Milvus is the store. PPR runs over the whole graph on every query.
  • Caveat: no chunking by default. One document becomes one passage. Long documents need your own preprocessor, or extraction quality and the passage-level ranking both suffer.
  • Caveat: quiet fallbacks. rerank_facts catches every exception and returns no facts, which silently turns the query into plain dense retrieval. difflib with cutoff=0.0 always maps the filter’s output to some candidate.
  • Caveat: benchmark-shaped defaults. QA prompts format passages as Wikipedia titles, and answer parsing expects an Answer: line. Fine for the papers; you will want your own prompt in production.

Sources: code at 2bfd831, verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit 2bfd831. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Chunking. Documents pass through a BaseTextPreprocessor before indexing. The default TextPreprocessor (src/hipporag/preprocessing.py:18–27) wraps each string as one Chunk without splitting. Configurable chunking is available via BaseConfig fields like preprocess_chunk_max_token_size, preprocess_chunk_overlap_token_size, and preprocess_chunk_func (src/hipporag/utils/config_utils.py:105–117) but the base class is abstract — users supply a custom preprocessor.

Entity and relation extraction. Two-stage OpenIE: first NER, then triple extraction. The NER prompt (src/hipporag/prompts/templates/ner.py:1–22) system-asks "extract named entities… respond with a JSON list of entities" with a one-shot example outputting {"named_entities": [...]}. The triple prompt (src/hipporag/prompts/templates/triple_extraction.py:4–50) system-instructs "construct an RDF graph… respond with a JSON list of triples", conditioned on the NER output. Both use the same LLM (default gpt-4o-mini). batch_openie (src/hipporag/information_extraction/openie_openai.py:186–286) runs NER and triple extraction concurrently via ThreadPoolExecutor with openie_max_workers (default 8) threads.

Entity resolution / deduplication. Entities are normalized by computing an MD5 hash of their lowercased, punctuation-stripped text (compute_mdhash_id(content, prefix="entity-")). This is a string-hash deduplication: two textually identical entity strings map to the same node. There is no fuzzy matching or LLM-based coreference; synonyms are discovered separately via embedding similarity (see synonymy edges).

Graph construction. From the extracted triples, extract_entity_nodes derives entity strings and flatten_facts collects all triples (HippoRAG.py:563–565). add_fact_edges (1179–1228) creates edges between entity nodes, weighted by frequency and tracked by source chunk via _fact_edge_source_counts. add_passage_edges (1230–1276) links each chunk node to the entity nodes appearing in it. add_synonymy_edges (1278–1352) runs KNN (synonymy_edge_topk=2047, threshold 0.8) over entity embeddings to link similar entities. The graph schema uses typed edges storing weight, edge_kind (fact/passage/synonym combinations), fact_source_counts, synonym_score, and passage_source (HippoRAG.py:1588–1596).

Editor's note. Correction: batch_openie does not run NER and triple extraction concurrently; it runs NER for all chunks in one thread pool, then triple extraction (conditioned on the NER output) in a second pool, and any failed chunk aborts the batch.

Where and how is the graph stored?

answered

Graph database vs files vs in-memory. The graph is stored as a serialized iGraph pickle file on disk (graph.pickle) and loaded into memory on init (HippoRAG.py:436–473). initialize_graph checks for an existing pickle; if found it calls ig.Graph.Read_Pickle and logs node/edge counts. Otherwise it creates an empty ig.Graph(directed=...). save_igraph (1679–1694) writes the graph back via graph.write_pickle using atomic file-replacement with a .tmp staging file and FileLock for concurrency safety.

Node/edge schema. The graph has two node types: phrase (entity) nodes (strings from NER, hashed as entity-<md5>) and passage (chunk) nodes (chunk content hashed as chunk-<md5>). Edges carry typed attributes: weight (float), edge_kind (string like "fact", "passage", "synonym", or "fact+synonym"), fact_source_counts (dict mapping chunk IDs to occurrence counts), synonym_score (float), passage_source (chunk ID), and source_key/target_key (HippoRAG.py:1561, 1588–1596). The graph is undirected by default (is_directed_graph=False in BaseConfig:176), but supports directed mode.

How embeddings sit next to the graph. Embeddings are stored separately in three EmbeddingStore instances — one per namespace (chunk, entity, fact) — by default as local Parquet files with columns hash_id, content, embedding (src/hipporag/embedding_store.py:183–187). The factory get_embedding_store (271–301) also supports Qdrant, ChromaDB, and Milvus backends. At retrieval time, prepare_retrieval_objects (HippoRAG.py:1750–1816) loads all embeddings into NumPy arrays indexed by node key. The graph and embeddings are linked by shared hash IDs: a phrase node in the graph has name = entity-<md5> which is the same key used in entity_embedding_store.

Are communities, summaries or hierarchies built over the graph?

answered

Not implemented. The repository does not implement community detection (Leiden, Louvain, or any other algorithm), hierarchical summaries over graph regions, or any form of graph partitioning. A grep for commun, leiden, louvain, hierarch, and summar across the entire src/ directory returned zero matches. The graph is used strictly as a flat, node-labeled structure where edges represent fact co-occurrence, passage–entity membership, and embedding-similarity (synonymy).

What happens instead. When querying, the system computes fact–query similarity scores, selects top facts, propagates their scores to the entities they contain, runs personalized PageRank (igraph.Graph.personalized_pagerank) over the entire entity+passage graph, and ranks passage nodes by their PPR score (HippoRAG.py:2177–2218). There is no hierarchical summarization step and no community-level retrieval. The graph is always treated as one monolithic component.

This is a deliberate architectural difference from approaches like GraphRAG (which uses Leiden clustering + community summaries). HippoRAG's PPR-based retrieval effectively lets relevance propagate along paths through the graph without needing a pre-computed community hierarchy.

How does query-time retrieval use the graph?

answered

No local/global/hybrid modes. HippoRAG has a single retrieval mode: it scores facts against the query, reranks, then runs PPR over the graph seeded by entity weights. The user-visible entry points are retrieve() (HippoRAG.py:704–794), retrieve_ircot() (804–857) and retrieve_dpr() (968–1039), plus QA helpers rag_qa() and answer_with_ircot().

Step 1 — Fact retrieval. Queries are encoded through the embedding model with a query_to_fact instruction (get_query_instruction, prompts/linking.py:5). Dot-product similarity scores are computed against all fact embeddings and min-max normalized (HippoRAG.py:1892–1928).

Step 2 — Recognition-memory reranking. The DSPyFilter (rerank.py:14–126) uses the extraction LLM to filter the top linking_top_k facts (default 5), passing each candidate fact list to the LLM with a prompt asking it to select only relevant facts. The LLM's selection is matched back by text similarity (difflib.get_close_matches).

Step 3 — Graph search with PPR. graph_search_with_fact_entities (HippoRAG.py:2008–2121) assigns each entity a weight from the average score of the facts containing it. Simultaneously, dense_passage_retrieval (1930–1967) computes passage–query dot-product scores, scaled by passage_node_weight (default 0.05). Entity and passage weights are combined into one node-weight vector, then run_ppr (2177–2218) calls igraph.Graph.personalized_pagerank with damping=0.5 (default), using the weight attribute on edges and an undirected projection (directed=False). Passage nodes are extracted from the PPR scores and sorted.

Step 4 — Context for the LLM. The top num_to_retrieve passages are collected via _build_retrieval_result (796–802). The QA step (qa, 1117–1177) passes the top qa_top_k (default 5) passages to the LLM with a prompt formatted as "Wikipedia Title: {passage}" + "Question: {query}", using dataset-specific prompt templates (e.g., rag_qa_musique).

Fallback. If no facts survive reranking, dense_passage_retrieval is used as a fallback (HippoRAG.py:762–764).

How are updates and incremental indexing handled?

answered

Document addition. index() (HippoRAG.py:499–593) is designed for incremental use. It calls load_existing_openie (1354–1404), which loads any prior OpenIE results from openie_state.json or the alternate openie_results_path. Only chunks whose hash IDs are not already in the persisted state (chunk_keys_to_save) are sent to batch_openie for NER+triple extraction. This means re-indexing the same documents is a no-op, and adding new documents only processes the new chunks. The graph is then augmented: add_fact_edges and add_passage_edges only process new chunks (they check current_graph_nodes at lines 1204 and 1256). add_synonymy_edges is optimized: when _pending_synonymy_entity_ids is set (only for new entities), it runs KNN only for the new entities against all existing entities (1329–1350), rather than recomputing all-vs-all (src/hipporag/HippoRAG.py:1301, 1316–1317).

Document deletion. delete() (595–682) finds the chunks to remove via hash, looks up their triples and entities, checks proc_triples_to_docs to see if a triple is referenced by other chunks (via remove_sources_from_mapping), and only removes unreferenced facts and entities. It then deletes vertices from the iGraph after cleaning up _fact_edge_source_counts and edge weights (684–702). Embedding store deletions are done separately per namespace.

State consistency. The index_manifest.json (317–344) binds the OpenIE provenance (model, endpoint, parameters) and embedding config to the persisted state. Loading with a different model/endpoint raises StateConsistencyError. The force_index_from_scratch and force_openie_from_scratch flags allow explicit rebuilds.

Caching of extraction results. OpenIE results (NER + triples) are saved to openie_state.json and optionally to a human-readable openie_results_path. On subsequent index calls, only missing chunk hashes are re-processed — the cached NER and triple outputs for unchanged chunks are reused at the graph-building stage (HippoRAG.py:544–553).

How are LLM cost and latency controlled during indexing and query?

answered

LLM response caching. The CacheOpenAI class (src/hipporag/llm/openai_gpt.py:93–220) wraps every infer() call with @cache_response (33–91), which SHA-256-hashes the serialized request (messages, model, endpoint, generation params) and checks a SQLite cache. On hit it returns (message, metadata, True) with zero API cost. The cache file is stored per-LLM name under save_dir/llm_cache/. Cache-aware token accounting (prompt/completion/billable vs logical) is displayed during batch OpenIE (src/hipporag/information_extraction/openie_openai.py:222–238), separating hit vs miss costs.

Batching. Embeddings are batched through embedding_batch_size (default 16, config_utils.py:140–141). OpenIE runs via ThreadPoolExecutor with openie_max_workers (default 8, config_utils.py:311), allowing concurrent NER and triple-extraction calls. KNN for synonymy edges is batched with configurable synonymy_edge_query_batch_size (default 1000) and synonymy_edge_key_batch_size (default 10000, config_utils.py:164–170).

Model choice per stage. The extraction/OpenIE LLM (extraction_llm) and the QA LLM (qa_llm) are independently configurable (HippoRAG.py:97–98, 220–229). By default they share the same model, but users can pass a cheaper model for extraction and a stronger one for QA. The embedding model is also independently configurable (embedding_model_name, default nvidia/NV-Embed-v2, config_utils.py:136).

Non-LLM shortcuts. The DSPy reranking filter (rerank.py) is the only LLM call at query time beyond QA — it uses the extraction LLM with max_new_tokens=512 to filter the candidate fact list. PPR is pure NumPy + iGraph C-extension (no LLM). Fact scoring is a simple np.dot (HippoRAG.py:1926). When no facts survive reranking, the system falls back to dense passage retrieval which is entirely embedding-based (1930–1967), avoiding any LLM call for that query.

Token budgets. openie_ner_max_tokens=2048 and openie_triple_max_tokens=4096 are configurable (config_utils.py:315–322). The DSPyFilter uses max_new_tokens=512 (rerank.py:36). QA inference uses max_new_tokens=2048 (config_utils.py:38). The model is configured with temperature=0 by default for deterministic extraction.