Where and how is the graph stored?
Graph database vs files vs in-memory; node/edge schema; how embeddings sit next to the graph.
Verdict
For a production database that holds both the graph and its vectors, LLM Graph Builder (Neo4j) and AutoFlow (TiDB) are the most integrated. For a portable file you can open and version, use GraphRAG’s tables or graphify’s graph.json.
Files and in-memory graphs. GraphRAG has no graph store. It writes entities, relationships and communities as Parquet or CSV tables and puts the embeddings in LanceDB. nano-graphrag keeps a NetworkX graph in a GraphML file, with JSON key-value files and nano-vectordb beside it. Neo4j is an option. graphify serialises NetworkX to node-link JSON and stores no embeddings at all. Its Neo4j and FalkorDB pushes are one-way exports. HippoRAG pickles an igraph graph of entity and passage nodes, written atomically under a file lock. Facts are not graph nodes. They live, with chunks and entities, in three Parquet embedding stores.
Pluggable backends. LightRAG has four storage roles. The default graph is a single-writer GraphML file, and Neo4j, Memgraph, PostgreSQL, MongoDB or OpenSearch can replace it. Entities and relations are also stored as vectors. In Semantica the graph is a plain Python dict unless you configure a GraphStore (Neo4j, FalkorDB, Apache AGE or Neptune). Vectors go to a separate store. TrustGraph stores RDF triples in Cassandra, Neo4j, Memgraph or FalkorDB, and entity vectors in Qdrant, Milvus or Pinecone. Its Neo4j writer drops the named-graph field, so the provenance features need Cassandra.
One store for graph and vectors. LLM Graph Builder keeps Document, Chunk, entity and community nodes in Neo4j, each with an embedding property and a vector index. AutoFlow creates entity and relationship tables per knowledge base, with HNSW vector columns in the same rows. TiDB is its only backend. Vector Graph RAG has no graph database at all. It uses three Milvus collections and stores adjacency as ID lists in dynamic fields. Each hop is an id in [...] query.
Pick: GraphRAG or HippoRAG for a batch artifact you rebuild and inspect offline. Pick: LightRAG, TrustGraph or LLM Graph Builder for a live graph in a real database. Pick: Vector Graph RAG or AutoFlow if you already run Milvus or TiDB and want no second system.
Per-project answers
Graphify-Labs/graphify
answeredIn-memory NetworkX. The graph lives in memory as a networkx.Graph (or nx.DiGraph when directed=True) throughout building, clustering, analysis, and query serving. build_from_json() (build.py:877–1038) constructs it: nodes are added with all attributes except id unpacked as kwargs, and edges are added with their metadata.
Serialised to JSON on disk. The canonical serialisation is graph.json, written by to_json() in export.py (line 272). It uses NetworkX's node_link_data format: a JSON object with top-level nodes (array of dicts) and links (array of edge dicts). Community IDs are stamped onto each node at write time. A safety check (export.py:278–331) refuses to overwrite an existing graph with a shrinking node count unless force=True. A backup is snapshotted before overwrite when the graph has semantic or curated content (backup_if_protected, export.py:42–104).
Optional graph database export. The exporters/graphdb.py module provides push_to_neo4j() (graphdb.py:22–98) which uses MERGE statements to upsert nodes and edges into a running Neo4j instance. Node labels are derived from file_type (capitalised). A similar push_to_falkordb exists. These are pure exports — there is no database-backed storage layer.
No vector store or embeddings. The README (line 38) states explicitly: "No embeddings, no vector store: a real graph you traverse." Node attributes include norm_label (diacritic-stripped lowercased label) for textual matching, but no embedding vectors are stored.
Global graph. A ~/.graphify/global-graph.json accumulates graphs from multiple repos via global_add() (global_graph.py:79). It is a flat union with repo-tagged nodes, not a federated index.
HKUDS/LightRAG
answeredGraph storage backends. The graph is stored via a pluggable BaseGraphStorage (lightrag/base.py:671). Implementations registered in lightrag/kg/__init__.py:12-23 include NetworkXStorage (default, file-based), Neo4JStorage (Neo4j graph DB), PGGraphStorage / PGTableGraphStorage (PostgreSQL), MongoGraphStorage, MemgraphStorage, and OpenSearchGraphStorage.
Default: NetworkX (in-memory + file). NetworkXStorage (lightrag/kg/networkx_impl.py:38-85) keeps the entire graph in process memory using networkx.Graph. It persists by rewriting one GraphML file (graph_<namespace>.graphml) per workspace via index_done_callback. It declares requires_single_writer = True, meaning only one process may mutate it at a time, enforced by the pipeline's busy reservation or LightRAG._admin_write_gate. Concurrent readers detect peer writes through a two-channel fence: file (st_mtime_ns, st_size) fingerprint plus a storage_updated flag. A commit publishes the whole namespace, so partial mutations from other in-flight writers can be published unintentionally — accepted as a documented residue.
Node/edge schema. Nodes and edges are NetworkX graph.add_node(id, **data) / graph.add_edge(src, tgt, **data) calls with arbitrary string-keyed dictionaries (lightrag/kg/networkx_impl.py:820-821, 844-845). Node data typically includes entity_name, entity_type, description, source_id, file_path, created_at. Edge data includes src_id, tgt_id, weight, source_id, keywords, description, file_path, created_at. XML attribute validation (validate_xml_attributes) is applied before every write since GraphML serialization can't handle nested structures.
Embeddings: separate vector stores. Graph nodes (entities) and edges (relations) are dually stored: once in the graph for structural traversal, and once in entity/relation vector databases (entities_vdb / relationships_vdb). The default NanoVectorDBStorage (lightrag/kg/nano_vector_db_impl.py:62-97) is an in-memory store serialized to a JSON file per workspace. Each entity vector record holds {content, entity_name, source_id, description, entity_type, file_path} (lightrag/operate.py:1805-1813). Each relation vector record holds {src_id, tgt_id, source_id, content, keywords, description, weight, file_path} (lightrag/operate.py:2324-2335). Other supported vector DBs: Milvus, PGVector, Faiss, Qdrant, Mongo, OpenSearch. NoopVectorDBStorage disables vector storage.
microsoft/graphrag
answeredThe graph is stored as tabular files (parquet/CSV) or in Azure CosmosDB — not as a graph database. The base Storage abstraction provides key-value file access with implementations for local disk, Azure Blob, Azure Cosmos DB, and in-memory. On top sits TableProvider for table-level operations: read_dataframe, write_dataframe, and open for streaming row operations.
Table schema. Six tables are produced with fixed column schemas in schemas.py:70-159. Entities: id, title, type, description, text_unit_ids, frequency, degree. Relationships: id, source, target, description, weight, combined_degree, text_unit_ids. Communities: id, community, level, parent, children, entity_ids, relationship_ids, text_unit_ids. TextUnits: id, text, n_tokens, document_id, entity_ids, relationship_ids.
Embeddings. The generate_text_embeddings workflow (at generate_text_embeddings.py:52) creates embeddings for three configured fields: text_unit text, entity title-description, and community report full content. The embed_text operation reads rows, batches them (configurable batch_size/batch_max_tokens), calls the embedding model, and loads vectors into a pluggable VectorStore. Optional snapshot tables store vectors alongside base tables.
No graph DB at query time. Query sessions load entities/relationships/communities into Python dataclass objects from serialized tables at startup.
semantica-agi/semantica
answeredThe graph is stored via a pluggable backend architecture. The core GraphStore class in semantica/graph_store/graph_store.py defines the interface (NodeManager, RelationshipManager, QueryEngine), and concrete implementations exist for Neo4j (Neo4jStore in neo4j_store.py, using Cypher queries via the neo4j driver), Apache AGE (ApacheAgeStore in age_store.py, wrapping openCypher through PostgreSQL psycopg2), FalkorDB (FalkorDBStore in falkordb_store.py), and Amazon Neptune. The builder's build() method optionally persists the graph to a store via self.graph_store.add_nodes(resolved_entities) and self.graph_store.add_edges(formatted_edges) (graph_builder.py:1012-1044). When no database backend is configured, the graph is returned as an in-memory Python dict {"entities": [...], "relationships": [...]} and held by the caller. Nodes are stored as property-graph entities with an id, optional labels (e.g. ["Person"]), and arbitrary properties. Edges carry source_id, target_id, type, and properties. The Neo4j driver (neo4j_store.py:96-100) normalizes Node/Relationship objects to plain dicts with underscored identity keys (_labels, _element_id, _type) so user properties are never shadowed. Embeddings sit alongside the graph in a separate vector store (semantica/vector_store/), supporting FAISS, Qdrant, Weaviate, Pinecone, Milvus, PgVector, and SQLiteVec. The ContextRetriever orchestrates hybrid retrieval: it queries the vector store via cosine similarity and the graph via semantic entity matching / BFS traversal, then merges results with configurable hybrid_alpha weighting — 0 = vector only, 1 = graph only, 0.5 = balanced (context_retriever.py:178). Embeddings for entities are generated via DecisionEmbeddingPipeline which uses the vector store's own embed method.
neo4j-labs/llm-graph-builder
answeredGraph database (Neo4j). All graph data is stored in a Neo4j 5.23+ instance with APOC. The langchain_neo4j.Neo4jGraph wrapper is used for writes (common_fn.py:281-295) and a raw neo4j.GraphDatabase.driver for reads (graph_query.py:26-37).
Node/edge schema. The core labels are:
:Document {fileName, fileSize, fileType, fileSource, status, model, url, createdAt, nodeCount, relationshipCount, ...}— one per uploaded file.:Chunk {id, text, position, length, fileName, content_offset, page_number?, start_time?, end_time?, embedding}— text fragments with vector embeddings.:__Entity__(base label viabaseEntityLabel=True,common_fn.py:285) plus user-defined sub-labels (e.g.:Person:__Entity__,:Organization:__Entity__).:__Community__ {id, level, summary, title, embedding, community_rank, weight}— hierarchical community summaries (communities.py).
Relationships include PART_OF (Chunk→Document), FIRST_CHUNK / NEXT_CHUNK (Document→Chunk / Chunk→Chunk), HAS_ENTITY (Chunk→Entity), SIMILAR (Chunk↔Chunk, KNN-based), IN_COMMUNITY (Entity→Community), and PARENT_COMMUNITY (Community→Community for hierarchy).
How embeddings sit next to the graph. Three Neo4j vector indexes exist (graphDB_dataAccess.py:551-583; post_processing.py:41-45): vector on Chunk.embedding, entity_vector on __Entity__.embedding, and community_vector on __Community__.embedding. All use cosine similarity. Embeddings are generated by a pluggable model (OpenAI, Gemini, Bedrock Titan, or sentence-transformers/all-MiniLM-L6-v2) and stored as array properties directly on the node, not in a separate vector store. A keyword full-text index on Chunk.text and a community_keyword full-text index on __Community__.summary support hybrid search. The KNN update (graphDB_dataAccess.py:184-195) places SIMILAR relationships between chunks whose embedding cosine similarity meets the KNN_MIN_SCORE threshold.
OSU-NLP-Group/HippoRAG
answeredGraph database vs files vs in-memory. The graph is stored as a serialized iGraph pickle file on disk (graph.pickle) and loaded into memory on init (HippoRAG.py:436–473). initialize_graph checks for an existing pickle; if found it calls ig.Graph.Read_Pickle and logs node/edge counts. Otherwise it creates an empty ig.Graph(directed=...). save_igraph (1679–1694) writes the graph back via graph.write_pickle using atomic file-replacement with a .tmp staging file and FileLock for concurrency safety.
Node/edge schema. The graph has two node types: phrase (entity) nodes (strings from NER, hashed as entity-<md5>) and passage (chunk) nodes (chunk content hashed as chunk-<md5>). Edges carry typed attributes: weight (float), edge_kind (string like "fact", "passage", "synonym", or "fact+synonym"), fact_source_counts (dict mapping chunk IDs to occurrence counts), synonym_score (float), passage_source (chunk ID), and source_key/target_key (HippoRAG.py:1561, 1588–1596). The graph is undirected by default (is_directed_graph=False in BaseConfig:176), but supports directed mode.
How embeddings sit next to the graph. Embeddings are stored separately in three EmbeddingStore instances — one per namespace (chunk, entity, fact) — by default as local Parquet files with columns hash_id, content, embedding (src/hipporag/embedding_store.py:183–187). The factory get_embedding_store (271–301) also supports Qdrant, ChromaDB, and Milvus backends. At retrieval time, prepare_retrieval_objects (HippoRAG.py:1750–1816) loads all embeddings into NumPy arrays indexed by node key. The graph and embeddings are linked by shared hash IDs: a phrase node in the graph has name = entity-<md5> which is the same key used in entity_embedding_store.
gusye1234/nano-graphrag
answeredGraph storage. The default graph backend is NetworkXStorage (_storage/gdb_networkx.py), an in-memory networkx.Graph that is serialized to a GraphML XML file on disk (graph_{namespace}.graphml) at index_done_callback time (gdb_networkx.py:80-98). On startup, if the GraphML file exists it is deserialized back. Nodes are stored as NetworkX node attributes: each upsert_node call adds a node with attributes dict containing entity_type, description, source_id, and after clustering a serialized JSON clusters field. Edges are NetworkX edge attributes with weight, description, source_id, order. An optional Neo4j backend (gdb_neo4j.py) mirrors the same operations via Cypher queries, storing all nodes under a single label derived from the working directory path with community IDs stored in a communityIds array property.
Key-value storage. JsonKVStorage (_storage/kv_json.py) stores Python dicts as flat JSON files under the working directory: kv_store_full_docs.json, kv_store_text_chunks.json, kv_store_llm_response_cache.json, kv_store_community_reports.json. These are loaded on init and flushed to disk at callbacks.
Vector storage. The default vector DB is NanoVectorDBStorage (_storage/vdb_nanovectordb.py), backed by the nano-vectordb library which persists to a JSON file (vdb_entities.json or vdb_chunks.json). Two vector DBs are created per GraphRAG instance: entities_vdb stores entity embeddings (content = entity_name + description, with entity_name as a meta field) for local-mode queries, and optionally chunks_vdb stores chunk embeddings for naive RAG mode. The embedding function (default OpenAI text-embedding-3-small, 1536-dim) is called in batches of embedding_batch_num (32), and concurrency is limited by embedding_func_max_async (16). Embeddings are concatenated into numpy arrays and upserted into NanoVectorDB, which supports cosine-similarity query with a better_than_threshold filter (default 0.2).
pingcap/autoflow
answeredThe knowledge graph is stored in TiDB (a distributed SQL database with vector support) across two per-knowledge-base tables.
Dynamic table creation: Each knowledge base gets dynamically-named tables (entities_{namespace}, relationships_{namespace}). The dynamic_create_models() function in core/autoflow/storage/graph_store/tidb_graph_store.py:40-149 generates SQLAlchemy model classes at runtime using pytidb for the core library. The backend's equivalent (backend/app/models/entity.py:38-96, backend/app/models/relationship.py:33-110) uses SQLModel with singleflight_cache-decorated factory functions so each (namespace, dimension) pair creates the model only once.
Entity schema (entities_{namespace}): id (UUID or auto-increment int), name (varchar 512), description (Text), meta (JSON), entity_type (enum: original/synopsis), embedding / description_vec (Vector column, dimension configurable per KB, with HNSW cosine-distance index), meta_vec (second vector column for metadata, also HNSW-indexed), created_at, updated_at. The synopsis type also stores a synopsis_info JSON column referencing groups of entity IDs (backend/app/models/entity.py:59-67).
Relationship schema (relationships_{namespace}): id, description (Text), meta (JSON), weight (int, default 0), source_entity_id (FK → entities), target_entity_id (FK → entities), description_vec (Vector column with HNSW index), chunk_id (UUID, nullable), document_id (int, nullable), last_modified_at (datetime). Foreign keys use SQLAlchemy relationships with lazy="joined" so source/target entities are eager-loaded (backend/app/models/relationship.py:58-68, 90-106).
Vector indexes: Both description_vec and meta_vec on entities, and description_vec on relationships, have HNSW indexes using TiDB Vector's VectorAdaptor with cosine distance (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:129-153). Embeddings are stored inline in the same row as the entity/relationship — there is no separate vector store for the graph.
Graph storage vs chunk storage: Chunks live in a separate chunks_{namespace} table with their own text, embedding, and metadata. The graph is cross-referenced to chunks via relationship.chunk_id and relationship.document_id. During retrieval, the get_chunks_by_relationships() method (tidb_graph_store.py:1052-1127) can jump from graph relationships back to the original chunk text.
trustgraph-ai/trustgraph
answeredGraph database backends. Triples are stored in one of four backends, all consuming Triples messages:
- Neo4j (
storage/triples/neo4j/write.py): Nodes stored as:Node {uri, workspace, collection}for URI entities and:Literal {value, workspace, collection}for literals. Relationships as:Rel {uri, workspace, collection}.MERGEprovides idempotent creation. Compound indexes on(workspace, collection, uri). - Memgraph (
storage/triples/memgraph/write.py): Identical Cypher-based schema. - FalkorDB (
storage/triples/falkordb/write.py): Same Cypher schema. - Cassandra (
storage/triples/cassandra/write.py:168-191):EntityCentricKnowledgeGraphwithinsert(collection, s, p, o, g, otype, dtype, lang). Each workspace is a separate keyspace.
Embedding stores alongside the graph. Entity vectors are stored separately in:
- Qdrant (
storage/graph_embeddings/qdrant/write.py:100-133): Collections namedt_{workspace}_{collection}_{dim}. Entity URI stored as payload, created lazily with cosine distance. - Milvus / Pinecone: Similar patterns.
Document/row embeddings go through storage/doc_embeddings/* and storage/row_embeddings/*. A keyword (BM25) index (storage/kw_index/fts5/) uses SQLite FTS5.
Document content is stored by the librarian service. Chunk provenance triples are emitted to GRAPH_SOURCE (chunking/recursive/chunker.py:203-223).
Isolation is per-workspace and per-collection. Deletion is collection-level only (storage/triples/neo4j/write.py:298-327).
zilliztech/vector-graph-rag
answeredThe graph is stored in Milvus, managed by MilvusStore (storage/milvus.py:35-1478). Three collections are created per named graph prefix: {prefix}_vgrag_entities, {prefix}_vgrag_relations, and {prefix}_vgrag_passages. There is no graph database.
Each collection has the same schema: a VARCHAR primary key id (string, UUID or user-provided), a FLOAT_VECTOR field named vector storing embeddings, and a VARCHAR field named text. Collections use the IP (inner product) metric type with configurable index type ("AUTOINDEX" default) (storage/milvus.py:209-225). Additional metadata fields are stored via Milvus's dynamic schema.
Entity records store relation_ids and passage_ids in metadata. Relation records store entity_ids (head and tail), passage_ids, plus structured triplet fields subject, predicate, object (storage/milvus.py:394-438). Passage records store entity_ids, relation_ids, and arbitrary user metadata (e.g. filterable fields like source). The graph adjacency is thus encoded entirely as ID lists in the metadata of each node — there are no explicit edge objects beyond what is captured in the relation collection.
Embeddings sit inside each record as the vector field, generated by EmbeddingModel (storage/embeddings.py) which supports multiple providers (OpenAI, HuggingFace, Voyage, Jina, Ollama, Mistral, Google). Entity, relation, and passage embeddings are all stored in the same vector index, albeit in separate collections. The default embedding model is text-embedding-3-large (3072 dimensions) in Settings, while the factory function defaults to text-embedding-3-small (as noted in CLAUDE.md).
← How is the knowledge graph extracted from documents? · Are communities, summaries or hierarchies built over the graph? →