# Where and how is the graph stored?

> Graph RAG — a good answer covers: Graph database vs files vs in-memory; node/edge schema; how embeddings sit next to the graph.

Canonical page: https://llms-technical-reviews.com/graph-rag/q/graph-storage/

## Verdict

For a production database that holds both the graph and its vectors, [LLM Graph Builder](/p/llm-graph-builder/) (Neo4j) and [AutoFlow](/p/autoflow/) (TiDB) are the most integrated. For a portable file you can open and version, use [GraphRAG](/p/graphrag/)'s tables or [graphify](/p/graphify/)'s `graph.json`.

**Files and in-memory graphs.** GraphRAG has no graph store. It writes entities, relationships and communities as Parquet or CSV tables and puts the embeddings in LanceDB. [nano-graphrag](/p/nano-graphrag/) keeps a NetworkX graph in a GraphML file, with JSON key-value files and nano-vectordb beside it. Neo4j is an option. graphify serialises NetworkX to node-link JSON and stores no embeddings at all. Its Neo4j and FalkorDB pushes are one-way exports. [HippoRAG](/p/hipporag/) pickles an igraph graph of entity and passage nodes, written atomically under a file lock. Facts are not graph nodes. They live, with chunks and entities, in three Parquet embedding stores.

**Pluggable backends.** [LightRAG](/p/lightrag/) has four storage roles. The default graph is a single-writer GraphML file, and Neo4j, Memgraph, PostgreSQL, MongoDB or OpenSearch can replace it. Entities and relations are also stored as vectors. In [Semantica](/p/semantica/) the graph is a plain Python dict unless you configure a `GraphStore` (Neo4j, FalkorDB, Apache AGE or Neptune). Vectors go to a separate store. [TrustGraph](/p/trustgraph/) stores RDF triples in Cassandra, Neo4j, Memgraph or FalkorDB, and entity vectors in Qdrant, Milvus or Pinecone. Its Neo4j writer drops the named-graph field, so the provenance features need Cassandra.

**One store for graph and vectors.** LLM Graph Builder keeps Document, Chunk, entity and community nodes in Neo4j, each with an embedding property and a vector index. AutoFlow creates entity and relationship tables per knowledge base, with HNSW vector columns in the same rows. TiDB is its only backend. [Vector Graph RAG](/p/vector-graph-rag/) has no graph database at all. It uses three Milvus collections and stores adjacency as ID lists in dynamic fields. Each hop is an `id in [...]` query.

Pick: GraphRAG or HippoRAG for a batch artifact you rebuild and inspect offline.
Pick: LightRAG, TrustGraph or LLM Graph Builder for a live graph in a real database.
Pick: Vector Graph RAG or AutoFlow if you already run Milvus or TiDB and want no second system.

## Per-project answers

### Graphify-Labs/graphify (answered)

**In-memory NetworkX.** The graph lives in memory as a `networkx.Graph` (or `nx.DiGraph` when directed=True) throughout building, clustering, analysis, and query serving. `build_from_json()` (build.py:877–1038) constructs it: nodes are added with all attributes except `id` unpacked as kwargs, and edges are added with their metadata.

**Serialised to JSON on disk.** The canonical serialisation is `graph.json`, written by `to_json()` in `export.py` (line 272). It uses NetworkX's `node_link_data` format: a JSON object with top-level `nodes` (array of dicts) and `links` (array of edge dicts). Community IDs are stamped onto each node at write time. A safety check (`export.py:278–331`) refuses to overwrite an existing graph with a shrinking node count unless `force=True`. A backup is snapshotted before overwrite when the graph has semantic or curated content (`backup_if_protected`, export.py:42–104).

**Optional graph database export.** The `exporters/graphdb.py` module provides `push_to_neo4j()` (graphdb.py:22–98) which uses MERGE statements to upsert nodes and edges into a running Neo4j instance. Node labels are derived from `file_type` (capitalised). A similar `push_to_falkordb` exists. These are pure exports — there is no database-backed storage layer.

**No vector store or embeddings.** The README (line 38) states explicitly: "No embeddings, no vector store: a real graph you traverse." Node attributes include `norm_label` (diacritic-stripped lowercased label) for textual matching, but no embedding vectors are stored.

**Global graph.** A `~/.graphify/global-graph.json` accumulates graphs from multiple repos via `global_add()` (global_graph.py:79). It is a flat union with repo-tagged nodes, not a federated index.


Citations: [graphify/build.py:877-1066](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/build.py#L877-L1066) · [graphify/export.py:272-331](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/export.py#L272-L331) · [graphify/exporters/graphdb.py:22-98](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/exporters/graphdb.py#L22-L98) · [graphify/global_graph.py:49-70](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/global_graph.py#L49-L70) · [graphify/build.py:1038-1066](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/build.py#L1038-L1066)

### HKUDS/LightRAG (answered)

**Graph storage backends.** The graph is stored via a pluggable `BaseGraphStorage` (`lightrag/base.py:671`). Implementations registered in `lightrag/kg/__init__.py:12-23` include NetworkXStorage (default, file-based), Neo4JStorage (Neo4j graph DB), PGGraphStorage / PGTableGraphStorage (PostgreSQL), MongoGraphStorage, MemgraphStorage, and OpenSearchGraphStorage.

**Default: NetworkX (in-memory + file).** `NetworkXStorage` (`lightrag/kg/networkx_impl.py:38-85`) keeps the entire graph in process memory using `networkx.Graph`. It persists by rewriting one GraphML file (`graph_<namespace>.graphml`) per workspace via `index_done_callback`. It declares `requires_single_writer = True`, meaning only one process may mutate it at a time, enforced by the pipeline's `busy` reservation or `LightRAG._admin_write_gate`. Concurrent readers detect peer writes through a two-channel fence: file `(st_mtime_ns, st_size)` fingerprint plus a `storage_updated` flag. A commit publishes the whole namespace, so partial mutations from other in-flight writers can be published unintentionally — accepted as a documented residue.

**Node/edge schema.** Nodes and edges are NetworkX `graph.add_node(id, **data)` / `graph.add_edge(src, tgt, **data)` calls with arbitrary string-keyed dictionaries (`lightrag/kg/networkx_impl.py:820-821, 844-845`). Node data typically includes `entity_name`, `entity_type`, `description`, `source_id`, `file_path`, `created_at`. Edge data includes `src_id`, `tgt_id`, `weight`, `source_id`, `keywords`, `description`, `file_path`, `created_at`. XML attribute validation (`validate_xml_attributes`) is applied before every write since GraphML serialization can't handle nested structures.

**Embeddings: separate vector stores.** Graph nodes (entities) and edges (relations) are **dually stored**: once in the graph for structural traversal, and once in entity/relation **vector databases** (`entities_vdb` / `relationships_vdb`). The default `NanoVectorDBStorage` (`lightrag/kg/nano_vector_db_impl.py:62-97`) is an in-memory store serialized to a JSON file per workspace. Each entity vector record holds `{content, entity_name, source_id, description, entity_type, file_path}` (`lightrag/operate.py:1805-1813`). Each relation vector record holds `{src_id, tgt_id, source_id, content, keywords, description, weight, file_path}` (`lightrag/operate.py:2324-2335`). Other supported vector DBs: Milvus, PGVector, Faiss, Qdrant, Mongo, OpenSearch. `NoopVectorDBStorage` disables vector storage.


Citations: [lightrag/kg/networkx_impl.py:38-85](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/networkx_impl.py#L38-L85) · [lightrag/kg/__init__.py:1-48](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/__init__.py#L1-L48) · [lightrag/base.py:671-714](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/base.py#L671-L714) · [lightrag/kg/nano_vector_db_impl.py:62-97](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/kg/nano_vector_db_impl.py#L62-L97) · [lightrag/operate.py:1800-1835](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L1800-L1835) · [lightrag/operate.py:2310-2340](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L2310-L2340)

### microsoft/graphrag (answered)

The graph is stored as **tabular files** (parquet/CSV) or in Azure CosmosDB — not as a graph database. The base `Storage` abstraction provides key-value file access with implementations for local disk, Azure Blob, Azure Cosmos DB, and in-memory. On top sits `TableProvider` for table-level operations: `read_dataframe`, `write_dataframe`, and `open` for streaming row operations.

**Table schema.** Six tables are produced with fixed column schemas in `schemas.py:70-159`. Entities: id, title, type, description, text_unit_ids, frequency, degree. Relationships: id, source, target, description, weight, combined_degree, text_unit_ids. Communities: id, community, level, parent, children, entity_ids, relationship_ids, text_unit_ids. TextUnits: id, text, n_tokens, document_id, entity_ids, relationship_ids.

**Embeddings.** The `generate_text_embeddings` workflow (at `generate_text_embeddings.py:52`) creates embeddings for three configured fields: text_unit text, entity title-description, and community report full content. The `embed_text` operation reads rows, batches them (configurable `batch_size`/`batch_max_tokens`), calls the embedding model, and loads vectors into a pluggable `VectorStore`. Optional snapshot tables store vectors alongside base tables.

**No graph DB at query time.** Query sessions load entities/relationships/communities into Python dataclass objects from serialized tables at startup.


Citations: [packages/graphrag-storage/graphrag_storage/storage.py:13-135](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-storage/graphrag_storage/storage.py#L13-L135) · [packages/graphrag/graphrag/data_model/schemas.py:70-159](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/data_model/schemas.py#L70-L159) · [packages/graphrag/graphrag/index/workflows/generate_text_embeddings.py:52-69](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/workflows/generate_text_embeddings.py#L52-L69) · [packages/graphrag/graphrag/index/operations/embed_text/embed_text.py:23-153](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/operations/embed_text/embed_text.py#L23-L153) · [packages/graphrag/graphrag/data_model/entity.py:12-69](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/data_model/entity.py#L12-L69)

### semantica-agi/semantica (answered)

The graph is stored via a pluggable backend architecture. The core `GraphStore` class in `semantica/graph_store/graph_store.py` defines the interface (NodeManager, RelationshipManager, QueryEngine), and concrete implementations exist for **Neo4j** (`Neo4jStore` in neo4j_store.py, using Cypher queries via the `neo4j` driver), **Apache AGE** (`ApacheAgeStore` in age_store.py, wrapping openCypher through PostgreSQL psycopg2), **FalkorDB** (`FalkorDBStore` in falkordb_store.py), and **Amazon Neptune**. The builder's `build()` method optionally persists the graph to a store via `self.graph_store.add_nodes(resolved_entities)` and `self.graph_store.add_edges(formatted_edges)` (graph_builder.py:1012-1044). When no database backend is configured, the graph is returned as an in-memory Python dict `{"entities": [...], "relationships": [...]}` and held by the caller. Nodes are stored as property-graph entities with an `id`, optional labels (e.g. `["Person"]`), and arbitrary `properties`. Edges carry `source_id`, `target_id`, `type`, and `properties`. The Neo4j driver (`neo4j_store.py:96-100`) normalizes Node/Relationship objects to plain dicts with underscored identity keys (`_labels`, `_element_id`, `_type`) so user properties are never shadowed. **Embeddings sit alongside the graph in a separate vector store** (`semantica/vector_store/`), supporting FAISS, Qdrant, Weaviate, Pinecone, Milvus, PgVector, and SQLiteVec. The `ContextRetriever` orchestrates hybrid retrieval: it queries the vector store via cosine similarity and the graph via semantic entity matching / BFS traversal, then merges results with configurable `hybrid_alpha` weighting — 0 = vector only, 1 = graph only, 0.5 = balanced (`context_retriever.py:178`). Embeddings for entities are generated via `DecisionEmbeddingPipeline` which uses the vector store's own `embed` method.


Citations: [semantica/graph_store/graph_store.py:1-100](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/graph_store/graph_store.py#L1-L100) · [semantica/graph_store/neo4j_store.py:88-103](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/graph_store/neo4j_store.py#L88-L103) · [semantica/kg/graph_builder.py:1010-1045](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/graph_builder.py#L1010-L1045) · [semantica/graph_store/age_store.py:1-60](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/graph_store/age_store.py#L1-L60) · [semantica/context/context_retriever.py:150-180](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/context/context_retriever.py#L150-L180) · [semantica/vector_store/__init__.py:1-50](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/vector_store/__init__.py#L1-L50)

### neo4j-labs/llm-graph-builder (answered)

**Graph database (Neo4j).** All graph data is stored in a Neo4j 5.23+ instance with APOC. The `langchain_neo4j.Neo4jGraph` wrapper is used for writes (`common_fn.py:281-295`) and a raw `neo4j.GraphDatabase.driver` for reads (`graph_query.py:26-37`).

**Node/edge schema.** The core labels are:
- `:Document {fileName, fileSize, fileType, fileSource, status, model, url, createdAt, nodeCount, relationshipCount, ...}` — one per uploaded file.
- `:Chunk {id, text, position, length, fileName, content_offset, page_number?, start_time?, end_time?, embedding}` — text fragments with vector embeddings.
- `:__Entity__` (base label via `baseEntityLabel=True`, `common_fn.py:285`) plus user-defined sub-labels (e.g. `:Person:__Entity__`, `:Organization:__Entity__`).
- `:__Community__ {id, level, summary, title, embedding, community_rank, weight}` — hierarchical community summaries (`communities.py`).

Relationships include `PART_OF` (Chunk→Document), `FIRST_CHUNK` / `NEXT_CHUNK` (Document→Chunk / Chunk→Chunk), `HAS_ENTITY` (Chunk→__Entity__), `SIMILAR` (Chunk↔Chunk, KNN-based), `IN_COMMUNITY` (__Entity__→__Community__), and `PARENT_COMMUNITY` (__Community__→__Community__ for hierarchy).

**How embeddings sit next to the graph.** Three Neo4j vector indexes exist (`graphDB_dataAccess.py:551-583`; `post_processing.py:41-45`): `vector` on `Chunk.embedding`, `entity_vector` on `__Entity__.embedding`, and `community_vector` on `__Community__.embedding`. All use cosine similarity. Embeddings are generated by a pluggable model (OpenAI, Gemini, Bedrock Titan, or sentence-transformers/all-MiniLM-L6-v2) and stored as array properties directly on the node, not in a separate vector store. A `keyword` full-text index on `Chunk.text` and a `community_keyword` full-text index on `__Community__.summary` support hybrid search. The KNN update (`graphDB_dataAccess.py:184-195`) places `SIMILAR` relationships between chunks whose embedding cosine similarity meets the `KNN_MIN_SCORE` threshold.


Citations: [backend/src/graphDB_dataAccess.py:540-585](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L540-L585) · [backend/src/post_processing.py:19-45](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/post_processing.py#L19-L45) · [backend/src/shared/constants.py:159-207](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L159-L207) · [backend/src/make_relationships.py:149-171](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L149-L171) · [backend/src/graphDB_dataAccess.py:151-196](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L151-L196)

### OSU-NLP-Group/HippoRAG (answered)

**Graph database vs files vs in-memory.** The graph is stored as a serialized **iGraph pickle file** on disk (`graph.pickle`) and loaded into memory on init (HippoRAG.py:436–473). `initialize_graph` checks for an existing pickle; if found it calls `ig.Graph.Read_Pickle` and logs node/edge counts. Otherwise it creates an empty `ig.Graph(directed=...)`. `save_igraph` (1679–1694) writes the graph back via `graph.write_pickle` using atomic file-replacement with a `.tmp` staging file and `FileLock` for concurrency safety.

**Node/edge schema.** The graph has two node types: **phrase (entity) nodes** (strings from NER, hashed as `entity-<md5>`) and **passage (chunk) nodes** (chunk content hashed as `chunk-<md5>`). Edges carry typed attributes: `weight` (float), `edge_kind` (string like `"fact"`, `"passage"`, `"synonym"`, or `"fact+synonym"`), `fact_source_counts` (dict mapping chunk IDs to occurrence counts), `synonym_score` (float), `passage_source` (chunk ID), and `source_key`/`target_key` (HippoRAG.py:1561, 1588–1596). The graph is undirected by default (`is_directed_graph=False` in BaseConfig:176), but supports directed mode.

**How embeddings sit next to the graph.** Embeddings are stored separately in **three `EmbeddingStore` instances** — one per namespace (`chunk`, `entity`, `fact`) — by default as local Parquet files with columns `hash_id`, `content`, `embedding` (src/hipporag/embedding_store.py:183–187). The factory `get_embedding_store` (271–301) also supports Qdrant, ChromaDB, and Milvus backends. At retrieval time, `prepare_retrieval_objects` (HippoRAG.py:1750–1816) loads all embeddings into NumPy arrays indexed by node key. The graph and embeddings are linked by shared hash IDs: a phrase node in the graph has `name = entity-<md5>` which is the same key used in `entity_embedding_store`.


Citations: [src/hipporag/HippoRAG.py:436-473](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L436-L473) · [src/hipporag/HippoRAG.py:1679-1694](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1679-L1694) · [src/hipporag/HippoRAG.py:1561-1596](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1561-L1596) · [src/hipporag/embedding_store.py:99-112](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/embedding_store.py#L99-L112) · [src/hipporag/embedding_store.py:271-301](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/embedding_store.py#L271-L301)

### gusye1234/nano-graphrag (answered)

**Graph storage.** The default graph backend is `NetworkXStorage` (`_storage/gdb_networkx.py`), an in-memory `networkx.Graph` that is serialized to a GraphML XML file on disk (`graph_{namespace}.graphml`) at `index_done_callback` time (`gdb_networkx.py:80-98`). On startup, if the GraphML file exists it is deserialized back. Nodes are stored as NetworkX node attributes: each `upsert_node` call adds a node with attributes dict containing `entity_type`, `description`, `source_id`, and after clustering a serialized JSON `clusters` field. Edges are NetworkX edge attributes with `weight`, `description`, `source_id`, `order`. An optional Neo4j backend (`gdb_neo4j.py`) mirrors the same operations via Cypher queries, storing all nodes under a single label derived from the working directory path with community IDs stored in a `communityIds` array property.

**Key-value storage.** `JsonKVStorage` (`_storage/kv_json.py`) stores Python dicts as flat JSON files under the working directory: `kv_store_full_docs.json`, `kv_store_text_chunks.json`, `kv_store_llm_response_cache.json`, `kv_store_community_reports.json`. These are loaded on init and flushed to disk at callbacks.

**Vector storage.** The default vector DB is `NanoVectorDBStorage` (`_storage/vdb_nanovectordb.py`), backed by the `nano-vectordb` library which persists to a JSON file (`vdb_entities.json` or `vdb_chunks.json`). Two vector DBs are created per GraphRAG instance: `entities_vdb` stores entity embeddings (content = entity_name + description, with `entity_name` as a meta field) for local-mode queries, and optionally `chunks_vdb` stores chunk embeddings for naive RAG mode. The embedding function (default OpenAI `text-embedding-3-small`, 1536-dim) is called in batches of `embedding_batch_num` (32), and concurrency is limited by `embedding_func_max_async` (16). Embeddings are concatenated into numpy arrays and upserted into NanoVectorDB, which supports cosine-similarity query with a `better_than_threshold` filter (default 0.2).


Citations: [nano_graphrag/_storage/gdb_networkx.py:19-98](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_storage/gdb_networkx.py#L19-L98) · [nano_graphrag/_storage/gdb_neo4j.py:19-92](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_storage/gdb_neo4j.py#L19-L92) · [nano_graphrag/_storage/vdb_nanovectordb.py:12-68](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_storage/vdb_nanovectordb.py#L12-L68) · [nano_graphrag/graphrag.py:196-217](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/graphrag.py#L196-L217)

### pingcap/autoflow (answered)

The knowledge graph is stored in **TiDB** (a distributed SQL database with vector support) across two per-knowledge-base tables.

**Dynamic table creation:** Each knowledge base gets dynamically-named tables (`entities_{namespace}`, `relationships_{namespace}`). The `dynamic_create_models()` function in `core/autoflow/storage/graph_store/tidb_graph_store.py:40-149` generates SQLAlchemy model classes at runtime using `pytidb` for the core library. The backend's equivalent (`backend/app/models/entity.py:38-96`, `backend/app/models/relationship.py:33-110`) uses SQLModel with `singleflight_cache`-decorated factory functions so each (namespace, dimension) pair creates the model only once.

**Entity schema (`entities_{namespace}`):** `id` (UUID or auto-increment int), `name` (varchar 512), `description` (Text), `meta` (JSON), `entity_type` (enum: `original`/`synopsis`), `embedding` / `description_vec` (Vector column, dimension configurable per KB, with HNSW cosine-distance index), `meta_vec` (second vector column for metadata, also HNSW-indexed), `created_at`, `updated_at`. The synopsis type also stores a `synopsis_info` JSON column referencing groups of entity IDs (`backend/app/models/entity.py:59-67`).

**Relationship schema (`relationships_{namespace}`):** `id`, `description` (Text), `meta` (JSON), `weight` (int, default 0), `source_entity_id` (FK → entities), `target_entity_id` (FK → entities), `description_vec` (Vector column with HNSW index), `chunk_id` (UUID, nullable), `document_id` (int, nullable), `last_modified_at` (datetime). Foreign keys use SQLAlchemy relationships with `lazy="joined"` so source/target entities are eager-loaded (`backend/app/models/relationship.py:58-68, 90-106`).

**Vector indexes:** Both `description_vec` and `meta_vec` on entities, and `description_vec` on relationships, have HNSW indexes using TiDB Vector's `VectorAdaptor` with cosine distance (`backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:129-153`). Embeddings are stored inline in the same row as the entity/relationship — there is no separate vector store for the graph.

**Graph storage vs chunk storage:** Chunks live in a separate `chunks_{namespace}` table with their own text, embedding, and metadata. The graph is cross-referenced to chunks via `relationship.chunk_id` and `relationship.document_id`. During retrieval, the `get_chunks_by_relationships()` method (`tidb_graph_store.py:1052-1127`) can jump from graph relationships back to the original chunk text.


Citations: [core/autoflow/storage/graph_store/tidb_graph_store.py:40-149](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/storage/graph_store/tidb_graph_store.py#L40-L149) · [backend/app/models/entity.py:38-96](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/models/entity.py#L38-L96) · [backend/app/models/relationship.py:33-110](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/models/relationship.py#L33-L110) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:117-161](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L117-L161)

### trustgraph-ai/trustgraph (answered)

**Graph database backends.** Triples are stored in one of four backends, all consuming `Triples` messages:

- **Neo4j** (`storage/triples/neo4j/write.py`): Nodes stored as `:Node {uri, workspace, collection}` for URI entities and `:Literal {value, workspace, collection}` for literals. Relationships as `:Rel {uri, workspace, collection}`. `MERGE` provides idempotent creation. Compound indexes on `(workspace, collection, uri)`.
- **Memgraph** (`storage/triples/memgraph/write.py`): Identical Cypher-based schema.
- **FalkorDB** (`storage/triples/falkordb/write.py`): Same Cypher schema.
- **Cassandra** (`storage/triples/cassandra/write.py:168-191`): `EntityCentricKnowledgeGraph` with `insert(collection, s, p, o, g, otype, dtype, lang)`. Each workspace is a separate keyspace.

**Embedding stores alongside the graph.** Entity vectors are stored separately in:
- **Qdrant** (`storage/graph_embeddings/qdrant/write.py:100-133`): Collections named `t_{workspace}_{collection}_{dim}`. Entity URI stored as payload, created lazily with cosine distance.
- **Milvus** / **Pinecone**: Similar patterns.

Document/row embeddings go through `storage/doc_embeddings/*` and `storage/row_embeddings/*`. A **keyword (BM25) index** (`storage/kw_index/fts5/`) uses SQLite FTS5.

**Document content** is stored by the **librarian** service. Chunk provenance triples are emitted to `GRAPH_SOURCE` (`chunking/recursive/chunker.py:203-223`).

**Isolation** is per-workspace and per-collection. Deletion is collection-level only (`storage/triples/neo4j/write.py:298-327`).


Citations: [trustgraph-flow/trustgraph/storage/triples/neo4j/write.py:143-200](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/storage/triples/neo4j/write.py#L143-L200) · [trustgraph-flow/trustgraph/storage/triples/neo4j/write.py:206-232](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/storage/triples/neo4j/write.py#L206-L232) · [trustgraph-flow/trustgraph/storage/triples/cassandra/write.py:168-191](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/storage/triples/cassandra/write.py#L168-L191) · [trustgraph-flow/trustgraph/storage/graph_embeddings/qdrant/write.py:90-133](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/storage/graph_embeddings/qdrant/write.py#L90-L133) · [trustgraph-flow/trustgraph/extract/kg/ontology/vector_store.py:23-52](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/vector_store.py#L23-L52) · [trustgraph-flow/trustgraph/chunking/recursive/chunker.py:203-223](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/chunking/recursive/chunker.py#L203-L223)

### zilliztech/vector-graph-rag (answered)

The graph is stored in **Milvus**, managed by `MilvusStore` (`storage/milvus.py:35-1478`). Three collections are created per named graph prefix: `{prefix}_vgrag_entities`, `{prefix}_vgrag_relations`, and `{prefix}_vgrag_passages`. There is no graph database.

Each collection has the same schema: a `VARCHAR` primary key `id` (string, UUID or user-provided), a `FLOAT_VECTOR` field named `vector` storing embeddings, and a `VARCHAR` field named `text`. Collections use the `IP` (inner product) metric type with configurable index type ("AUTOINDEX" default) (`storage/milvus.py:209-225`). Additional metadata fields are stored via Milvus's dynamic schema.

**Entity records** store `relation_ids` and `passage_ids` in metadata. **Relation records** store `entity_ids` (head and tail), `passage_ids`, plus structured triplet fields `subject`, `predicate`, `object` (`storage/milvus.py:394-438`). **Passage records** store `entity_ids`, `relation_ids`, and arbitrary user metadata (e.g. filterable fields like `source`). The graph adjacency is thus encoded entirely as **ID lists in the metadata** of each node — there are no explicit edge objects beyond what is captured in the relation collection.

Embeddings sit **inside each record** as the `vector` field, generated by `EmbeddingModel` (`storage/embeddings.py`) which supports multiple providers (OpenAI, HuggingFace, Voyage, Jina, Ollama, Mistral, Google). Entity, relation, and passage embeddings are all stored in the same vector index, albeit in separate collections. The default embedding model is text-embedding-3-large (3072 dimensions) in Settings, while the factory function defaults to text-embedding-3-small (as noted in CLAUDE.md).


Citations: [src/vector_graph_rag/storage/milvus.py:35-80](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/storage/milvus.py#L35-L80) · [src/vector_graph_rag/storage/milvus.py:191-243](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/storage/milvus.py#L191-L243) · [src/vector_graph_rag/storage/milvus.py:285-340](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/storage/milvus.py#L285-L340) · [src/vector_graph_rag/storage/embeddings.py:1-80](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/storage/embeddings.py#L1-L80)
