# pingcap/autoflow

> Self-hosted Graph RAG chat app on TiDB that extracts a DSPy knowledge graph per chunk and fuses it with vector search.

- Category: [Graph RAG](https://llms-technical-reviews.com/graph-rag/)
- Repository: https://github.com/pingcap/autoflow (reviewed at commit `c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2`, 2026-04-27)
- Stars: 2978 · Language: TypeScript · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/autoflow/

## Overview

AutoFlow is PingCAP's open-source conversational knowledge base, the code behind tidb.ai. It is a full application, not a library: a Next.js frontend, a FastAPI backend, Celery workers, Redis, and TiDB as the only database. TiDB holds the relational data, the chunk vectors and the knowledge graph. Admins create knowledge bases, attach data sources (file uploads, single web pages, sitemaps), pick LLM, embedding and reranker models in the UI, and configure one or more chat engines. Users then chat with a streaming answer that cites its source documents.

The Graph RAG part is LlamaIndex plumbing plus DSPy programs. Every chunk goes through a DSPy extraction program that returns entities with descriptions, a JSON "covariates" tree per entity, and directed relationships. Entities and relationships are rows in per-knowledge-base TiDB tables with cosine HNSW vector indexes. There is no graph database, no Cypher and no community detection. At query time the graph is used as extra context: a weighted, vector-guided walk over relationship embeddings produces a small subgraph that is rendered into the prompt alongside the vector-retrieved chunks.

The prompts are written for database documentation ("identify all entities related to database technologies"), which shows its origin as TiDB's docs assistant. The repo also contains `core/`, an early `autoflow-ai` Python package built on `pytidb` (version 0.0.2.dev5); the deployed product is `backend/` plus `frontend/`, and that is what this page describes.

## Architecture

```mermaid
flowchart LR
  DS["Data sources: files, URLs, sitemaps"] --> CW["Celery: build_index_for_document"]
  CW --> CH["Chunking: SentenceSplitter / MarkdownNodeParser"]
  CH --> VT["TiDB chunks_ns (vector)"]
  CH --> KW["Celery: build_kg_index_for_chunk"]
  KW --> EX["DSPy Extractor"]
  EX --> GS["TiDBGraphStore.save"]
  GS --> GT["TiDB entities_ns / relationships_ns"]
  UI["Next.js chat UI"] --> API["FastAPI /chats"]
  API --> CF["ChatFlow"]
  CF --> KG["KnowledgeGraphFusionRetriever"]
  KG --> GT
  CF --> CR["ChunkFusionRetriever + reranker"]
  CR --> VT
  CF --> LLM["LLM answer (streamed)"]
```

| Component | Path | Role |
|---|---|---|
| Index tasks | `backend/app/tasks/build_index.py` | Celery tasks: vector index per document, then one KG task per chunk |
| Index service | `backend/app/rag/build_index.py` | Chunking rules per MIME type, LlamaIndex `VectorStoreIndex` insert, KG insert |
| Extractor | `backend/app/rag/indices/knowledge_graph/extractor.py` | DSPy `ExtractGraphTriplet` and `ExtractCovariate` signatures |
| Graph store | `backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py` | Entity resolution, LLM merge, weighted relationship search |
| KG retrievers | `backend/app/rag/retrievers/knowledge_graph/` | Per-KB retriever, multi-KB fusion with optional query decomposition |
| Chunk retrievers | `backend/app/rag/retrievers/chunk/` | TiDB vector search, metadata post-filter, reranker |
| Chat flow | `backend/app/rag/chat/chat_flow.py` | KG search, question refinement, clarification, answer streaming |
| Chat engine config | `backend/app/rag/chat/config.py` | Per-engine prompts, KG depth, intent search, reranker, external engine |
| Frontend | `frontend/app/` | Next.js 15 chat UI and admin console (KBs, models, graph editor, evaluations) |

## How a request flows

Indexing a document:

1. `import_documents_from_kb_datasource` loads documents from a data source, commits each one and queues `build_index_for_document` ([tasks/knowledge_base.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/tasks/knowledge_base.py#L47-L83)).
2. The task chunks the document with LlamaIndex's `SentenceSplitter` (plain text) or a custom `MarkdownNodeParser`, embeds the chunks and inserts them into the KB's `chunks_<namespace>` table ([build_index.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/build_index.py#L51-L136)). If the KB has the knowledge-graph index method, one `build_kg_index_for_chunk` task is queued per chunk ([tasks/build_index.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/tasks/build_index.py#L61-L92)).
3. The DSPy `Extractor` runs two LLM calls per chunk: one for entities and relationships, one for per-entity covariates ([extractor.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/extractor.py#L97-L128)). Relationship endpoints missing from the entity list become stub entities marked `need-revised`.
4. `TiDBGraphStore.save` skips the chunk if any relationship already carries its `chunk_id`. Otherwise each entity goes through `get_or_create_entity`, and each relationship is stored with an embedding of "source(description) -> relation -> target(description)" ([tidb_graph_store.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L198-L282)).

Answering a chat message ([chat_flow.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/chat_flow.py#L200-L258)):

1. **Graph search.** `RetrieveFlow.search_knowledge_graph` builds a `KnowledgeGraphFusionRetriever` over the engine's knowledge bases ([retrieve_flow.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/retrieve/retrieve_flow.py#L81-L99)). With `using_intent_search` (on by default), a DSPy `QueryDecomposer` first splits the question into sub-questions. Each sub-question runs against each KB in parallel, and results are merged by relationship text, summing weights ([fusion_retriever.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py#L97-L137)).
2. **Weighted walk.** `retrieve_with_weight` takes the 10 relationships closest to the query embedding, then for `depth - 1` rounds (default depth 2) searches relationships whose source is an already visited entity, quota-filling from cosine-distance bands 0-0.25, 0.25-0.35, 0.35-0.45 and 0.45-0.55. It also adds the two most similar "synopsis" entities ([tidb_graph_store.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L478-L589)).
3. **Context.** The subgraph is rendered through `intent_graph_knowledge` or `normal_graph_knowledge` Jinja prompts, which tell the model to prefer higher-weight and more recently modified relationships.
4. **Refine and clarify.** The fast LLM rewrites the question using the graph context and chat history. Optionally a clarification prompt can stop the turn and ask the user a question.
5. **Chunks.** `ChunkFusionRetriever` runs TiDB vector search on the refined question, applies metadata filters and the configured reranker, and keeps `top_k`.
6. **Answer.** LlamaIndex's response synthesizer streams the answer from the main LLM with graph context, chunks and the original question in `text_qa_prompt`. Source documents and the retrieved graph are saved with the message.

## Key components

### Entity resolution

`get_or_create_entity` embeds "name: description", finds the closest existing entity with the same name, and treats it as a match when cosine distance is under 0.1 ([tidb_graph_store.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L367-L466)). An identical match is reused. A close but different one goes to a DSPy `MergeEntities` program that either returns a merged description and metadata (the row is updated and re-embedded) or declines, in which case a new entity with the same name is created. Same-name entities with different meanings can therefore coexist. Note that the `entity_type` condition in that filter is joined with Python `and` instead of SQLAlchemy `and_`, so only the name condition reaches SQL.

### Relationship scoring

`search_relationships_weight` pre-selects the `limit * 10` nearest relationships by cosine distance, applies meta filters and the visited-entity constraint, then scores `1 / distance` plus a piecewise weight bonus and, optionally, an in/out-degree term ([tidb_graph_store.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L632-L753), [helpers.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/helpers.py#L9-L67)). Relationships with negative weight are excluded, which gives admins a way to suppress bad edges. Nothing in the backend raises weights automatically; they change only through the graph editor.

### Admin graph editing and synopsis entities

`TiDBGraphEditor` lets staff edit entities and relationships (re-embedding on save, with a staff action log) and create "synopsis" entities that summarise a group of related entities. Synopsis entities are the closest thing AutoFlow has to a summary layer, and they are manual.

### Deletion and re-indexing

Deleting a document deletes relationships whose `chunk_id` belongs to it, then entities with no remaining relationships, then the chunks ([repositories/graph.py](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/repositories/graph.py#L29-L64)). Entities that survive keep any description merged in from the deleted text. Re-indexing a document re-queues the Celery tasks, but because `save` skips chunks that already have relationships, a completed chunk is not re-extracted unless its edges were removed first.

## Extending it

- **Models.** LLM providers include OpenAI, OpenAI-compatible, Azure OpenAI, Gemini, Vertex AI, Bedrock, Ollama and Gitee AI; embedding and reranker models (including a bundled local embedding/reranker service) are configured as database rows from the admin UI.
- **Prompts.** Every chat engine stores its own extraction-independent prompts (graph context, condense, clarify, QA, follow-up questions, goal generation).
- **Compiled DSPy programs.** `SimpleGraphExtractor` and `QueryDecomposer` can load compiled DSPy programs from disk, so the extraction and decomposition prompts can be optimised offline.
- **External engine.** A chat engine can delegate to an external streaming API with goal generation, and a post-verification URL can be called after each answer.
- **Retrieval API.** `POST /retrieve/chunks` and `POST /retrieve/knowledge_graph` expose the two retrieval paths separately for use outside the chat UI.

## Running it

The supported path is Docker Compose: `backend` (FastAPI), `background` (supervisord running two Celery workers and Flower), `frontend` (Next.js), `redis`, and an optional `local-embedding-reranker`. TiDB is not in the compose file; you point `.env` at a TiDB Cloud Serverless cluster or a self-managed TiDB with vector search. Alembic migrations create the shared tables, and per-KB `chunks_*`, `entities_*` and `relationships_*` tables are created with vector dimensions taken from the KB's embedding model. Langfuse tracing is built in.

## Strengths and caveats

- **Strength: one database.** Chunks, vectors, graph, chat history and admin data all live in TiDB, with HNSW cosine indexes on entity and relationship embeddings.
- **Strength: a real product.** Multi-KB chat engines, streaming with stage annotations, citations, a graph editor, evaluations and per-chunk index status with retry make it usable as-is.
- **Strength: entity merging with judgement.** Name plus description similarity plus an LLM merge decision is more careful than exact-string deduplication.
- **Caveat: TiDB lock-in.** The graph store, vector store and migrations are TiDB-specific; there is no other backend.
- **Caveat: expensive indexing.** Two LLM calls per chunk for extraction, plus possible merge calls per entity, plus embeddings for every entity, metadata tree and relationship.
- **Caveat: no global queries.** No communities, no hierarchical summaries; the graph contributes up to a few dozen relationships of context per sub-question.
- **Caveat: domain-tuned prompts.** The extraction and merge prompts are written for database documentation and need rewriting for other domains.
- **Caveat: weak re-indexing.** Re-indexing does not refresh the graph for already-extracted chunks, and deleting documents leaves merged entity descriptions behind.

*Sources: code at c4cb19d, deepwiki-open wiki (12 pages), verified Q&A.*

## How pingcap/autoflow answers the Graph RAG questions

### How is the knowledge graph extracted from documents? (answered)

Graph construction is a two-phase DSPy-based LLM extraction pipeline operating per chunk.

**Chunking:** Documents are split via LlamaIndex's `SentenceSplitter` (plain text, default 1024 tokens with 20-token overlap) or a custom `MarkdownNodeParser` (markdown), configured per knowledge base (`backend/app/rag/build_index.py:83-133`). Each chunk becomes a `TextNode` stored in a per-KB `chunks_{namespace}` TiDB table (`backend/app/models/chunk.py:53-61`, fields: `text`, `meta`, `embedding`, `document_id`, `index_status`).

**Extraction (core library — `core/autoflow/knowledge_graph/`):** The `SimpleKGExtractor` (`core/autoflow/knowledge_graph/extractors/simple.py:11-23`) runs two DSPy programs in sequence. First, `KnowledgeGraphExtractor` (`core/autoflow/knowledge_graph/programs/extract_graph.py:118-147`) calls an LLM with the `ExtractKnowledgeGraph` signature (`extract_graph.py:85-115`) — a structured DSPy Signature whose prompt asks to identify "all meaningful entities and relationships" from database documentation, consolidating similar entities and capturing directionality. The LLM returns a `PredictKnowledgeGraph` (lists of `PredictEntity` and `PredictRelationship`). Second, `EntityCovariateExtractor` (`core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85`) re-reads the text with the entity list and enriches each entity's `.meta` with a covariate JSON tree (topic + attributes).

**Extraction (backend — `backend/app/rag/indices/knowledge_graph/extractor.py`):** The `SimpleGraphExtractor` does the same but additionally creates fallback entities for any source/target mentioned in a relationship that wasn't in the entity list, marking them `"need-revised"` (`extractor.py:173-208`).

**Entity resolution / deduplication:** In `TiDBGraphStore.get_or_create_entity()` (`backend/.../tidb_graph_store.py:367-466`), before inserting, the store computes the cosine distance between the new entity's description embedding and existing entities with the same name. If distance is below a threshold (`description_similarity_threshold`, default 0.9 → max cosine distance 0.1), it considers them the same. If the descriptions/metadata differ, it calls `MergeEntities` — a DSPy program (`tidb_graph_store.py:53-81`) that asks the LLM to decide whether two same-named entities are genuinely identical and, if so, merge their descriptions and metadata. Otherwise a new row is created. Relationships also use description-vector similarity to resolve source/target entities (`tidb_graph_store.py:231-249`).

**Schema/Ontology:** There is no fixed ontology. The graph is an open, untyped property graph. Entities have `name`, `description`, `meta` (covariates dict), and an optional `entity_type` (`original` or `synopsis`). Relationships have `source_entity`, `target_entity`, `description`, `weight`, and `meta` (including `chunk_id`/`document_id`).


Citations: [core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/programs/extract_covariates.py#L56-L85) · [backend/app/rag/indices/knowledge_graph/extractor.py:83-225](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/extractor.py#L83-L225) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:367-477](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L367-L477) · [backend/app/rag/build_index.py:82-158](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/build_index.py#L82-L158) · [core/autoflow/configs/chunkers/text.py:1-12](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/configs/chunkers/text.py#L1-L12)

### Where and how is the graph stored? (answered)

The knowledge graph is stored in **TiDB** (a distributed SQL database with vector support) across two per-knowledge-base tables.

**Dynamic table creation:** Each knowledge base gets dynamically-named tables (`entities_{namespace}`, `relationships_{namespace}`). The `dynamic_create_models()` function in `core/autoflow/storage/graph_store/tidb_graph_store.py:40-149` generates SQLAlchemy model classes at runtime using `pytidb` for the core library. The backend's equivalent (`backend/app/models/entity.py:38-96`, `backend/app/models/relationship.py:33-110`) uses SQLModel with `singleflight_cache`-decorated factory functions so each (namespace, dimension) pair creates the model only once.

**Entity schema (`entities_{namespace}`):** `id` (UUID or auto-increment int), `name` (varchar 512), `description` (Text), `meta` (JSON), `entity_type` (enum: `original`/`synopsis`), `embedding` / `description_vec` (Vector column, dimension configurable per KB, with HNSW cosine-distance index), `meta_vec` (second vector column for metadata, also HNSW-indexed), `created_at`, `updated_at`. The synopsis type also stores a `synopsis_info` JSON column referencing groups of entity IDs (`backend/app/models/entity.py:59-67`).

**Relationship schema (`relationships_{namespace}`):** `id`, `description` (Text), `meta` (JSON), `weight` (int, default 0), `source_entity_id` (FK → entities), `target_entity_id` (FK → entities), `description_vec` (Vector column with HNSW index), `chunk_id` (UUID, nullable), `document_id` (int, nullable), `last_modified_at` (datetime). Foreign keys use SQLAlchemy relationships with `lazy="joined"` so source/target entities are eager-loaded (`backend/app/models/relationship.py:58-68, 90-106`).

**Vector indexes:** Both `description_vec` and `meta_vec` on entities, and `description_vec` on relationships, have HNSW indexes using TiDB Vector's `VectorAdaptor` with cosine distance (`backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:129-153`). Embeddings are stored inline in the same row as the entity/relationship — there is no separate vector store for the graph.

**Graph storage vs chunk storage:** Chunks live in a separate `chunks_{namespace}` table with their own text, embedding, and metadata. The graph is cross-referenced to chunks via `relationship.chunk_id` and `relationship.document_id`. During retrieval, the `get_chunks_by_relationships()` method (`tidb_graph_store.py:1052-1127`) can jump from graph relationships back to the original chunk text.


Citations: [core/autoflow/storage/graph_store/tidb_graph_store.py:40-149](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/storage/graph_store/tidb_graph_store.py#L40-L149) · [backend/app/models/entity.py:38-96](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/models/entity.py#L38-L96) · [backend/app/models/relationship.py:33-110](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/models/relationship.py#L33-L110) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:117-161](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L117-L161)

### Are communities, summaries or hierarchies built over the graph? (insufficient evidence)

AutoFlow does **not** implement automated community detection (e.g., Leiden algorithm), hierarchical summarization, or any Graph ML-based partitioning of the graph. A search of all Python files in both `core/` and `backend/` for terms like "community", "leiden", "hierarch", "cluster", and "partition" — none appear in the graph-related code. 

What exists instead is a **synopsis entity** mechanism, exposed as an admin API (`POST /admin/knowledge_bases/{kb_id}/graph/entities/synopsis`, `backend/app/api/admin_routes/knowledge_base/graph/routes.py:56-78`). The `TiDBGraphEditor.create_synopsis_entity()` method (`backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py:163-222`) allows a **user (or an external process) to manually create** an entity of type `synopsis`, giving it a name, description, topic, and a list of related entity IDs. It then creates "is a part of" relationships linking the synopsis entity to each listed entity. During graph retrieval, synopsis entities are fetched via `fetch_similar_entities(..., entity_type=EntityType.synopsis)` and added to the result set (`tidb_graph_store.py:551-554`). This provides a form of user-driven grouping, but it is not an automatic community-detection pass.

There is also no evidence of hierarchical summaries being computed over entity groups. The graph is flat — entities of type `original` are the only automatically-created kind.


Citations: [backend/app/api/admin_routes/knowledge_base/graph/routes.py:56-78](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/api/admin_routes/knowledge_base/graph/routes.py#L56-L78) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py:163-222](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py#L163-L222) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:547-554](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L547-L554)

### How does query-time retrieval use the graph? (answered)

Query-time retrieval is a **hybrid** process that combines graph traversal, vector similarity, and chunk retrieval, assembled into context for the LLM.

**How it's invoked (chat flow):** In `ChatFlow._builtin_chat()` (`backend/app/rag/chat/chat_flow.py:200-258`), the pipeline is: (1) search the knowledge graph, (2) refine the user question using graph context, (3) optionally clarify, (4) search for relevant chunks, (5) generate the answer with both graph context and chunks.

**Graph retrieval modes:** The `KnowledgeGraphOption` config has two modes controlled by `using_intent_search` (`backend/app/rag/chat/config.py:53-55`):
- **Normal mode** (`using_intent_search=False`): The `KnowledgeGraphSimpleRetriever` wraps `TiDBGraphStore.retrieve_with_weight()` which performs a multi-depth traversal. At depth 0, it vector-search-ranks relationships against the query embedding using a weighted score: `alpha * (1/embedding_distance) + weight_score + degree_score`. Then for subsequent depths, it progressively explores neighbors, breaking the search into distance-range bands (`[(0,0.25), (0.25,0.35), (0.35,0.45), (0.45,0.55)]`) with decreasing search ratios, ensuring broad coverage. Synopsis entities are appended via `fetch_similar_entities(entity_type=synopsis)`. Results are rendered into a prompt template (`normal_graph_knowledge`) listing entities and relationships.
- **Intent mode** (`using_intent_search=True`): The `KnowledgeGraphFusionRetriever` first decomposes the question into sub-queries (via `DecomposedFactors` schema, found in `schema.py:113-115`), retrieves per sub-query, then fuses the results — merging entities by set union and relationships by `rag_description` key (accumulating weight). Each sub-graph is rendered via `intent_graph_knowledge` prompt template.

**Combining graph + vector hits:** After graph retrieval, the refined question (enriched with graph context) is used by `ChunkFusionRetriever` (`retrieve_flow.py:134-142`) to search the vector index for relevant chunks. Both graph context (as a formatted string `knowledge_graph_context`) and chunks (as `NodeWithScore[]`) are passed to the final LLM prompt (`_generate_answer` at `chat_flow.py:460-524`) via the `text_qa_template` prompt, which includes `{graph_knowledges}` as a partial format variable.

**Fusion across multiple knowledge bases:** The `KnowledgeGraphFusionRetriever._knowledge_graph_fusion()` (`backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py:97-137`) merges entities/relationships from multiple KB retrievers into a single result with child subgraph metadata.

**No true local/global modes:** Unlike Microsoft's GraphRAG, AutoFlow does not have a separate "global" search over community summaries. It always retrieves graph context by traversing from query-similar relationships.


Citations: [backend/app/rag/chat/chat_flow.py:200-258](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/chat_flow.py#L200-L258) · [backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py:27-137](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py#L27-L137) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:488-588](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L488-L588) · [backend/app/rag/chat/config.py:53-57](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/config.py#L53-L57) · [core/autoflow/knowledge_graph/retrievers/weighted.py:68-168](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/retrievers/weighted.py#L68-L168)

### How are updates and incremental indexing handled? (answered)

AutoFlow supports **per-document incremental indexing** and **re-indexing of failed tasks**, but has no incremental graph update for individual entity/relationship changes at the chunk level.

**Adding new documents:** Documents are added per knowledge base. The `import_documents_for_knowledge_base` Celery task feeds documents through `IndexService.build_vector_index_for_document()` and `IndexService.build_kg_index_for_chunk()` (`backend/app/rag/build_index.py:51-159`). For vector indexing, each document's content is chunked by LlamaIndex's `SentenceSplitter`/`MarkdownNodeParser`, embedded, and inserted into the `chunks_{namespace}` table. For the KG index, each chunk becomes a `TextNode`, entities and relationships are extracted via the LLM extraction pipeline, and saved to the `entities`/`relationships` tables.

**Re-indexing (update):** Documents can be reindexed via the admin API (`POST /admin/knowledge_bases/{kb_id}/documents/reindex`, `backend/app/api/admin_routes/knowledge_base/document/routes.py:134-222`). The process sets the document's `index_status` to `PENDING` and enqueues `build_index_for_document` (vector) and `build_kg_index_for_chunk` (KG) Celery tasks. It only reindexes documents whose status is `FAILED` (or when `reindex_completed_task=True`). For chunks that already have `COMPLETED` KG index, the reindex is skipped (`routes.py:201-215`).

**Deletion:** Deleting a document (`DELETE /admin/knowledge_bases/{kb_id}/documents/{document_id}`, `routes.py:91-131`) first calls `graph_repo.delete_document_relationships()` which deletes all relationships whose `chunk_id` belongs to that document's chunks (`backend/app/repositories/graph.py:57-64`). Then `delete_orphaned_entities()` removes any entity left with no relationships. Finally the chunk entries and document itself are deleted.

**Idempotency / duplicate detection:** `KnowledgeGraphIndex.add_chunk()` (`core/autoflow/knowledge_graph/index.py:38-51`) checks `list_relationships(chunk_id=...)` before extraction — if any relationships already exist for that chunk, it skips. The backend `TiDBGraphStore.save()` similarly checks if any relationship with the same `chunk_id` meta already exists (`tidb_graph_store.py:206-213`). This prevents double-extraction.

**No true incremental graph merge:** When a new chunk is added and its entities overlap with existing ones, they are resolved via the embedding-similarity + LLM-merge mechanism in `get_or_create_entity()` — existing entities are updated (merged) rather than duplicated. But there is no mechanism to re-extract or update relationships for previously-indexed chunks when a new document changes the overall understanding of an entity.

**Namespace isolation:** Each knowledge base gets its own set of tables (`chunks_{kb_id}`, `entities_{kb_id}`, `relationships_{kb_id}`), so operations on one KB do not affect others.


Citations: [backend/app/api/admin_routes/knowledge_base/document/routes.py:91-131](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/api/admin_routes/knowledge_base/document/routes.py#L91-L131) · [backend/app/api/admin_routes/knowledge_base/document/routes.py:173-222](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/api/admin_routes/knowledge_base/document/routes.py#L173-L222) · [backend/app/repositories/graph.py:45-64](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/repositories/graph.py#L45-L64) · [core/autoflow/knowledge_graph/index.py:38-51](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/index.py#L38-L51) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:198-217](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L198-L217)

### How are LLM cost and latency controlled during indexing and query? (answered)

AutoFlow employs several strategies to control LLM cost and latency.

**Two-tier LLM model:** The `ChatEngineConfig` distinguishes between a primary `llm` (used for the final answer synthesis) and a `fast_llm` (used for cheaper operations like question refinement, clarification, and intent generation) (`backend/app/rag/chat/config.py:133-158`). The fast LLM is typically a smaller/cheaper model, while the primary LLM handles the expensive QA generation.

**Embedding model for similarity, not LLM for most retrieval:** The core graph retrieval is driven by **vector similarity** (cosine distance on TiDB HNSW indexes), not by LLM calls. The `WeightedGraphRetriever` (`core/autoflow/knowledge_graph/retrievers/weighted.py`) and `TiDBGraphStore.search_relationships_weight()` (`backend/.../tidb_graph_store.py:632-753`) use embedding distance as the primary signal, only ranking results with the weighted formula. The LLM is only invoked during extraction (indexing time) and during the final response generation (query time).

**Entity deduplication with conditional LLM merge:** The `get_or_create_entity()` method first checks cosine similarity of description embeddings. Only when an entity with the same name is found *and* its description/metadata are different does it call the `MergeEntities` DSPy program (an LLM call) — otherwise it simply reuses the existing row (`tidb_graph_store.py:397-442`). This avoids an LLM call for duplicate entities.

**Singleflight cache for model factories:** The `@singleflight_cache` decorator on `get_dynamic_entity_model`, `get_dynamic_relationship_model`, and `get_dynamic_chunk_model` ensures that per-namespace model creation (which involves schema reflection) runs only once across concurrent threads (`backend/app/utils/singleflight_cache.py:5-45`).

**Semantic cache for external engines:** The `SemanticCacheManager` (`backend/app/rag/semantic_cache/base.py:91-199`) stores question-answer pairs in TiDB with vector embeddings. Before invoking the external chat engine, the system checks if a semantically similar question has been answered before, using cosine-distance search (threshold 0.5) followed by a DSPy `QASemanticSearchModule` to classify exact/similar/no-match. This avoids redundant LLM calls for repeated queries.

**Indexing batching via Celery:** Indexing tasks are dispatched to Celery workers (`build_index_for_document.delay(...)`, `build_kg_index_for_chunk.delay(...)`), enabling parallel processing of documents/chunks across workers (`backend/app/tasks/build_index.py`).

**No per-chunk LLM call deduplication:** Every chunk that has not been indexed (no existing relationships for its `chunk_id`) gets a fresh extraction call, which is the main cost driver at indexing time.

**Configurable model choice:** The system uses OpenAI by default (GPT-4o for indexing extraction, text-embedding-3-small for embeddings), but the `LLMConfig` and `EmbeddingModelConfig` abstractions support pluggable providers (`core/autoflow/configs/models/llms/base.py:6-13`, `embeddings/base.py`), so operators can choose cheaper models.


Citations: [backend/app/rag/chat/config.py:147-158](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/chat/config.py#L147-L158) · [backend/app/rag/semantic_cache/base.py:91-199](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/semantic_cache/base.py#L91-L199) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:367-477](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L367-L477) · [core/autoflow/knowledge_graph/retrievers/weighted.py:45-68](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/retrievers/weighted.py#L45-L68) · [backend/app/utils/singleflight_cache.py:5-45](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/utils/singleflight_cache.py#L5-L45)
