pingcap/autoflow
Self-hosted Graph RAG chat app on TiDB that extracts a DSPy knowledge graph per chunk and fuses it with vector search.
Overview
AutoFlow is PingCAP’s open-source conversational knowledge base, the code behind tidb.ai. It is a full application, not a library: a Next.js frontend, a FastAPI backend, Celery workers, Redis, and TiDB as the only database. TiDB holds the relational data, the chunk vectors and the knowledge graph. Admins create knowledge bases, attach data sources (file uploads, single web pages, sitemaps), pick LLM, embedding and reranker models in the UI, and configure one or more chat engines. Users then chat with a streaming answer that cites its source documents.
The Graph RAG part is LlamaIndex plumbing plus DSPy programs. Every chunk goes through a DSPy extraction program that returns entities with descriptions, a JSON “covariates” tree per entity, and directed relationships. Entities and relationships are rows in per-knowledge-base TiDB tables with cosine HNSW vector indexes. There is no graph database, no Cypher and no community detection. At query time the graph is used as extra context: a weighted, vector-guided walk over relationship embeddings produces a small subgraph that is rendered into the prompt alongside the vector-retrieved chunks.
The prompts are written for database documentation (“identify all entities related to database technologies”), which shows its origin as TiDB’s docs assistant. The repo also contains core/, an early autoflow-ai Python package built on pytidb (version 0.0.2.dev5); the deployed product is backend/ plus frontend/, and that is what this page describes.
Architecture
flowchart LR
DS["Data sources: files, URLs, sitemaps"] --> CW["Celery: build_index_for_document"]
CW --> CH["Chunking: SentenceSplitter / MarkdownNodeParser"]
CH --> VT["TiDB chunks_ns (vector)"]
CH --> KW["Celery: build_kg_index_for_chunk"]
KW --> EX["DSPy Extractor"]
EX --> GS["TiDBGraphStore.save"]
GS --> GT["TiDB entities_ns / relationships_ns"]
UI["Next.js chat UI"] --> API["FastAPI /chats"]
API --> CF["ChatFlow"]
CF --> KG["KnowledgeGraphFusionRetriever"]
KG --> GT
CF --> CR["ChunkFusionRetriever + reranker"]
CR --> VT
CF --> LLM["LLM answer (streamed)"]
| Component | Path | Role |
|---|---|---|
| Index tasks | backend/app/tasks/build_index.py |
Celery tasks: vector index per document, then one KG task per chunk |
| Index service | backend/app/rag/build_index.py |
Chunking rules per MIME type, LlamaIndex VectorStoreIndex insert, KG insert |
| Extractor | backend/app/rag/indices/knowledge_graph/extractor.py |
DSPy ExtractGraphTriplet and ExtractCovariate signatures |
| Graph store | backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py |
Entity resolution, LLM merge, weighted relationship search |
| KG retrievers | backend/app/rag/retrievers/knowledge_graph/ |
Per-KB retriever, multi-KB fusion with optional query decomposition |
| Chunk retrievers | backend/app/rag/retrievers/chunk/ |
TiDB vector search, metadata post-filter, reranker |
| Chat flow | backend/app/rag/chat/chat_flow.py |
KG search, question refinement, clarification, answer streaming |
| Chat engine config | backend/app/rag/chat/config.py |
Per-engine prompts, KG depth, intent search, reranker, external engine |
| Frontend | frontend/app/ |
Next.js 15 chat UI and admin console (KBs, models, graph editor, evaluations) |
How a request flows
Indexing a document:
import_documents_from_kb_datasourceloads documents from a data source, commits each one and queuesbuild_index_for_document(tasks/knowledge_base.py).- The task chunks the document with LlamaIndex’s
SentenceSplitter(plain text) or a customMarkdownNodeParser, embeds the chunks and inserts them into the KB’schunks_<namespace>table (build_index.py). If the KB has the knowledge-graph index method, onebuild_kg_index_for_chunktask is queued per chunk (tasks/build_index.py). - The DSPy
Extractorruns two LLM calls per chunk: one for entities and relationships, one for per-entity covariates (extractor.py). Relationship endpoints missing from the entity list become stub entities markedneed-revised. TiDBGraphStore.saveskips the chunk if any relationship already carries itschunk_id. Otherwise each entity goes throughget_or_create_entity, and each relationship is stored with an embedding of “source(description) -> relation -> target(description)” (tidb_graph_store.py).
Answering a chat message (chat_flow.py):
- Graph search.
RetrieveFlow.search_knowledge_graphbuilds aKnowledgeGraphFusionRetrieverover the engine’s knowledge bases (retrieve_flow.py). Withusing_intent_search(on by default), a DSPyQueryDecomposerfirst splits the question into sub-questions. Each sub-question runs against each KB in parallel, and results are merged by relationship text, summing weights (fusion_retriever.py). - Weighted walk.
retrieve_with_weighttakes the 10 relationships closest to the query embedding, then fordepth - 1rounds (default depth 2) searches relationships whose source is an already visited entity, quota-filling from cosine-distance bands 0-0.25, 0.25-0.35, 0.35-0.45 and 0.45-0.55. It also adds the two most similar “synopsis” entities (tidb_graph_store.py). - Context. The subgraph is rendered through
intent_graph_knowledgeornormal_graph_knowledgeJinja prompts, which tell the model to prefer higher-weight and more recently modified relationships. - Refine and clarify. The fast LLM rewrites the question using the graph context and chat history. Optionally a clarification prompt can stop the turn and ask the user a question.
- Chunks.
ChunkFusionRetrieverruns TiDB vector search on the refined question, applies metadata filters and the configured reranker, and keepstop_k. - Answer. LlamaIndex’s response synthesizer streams the answer from the main LLM with graph context, chunks and the original question in
text_qa_prompt. Source documents and the retrieved graph are saved with the message.
Key components
Entity resolution
get_or_create_entity embeds “name: description”, finds the closest existing entity with the same name, and treats it as a match when cosine distance is under 0.1 (tidb_graph_store.py). An identical match is reused. A close but different one goes to a DSPy MergeEntities program that either returns a merged description and metadata (the row is updated and re-embedded) or declines, in which case a new entity with the same name is created. Same-name entities with different meanings can therefore coexist. Note that the entity_type condition in that filter is joined with Python and instead of SQLAlchemy and_, so only the name condition reaches SQL.
Relationship scoring
search_relationships_weight pre-selects the limit * 10 nearest relationships by cosine distance, applies meta filters and the visited-entity constraint, then scores 1 / distance plus a piecewise weight bonus and, optionally, an in/out-degree term (tidb_graph_store.py, helpers.py). Relationships with negative weight are excluded, which gives admins a way to suppress bad edges. Nothing in the backend raises weights automatically; they change only through the graph editor.
Admin graph editing and synopsis entities
TiDBGraphEditor lets staff edit entities and relationships (re-embedding on save, with a staff action log) and create “synopsis” entities that summarise a group of related entities. Synopsis entities are the closest thing AutoFlow has to a summary layer, and they are manual.
Deletion and re-indexing
Deleting a document deletes relationships whose chunk_id belongs to it, then entities with no remaining relationships, then the chunks (repositories/graph.py). Entities that survive keep any description merged in from the deleted text. Re-indexing a document re-queues the Celery tasks, but because save skips chunks that already have relationships, a completed chunk is not re-extracted unless its edges were removed first.
Extending it
- Models. LLM providers include OpenAI, OpenAI-compatible, Azure OpenAI, Gemini, Vertex AI, Bedrock, Ollama and Gitee AI; embedding and reranker models (including a bundled local embedding/reranker service) are configured as database rows from the admin UI.
- Prompts. Every chat engine stores its own extraction-independent prompts (graph context, condense, clarify, QA, follow-up questions, goal generation).
- Compiled DSPy programs.
SimpleGraphExtractorandQueryDecomposercan load compiled DSPy programs from disk, so the extraction and decomposition prompts can be optimised offline. - External engine. A chat engine can delegate to an external streaming API with goal generation, and a post-verification URL can be called after each answer.
- Retrieval API.
POST /retrieve/chunksandPOST /retrieve/knowledge_graphexpose the two retrieval paths separately for use outside the chat UI.
Running it
The supported path is Docker Compose: backend (FastAPI), background (supervisord running two Celery workers and Flower), frontend (Next.js), redis, and an optional local-embedding-reranker. TiDB is not in the compose file; you point .env at a TiDB Cloud Serverless cluster or a self-managed TiDB with vector search. Alembic migrations create the shared tables, and per-KB chunks_*, entities_* and relationships_* tables are created with vector dimensions taken from the KB’s embedding model. Langfuse tracing is built in.
Strengths and caveats
- Strength: one database. Chunks, vectors, graph, chat history and admin data all live in TiDB, with HNSW cosine indexes on entity and relationship embeddings.
- Strength: a real product. Multi-KB chat engines, streaming with stage annotations, citations, a graph editor, evaluations and per-chunk index status with retry make it usable as-is.
- Strength: entity merging with judgement. Name plus description similarity plus an LLM merge decision is more careful than exact-string deduplication.
- Caveat: TiDB lock-in. The graph store, vector store and migrations are TiDB-specific; there is no other backend.
- Caveat: expensive indexing. Two LLM calls per chunk for extraction, plus possible merge calls per entity, plus embeddings for every entity, metadata tree and relationship.
- Caveat: no global queries. No communities, no hierarchical summaries; the graph contributes up to a few dozen relationships of context per sub-question.
- Caveat: domain-tuned prompts. The extraction and merge prompts are written for database documentation and need rewriting for other domains.
- Caveat: weak re-indexing. Re-indexing does not refresh the graph for already-extracted chunks, and deleting documents leaves merged entity descriptions behind.
Sources: code at c4cb19d, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit c4cb19d. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredGraph construction is a two-phase DSPy-based LLM extraction pipeline operating per chunk.
Chunking: Documents are split via LlamaIndex's SentenceSplitter (plain text, default 1024 tokens with 20-token overlap) or a custom MarkdownNodeParser (markdown), configured per knowledge base (backend/app/rag/build_index.py:83-133). Each chunk becomes a TextNode stored in a per-KB chunks_{namespace} TiDB table (backend/app/models/chunk.py:53-61, fields: text, meta, embedding, document_id, index_status).
Extraction (core library — core/autoflow/knowledge_graph/): The SimpleKGExtractor (core/autoflow/knowledge_graph/extractors/simple.py:11-23) runs two DSPy programs in sequence. First, KnowledgeGraphExtractor (core/autoflow/knowledge_graph/programs/extract_graph.py:118-147) calls an LLM with the ExtractKnowledgeGraph signature (extract_graph.py:85-115) — a structured DSPy Signature whose prompt asks to identify "all meaningful entities and relationships" from database documentation, consolidating similar entities and capturing directionality. The LLM returns a PredictKnowledgeGraph (lists of PredictEntity and PredictRelationship). Second, EntityCovariateExtractor (core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85) re-reads the text with the entity list and enriches each entity's .meta with a covariate JSON tree (topic + attributes).
Extraction (backend — backend/app/rag/indices/knowledge_graph/extractor.py): The SimpleGraphExtractor does the same but additionally creates fallback entities for any source/target mentioned in a relationship that wasn't in the entity list, marking them "need-revised" (extractor.py:173-208).
Entity resolution / deduplication: In TiDBGraphStore.get_or_create_entity() (backend/.../tidb_graph_store.py:367-466), before inserting, the store computes the cosine distance between the new entity's description embedding and existing entities with the same name. If distance is below a threshold (description_similarity_threshold, default 0.9 → max cosine distance 0.1), it considers them the same. If the descriptions/metadata differ, it calls MergeEntities — a DSPy program (tidb_graph_store.py:53-81) that asks the LLM to decide whether two same-named entities are genuinely identical and, if so, merge their descriptions and metadata. Otherwise a new row is created. Relationships also use description-vector similarity to resolve source/target entities (tidb_graph_store.py:231-249).
Schema/Ontology: There is no fixed ontology. The graph is an open, untyped property graph. Entities have name, description, meta (covariates dict), and an optional entity_type (original or synopsis). Relationships have source_entity, target_entity, description, weight, and meta (including chunk_id/document_id).
Where and how is the graph stored?
answeredThe knowledge graph is stored in TiDB (a distributed SQL database with vector support) across two per-knowledge-base tables.
Dynamic table creation: Each knowledge base gets dynamically-named tables (entities_{namespace}, relationships_{namespace}). The dynamic_create_models() function in core/autoflow/storage/graph_store/tidb_graph_store.py:40-149 generates SQLAlchemy model classes at runtime using pytidb for the core library. The backend's equivalent (backend/app/models/entity.py:38-96, backend/app/models/relationship.py:33-110) uses SQLModel with singleflight_cache-decorated factory functions so each (namespace, dimension) pair creates the model only once.
Entity schema (entities_{namespace}): id (UUID or auto-increment int), name (varchar 512), description (Text), meta (JSON), entity_type (enum: original/synopsis), embedding / description_vec (Vector column, dimension configurable per KB, with HNSW cosine-distance index), meta_vec (second vector column for metadata, also HNSW-indexed), created_at, updated_at. The synopsis type also stores a synopsis_info JSON column referencing groups of entity IDs (backend/app/models/entity.py:59-67).
Relationship schema (relationships_{namespace}): id, description (Text), meta (JSON), weight (int, default 0), source_entity_id (FK → entities), target_entity_id (FK → entities), description_vec (Vector column with HNSW index), chunk_id (UUID, nullable), document_id (int, nullable), last_modified_at (datetime). Foreign keys use SQLAlchemy relationships with lazy="joined" so source/target entities are eager-loaded (backend/app/models/relationship.py:58-68, 90-106).
Vector indexes: Both description_vec and meta_vec on entities, and description_vec on relationships, have HNSW indexes using TiDB Vector's VectorAdaptor with cosine distance (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:129-153). Embeddings are stored inline in the same row as the entity/relationship — there is no separate vector store for the graph.
Graph storage vs chunk storage: Chunks live in a separate chunks_{namespace} table with their own text, embedding, and metadata. The graph is cross-referenced to chunks via relationship.chunk_id and relationship.document_id. During retrieval, the get_chunks_by_relationships() method (tidb_graph_store.py:1052-1127) can jump from graph relationships back to the original chunk text.
Are communities, summaries or hierarchies built over the graph?
insufficient evidenceAutoFlow does not implement automated community detection (e.g., Leiden algorithm), hierarchical summarization, or any Graph ML-based partitioning of the graph. A search of all Python files in both core/ and backend/ for terms like "community", "leiden", "hierarch", "cluster", and "partition" — none appear in the graph-related code.
What exists instead is a synopsis entity mechanism, exposed as an admin API (POST /admin/knowledge_bases/{kb_id}/graph/entities/synopsis, backend/app/api/admin_routes/knowledge_base/graph/routes.py:56-78). The TiDBGraphEditor.create_synopsis_entity() method (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py:163-222) allows a user (or an external process) to manually create an entity of type synopsis, giving it a name, description, topic, and a list of related entity IDs. It then creates "is a part of" relationships linking the synopsis entity to each listed entity. During graph retrieval, synopsis entities are fetched via fetch_similar_entities(..., entity_type=EntityType.synopsis) and added to the result set (tidb_graph_store.py:551-554). This provides a form of user-driven grouping, but it is not an automatic community-detection pass.
There is also no evidence of hierarchical summaries being computed over entity groups. The graph is flat — entities of type original are the only automatically-created kind.
How does query-time retrieval use the graph?
answeredQuery-time retrieval is a hybrid process that combines graph traversal, vector similarity, and chunk retrieval, assembled into context for the LLM.
How it's invoked (chat flow): In ChatFlow._builtin_chat() (backend/app/rag/chat/chat_flow.py:200-258), the pipeline is: (1) search the knowledge graph, (2) refine the user question using graph context, (3) optionally clarify, (4) search for relevant chunks, (5) generate the answer with both graph context and chunks.
Graph retrieval modes: The KnowledgeGraphOption config has two modes controlled by using_intent_search (backend/app/rag/chat/config.py:53-55):
- Normal mode (
using_intent_search=False): TheKnowledgeGraphSimpleRetrieverwrapsTiDBGraphStore.retrieve_with_weight()which performs a multi-depth traversal. At depth 0, it vector-search-ranks relationships against the query embedding using a weighted score:alpha * (1/embedding_distance) + weight_score + degree_score. Then for subsequent depths, it progressively explores neighbors, breaking the search into distance-range bands ([(0,0.25), (0.25,0.35), (0.35,0.45), (0.45,0.55)]) with decreasing search ratios, ensuring broad coverage. Synopsis entities are appended viafetch_similar_entities(entity_type=synopsis). Results are rendered into a prompt template (normal_graph_knowledge) listing entities and relationships. - Intent mode (
using_intent_search=True): TheKnowledgeGraphFusionRetrieverfirst decomposes the question into sub-queries (viaDecomposedFactorsschema, found inschema.py:113-115), retrieves per sub-query, then fuses the results — merging entities by set union and relationships byrag_descriptionkey (accumulating weight). Each sub-graph is rendered viaintent_graph_knowledgeprompt template.
Combining graph + vector hits: After graph retrieval, the refined question (enriched with graph context) is used by ChunkFusionRetriever (retrieve_flow.py:134-142) to search the vector index for relevant chunks. Both graph context (as a formatted string knowledge_graph_context) and chunks (as NodeWithScore[]) are passed to the final LLM prompt (_generate_answer at chat_flow.py:460-524) via the text_qa_template prompt, which includes {graph_knowledges} as a partial format variable.
Fusion across multiple knowledge bases: The KnowledgeGraphFusionRetriever._knowledge_graph_fusion() (backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py:97-137) merges entities/relationships from multiple KB retrievers into a single result with child subgraph metadata.
No true local/global modes: Unlike Microsoft's GraphRAG, AutoFlow does not have a separate "global" search over community summaries. It always retrieves graph context by traversing from query-similar relationships.
How are updates and incremental indexing handled?
answeredAutoFlow supports per-document incremental indexing and re-indexing of failed tasks, but has no incremental graph update for individual entity/relationship changes at the chunk level.
Adding new documents: Documents are added per knowledge base. The import_documents_for_knowledge_base Celery task feeds documents through IndexService.build_vector_index_for_document() and IndexService.build_kg_index_for_chunk() (backend/app/rag/build_index.py:51-159). For vector indexing, each document's content is chunked by LlamaIndex's SentenceSplitter/MarkdownNodeParser, embedded, and inserted into the chunks_{namespace} table. For the KG index, each chunk becomes a TextNode, entities and relationships are extracted via the LLM extraction pipeline, and saved to the entities/relationships tables.
Re-indexing (update): Documents can be reindexed via the admin API (POST /admin/knowledge_bases/{kb_id}/documents/reindex, backend/app/api/admin_routes/knowledge_base/document/routes.py:134-222). The process sets the document's index_status to PENDING and enqueues build_index_for_document (vector) and build_kg_index_for_chunk (KG) Celery tasks. It only reindexes documents whose status is FAILED (or when reindex_completed_task=True). For chunks that already have COMPLETED KG index, the reindex is skipped (routes.py:201-215).
Deletion: Deleting a document (DELETE /admin/knowledge_bases/{kb_id}/documents/{document_id}, routes.py:91-131) first calls graph_repo.delete_document_relationships() which deletes all relationships whose chunk_id belongs to that document's chunks (backend/app/repositories/graph.py:57-64). Then delete_orphaned_entities() removes any entity left with no relationships. Finally the chunk entries and document itself are deleted.
Idempotency / duplicate detection: KnowledgeGraphIndex.add_chunk() (core/autoflow/knowledge_graph/index.py:38-51) checks list_relationships(chunk_id=...) before extraction — if any relationships already exist for that chunk, it skips. The backend TiDBGraphStore.save() similarly checks if any relationship with the same chunk_id meta already exists (tidb_graph_store.py:206-213). This prevents double-extraction.
No true incremental graph merge: When a new chunk is added and its entities overlap with existing ones, they are resolved via the embedding-similarity + LLM-merge mechanism in get_or_create_entity() — existing entities are updated (merged) rather than duplicated. But there is no mechanism to re-extract or update relationships for previously-indexed chunks when a new document changes the overall understanding of an entity.
Namespace isolation: Each knowledge base gets its own set of tables (chunks_{kb_id}, entities_{kb_id}, relationships_{kb_id}), so operations on one KB do not affect others.
How are LLM cost and latency controlled during indexing and query?
answeredAutoFlow employs several strategies to control LLM cost and latency.
Two-tier LLM model: The ChatEngineConfig distinguishes between a primary llm (used for the final answer synthesis) and a fast_llm (used for cheaper operations like question refinement, clarification, and intent generation) (backend/app/rag/chat/config.py:133-158). The fast LLM is typically a smaller/cheaper model, while the primary LLM handles the expensive QA generation.
Embedding model for similarity, not LLM for most retrieval: The core graph retrieval is driven by vector similarity (cosine distance on TiDB HNSW indexes), not by LLM calls. The WeightedGraphRetriever (core/autoflow/knowledge_graph/retrievers/weighted.py) and TiDBGraphStore.search_relationships_weight() (backend/.../tidb_graph_store.py:632-753) use embedding distance as the primary signal, only ranking results with the weighted formula. The LLM is only invoked during extraction (indexing time) and during the final response generation (query time).
Entity deduplication with conditional LLM merge: The get_or_create_entity() method first checks cosine similarity of description embeddings. Only when an entity with the same name is found and its description/metadata are different does it call the MergeEntities DSPy program (an LLM call) — otherwise it simply reuses the existing row (tidb_graph_store.py:397-442). This avoids an LLM call for duplicate entities.
Singleflight cache for model factories: The @singleflight_cache decorator on get_dynamic_entity_model, get_dynamic_relationship_model, and get_dynamic_chunk_model ensures that per-namespace model creation (which involves schema reflection) runs only once across concurrent threads (backend/app/utils/singleflight_cache.py:5-45).
Semantic cache for external engines: The SemanticCacheManager (backend/app/rag/semantic_cache/base.py:91-199) stores question-answer pairs in TiDB with vector embeddings. Before invoking the external chat engine, the system checks if a semantically similar question has been answered before, using cosine-distance search (threshold 0.5) followed by a DSPy QASemanticSearchModule to classify exact/similar/no-match. This avoids redundant LLM calls for repeated queries.
Indexing batching via Celery: Indexing tasks are dispatched to Celery workers (build_index_for_document.delay(...), build_kg_index_for_chunk.delay(...)), enabling parallel processing of documents/chunks across workers (backend/app/tasks/build_index.py).
No per-chunk LLM call deduplication: Every chunk that has not been indexed (no existing relationships for its chunk_id) gets a fresh extraction call, which is the main cost driver at indexing time.
Configurable model choice: The system uses OpenAI by default (GPT-4o for indexing extraction, text-embedding-3-small for embeddings), but the LLMConfig and EmbeddingModelConfig abstractions support pluggable providers (core/autoflow/configs/models/llms/base.py:6-13, embeddings/base.py), so operators can choose cheaper models.