LLMs Technical Reviews
Home / Graph RAG / autoflow

pingcap/autoflow

Self-hosted Graph RAG chat app on TiDB that extracts a DSPy knowledge graph per chunk and fuses it with vector search.

GitHub ↗★ 3.0kTypeScriptApache-2.0commit c4cb19d · 2026-04-27homepage ↗

Overview

AutoFlow is PingCAP’s open-source conversational knowledge base, the code behind tidb.ai. It is a full application, not a library: a Next.js frontend, a FastAPI backend, Celery workers, Redis, and TiDB as the only database. TiDB holds the relational data, the chunk vectors and the knowledge graph. Admins create knowledge bases, attach data sources (file uploads, single web pages, sitemaps), pick LLM, embedding and reranker models in the UI, and configure one or more chat engines. Users then chat with a streaming answer that cites its source documents.

The Graph RAG part is LlamaIndex plumbing plus DSPy programs. Every chunk goes through a DSPy extraction program that returns entities with descriptions, a JSON “covariates” tree per entity, and directed relationships. Entities and relationships are rows in per-knowledge-base TiDB tables with cosine HNSW vector indexes. There is no graph database, no Cypher and no community detection. At query time the graph is used as extra context: a weighted, vector-guided walk over relationship embeddings produces a small subgraph that is rendered into the prompt alongside the vector-retrieved chunks.

The prompts are written for database documentation (“identify all entities related to database technologies”), which shows its origin as TiDB’s docs assistant. The repo also contains core/, an early autoflow-ai Python package built on pytidb (version 0.0.2.dev5); the deployed product is backend/ plus frontend/, and that is what this page describes.

Architecture

flowchart LR
  DS["Data sources: files, URLs, sitemaps"] --> CW["Celery: build_index_for_document"]
  CW --> CH["Chunking: SentenceSplitter / MarkdownNodeParser"]
  CH --> VT["TiDB chunks_ns (vector)"]
  CH --> KW["Celery: build_kg_index_for_chunk"]
  KW --> EX["DSPy Extractor"]
  EX --> GS["TiDBGraphStore.save"]
  GS --> GT["TiDB entities_ns / relationships_ns"]
  UI["Next.js chat UI"] --> API["FastAPI /chats"]
  API --> CF["ChatFlow"]
  CF --> KG["KnowledgeGraphFusionRetriever"]
  KG --> GT
  CF --> CR["ChunkFusionRetriever + reranker"]
  CR --> VT
  CF --> LLM["LLM answer (streamed)"]
Component Path Role
Index tasks backend/app/tasks/build_index.py Celery tasks: vector index per document, then one KG task per chunk
Index service backend/app/rag/build_index.py Chunking rules per MIME type, LlamaIndex VectorStoreIndex insert, KG insert
Extractor backend/app/rag/indices/knowledge_graph/extractor.py DSPy ExtractGraphTriplet and ExtractCovariate signatures
Graph store backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py Entity resolution, LLM merge, weighted relationship search
KG retrievers backend/app/rag/retrievers/knowledge_graph/ Per-KB retriever, multi-KB fusion with optional query decomposition
Chunk retrievers backend/app/rag/retrievers/chunk/ TiDB vector search, metadata post-filter, reranker
Chat flow backend/app/rag/chat/chat_flow.py KG search, question refinement, clarification, answer streaming
Chat engine config backend/app/rag/chat/config.py Per-engine prompts, KG depth, intent search, reranker, external engine
Frontend frontend/app/ Next.js 15 chat UI and admin console (KBs, models, graph editor, evaluations)

How a request flows

Indexing a document:

  1. import_documents_from_kb_datasource loads documents from a data source, commits each one and queues build_index_for_document (tasks/knowledge_base.py).
  2. The task chunks the document with LlamaIndex’s SentenceSplitter (plain text) or a custom MarkdownNodeParser, embeds the chunks and inserts them into the KB’s chunks_<namespace> table (build_index.py). If the KB has the knowledge-graph index method, one build_kg_index_for_chunk task is queued per chunk (tasks/build_index.py).
  3. The DSPy Extractor runs two LLM calls per chunk: one for entities and relationships, one for per-entity covariates (extractor.py). Relationship endpoints missing from the entity list become stub entities marked need-revised.
  4. TiDBGraphStore.save skips the chunk if any relationship already carries its chunk_id. Otherwise each entity goes through get_or_create_entity, and each relationship is stored with an embedding of “source(description) -> relation -> target(description)” (tidb_graph_store.py).

Answering a chat message (chat_flow.py):

  1. Graph search. RetrieveFlow.search_knowledge_graph builds a KnowledgeGraphFusionRetriever over the engine’s knowledge bases (retrieve_flow.py). With using_intent_search (on by default), a DSPy QueryDecomposer first splits the question into sub-questions. Each sub-question runs against each KB in parallel, and results are merged by relationship text, summing weights (fusion_retriever.py).
  2. Weighted walk. retrieve_with_weight takes the 10 relationships closest to the query embedding, then for depth - 1 rounds (default depth 2) searches relationships whose source is an already visited entity, quota-filling from cosine-distance bands 0-0.25, 0.25-0.35, 0.35-0.45 and 0.45-0.55. It also adds the two most similar “synopsis” entities (tidb_graph_store.py).
  3. Context. The subgraph is rendered through intent_graph_knowledge or normal_graph_knowledge Jinja prompts, which tell the model to prefer higher-weight and more recently modified relationships.
  4. Refine and clarify. The fast LLM rewrites the question using the graph context and chat history. Optionally a clarification prompt can stop the turn and ask the user a question.
  5. Chunks. ChunkFusionRetriever runs TiDB vector search on the refined question, applies metadata filters and the configured reranker, and keeps top_k.
  6. Answer. LlamaIndex’s response synthesizer streams the answer from the main LLM with graph context, chunks and the original question in text_qa_prompt. Source documents and the retrieved graph are saved with the message.

Key components

Entity resolution

get_or_create_entity embeds “name: description”, finds the closest existing entity with the same name, and treats it as a match when cosine distance is under 0.1 (tidb_graph_store.py). An identical match is reused. A close but different one goes to a DSPy MergeEntities program that either returns a merged description and metadata (the row is updated and re-embedded) or declines, in which case a new entity with the same name is created. Same-name entities with different meanings can therefore coexist. Note that the entity_type condition in that filter is joined with Python and instead of SQLAlchemy and_, so only the name condition reaches SQL.

Relationship scoring

search_relationships_weight pre-selects the limit * 10 nearest relationships by cosine distance, applies meta filters and the visited-entity constraint, then scores 1 / distance plus a piecewise weight bonus and, optionally, an in/out-degree term (tidb_graph_store.py, helpers.py). Relationships with negative weight are excluded, which gives admins a way to suppress bad edges. Nothing in the backend raises weights automatically; they change only through the graph editor.

Admin graph editing and synopsis entities

TiDBGraphEditor lets staff edit entities and relationships (re-embedding on save, with a staff action log) and create “synopsis” entities that summarise a group of related entities. Synopsis entities are the closest thing AutoFlow has to a summary layer, and they are manual.

Deletion and re-indexing

Deleting a document deletes relationships whose chunk_id belongs to it, then entities with no remaining relationships, then the chunks (repositories/graph.py). Entities that survive keep any description merged in from the deleted text. Re-indexing a document re-queues the Celery tasks, but because save skips chunks that already have relationships, a completed chunk is not re-extracted unless its edges were removed first.

Extending it

  • Models. LLM providers include OpenAI, OpenAI-compatible, Azure OpenAI, Gemini, Vertex AI, Bedrock, Ollama and Gitee AI; embedding and reranker models (including a bundled local embedding/reranker service) are configured as database rows from the admin UI.
  • Prompts. Every chat engine stores its own extraction-independent prompts (graph context, condense, clarify, QA, follow-up questions, goal generation).
  • Compiled DSPy programs. SimpleGraphExtractor and QueryDecomposer can load compiled DSPy programs from disk, so the extraction and decomposition prompts can be optimised offline.
  • External engine. A chat engine can delegate to an external streaming API with goal generation, and a post-verification URL can be called after each answer.
  • Retrieval API. POST /retrieve/chunks and POST /retrieve/knowledge_graph expose the two retrieval paths separately for use outside the chat UI.

Running it

The supported path is Docker Compose: backend (FastAPI), background (supervisord running two Celery workers and Flower), frontend (Next.js), redis, and an optional local-embedding-reranker. TiDB is not in the compose file; you point .env at a TiDB Cloud Serverless cluster or a self-managed TiDB with vector search. Alembic migrations create the shared tables, and per-KB chunks_*, entities_* and relationships_* tables are created with vector dimensions taken from the KB’s embedding model. Langfuse tracing is built in.

Strengths and caveats

  • Strength: one database. Chunks, vectors, graph, chat history and admin data all live in TiDB, with HNSW cosine indexes on entity and relationship embeddings.
  • Strength: a real product. Multi-KB chat engines, streaming with stage annotations, citations, a graph editor, evaluations and per-chunk index status with retry make it usable as-is.
  • Strength: entity merging with judgement. Name plus description similarity plus an LLM merge decision is more careful than exact-string deduplication.
  • Caveat: TiDB lock-in. The graph store, vector store and migrations are TiDB-specific; there is no other backend.
  • Caveat: expensive indexing. Two LLM calls per chunk for extraction, plus possible merge calls per entity, plus embeddings for every entity, metadata tree and relationship.
  • Caveat: no global queries. No communities, no hierarchical summaries; the graph contributes up to a few dozen relationships of context per sub-question.
  • Caveat: domain-tuned prompts. The extraction and merge prompts are written for database documentation and need rewriting for other domains.
  • Caveat: weak re-indexing. Re-indexing does not refresh the graph for already-extracted chunks, and deleting documents leaves merged entity descriptions behind.

Sources: code at c4cb19d, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit c4cb19d. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Graph construction is a two-phase DSPy-based LLM extraction pipeline operating per chunk.

Chunking: Documents are split via LlamaIndex's SentenceSplitter (plain text, default 1024 tokens with 20-token overlap) or a custom MarkdownNodeParser (markdown), configured per knowledge base (backend/app/rag/build_index.py:83-133). Each chunk becomes a TextNode stored in a per-KB chunks_{namespace} TiDB table (backend/app/models/chunk.py:53-61, fields: text, meta, embedding, document_id, index_status).

Extraction (core library — core/autoflow/knowledge_graph/): The SimpleKGExtractor (core/autoflow/knowledge_graph/extractors/simple.py:11-23) runs two DSPy programs in sequence. First, KnowledgeGraphExtractor (core/autoflow/knowledge_graph/programs/extract_graph.py:118-147) calls an LLM with the ExtractKnowledgeGraph signature (extract_graph.py:85-115) — a structured DSPy Signature whose prompt asks to identify "all meaningful entities and relationships" from database documentation, consolidating similar entities and capturing directionality. The LLM returns a PredictKnowledgeGraph (lists of PredictEntity and PredictRelationship). Second, EntityCovariateExtractor (core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85) re-reads the text with the entity list and enriches each entity's .meta with a covariate JSON tree (topic + attributes).

Extraction (backend — backend/app/rag/indices/knowledge_graph/extractor.py): The SimpleGraphExtractor does the same but additionally creates fallback entities for any source/target mentioned in a relationship that wasn't in the entity list, marking them "need-revised" (extractor.py:173-208).

Entity resolution / deduplication: In TiDBGraphStore.get_or_create_entity() (backend/.../tidb_graph_store.py:367-466), before inserting, the store computes the cosine distance between the new entity's description embedding and existing entities with the same name. If distance is below a threshold (description_similarity_threshold, default 0.9 → max cosine distance 0.1), it considers them the same. If the descriptions/metadata differ, it calls MergeEntities — a DSPy program (tidb_graph_store.py:53-81) that asks the LLM to decide whether two same-named entities are genuinely identical and, if so, merge their descriptions and metadata. Otherwise a new row is created. Relationships also use description-vector similarity to resolve source/target entities (tidb_graph_store.py:231-249).

Schema/Ontology: There is no fixed ontology. The graph is an open, untyped property graph. Entities have name, description, meta (covariates dict), and an optional entity_type (original or synopsis). Relationships have source_entity, target_entity, description, weight, and meta (including chunk_id/document_id).

Where and how is the graph stored?

answered

The knowledge graph is stored in TiDB (a distributed SQL database with vector support) across two per-knowledge-base tables.

Dynamic table creation: Each knowledge base gets dynamically-named tables (entities_{namespace}, relationships_{namespace}). The dynamic_create_models() function in core/autoflow/storage/graph_store/tidb_graph_store.py:40-149 generates SQLAlchemy model classes at runtime using pytidb for the core library. The backend's equivalent (backend/app/models/entity.py:38-96, backend/app/models/relationship.py:33-110) uses SQLModel with singleflight_cache-decorated factory functions so each (namespace, dimension) pair creates the model only once.

Entity schema (entities_{namespace}): id (UUID or auto-increment int), name (varchar 512), description (Text), meta (JSON), entity_type (enum: original/synopsis), embedding / description_vec (Vector column, dimension configurable per KB, with HNSW cosine-distance index), meta_vec (second vector column for metadata, also HNSW-indexed), created_at, updated_at. The synopsis type also stores a synopsis_info JSON column referencing groups of entity IDs (backend/app/models/entity.py:59-67).

Relationship schema (relationships_{namespace}): id, description (Text), meta (JSON), weight (int, default 0), source_entity_id (FK → entities), target_entity_id (FK → entities), description_vec (Vector column with HNSW index), chunk_id (UUID, nullable), document_id (int, nullable), last_modified_at (datetime). Foreign keys use SQLAlchemy relationships with lazy="joined" so source/target entities are eager-loaded (backend/app/models/relationship.py:58-68, 90-106).

Vector indexes: Both description_vec and meta_vec on entities, and description_vec on relationships, have HNSW indexes using TiDB Vector's VectorAdaptor with cosine distance (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:129-153). Embeddings are stored inline in the same row as the entity/relationship — there is no separate vector store for the graph.

Graph storage vs chunk storage: Chunks live in a separate chunks_{namespace} table with their own text, embedding, and metadata. The graph is cross-referenced to chunks via relationship.chunk_id and relationship.document_id. During retrieval, the get_chunks_by_relationships() method (tidb_graph_store.py:1052-1127) can jump from graph relationships back to the original chunk text.

Are communities, summaries or hierarchies built over the graph?

insufficient evidence

AutoFlow does not implement automated community detection (e.g., Leiden algorithm), hierarchical summarization, or any Graph ML-based partitioning of the graph. A search of all Python files in both core/ and backend/ for terms like "community", "leiden", "hierarch", "cluster", and "partition" — none appear in the graph-related code.

What exists instead is a synopsis entity mechanism, exposed as an admin API (POST /admin/knowledge_bases/{kb_id}/graph/entities/synopsis, backend/app/api/admin_routes/knowledge_base/graph/routes.py:56-78). The TiDBGraphEditor.create_synopsis_entity() method (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py:163-222) allows a user (or an external process) to manually create an entity of type synopsis, giving it a name, description, topic, and a list of related entity IDs. It then creates "is a part of" relationships linking the synopsis entity to each listed entity. During graph retrieval, synopsis entities are fetched via fetch_similar_entities(..., entity_type=EntityType.synopsis) and added to the result set (tidb_graph_store.py:551-554). This provides a form of user-driven grouping, but it is not an automatic community-detection pass.

There is also no evidence of hierarchical summaries being computed over entity groups. The graph is flat — entities of type original are the only automatically-created kind.

How does query-time retrieval use the graph?

answered

Query-time retrieval is a hybrid process that combines graph traversal, vector similarity, and chunk retrieval, assembled into context for the LLM.

How it's invoked (chat flow): In ChatFlow._builtin_chat() (backend/app/rag/chat/chat_flow.py:200-258), the pipeline is: (1) search the knowledge graph, (2) refine the user question using graph context, (3) optionally clarify, (4) search for relevant chunks, (5) generate the answer with both graph context and chunks.

Graph retrieval modes: The KnowledgeGraphOption config has two modes controlled by using_intent_search (backend/app/rag/chat/config.py:53-55):

  • Normal mode (using_intent_search=False): The KnowledgeGraphSimpleRetriever wraps TiDBGraphStore.retrieve_with_weight() which performs a multi-depth traversal. At depth 0, it vector-search-ranks relationships against the query embedding using a weighted score: alpha * (1/embedding_distance) + weight_score + degree_score. Then for subsequent depths, it progressively explores neighbors, breaking the search into distance-range bands ([(0,0.25), (0.25,0.35), (0.35,0.45), (0.45,0.55)]) with decreasing search ratios, ensuring broad coverage. Synopsis entities are appended via fetch_similar_entities(entity_type=synopsis). Results are rendered into a prompt template (normal_graph_knowledge) listing entities and relationships.
  • Intent mode (using_intent_search=True): The KnowledgeGraphFusionRetriever first decomposes the question into sub-queries (via DecomposedFactors schema, found in schema.py:113-115), retrieves per sub-query, then fuses the results — merging entities by set union and relationships by rag_description key (accumulating weight). Each sub-graph is rendered via intent_graph_knowledge prompt template.

Combining graph + vector hits: After graph retrieval, the refined question (enriched with graph context) is used by ChunkFusionRetriever (retrieve_flow.py:134-142) to search the vector index for relevant chunks. Both graph context (as a formatted string knowledge_graph_context) and chunks (as NodeWithScore[]) are passed to the final LLM prompt (_generate_answer at chat_flow.py:460-524) via the text_qa_template prompt, which includes {graph_knowledges} as a partial format variable.

Fusion across multiple knowledge bases: The KnowledgeGraphFusionRetriever._knowledge_graph_fusion() (backend/app/rag/retrievers/knowledge_graph/fusion_retriever.py:97-137) merges entities/relationships from multiple KB retrievers into a single result with child subgraph metadata.

No true local/global modes: Unlike Microsoft's GraphRAG, AutoFlow does not have a separate "global" search over community summaries. It always retrieves graph context by traversing from query-similar relationships.

How are updates and incremental indexing handled?

answered

AutoFlow supports per-document incremental indexing and re-indexing of failed tasks, but has no incremental graph update for individual entity/relationship changes at the chunk level.

Adding new documents: Documents are added per knowledge base. The import_documents_for_knowledge_base Celery task feeds documents through IndexService.build_vector_index_for_document() and IndexService.build_kg_index_for_chunk() (backend/app/rag/build_index.py:51-159). For vector indexing, each document's content is chunked by LlamaIndex's SentenceSplitter/MarkdownNodeParser, embedded, and inserted into the chunks_{namespace} table. For the KG index, each chunk becomes a TextNode, entities and relationships are extracted via the LLM extraction pipeline, and saved to the entities/relationships tables.

Re-indexing (update): Documents can be reindexed via the admin API (POST /admin/knowledge_bases/{kb_id}/documents/reindex, backend/app/api/admin_routes/knowledge_base/document/routes.py:134-222). The process sets the document's index_status to PENDING and enqueues build_index_for_document (vector) and build_kg_index_for_chunk (KG) Celery tasks. It only reindexes documents whose status is FAILED (or when reindex_completed_task=True). For chunks that already have COMPLETED KG index, the reindex is skipped (routes.py:201-215).

Deletion: Deleting a document (DELETE /admin/knowledge_bases/{kb_id}/documents/{document_id}, routes.py:91-131) first calls graph_repo.delete_document_relationships() which deletes all relationships whose chunk_id belongs to that document's chunks (backend/app/repositories/graph.py:57-64). Then delete_orphaned_entities() removes any entity left with no relationships. Finally the chunk entries and document itself are deleted.

Idempotency / duplicate detection: KnowledgeGraphIndex.add_chunk() (core/autoflow/knowledge_graph/index.py:38-51) checks list_relationships(chunk_id=...) before extraction — if any relationships already exist for that chunk, it skips. The backend TiDBGraphStore.save() similarly checks if any relationship with the same chunk_id meta already exists (tidb_graph_store.py:206-213). This prevents double-extraction.

No true incremental graph merge: When a new chunk is added and its entities overlap with existing ones, they are resolved via the embedding-similarity + LLM-merge mechanism in get_or_create_entity() — existing entities are updated (merged) rather than duplicated. But there is no mechanism to re-extract or update relationships for previously-indexed chunks when a new document changes the overall understanding of an entity.

Namespace isolation: Each knowledge base gets its own set of tables (chunks_{kb_id}, entities_{kb_id}, relationships_{kb_id}), so operations on one KB do not affect others.

How are LLM cost and latency controlled during indexing and query?

answered

AutoFlow employs several strategies to control LLM cost and latency.

Two-tier LLM model: The ChatEngineConfig distinguishes between a primary llm (used for the final answer synthesis) and a fast_llm (used for cheaper operations like question refinement, clarification, and intent generation) (backend/app/rag/chat/config.py:133-158). The fast LLM is typically a smaller/cheaper model, while the primary LLM handles the expensive QA generation.

Embedding model for similarity, not LLM for most retrieval: The core graph retrieval is driven by vector similarity (cosine distance on TiDB HNSW indexes), not by LLM calls. The WeightedGraphRetriever (core/autoflow/knowledge_graph/retrievers/weighted.py) and TiDBGraphStore.search_relationships_weight() (backend/.../tidb_graph_store.py:632-753) use embedding distance as the primary signal, only ranking results with the weighted formula. The LLM is only invoked during extraction (indexing time) and during the final response generation (query time).

Entity deduplication with conditional LLM merge: The get_or_create_entity() method first checks cosine similarity of description embeddings. Only when an entity with the same name is found and its description/metadata are different does it call the MergeEntities DSPy program (an LLM call) — otherwise it simply reuses the existing row (tidb_graph_store.py:397-442). This avoids an LLM call for duplicate entities.

Singleflight cache for model factories: The @singleflight_cache decorator on get_dynamic_entity_model, get_dynamic_relationship_model, and get_dynamic_chunk_model ensures that per-namespace model creation (which involves schema reflection) runs only once across concurrent threads (backend/app/utils/singleflight_cache.py:5-45).

Semantic cache for external engines: The SemanticCacheManager (backend/app/rag/semantic_cache/base.py:91-199) stores question-answer pairs in TiDB with vector embeddings. Before invoking the external chat engine, the system checks if a semantically similar question has been answered before, using cosine-distance search (threshold 0.5) followed by a DSPy QASemanticSearchModule to classify exact/similar/no-match. This avoids redundant LLM calls for repeated queries.

Indexing batching via Celery: Indexing tasks are dispatched to Celery workers (build_index_for_document.delay(...), build_kg_index_for_chunk.delay(...)), enabling parallel processing of documents/chunks across workers (backend/app/tasks/build_index.py).

No per-chunk LLM call deduplication: Every chunk that has not been indexed (no existing relationships for its chunk_id) gets a fresh extraction call, which is the main cost driver at indexing time.

Configurable model choice: The system uses OpenAI by default (GPT-4o for indexing extraction, text-embedding-3-small for embeddings), but the LLMConfig and EmbeddingModelConfig abstractions support pluggable providers (core/autoflow/configs/models/llms/base.py:6-13, embeddings/base.py), so operators can choose cheaper models.