# infiniflow/ragflow

> Go RAG server with DeepDoc layout parsing, hybrid BM25 + vector search over pluggable engines, and agentic chat modes.

- Category: [RAG engines](https://llms-technical-reviews.com/rag/)
- Repository: https://github.com/infiniflow/ragflow (reviewed at commit `cc72ecb0af18ade5d58f84d107af7cd59a88c39f`, 2026-10-06)
- Stars: 91730 · Language: Go · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/ragflow/

## Overview

RAGFlow is a self-hosted RAG server. You upload documents into datasets ("knowledge bases"). RAGFlow parses them with its own layout and OCR models, splits them into chunks, embeds them, and stores them in a search engine. Chat assistants, search apps, an agent canvas and an OpenAI-compatible endpoint then answer questions from those chunks, with citations.

At this commit the backend is Go. One binary, `bin/ragflow_server`, runs in five modes: `--api`, `--admin`, `--ingestor`, `--syncer` and `--deepdoc`. The React + Vite UI is in `web/`. The repository still has some Python (the client SDK in `sdk/python`, tests and helper scripts), but the Docker entrypoint starts only the Go server and the Go ingestor. Its own comment says the image "ships no Python task executor". The repo's `AGENTS.md` still describes a Python API that is the default, so trust the entrypoint over that file.

RAGFlow fits teams that want a complete product: a UI, multi-tenant accounts, many model providers and strong PDF parsing. The cost is weight. A default deployment runs MySQL, MinIO, Kvrocks, NATS, ClickHouse and a search engine next to the server.

## Architecture

```mermaid
flowchart LR
  UI["Web UI (React)"] --> NG["nginx"]
  NG --> API["ragflow_server --api"]
  API --> CP["ChatPipelineService"]
  CP --> RS["RetrievalService"]
  CP --> AG["Agentic RAG (eino)"]
  RS --> DE["DocEngine"]
  AG --> DE
  DE --> ES["ES / Infinity / OceanBase / SereneDB"]
  API --> MQ["NATS JetStream"]
  MQ --> ING["ragflow_server --ingestor"]
  ING --> PL["Pipeline: Parser, Chunker, Tokenizer, Extractor"]
  PL --> DD["DeepDoc (ONNX OCR + layout)"]
  PL --> DE
  API --> DB["MySQL"]
  ING --> S3["MinIO"]
  CP --> LLM["Model providers"]
```

| Component | Path | Role |
|---|---|---|
| Server entrypoint | `cmd/ragflow_server.go` | Parses the mode flag and wires services, engines and the agentic retriever |
| HTTP routes | `internal/router/router.go` | Gin routes under `/api/v1`, including `/chat/completions`, `/openai/:chat_id/chat/completions` and `/retrieval` |
| Chat pipeline | `internal/service/chat_pipeline.go` | 11-phase `AsyncChat`: model binding, query refinement, retrieval, prompt, LLM call, citations |
| Retrieval | `internal/service/nlp/retrieval.go`, `reranker.go` | Hybrid search, retries, score fusion and reranking |
| Document engines | `internal/engine/` | `DocEngine` interface with Elasticsearch, Infinity, OceanBase/SeekDB and SereneDB backends |
| Ingestor | `internal/ingestion/service/ingestion_service.go` | Pulls tasks from NATS and runs the pipeline per document |
| Ingestion pipeline | `internal/ingestion/pipeline/`, `component/` | DSL-driven components; templates in `pipeline/template/*.json` |
| DeepDoc | `internal/deepdoc/parser/pdf/` | Native-backed PDF parsing, OCR, layout and table structure |
| Agentic RAG | `internal/agentic_rag/`, `internal/rag/agentic-rag/` | Tool-using agent (level 5) and the reasoning graph (levels 1-4) |
| Model drivers | `internal/entity/models/` | About 90 provider drivers for chat, embedding, rerank, TTS and ASR |
| Syncer | `internal/syncer/` | Scheduled sync of external data sources into datasets |

## How a request flows

1. A chat request reaches `POST /api/v1/chat/completions`, or the OpenAI-compatible route ([router.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/router/router.go#L320-L331)). Both call `ChatPipelineService.AsyncChat`.
2. `AsyncChat` checks that the last message is from the user. It then picks one of three engines. Reasoning level 5 (or an `agent_mode` kwarg) goes to the smart-reasoning agent. A chat with no datasets and no web search goes to a plain LLM call (`AsyncChatSolo`). Everything else runs the pipeline in a goroutine ([chat_pipeline.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L204-L290)).
3. Phases 2-8 resolve the model config, start a Langfuse trace when the tenant has keys, bind the embedding, rerank, chat and TTS models, and can answer from SQL for table datasets. They then refine the query with the LLM: multi-turn rewrite, cross-language translation, metadata filters and keyword extraction.
4. Phase 9 retrieves. At reasoning levels 1-4 the harness (`retrieveViaHarness`) runs the agentic graph from `internal/rag/agentic-rag`. At level 0 it calls `RetrievalService.Retrieval`, then optionally adds TOC enhancement, parent chunks for child fragments, web search (Tavily) and knowledge-graph results ([chat_pipeline.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L685-L725)).
5. `Search` builds a BM25 `matchText` expression, a dense `matchDense` expression and a fusion expression. If nothing comes back, it retries with looser thresholds, then falls back to dense-only search ([retrieval.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L671-L880)).
6. Phase 10 packs the chunks into the prompt and trims the history to fit the model window. Phase 11 streams the LLM answer and calls `decorateAnswer`. If the model did not write `[ID:n]` markers itself, `InsertCitations` embeds each answer sentence and attaches up to 4 chunks per sentence by cosine similarity ([chat_pipeline.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L3127-L3230), [citation.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/citation.go#L249-L274)).

On the write side, the API publishes an ingestion task to NATS. An `--ingestor` worker pulls it in `consumeLoop` ([ingestion_service.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/ingestion/service/ingestion_service.go#L302-L330)) and runs the dataset's pipeline DSL. Chunks are written in batches without an index refresh. The last batch refreshes, so the document is searchable before the task is acknowledged ([pipeline_executor.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/ingestion/task/pipeline_executor.go#L104-L135)).

## Key components

### Ingestion pipeline

Each parser template (General, Book, Paper, Laws, Manual, QA, Table, Presentation, Picture, Audio, Email, One) is a JSON DSL. It chains `File`, `Parser`, a chunker, `Tokenizer` and usually `Extractor`. The General template uses `GeneralChunker` with `chunk_token_size: 512` and sentence delimiters ([ingestion_pipeline_general.json](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/ingestion/pipeline/template/ingestion_pipeline_general.json)). PDFs default to `parse_method: DeepDOC`, which runs the native OCR and layout models. Other PDF backends, including MinerU, PaddleOCR and Docling, are selected by `parse_method`. `Tokenizer` writes `content_ltks` for BM25 and sends all of a document's chunks to the embedding model in one batched `Encode` call.

### Knowledge compiler

Ingestion can also run a knowledge compiler (`internal/ingestion/component/knowledge_compiler/`). Its variants are wiki, mind map, tree and structure. The wiki variant makes LLM passes that extract entities and claims, plan pages, reconcile them with existing pages and merge new versions, using the prompts in `wiki/prompt.go`. It stores the resulting pages as searchable chunks. The ingestor also runs a dataset-level compile consumer with NATS leases, so pages can be rebuilt across documents. This is an add-on to RAG over chunks, not the core of the product.

### Document engines

`DocEngine` is the storage contract. It covers chunk CRUD, `Search`, a per-tenant metadata store, `RunSQL`, KNN scoring and metadata filter push-down ([engine.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/engine.go#L29-L94)). It has five engine types: Elasticsearch, Infinity, OceanBase, SeekDB (which shares the OceanBase code) and SereneDB. Compose also has an OpenSearch profile, but there is no OpenSearch `DocEngine` package in `internal/engine/`.

### Scoring and reranking

On every backend except Infinity, the first search uses fixed fusion weights of `0.001,1`, so it is in effect a vector recall pass. The caller's `vector_similarity_weight` (default 0.3) is applied afterwards, as `tkWeight*tsim + vtWeight*vsim` ([retrieval.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L601-L650)). When the chat has a rerank model, `RerankByModel` replaces that step. Infinity uses its own fused score ([retrieval.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L450-L490)). A `pagerank_fea` rank feature (weight 10) adds a per-dataset boost.

### Agentic modes

RAGFlow has two separate agentic paths. Levels 1-4 (low to ultra) run the planner graph in `internal/rag/agentic-rag`. `cmd/ragflow_server.go` wires it in through `retrievalbridge.NewHarnessRetriever`. Level 5 runs a cloudwego/eino `ChatModelAgent` named `smart-reasoning` ([agent.go](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/agentic_rag/agent.go#L291-L375)). Its tools come from `conf/agentic_rag.yaml`: `grep_chunks`, `search_bm25_chunks`, `search_semantic_chunks`, `list_chunks`, `think`, `todo_write` and a goja `run_javascript` sandbox ([agentic_rag.yaml](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/conf/agentic_rag.yaml#L1-L19)). The agent ends each factual line with `chunk_id: <id>`, and the pipeline turns that into citation markers. `internal/service/deep_researcher.go` defines a recursive `DeepResearcher`, but no production code constructs it.

## Extending it

- **Model providers:** add a driver under `internal/entity/models/` and an entry in `conf/llm_factories.json` (78 factories at this commit).
- **Ingestion:** chunkers register with `MustRegisterChunker`, and pipelines are editable JSON DSL, so a dataset can run a custom canvas.
- **Agentic RAG:** tool lists and prompts live in `conf/agentic_rag.yaml`, which is reloaded when the file changes.
- **Agents and integrations:** `internal/agent/component/` holds the canvas components (LLM, categorize, invoke, SQL, browser, MCP and others). `internal/syncer/connector/` holds the data-source connectors. An optional MCP listener and a Dify-compatible retrieval endpoint are also available.

## Running it

The supported path is Docker Compose in `docker/`. `DOC_ENGINE` (default `elasticsearch`) selects the search-engine profile. The default `COMPOSE_PROFILES` also brings up MySQL, MinIO, Kvrocks, NATS and ClickHouse. The entrypoint runs Go migrations, then starts the syncer, nginx with `ragflow_server --api`, and one or more `--ingestor` processes (`--workers=N`) ([entrypoint.sh](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/docker/entrypoint.sh#L195-L236)). Ingestor concurrency (`max_concurrent_workers`, 1-256) and PDF page concurrency (`page_concurrency`, 1-16) are set in `service_conf.yaml` or by CLI flag ([service_conf.yaml.template](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/docker/service_conf.yaml.template#L115-L127)). A source build needs Go 1.27, plus cgo and native libraries for DeepDoc and the C++ tokenizer (`build.sh`).

## Strengths and caveats

- **Strength: parsing depth.** It has a native OCR and layout pipeline with table structure recognition and twelve template pipelines, so most document types get a purpose-built chunker.
- **Strength: retrieval is not left to luck.** Empty results trigger staged retries, a dense-only fallback and a filter-only fallback when a metadata filter is active.
- **Strength: product surface.** It has multi-tenancy, an OpenAI-compatible endpoint, many model providers, an agent canvas and data-source sync in one deployment.
- **Caveat: an evaluation schema with no runner.** `internal/entity/evaluation.go` defines datasets, cases, runs and results tables. No Go service, route or UI uses them. The observability you can use today is per-tenant Langfuse tracing and OpenTelemetry in the harness.
- **Caveat: a heavy stack.** Six or more stateful services, and Elasticsearch needs large memory and ulimit settings.
- **Caveat: a migration in progress.** Many Go files cite the Python code they replace ("mirrors dialog_service.py"). `AGENTS.md` is out of date, and `deep_researcher.go` is unused. Expect fast change in `internal/ingestion` and `internal/parser`.
- **Caveat: the consumer range is cosmetic.** In the entrypoint, `--consumer-no-beg/--consumer-no-end` start a single ingestor and do not shard work.

*Sources: code at cc72ecb, deepwiki-open wiki (11 pages), verified Q&A.*

## How infiniflow/ragflow answers the RAG engines questions

### How are documents parsed and chunked? (answered)

**Formats:** RAGFlow supports PDF, DOCX, TXT, Markdown, HTML, spreadsheets, slides, images, audio, email, and EPUB. Parser type is selected per-document via a `ParserConfig.parser_id` field, with known types including `general`, `naive`, `book`, `paper`, `laws`, `manual`, `presentation`, `qa`, `table`, `picture`, `one`, `audio`, `email`, and `knowledge_graph` (`entity/dataset.go:48-62`).

**PDF handling:** The PDF parser (`deedoc/parser/pdf/parser.go`) is the most sophisticated. It runs a multi-stage pipeline: per-page render → text extraction via both embedded PDF characters and OCR (`parser_ocr.go` uses a DLA detection model + ONNX text recognizer at `recBatchNum=16`), → DLA (Document Layout Analysis) region detection → table extraction (`grid_structure_test.go`) → section assembly. The system can classify text boxes into layout types (title, text, figure, table) and handles de-skewing, rotation candidates, and coordinate-space normalization for all downstream layout algorithms.

**Office documents:** The DOCX parser (`deedoc/parser/office/parser.go`) converts raw blocks to Sections: headings get `layoutType="title"`, tables get `DocTypeKwd="table"` with an HTML row representation, and images get `DocTypeKwd="image"`. The Go port uses `office_oxide` for raw block extraction.

**OCR pipeline:** The OCR pipeline (`parser_ocr.go`) detects text boxes via an ONNX inference model, de-skews each box using `WarpCrop`, generates rotation candidates (0°, ±90°), and runs batched recognition (`recBatchNum=16`). Coordinate scaling divides by zoom factor so downstream layout receives PDF-point coordinates.

**Chunking strategy:** After parsing, the chunker (`ingestion/component/chunker/`) has multiple strategies: token-based (`token.go`), delimiter-based (`delimiter_case_sensitive_test.go`), page-based (`page.go`), section-based (`positions_slice.go`), QA pair (`qa.go`), table-aware (`table.go`, `table_split.go`), and hierarchy-aware (`hierarchy.go`). The `general.go` chunker merges text segments from parsed sections with configurable token caps. Chunk sizes are controlled by parser configuration rather than fixed lengths. The tokenizer component (`tokenizer.go`) computes per-chunk `content_ltks` (tokenized string for BM25) and `content_sm_ltks` (fine-grained variant), plus embedding vectors via the tenant's embedding model.

> **Editor's note.** Correction: the PDF and Office parsers live under `internal/deepdoc/parser/` (not `deedoc/`). `grid_structure_test.go` and `delimiter_case_sensitive_test.go` are tests, not the table or delimiter implementations. The General template chunks at `chunk_token_size: 512` (`internal/ingestion/pipeline/template/ingestion_pipeline_general.json`).

Citations: [internal/entity/dataset.go:44-62](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/entity/dataset.go#L44-L62) · [internal/ingestion/component/tokenizer.go:1-60](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/ingestion/component/tokenizer.go#L1-L60)

### How are embeddings and indexes built and stored? (answered)

**Embedding models:** RAGFlow supports a broad range of embedding providers through a plugin-style model architecture (`internal/entity/models/`). The factory (`factory.go`) resolves model requests to provider-specific drivers — there are 80+ model provider files for OpenAI, Anthropic, Google/Gemini, DeepSeek, Ollama, Cohere, Jina, Voyage, vLLM, Xinference, HuggingFace, and many Chinese providers (e.g. Baichuan, ZhipuAI, Qwen, Hunyuan). Each supports chat, text-embedding, reranking, or TTS modalities. Models are registered per-tenant in the `tenant_model_instance` table and resolved at runtime by `ModelSolver` (`service/model_solver.go`).

**Vector stores:** Four backends are supported (`engine/engine.go:33-38`): Elasticsearch, Infinity, OceanBase, and SereneDB, plus OpenSearch (via the OS config block in `service_conf.yaml.template`). Each implements the `DocEngine` interface (`engine/engine.go:41-93`) with chunk CRUD, search, metadata operations, and SQL execution. The Go code has Elasticsearch (`engine/elasticsearch/chunk.go`) and Infinity (`engine/infinity/`) implementations; OceanBase and SereneDB share a PostgreSQL-compatible layer. Chunks are created via `CreateChunkStore()` which sets up an ES index with `number_of_shards=1`, `number_of_replicas=0` and appropriate mappings including `tag_feas` as `rank_features` for tag-based ranking. Bulk writes are batched at 500 chunks or 20MB.

**Hybrid / keyword (BM25) indexes:** Chunks carry both `content_ltks` (tokenized text for BM25 matching) and embedding vectors. BM25 is implemented via the QueryBuilder (`service/nlp/query_builder.go`) that produces `matchText` expressions for the engine. OpenSearch has dedicated hybrid search pipeline configuration (`service_conf.yaml.template:59`).

**Metadata:** The `DocEngine` interface defines `CreateMetadataStore`, `InsertMetadata`, `UpdateMetadata`, `SearchMetadata` operations (`engine/engine.go:54-61`). Metadata fields per chunk include `docnm_kwd` (document name), `kb_id`, `img_id`, `title_tks`, `position_int`, `page_num_int`, `chunk_order_int`, `content_with_weight`, `doc_type_kwd` (text/image/table), `mom_id` (parent chunk for child fragments), and user-defined metadata values set via filter conditions.

> **Editor's note.** Correction: `DocEngine` has five engine types: Elasticsearch, Infinity, OceanBase, SeekDB and SereneDB. OceanBase and SeekDB share one implementation (`IsOceanBaseFamily`), and SereneDB is a separate package. There is no OpenSearch engine under `internal/engine/`, only a Compose profile and a mapping file.

Citations: [internal/engine/engine.go:30-93](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/engine.go#L30-L93) · [internal/engine/elasticsearch/chunk.go:54-151](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/elasticsearch/chunk.go#L54-L151) · [internal/service/nlp/retrieval.go:697-710](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L697-L710)

### How is retrieval performed? (answered)

**Hybrid search:** The core retrieval is implemented in `service/nlp/retrieval.go`. The `RetrievalService.Search()` method (line 671) builds parallel BM25 (`matchText`) and dense vector (`matchDense`) query expressions and optionally a `fusionExpr` that combines them. Defaults: KNN topK=1024, numCandidates=2048, similarityThreshold=0.2, vectorSimilarityWeight=0.3 (1.0 for vector-only), rankFeature with pagerank_fea=10.0, rerankCandidatesCount=64. When no embedding model is available, it falls back to keyword-only search. The fusion expression weight is controlled by `VectorSimilarityWeight`, and if empty results are returned the system retries with progressively relaxed thresholds (min_match=0.1, similarity=0.17) and optionally pure dense-only fallback.

**Dense / sparse / fusion:** The `types/SearchRequest` struct (`engine/types/types.go:34-56`) carries `MatchExprs` — a list of text match, dense vector match, and fusion expressions. The vector expression is built via `GetVector()` which calls the embedding model. For hybrid, all three are passed to the engine: BM25 text match + vector kNN + fusion (weighted sum).

**Reranking:** Reranking (`service/nlp/reranker.go`) supports model-based reranking (via a `RerankModel` provider, e.g., Jina, Cohere, BGE-reranker) and engine-dependent fallback — Infinity uses pre-normalized scores, Elasticsearch recomputes via `RerankStandard()` which combines token cosine similarity `tkSim` and vector cosine similarity `vtSim` weighted by `tkWeight` (0.7) and `vtWeight` (0.3).

**Query rewriting:** The chat pipeline (`chat_pipeline.go` phase 8, lines 188-189) applies multi-turn refinement (`refine_multiturn`), cross-language translation (`cross_languages`), metadata filtering (`meta_data_filter`), and keyword extraction via LLM (remaining `service/generator.go:51-95` uses a configured ChatModel with a `keyword_prompt` template).

**DeepResearcher (agentic retrieval):** When `reasoning=true` and agentic mode is enabled (`chat_pipeline.go:267-268`), the system runs a recursive DeepResearcher (up to depth=3) that iterates between KB search, web search and optional Knowledge Graph retrieval, with sufficiency checking and multi-query generation at each layer.

**Filters:** The `RetrievalRequest` carries `Filter` (arbitrary key-value metadata conditions) and `DocIDs` for document-scoped search. The `available_int=1` filter defaults to exclude unavailable chunks. Metadata pushdown filtering is supported via `FilterDocIdsByMetaPushdown()` on the engine interface.

> **Editor's note.** Correction: `DeepResearcher` (`service/deep_researcher.go`) has no production caller. Reasoning levels 1-4 run the agentic graph in `internal/rag/agentic-rag` through `retrievalbridge.NewHarnessRetriever`, and level 5 runs the `smart-reasoning` eino agent (`internal/agentic_rag`). On non-Infinity engines the first search uses fixed fusion weights `0.001,1`. `vector_similarity_weight` is applied afterwards (tkWeight = 1 - weight), so 0.7/0.3 is only the default.

Citations: [internal/service/nlp/retrieval.go:670-870](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L670-L870) · [internal/service/nlp/reranker.go:40-89](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/reranker.go#L40-L89) · [internal/engine/types/types.go:34-56](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/types/types.go#L34-L56) · [internal/service/chat_pipeline.go:146-268](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L146-L268) · [internal/service/generator.go:51-95](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/generator.go#L51-L95)

### How are answers generated and grounded? (answered)

**Pipeline architecture:** The `ChatPipelineService.AsyncChat()` method (`service/chat_pipeline.go:204-280`) implements the full generation pipeline. After entry validation (non-empty messages, last role=user), it resolves the LLM model config with max_tokens, sets up Langfuse tracing, binds models (embedding, rerank, chat, TTS), and performs optional tool-call session binding.

**Prompt assembly:** Phase 7 (line 188-189) resolves prompt parameters from `chat.PromptConfig` — including the `{knowledge}` placeholder auto-fill, system prompt templates, empty-response handling, and citation prompt formatting. Prompts are loaded from `rag/prompts/` markdown files via `LoadPrompt()` (`service/load_prompt.go`), which caches them from the filesystem. Phase 8 performs LLM-based query refinement (multi-turn context, cross-language translation, metadata filtering, keyword extraction).

**Citations / source attribution:** The `InsertCitations` function (`service/citation.go:68-80`) implements the citation decoration algorithm: it splits the answer into sentences, encodes each into a vector via the embedding model, computes cosine similarity with chunk vectors, and applies threshold descent (0.63→0.3×0.8 per round). Up to 4 chunks can be cited per sentence, with `[ID:n]` markers inserted into the output. The `DoRefer` flag on the Chat entity controls whether citation is enabled. The citation prompt (`citation_prompt.md` in `rag/prompts/`) instructs the LLM to reference sources by chunk ID. The agentic RAG mode additionally normalizes malformed citation formats (`citation.go:43-53`).

**Streaming:** The pipeline supports both streaming and non-streaming modes (`stream bool` parameter). In streaming mode, it yields `AsyncChatResult` deltas containing `Answer`, optional `Reasoning` (chain-of-thought routed to `delta.reasoning_content`), `Reference` metadata, and structured `ThinkEvent` steps for agentic RAG reasoning traceability.

**Agentic / multi-step answering:** When the reasoning level is set to `agentic` and knowledge bases are configured (`chat_pipeline.go:267-268`), the system runs agentic RAG (`internal/rag/agentic-rag/agentic_rag.go`). This uses tools defined in `config/agentic_rag.yaml` (think, grep_chunks, search_bm25_chunks, search_semantic_chunks, list_chunks, run_javascript, todo_write) following an "Evidence-First" philosophy. The DeepResearcher mode (`rag/agentic-rag/runtime/`) can recursively search, read, and verify across multiple rounds.

**Structured outputs:** The `structured_output_prompt.md` template supports forcing JSON-formatted outputs from the LLM, and the generation pipeline includes a `messageFitIn` check that truncates context to stay within 95% of the model's token budget before the LLM call.

> **Editor's note.** Correction: the `conf/agentic_rag.yaml` tools (grep_chunks, search_bm25_chunks, …) belong to the level-5 agent in `internal/agentic_rag/`, not to `internal/rag/agentic-rag/agentic_rag.go`, which is the levels 1-4 graph. Embedding-based `InsertCitations` runs only when the LLM did not already write citation markers (`decorateAnswer`). DeepResearcher is not wired into the chat path.

Citations: [internal/service/chat_pipeline.go:204-280](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L204-L280) · [internal/service/citation.go:38-100](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/citation.go#L38-L100) · [internal/service/generator.go:51-100](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/generator.go#L51-L100)

### How is quality evaluated or observed? (answered)

**Database-backed evaluation framework:** RAGFlow has a built-in evaluation system with four entity types (`entity/evaluation.go`): `EvaluationDataset` (named groups of test cases with KB bindings), `EvaluationCase` (question + reference_answer + relevant_doc_ids/chunk_ids per case), `EvaluationRun` (ties a dataset to a dialog/chat config with a `MetricsSummary` JSON blob), and `EvaluationResult` (per-case generated_answer, retrieved_chunks, Metrics, execution_time, token_usage). Runs have a status lifecycle (PENDING → ...) and store a `ConfigSnapshot` of the chat configuration at run time.

**Metrics:** Individual `EvaluationResult` records carry a `Metrics` JSONMap that stores arbitrary evaluation metrics per case. The `EvaluationRun` aggregates these into a `MetricsSummary` JSONMap. The evaluation entity types include both `execution_time` and `token_usage` tracking for cost monitoring. This provides an offline benchmark framework rather than continuous production metrics — users create datasets of question-answer pairs, run the chat against them, and inspect results.

**Langfuse observability:** Beyond the built-in evaluation tables, RAGFlow integrates with Langfuse for production tracing (`service/langfuse.go:45-57`). Per-tenant Langfuse credentials are stored in `TenantLangfuse` and resolved at request time. The chat pipeline (phase 3, `chat_pipeline.go:311-326`) creates a Langfuse trace on each chat invocation when configured, posting trace metadata including stream mode, KB count, and session/user IDs. Observations are posted async via a background worker.

**No built-in metrics library:** The system does not include an implementation of standard RAG evaluation metrics like Faithfulness, Answer Relevancy, or Context Precision. The evaluation framework stores whatever metrics are computed externally or via custom scripts — it provides the storage infrastructure (datasets, cases, runs, results) and the mechanism to execute a dialog against a set of test questions, but the metric computation itself is not shipped in this Go codebase. What it stores in the JSONMap fields is up to the operator.

> **Editor's note.** Correction: the evaluation entities in `entity/evaluation.go` are schema only. No Go service, route or web UI creates or runs evaluations at this commit. Observability also includes OpenTelemetry spans in the harness (`internal/harness/graph/pregel/otel_telemetry.go`).

Citations: [internal/entity/evaluation.go:1-97](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/entity/evaluation.go#L1-L97) · [internal/service/langfuse.go:38-80](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/langfuse.go#L38-L80) · [internal/service/chat_pipeline.go:309-327](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L309-L327)

### How is it deployed and operated? (answered)

**Service architecture:** RAGFlow is a Docker-based service with multiple profiles. The main API server runs on port 9380 (Go, via `cmd/ragflow_server.go`), an optional admin server on 9381, and nginx serves the web UI on ports 80/443. Task executors are pluggable workers that consume ingestion/parsing tasks from a message queue.

**Infrastructure requirements:** The minimum stack requires MySQL (metadata/settings), MinIO (document storage), Kvrocks (cache/task queue), and a vector store (Elasticsearch, Infinity, OceanBase, or SereneDB). Optional components include NATS (message queue), ClickHouse (analytics), and OpenSearch (hybrid search). The docker-compose (`docker/docker-compose.yml`, `docker-compose-base.yml`) configures all services including ES, Infinity, OceanBase, etc., gated behind profiles (`elasticsearch`, `infinity`, `oceanbase`, `opensearch`). Memory limits, ulimits (nofile=65535), and Docker healthchecks are preconfigured.

**UI:** The web frontend (`web/`) is a React/TypeScript application built with Vite, using Ant Design icons, AntV G2/G6 for graphs, and Lexical for rich text. It connects to the Go API server over HTTP. The UI includes knowledge base management, chat/dialog management, search app configuration, document upload, evaluation dashboards, and agent canvas workflows.

**API:** The system exposes REST APIs under `/api/v1/` and `/api/v1/openai/<chat_id>/chat/completions` for OpenAI-compatible endpoints. The Go router (`internal/router/router.go`) handles all routes. There's also an MCP (Model Context Protocol) server endpoint on port 9382, supporting SSE and Streamable HTTP transports.

**Multi-tenancy:** Multi-tenancy is built in at the model layer — all entities carry `TenantID`/`CreatedBy`, and the evaluation/chat/knowledgebase tables all filter by tenant. Models are per-tenant via `tenant_model_instance`. The `tenant_model_provider.go` and `tenant_llm.go` tables store provider credentials and LLM configurations per tenant. The Search/Share system (`service/search.go`) supports access control with "me" vs "team" permission levels.

**Scaling:** Task executors can be horizontally scaled using `--workers=N` or range-based `--consumer-no-beg/--consumer-no-end` flags. The entrypoint.sh (`docker/entrypoint.sh`) allows selectively disabling the web server, task executor, or data sync for separate container roles. Redis/Kvrocks serves as the task queue backend, and NATS can be used for high-throughput messaging, both naturally supporting distributed worker pools.

> **Editor's note.** Correction: the Go ingestor's queue is NATS JetStream (`ingestor.mq_type: 'nats'`), and NATS is in the default `ragflow-go` Compose profile, so it is not optional. Kvrocks is the cache. The MCP listener on 9382 is commented out by default. In the entrypoint, the `--consumer-no-beg/--consumer-no-end` range starts a single ingestor, and only `--workers=N` starts N of them.

Citations: [internal/entity/evaluation.go:19-32](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/entity/evaluation.go#L19-L32) · [internal/service/search.go:136-167](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/search.go#L136-L167)
