# How are answers generated and grounded?

> RAG engines — a good answer covers: Prompt assembly; citations / source attribution; streaming; agentic or multi-step answering.

Canonical page: https://llms-technical-reviews.com/rag/q/generation/

## Verdict

[RAGFlow](/p/ragflow/) does the most to ground answers, because it attaches citations even when the model leaves them out. [Kotaemon](/p/kotaemon/) shows evidence most clearly. [R2R](/p/r2r/) has the cleanest citation stream for building your own product.

**The engine enforces citations.** If the model writes no `[ID:n]` markers, RAGFlow embeds each answer sentence and attaches up to four chunks by cosine similarity. Its reasoning levels add a planner graph or a tool-using agent. R2R turns bracketed short ids into `Citation` objects. When streaming, it emits a `citation` event the first time each id appears. [Onyx](/p/onyx/) gives the model numbered documents inside an agent loop of up to six cycles. A stream processor turns `[1]` markers into links. In highlight mode, Kotaemon runs a separate function-calling pass that extracts verbatim quotes and fuzzy-matches them to chunk spans in its PDF viewer. [PageIndex](/p/pageindex/) has the model write `<cite doc page/>` tags. These are page-level in local mode.

**Sources attached, citation left to you.** [AnythingLLM](/p/anything-llm/) returns retrieved chunks as `sources`. To fit the context window, it cuts text out of the middle of the system prompt, history or user message. [LlamaIndex](/p/llama_index/) returns `source_nodes` with every response and uses compact-and-refine synthesis by default. `CitationQueryEngine` re-splits sources into numbered 512-token chunks. [Haystack](/p/haystack/)'s `AnswerBuilder` can map `[n]` references back to documents through `reference_pattern`. Its `Agent` handles multi-step answering, with hooks.

**Delegated or absent.** [RAG-Anything](/p/rag-anything/) leaves the answer to LightRAG and does not render citations. Its VLM-enhanced query puts retrieved images into the prompt. [Quivr](/p/quivr/) does not generate answers at all. It returns segments with Unicode code-point offsets, and its MCP tools tell an external agent how to cite them.

Pick: RAGFlow or Onyx when every answer must be traceable without depending on the model.
Pick: Kotaemon when users need to see the highlighted evidence.
Pick: R2R or Quivr when you build your own front end on top of an API.

## Per-project answers

### infiniflow/ragflow (answered)

**Pipeline architecture:** The `ChatPipelineService.AsyncChat()` method (`service/chat_pipeline.go:204-280`) implements the full generation pipeline. After entry validation (non-empty messages, last role=user), it resolves the LLM model config with max_tokens, sets up Langfuse tracing, binds models (embedding, rerank, chat, TTS), and performs optional tool-call session binding.

**Prompt assembly:** Phase 7 (line 188-189) resolves prompt parameters from `chat.PromptConfig` — including the `{knowledge}` placeholder auto-fill, system prompt templates, empty-response handling, and citation prompt formatting. Prompts are loaded from `rag/prompts/` markdown files via `LoadPrompt()` (`service/load_prompt.go`), which caches them from the filesystem. Phase 8 performs LLM-based query refinement (multi-turn context, cross-language translation, metadata filtering, keyword extraction).

**Citations / source attribution:** The `InsertCitations` function (`service/citation.go:68-80`) implements the citation decoration algorithm: it splits the answer into sentences, encodes each into a vector via the embedding model, computes cosine similarity with chunk vectors, and applies threshold descent (0.63→0.3×0.8 per round). Up to 4 chunks can be cited per sentence, with `[ID:n]` markers inserted into the output. The `DoRefer` flag on the Chat entity controls whether citation is enabled. The citation prompt (`citation_prompt.md` in `rag/prompts/`) instructs the LLM to reference sources by chunk ID. The agentic RAG mode additionally normalizes malformed citation formats (`citation.go:43-53`).

**Streaming:** The pipeline supports both streaming and non-streaming modes (`stream bool` parameter). In streaming mode, it yields `AsyncChatResult` deltas containing `Answer`, optional `Reasoning` (chain-of-thought routed to `delta.reasoning_content`), `Reference` metadata, and structured `ThinkEvent` steps for agentic RAG reasoning traceability.

**Agentic / multi-step answering:** When the reasoning level is set to `agentic` and knowledge bases are configured (`chat_pipeline.go:267-268`), the system runs agentic RAG (`internal/rag/agentic-rag/agentic_rag.go`). This uses tools defined in `config/agentic_rag.yaml` (think, grep_chunks, search_bm25_chunks, search_semantic_chunks, list_chunks, run_javascript, todo_write) following an "Evidence-First" philosophy. The DeepResearcher mode (`rag/agentic-rag/runtime/`) can recursively search, read, and verify across multiple rounds.

**Structured outputs:** The `structured_output_prompt.md` template supports forcing JSON-formatted outputs from the LLM, and the generation pipeline includes a `messageFitIn` check that truncates context to stay within 95% of the model's token budget before the LLM call.

> **Editor's note.** Correction: the `conf/agentic_rag.yaml` tools (grep_chunks, search_bm25_chunks, …) belong to the level-5 agent in `internal/agentic_rag/`, not to `internal/rag/agentic-rag/agentic_rag.go`, which is the levels 1-4 graph. Embedding-based `InsertCitations` runs only when the LLM did not already write citation markers (`decorateAnswer`). DeepResearcher is not wired into the chat path.

Citations: [internal/service/chat_pipeline.go:204-280](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L204-L280) · [internal/service/citation.go:38-100](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/citation.go#L38-L100) · [internal/service/generator.go:51-100](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/generator.go#L51-L100)

### Mintplex-Labs/anything-llm (answered)

**Prompt assembly.** The `streamChatWithWorkspace()` function builds a system prompt by calling `chatPrompt()` which takes the workspace's configured prompt template, expands system-prompt variables, and appends user memories (if enabled) as a `## Things I Remember About You` section (scoped per user, reranked via the native cross-encoder when there are many memories). Context from RAG is appended as numbered `[CONTEXT N]`/`[END CONTEXT N]` blocks via `#appendContext()`. The chat history is formatted as alternating user/assistant messages. The final message array is built by each provider's `constructPrompt()` method: a system role message, the history messages, and a user role message.

**Compression.** The `messageArrayCompressor` (or `messageStringCompressor` for string-mode models) enforces token-budget proportions: system 15%, history 15%, user 70%. It uses a "cannonball" strategy — when a component exceeds its limit, it truncates from the middle outwards, inserting a `--prompt truncated for brevity--` marker. History is the most aggressively compressed component, with a preference for keeping the 3 most recent exchanges even if they must be cannonballed individually.

**Citation / source attribution.** Vector search results carry their chunk metadata (title, source URL, published date, score) which are returned alongside the text response as `sources`. The chunk metadata is included in the context block sent to the LLM, and the UI renders these sources as clickable citations. The `fillSourceWindow` mechanism ensures the context window has consistently `topN` sources for follow-up questions by backfilling from previous turns.

**Streaming.** Every LLM provider implements `streamGetChatCompletion()` and `handleStream()`. The response is delivered as SSE (Server-Sent Events) with chunk types: `textResponseChunk`, `textResponse`, `abort`, `finalizeResponseStream`. OpenAI's Responses API streams reasoning summary as thinking blocks. Cost metrics (prompt tokens, completion tokens, duration, tokens/sec) are computed per turn.

**Multi-step / agentic.** When the workspace has agent skills enabled, `grepAgents()` diverts to the agent system (aibitat) which supports multi-turn tool-using agents with MCP server integration, Gmail/Google Calendar/Outlook skills, SQL agents, file creation, web browsing, and scheduled-job creation. These run as independent subagents with their own provider connections.


Citations: [server/utils/chats/stream.js:273-340](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L273-L340) · [server/utils/chats/index.js:122-144](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/index.js#L122-L144) · [server/utils/memories/index.js:1-139](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/memories/index.js#L1-L139) · [server/utils/helpers/chat/index.js:49-192](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/index.js#L49-L192) · [server/utils/AiProviders/openAi/index.js:49-135](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/AiProviders/openAi/index.js#L49-L135)

### run-llama/llama_index (answered)

**Prompt assembly**: The `BaseSynthesizer` (`llama_index/core/response_synthesizers/base.py:64-107`) holds a `_text_qa_template` (default `DEFAULT_TEXT_QA_PROMPT`) and `_refine_template`. The QA prompt receives `context_str` (concatenated node text at `MetadataMode.LLM`) and `query_str`. For chat models, parallel `_chat_content_qa_template` and `_chat_content_refine_template` variants use `ChatMessage` blocks.

**Response modes** (`llama_index/core/response_synthesizers/type.py:4-58`): `REFINE` iterates through nodes, building an initial answer then refining with each subsequent chunk. `COMPACT` (via `CompactAndRefine` at `response_synthesizers/compact_and_refine.py:13-60`) packs multiple chunks into the context window before refining. `TREE_SUMMARIZE` builds a bottom-up summary tree. `GENERATION` (`llama_index/core/response_synthesizers/generation.py:32-50`) ignores context entirely. `SIMPLE_SUMMARIZE` stuffs all text into one prompt. `NO_TEXT` and `CONTEXT_ONLY` return nodes as-is.

**Source attribution**: Every `Response` object (`llama_index/core/base/response/schema.py:14-42`) carries `source_nodes: List[NodeWithScore]` alongside the response text. `Response.get_formatted_sources()` renders truncated node content with node IDs. Source nodes flow from retrieval through synthesis unchanged.

**Streaming**: When `streaming=True` is passed, calls return `StreamingResponse` or `AsyncStreamingResponse` (`llama_index/core/base/response/schema.py:108-200`), which wrap a generator/async generator yielding tokens. The response also carries `source_nodes` for display.

**Agentic/multi-step answering**: The `AgentWorkflow` and `FunctionAgent` (`llama_index/core/agent/`) support tool-calling agents that can iterate, reflect, and call retrieval tools. `SubQuestionQueryEngine` decomposes queries into parallel sub-questions. `ReActAgent` implements the ReAct pattern with chat formatters and output parsers.


Citations: [llama-index-core/llama_index/core/response_synthesizers/base.py:64-107](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/base.py#L64-L107) · [llama-index-core/llama_index/core/response_synthesizers/compact_and_refine.py:13-60](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/compact_and_refine.py#L13-L60) · [llama-index-core/llama_index/core/response_synthesizers/generation.py:32-50](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/generation.py#L32-L50) · [llama-index-core/llama_index/core/base/response/schema.py:14-42](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/response/schema.py#L14-L42) · [llama-index-core/llama_index/core/base/response/schema.py:108-200](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/response/schema.py#L108-L200)

### The-Vibe-Company/quivr (answered)

Quivr **does not implement answer generation**. It is a retrieval engine, not an LLM orchestrator — there is no prompt assembly, no LLM call for answer synthesis, no citation formatting, and no streaming of generated text. Instead, it returns ranked segments with **precise provenance** suitable for an external generator. Every search hit includes `record_id`, `version_id`, `part_key`, `segment_id`, and an `excerpt` with exact Unicode code-point `[start, end)` offsets into the source Part text (transport/httpapi/search.go:75-77). This allows an external LLM frontend to cite sources with verifiable precision. The MCP interface (`quivr mcp` command) exposes these same capabilities to AI agents — the `read` profile provides search and read-record tools, and its instructions explicitly tell agents how to cite hits by record_id, version_id, part_key and offset ranges (online/mcp_catalogue.go:41-44). The `ingest` profile adds document ingestion. The monitoring/evaluation subsystem (monitoring/monitoring.go) supports Saved Queries and Subscriptions that evaluate new documents against persistent queries, producing Matches delivered via webhook — but this is alerting, not generation. For answer generation, users are expected to call Quivr's search API from a separate application that handles prompt construction, context assembly, LLM inference, and citation rendering.


Citations: [internal/transport/httpapi/search.go:18-80](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/transport/httpapi/search.go#L18-L80) · [internal/online/mcp_catalogue.go:37-84](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/online/mcp_catalogue.go#L37-L84) · [internal/online/mcp.go:1-36](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/online/mcp.go#L1-L36) · [internal/monitoring/monitoring.go:1-50](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/monitoring/monitoring.go#L1-L50)

### VectifyAI/PageIndex (answered)

**Prompt assembly.** Generation uses the OpenAI Agents SDK or Anthropic SDK tool runner (`local_chat.py`). The system prompt is assembled from three parts: `CHAT_HEADER` (`"You are PageIndex by Vectify AI…"`), the `_base_instructions` (tool descriptions and discovery workflow), and the client's `instructions` plus any system/developer messages from the conversation history. The `_managed_instructions` function combines these (`local_chat.py:28-33`). For the chat_completions lane, user messages are validated and system rows are extracted into the instructions (`_split_chat_messages`, line 50-85).

**Document grounding.** A **targeting block** is prepended as the first user message — it lists the document(s) with metadata and directs the agent to use the tools (`local_chat.py:645-647`). The agent must use tools (`get_document_structure`, `get_page_content`) to read from actual documents; it cannot answer from general knowledge. The tool descriptions instruct the agent: "Cite only statements supported by tool outputs" and "Never fill a gap from general knowledge" (`agent_tools.py:1636-1644`).

**Citations.** When `citations=True`, a citation prompt (the `cited_answer` MCP server prompt or its local frozen copy) is appended to instructions, teaching the model to emit `<cite doc="..." page="..."/>` tags after each claim (`agent_tools.py:1634-1655`, `client.py:1391-1402`). The client parses these tags into structured citation objects (`client.py:48-88`). Cloud mode supports block-level citations (`block_id`); local mode is page-level only.

**Streaming.** `chat(stream=True)` returns a `ChatStream` — iterating yields text chunks with the agent's process woven in (thinking, tool calls, results) when `show_process=True` (default). The raw event stream via `.events` yields typed dicts: `thinking`, `answer`, `tool_call`, `tool_result` (`local_chat.py:721-773`). A sync pump drives an async generator on a background thread (`_stream_sync`, line 96-160). The managed endpoint's stream uses its own chunk parsing (`_cloud_chunk_events`, line 845-888).

**Protocol lanes.** Three wire protocols: (1) The default **answer lane** — simplified Chat Completions envelope; (2) `protocol="responses"` — native OpenAI Responses API with full transcript (tool calls/results as native items), enabling prompt-cache continuity across turns; (3) `protocol="messages"` — native Anthropic Messages API with the SDK's tool runner, cache-control markings, and cross-turn aggregated usage. All three support multi-turn via tool_use/tool_result round-trips.

**Agentic answering.** The answering agent uses `max_turns` (default 10) to limit the reasoning loop (`local_chat.py:404-408`). It can iteratively explore the document structure, read specific pages, and refine its understanding. The OpenAI Agents SDK `Runner` handles the turn loop; the Anthropic SDK's own `tool_runner` does the same for the Messages lane.


Citations: [pageindex/local_chat.py:28-33](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L28-L33) · [pageindex/agent_tools.py:1634-1655](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/agent_tools.py#L1634-L1655) · [pageindex/local_chat.py:721-773](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L721-L773) · [pageindex/local_chat.py:637-659](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L637-L659)

### onyx-dot-app/onyx (answered)

**Answer generation entry point:** `_stream_chat_turn` (`process_message.py:1650–1750`) orchestrates the entire streaming answer generation flow. It calls `build_chat_turn` for setup, then `_run_models` which invokes the LLM.

**LLM loop with multi-turn tool calling:** `run_llm_loop` (`chat/llm_loop.py`, lines 787–1467) runs up to `MAX_LLM_CYCLES` (default 6, configurable) iterations. Each cycle: (1) constructs the message history including system prompt, prior conversation, context files, and tool definitions; (2) calls the LLM via `run_llm_step`; (3) if the LLM returns tool calls, executes them via `run_tool_calls`; (4) feeds tool responses back as additional history; (5) breaks when the LLM returns a final answer with no tool calls. Available tools include `SearchTool` (internal and web search), `OpenURLTool`, `ImageGenerationTool`, `PythonTool` (code interpreter), `FileReaderTool`, and user-defined custom tools.

**Prompt assembly:** `construct_message_history` (llm_loop.py, lines 355–637) assembles the message list in this order: [system prompt] [truncated chat history] [custom agent prompt] [context files / project docs] [last user message] [prior tool rounds] [oversized-file metadata] [reminder]. A token budget is computed from the LLM's context window minus tool definitions, then history is truncated from the top to fit. Prompt caching markers are set on the static prefix (system prompt, context files, cacheable history).

**Citations and source attribution:** The `DynamicCitationProcessor` (`chat/citation_processor.py`) processes the LLM stream in real-time: in `HYPERLINK` mode, it converts citation markers like `[1]` into hyperlinks formatted as `[[1]](url)`, emitting `CitationInfo` objects. In `REMOVE` mode it strips citations from the answer. Citation mappings come from search tool responses (SearchDoc → URL). The `run_llm_loop` tracks `gathered_documents` from search tools and the `citation_processor` is updated with `update_citation_processor_from_tool_response` after each tool cycle.

**Streaming:** The entire answer is streamed as `Packet` objects via an `Emitter`. Packets include answer tokens (`AgentResponseDelta`), citations (`CitationInfo`), tool call metadata (`ToolCallDebug`), and stop signals (`OverallStop`). Multiple models can be compared side-by-side via `handle_multi_model_stream`.

**Query rewriting:** Before retrieval/generation, `semantic_query_rephrase` and `keyword_query_expansion` use an LLM to rewrite the user's query for better search recall, incorporating chat history, user memories, and user information.

**Agentic / multi-step answering:** The LLM can autonomously decide tool call ordering across multiple cycles. The system supports deep research (`run_deep_research_llm_loop`) for complex multi-step research queries.


Citations: [backend/onyx/chat/process_message.py:1650-1750](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/chat/process_message.py#L1650-L1750) · [backend/onyx/chat/llm_loop.py:355-637](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/chat/llm_loop.py#L355-L637) · [backend/onyx/chat/citation_processor.py:1-50](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/chat/citation_processor.py#L1-L50)

### deepset-ai/haystack (answered)

**Prompt assembly** – Haystack uses `ChatPromptBuilder` and `PromptBuilder` (`haystack/components/builders/`) with Jinja2 templates to assemble prompts. The `OpenAIChatGenerator` (`haystack/components/generators/chat/openai.py:57-409`) works with `ChatMessage` objects (with `user`, `assistant`, `system`, `tool` roles). Generation parameters like `temperature`, `max_completion_tokens`, `response_format`, and `stop` sequences can be set at init or merged per-key at run time.

**Streaming** – The `OpenAIChatGenerator` accepts a `streaming_callback: StreamingCallbackT` parameter both in `__init__` and at `run()` time (line 348). When a callback is provided, the generator streams `ChatCompletionChunk` objects through the callback function via OpenAI's `Stream` response, assembling the full message from chunks (`_handle_stream_response`). The `StreamingChunk` dataclass carries each delta: `content`, `delta`, and `tool_call` data.

**Tool calling / agentic answering** – The `Agent` component (`haystack/components/agents/agent.py:831-921`) wraps a ChatGenerator with tools and runs a multi-step loop: call LLM, execute requested tools, append results, repeat until the model replies with text (no tool calls) or `max_agent_steps` is reached. Each step runs the LLM once and executes all tool calls the model requested in that turn. The Agent returns full conversation history, token usage, step count, and an `exit_reason` (`"text"`, `"max_agent_steps"`, tool name, etc.). `State` management with user-defined schemas, hooks (`BEFORE_RUN`, `BEFORE_TOOL`, `AFTER_TOOL`, etc.), and `Toolset` subclassing provide extensible agent orchestration.

**Citations / source attribution** – There is no built-in citation-formatting mechanism in the generator components themselves. Source attribution must be implemented by the application pipeline — typically by passing retrieved documents alongside the query in the prompt template. The `FaithfulnessEvaluator` can verify whether generated statements are inferable from provided contexts.

**Structured output** – The `OpenAIChatGenerator` supports OpenAI's `response_format` parameter including Pydantic model or JSON schema validation (OpenAI Structured Outputs).

> **Editor's note.** Correction: citations are partly built in. AnswerBuilder accepts a reference_pattern (e.g. \[(\d+)\]) that maps numbered references in the reply back to the input documents and flags them as referenced.

Citations: [haystack/components/generators/chat/openai.py:57-409](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/generators/chat/openai.py#L57-L409) · [haystack/components/agents/agent.py:831-921](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/agents/agent.py#L831-L921) · [haystack/components/evaluators/faithfulness.py:54-239](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/evaluators/faithfulness.py#L54-L239) · [haystack/components/generators/chat/openai.py:156-179](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/generators/chat/openai.py#L156-L179)

### Cinnamon/kotaemon (answered)

Answers are generated via `AnswerWithContextPipeline` (or `AnswerWithInlineCitation`) which assembles the retrieved evidence and user question into a prompt (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, lines 40–80). Four prompt templates exist: text, table, figure, and chatbot. The chosen template is populated with `{context}`, `{question}`, and `{lang}` variables. The conversation history (up to `n_last_interactions`) is included as message turn pairs (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, lines 234–240). The system prompt is configurable via UI.

Streaming is the default path: the LLM's `.stream()` method yields `Document(channel="chat")` tokens in real-time. If streaming is unsupported, it falls back to `.invoke()` (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, lines 262–272).

Citation/attribution works two ways:
- **Highlight citation**: An LLM function-calling pipeline (`CitationPipeline`) extracts direct quotes from the evidence using `CiteEvidence` schema, then `match_evidence_with_context()` uses SequenceMatcher to find spans in original docs for PDF viewer highlighting (`libs/kotaemon/kotaemon/indices/qa/citation.py`, lines 22–95).
- **Inline citation**: `AnswerWithInlineCitation` prompts the LLM to output `【N】` markers in the answer alongside START_PHRASE/END_PHRASE delimiting exact spans, then parses and links them to the source documents (`libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py`, lines 70–361).

Multimodal QA is supported: when `evidence_mode` is figure and `use_multimodal` is true, image URLs are injected into the HumanMessage (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, lines 242–256). An optional LLMTrulensScoring relevance score is computed per document in a background thread (`libs/ktem/ktem/reasoning/simple.py`, lines 297–306). Mindmap generation and citation embedding visualizations run as optional background threads.

Agentic/multi-step answering is implemented via **ReAct** (`libs/ktem/ktem/reasoning/react.py`) and **ReWOO** (`libs/ktem/ktem/reasoning/rewoo/`) agents. The ReAct agent uses the Think-Action-Observation loop with tools (doc search, Wikipedia, Google, LLM, MCP tools).


Citations: [libs/kotaemon/kotaemon/indices/qa/citation_qa.py:190-294](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/qa/citation_qa.py#L190-L294) · [libs/kotaemon/kotaemon/indices/qa/citation.py:22-95](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/qa/citation.py#L22-L95) · [libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py:86-105](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py#L86-L105) · [libs/ktem/ktem/reasoning/simple.py:281-331](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/ktem/ktem/reasoning/simple.py#L281-L331)

### HKUDS/RAG-Anything (answered)

**Prompt assembly.** Generation is delegated to LightRAG's `aquery()` method, which takes a `system_prompt` parameter (`query.py:129-131`). RAG-Anything passes through any `system_prompt` the user provides. For VLM-enhanced queries, a hardcoded system prompt is used: `"You are a helpful assistant that can analyze both text and image content to provide comprehensive answers."` (`query.py:817`), optionally extended by user-provided prompts. For multimodal `aquery_with_multimodal()`, the system uses extensive prompt templates defined in `prompt.py` — separate prompts for `IMAGE_ANALYSIS_SYSTEM`, `TABLE_ANALYSIS_SYSTEM`, `EQUATION_ANALYSIS_SYSTEM`, and `GENERIC_ANALYSIS_SYSTEM` — to generate structured JSON descriptions (with `detailed_description`, `entity_info` fields) that are then prepended to the user query.

**Citations / source attribution.** The processed chunks and entities are stored by LightRAG with provenance: each text chunk records `page_idx` and `page_idx_end` via the page-provenance system (`processor.py:251-385`). Multimodal chunks are associated with a `doc_id` and file path. The `QueryParam` supports document-level source tracking, but RAG-Anything does not implement explicit citation rendering in generated answers.

**Streaming.** The `aquery()` method accepts `stream` in `**kwargs` (`query.py:284`), which is passed through to LightRAG's `QueryParam`. However, RAG-Anything's own callback/response handling does not explicitly support streaming responses — the flag is mainly used to disable caching (`use_cache = not kwargs.get("stream", False)`).

**Agentic or multi-step answering.** RAG-Anything does not implement agentic reasoning loops or multi-step answering. It is a single-turn retrieval-and-generate pipeline. The multimodal query pipeline (`aquery_with_multimodal`) does have a two-step flow (pre-process multimodal content → enhanced query → retrieve → generate) but this is not adaptive or iterative.


Citations: [raganything/query.py:129-197](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L129-L197) · [raganything/query.py:284-285](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L284-L285) · [raganything/query.py:751-830](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L751-L830) · [raganything/prompt.py:67-100](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/prompt.py#L67-L100) · [raganything/processor.py:251-265](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/processor.py#L251-L265)

### SciPhi-AI/R2R (answered)

**Prompt assembly.** System prompts are stored in the database via `PostgresPromptsHandler`. The `Agent` base class (`py/core/base/agent/agent.py:81-97`) loads a static prompt (e.g., `static_rag_agent`) and a dynamic reasoning prompt at startup. The `R2RAgent.arun()` (`py/core/agent/base.py:100-151`) assembles messages from the conversation history, passes them to `llm_provider.aget_completion()`, and iterates up to `max_iterations` (default 10). `GenerationConfig` (`py/shared/abstractions/llm.py`) controls model, temperature, max_tokens, streaming, and extended thinking.

**Citations / source attribution.** During streaming (`py/core/agent/base.py:351-638`), the `R2RStreamingAgent` tracks a `SearchResultsCollector` and `CitationTracker`. As the LLM streams tokens, `find_new_citation_spans()` detects bracket references like `[abc1234]` in the text. Each citation emits an SSE `citation_event` with the short ID, span boundaries, and full source payload on first occurrence. On `finish_reason=stop`, consolidated citations with all spans and payloads are emitted in a `final_answer_event` (`py/core/agent/base.py:599-633`). Citations are stored in `Message.metadata` for persistence.

**Streaming.** The `R2RStreamingAgent` (`py/core/agent/base.py:351-638`) yields SSE-formatted events: `thinking` (model chain-of-thought), `message` (partial text tokens), `tool_call` (tool name/arguments), `citation` (source reference), and `final_answer` (complete answer with structured citations). The endpoint returns `text/event-stream` content type via FastAPI's `StreamingResponse` (`py/core/main/api/v3/retrieval_router.py:408-427`).

**Agentic / multi-step.** `AgentFactory` (`py/core/main/services/retrieval_service.py:60-245`) creates one of 10 agent types based on mode (RAG or research), streaming preference, and XML tool format. `R2RRAGAgent` and `R2RResearchAgent` support tool usage via `ToolRegistry` — tools include `web_search`, `web_scrape`, `search_file_knowledge`, `get_file_content`, `reasoning`, `critique`, `python_executor`. The research agent adds its own reasoning/critique/Python tools (`py/core/agent/research.py:35-80`).


Citations: [py/core/base/agent/agent.py:81-98](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/base/agent/agent.py#L81-L98) · [py/core/agent/base.py:100-151](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/agent/base.py#L100-L151) · [py/core/agent/base.py:351-638](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/agent/base.py#L351-L638) · [py/core/main/services/retrieval_service.py:60-245](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/main/services/retrieval_service.py#L60-L245) · [py/core/main/api/v3/retrieval_router.py:408-427](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/main/api/v3/retrieval_router.py#L408-L427) · [py/core/agent/research.py:35-80](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/agent/research.py#L35-L80)
