LLMs Technical Reviews

How are answers generated and grounded?

Prompt assembly; citations / source attribution; streaming; agentic or multi-step answering.

Verdict

RAGFlow does the most to ground answers, because it attaches citations even when the model leaves them out. Kotaemon shows evidence most clearly. R2R has the cleanest citation stream for building your own product.

The engine enforces citations. If the model writes no [ID:n] markers, RAGFlow embeds each answer sentence and attaches up to four chunks by cosine similarity. Its reasoning levels add a planner graph or a tool-using agent. R2R turns bracketed short ids into Citation objects. When streaming, it emits a citation event the first time each id appears. Onyx gives the model numbered documents inside an agent loop of up to six cycles. A stream processor turns [1] markers into links. In highlight mode, Kotaemon runs a separate function-calling pass that extracts verbatim quotes and fuzzy-matches them to chunk spans in its PDF viewer. PageIndex has the model write <cite doc page/> tags. These are page-level in local mode.

Sources attached, citation left to you. AnythingLLM returns retrieved chunks as sources. To fit the context window, it cuts text out of the middle of the system prompt, history or user message. LlamaIndex returns source_nodes with every response and uses compact-and-refine synthesis by default. CitationQueryEngine re-splits sources into numbered 512-token chunks. Haystack’s AnswerBuilder can map [n] references back to documents through reference_pattern. Its Agent handles multi-step answering, with hooks.

Delegated or absent. RAG-Anything leaves the answer to LightRAG and does not render citations. Its VLM-enhanced query puts retrieved images into the prompt. Quivr does not generate answers at all. It returns segments with Unicode code-point offsets, and its MCP tools tell an external agent how to cite them.

Pick: RAGFlow or Onyx when every answer must be traceable without depending on the model. Pick: Kotaemon when users need to see the highlighted evidence. Pick: R2R or Quivr when you build your own front end on top of an API.

Per-project answers

infiniflow/ragflow

answered

Pipeline architecture: The ChatPipelineService.AsyncChat() method (service/chat_pipeline.go:204-280) implements the full generation pipeline. After entry validation (non-empty messages, last role=user), it resolves the LLM model config with max_tokens, sets up Langfuse tracing, binds models (embedding, rerank, chat, TTS), and performs optional tool-call session binding.

Prompt assembly: Phase 7 (line 188-189) resolves prompt parameters from chat.PromptConfig — including the {knowledge} placeholder auto-fill, system prompt templates, empty-response handling, and citation prompt formatting. Prompts are loaded from rag/prompts/ markdown files via LoadPrompt() (service/load_prompt.go), which caches them from the filesystem. Phase 8 performs LLM-based query refinement (multi-turn context, cross-language translation, metadata filtering, keyword extraction).

Citations / source attribution: The InsertCitations function (service/citation.go:68-80) implements the citation decoration algorithm: it splits the answer into sentences, encodes each into a vector via the embedding model, computes cosine similarity with chunk vectors, and applies threshold descent (0.63→0.3×0.8 per round). Up to 4 chunks can be cited per sentence, with [ID:n] markers inserted into the output. The DoRefer flag on the Chat entity controls whether citation is enabled. The citation prompt (citation_prompt.md in rag/prompts/) instructs the LLM to reference sources by chunk ID. The agentic RAG mode additionally normalizes malformed citation formats (citation.go:43-53).

Streaming: The pipeline supports both streaming and non-streaming modes (stream bool parameter). In streaming mode, it yields AsyncChatResult deltas containing Answer, optional Reasoning (chain-of-thought routed to delta.reasoning_content), Reference metadata, and structured ThinkEvent steps for agentic RAG reasoning traceability.

Agentic / multi-step answering: When the reasoning level is set to agentic and knowledge bases are configured (chat_pipeline.go:267-268), the system runs agentic RAG (internal/rag/agentic-rag/agentic_rag.go). This uses tools defined in config/agentic_rag.yaml (think, grep_chunks, search_bm25_chunks, search_semantic_chunks, list_chunks, run_javascript, todo_write) following an "Evidence-First" philosophy. The DeepResearcher mode (rag/agentic-rag/runtime/) can recursively search, read, and verify across multiple rounds.

Structured outputs: The structured_output_prompt.md template supports forcing JSON-formatted outputs from the LLM, and the generation pipeline includes a messageFitIn check that truncates context to stay within 95% of the model's token budget before the LLM call.

Editor's note. Correction: the conf/agentic_rag.yaml tools (grep_chunks, search_bm25_chunks, …) belong to the level-5 agent in internal/agentic_rag/, not to internal/rag/agentic-rag/agentic_rag.go, which is the levels 1-4 graph. Embedding-based InsertCitations runs only when the LLM did not already write citation markers (decorateAnswer). DeepResearcher is not wired into the chat path.

Mintplex-Labs/anything-llm

answered

Prompt assembly. The streamChatWithWorkspace() function builds a system prompt by calling chatPrompt() which takes the workspace's configured prompt template, expands system-prompt variables, and appends user memories (if enabled) as a ## Things I Remember About You section (scoped per user, reranked via the native cross-encoder when there are many memories). Context from RAG is appended as numbered [CONTEXT N]/[END CONTEXT N] blocks via #appendContext(). The chat history is formatted as alternating user/assistant messages. The final message array is built by each provider's constructPrompt() method: a system role message, the history messages, and a user role message.

Compression. The messageArrayCompressor (or messageStringCompressor for string-mode models) enforces token-budget proportions: system 15%, history 15%, user 70%. It uses a "cannonball" strategy — when a component exceeds its limit, it truncates from the middle outwards, inserting a --prompt truncated for brevity-- marker. History is the most aggressively compressed component, with a preference for keeping the 3 most recent exchanges even if they must be cannonballed individually.

Citation / source attribution. Vector search results carry their chunk metadata (title, source URL, published date, score) which are returned alongside the text response as sources. The chunk metadata is included in the context block sent to the LLM, and the UI renders these sources as clickable citations. The fillSourceWindow mechanism ensures the context window has consistently topN sources for follow-up questions by backfilling from previous turns.

Streaming. Every LLM provider implements streamGetChatCompletion() and handleStream(). The response is delivered as SSE (Server-Sent Events) with chunk types: textResponseChunk, textResponse, abort, finalizeResponseStream. OpenAI's Responses API streams reasoning summary as thinking blocks. Cost metrics (prompt tokens, completion tokens, duration, tokens/sec) are computed per turn.

Multi-step / agentic. When the workspace has agent skills enabled, grepAgents() diverts to the agent system (aibitat) which supports multi-turn tool-using agents with MCP server integration, Gmail/Google Calendar/Outlook skills, SQL agents, file creation, web browsing, and scheduled-job creation. These run as independent subagents with their own provider connections.

run-llama/llama_index

answered

Prompt assembly: The BaseSynthesizer (llama_index/core/response_synthesizers/base.py:64-107) holds a _text_qa_template (default DEFAULT_TEXT_QA_PROMPT) and _refine_template. The QA prompt receives context_str (concatenated node text at MetadataMode.LLM) and query_str. For chat models, parallel _chat_content_qa_template and _chat_content_refine_template variants use ChatMessage blocks.

Response modes (llama_index/core/response_synthesizers/type.py:4-58): REFINE iterates through nodes, building an initial answer then refining with each subsequent chunk. COMPACT (via CompactAndRefine at response_synthesizers/compact_and_refine.py:13-60) packs multiple chunks into the context window before refining. TREE_SUMMARIZE builds a bottom-up summary tree. GENERATION (llama_index/core/response_synthesizers/generation.py:32-50) ignores context entirely. SIMPLE_SUMMARIZE stuffs all text into one prompt. NO_TEXT and CONTEXT_ONLY return nodes as-is.

Source attribution: Every Response object (llama_index/core/base/response/schema.py:14-42) carries source_nodes: List[NodeWithScore] alongside the response text. Response.get_formatted_sources() renders truncated node content with node IDs. Source nodes flow from retrieval through synthesis unchanged.

Streaming: When streaming=True is passed, calls return StreamingResponse or AsyncStreamingResponse (llama_index/core/base/response/schema.py:108-200), which wrap a generator/async generator yielding tokens. The response also carries source_nodes for display.

Agentic/multi-step answering: The AgentWorkflow and FunctionAgent (llama_index/core/agent/) support tool-calling agents that can iterate, reflect, and call retrieval tools. SubQuestionQueryEngine decomposes queries into parallel sub-questions. ReActAgent implements the ReAct pattern with chat formatters and output parsers.

The-Vibe-Company/quivr

answered

Quivr does not implement answer generation. It is a retrieval engine, not an LLM orchestrator — there is no prompt assembly, no LLM call for answer synthesis, no citation formatting, and no streaming of generated text. Instead, it returns ranked segments with precise provenance suitable for an external generator. Every search hit includes record_id, version_id, part_key, segment_id, and an excerpt with exact Unicode code-point [start, end) offsets into the source Part text (transport/httpapi/search.go:75-77). This allows an external LLM frontend to cite sources with verifiable precision. The MCP interface (quivr mcp command) exposes these same capabilities to AI agents — the read profile provides search and read-record tools, and its instructions explicitly tell agents how to cite hits by record_id, version_id, part_key and offset ranges (online/mcp_catalogue.go:41-44). The ingest profile adds document ingestion. The monitoring/evaluation subsystem (monitoring/monitoring.go) supports Saved Queries and Subscriptions that evaluate new documents against persistent queries, producing Matches delivered via webhook — but this is alerting, not generation. For answer generation, users are expected to call Quivr's search API from a separate application that handles prompt construction, context assembly, LLM inference, and citation rendering.

VectifyAI/PageIndex

answered

Prompt assembly. Generation uses the OpenAI Agents SDK or Anthropic SDK tool runner (local_chat.py). The system prompt is assembled from three parts: CHAT_HEADER ("You are PageIndex by Vectify AI…"), the _base_instructions (tool descriptions and discovery workflow), and the client's instructions plus any system/developer messages from the conversation history. The _managed_instructions function combines these (local_chat.py:28-33). For the chat_completions lane, user messages are validated and system rows are extracted into the instructions (_split_chat_messages, line 50-85).

Document grounding. A targeting block is prepended as the first user message — it lists the document(s) with metadata and directs the agent to use the tools (local_chat.py:645-647). The agent must use tools (get_document_structure, get_page_content) to read from actual documents; it cannot answer from general knowledge. The tool descriptions instruct the agent: "Cite only statements supported by tool outputs" and "Never fill a gap from general knowledge" (agent_tools.py:1636-1644).

Citations. When citations=True, a citation prompt (the cited_answer MCP server prompt or its local frozen copy) is appended to instructions, teaching the model to emit <cite doc="..." page="..."/> tags after each claim (agent_tools.py:1634-1655, client.py:1391-1402). The client parses these tags into structured citation objects (client.py:48-88). Cloud mode supports block-level citations (block_id); local mode is page-level only.

Streaming. chat(stream=True) returns a ChatStream — iterating yields text chunks with the agent's process woven in (thinking, tool calls, results) when show_process=True (default). The raw event stream via .events yields typed dicts: thinking, answer, tool_call, tool_result (local_chat.py:721-773). A sync pump drives an async generator on a background thread (_stream_sync, line 96-160). The managed endpoint's stream uses its own chunk parsing (_cloud_chunk_events, line 845-888).

Protocol lanes. Three wire protocols: (1) The default answer lane — simplified Chat Completions envelope; (2) protocol="responses" — native OpenAI Responses API with full transcript (tool calls/results as native items), enabling prompt-cache continuity across turns; (3) protocol="messages" — native Anthropic Messages API with the SDK's tool runner, cache-control markings, and cross-turn aggregated usage. All three support multi-turn via tool_use/tool_result round-trips.

Agentic answering. The answering agent uses max_turns (default 10) to limit the reasoning loop (local_chat.py:404-408). It can iteratively explore the document structure, read specific pages, and refine its understanding. The OpenAI Agents SDK Runner handles the turn loop; the Anthropic SDK's own tool_runner does the same for the Messages lane.

onyx-dot-app/onyx

answered

Answer generation entry point: _stream_chat_turn (process_message.py:1650–1750) orchestrates the entire streaming answer generation flow. It calls build_chat_turn for setup, then _run_models which invokes the LLM.

LLM loop with multi-turn tool calling: run_llm_loop (chat/llm_loop.py, lines 787–1467) runs up to MAX_LLM_CYCLES (default 6, configurable) iterations. Each cycle: (1) constructs the message history including system prompt, prior conversation, context files, and tool definitions; (2) calls the LLM via run_llm_step; (3) if the LLM returns tool calls, executes them via run_tool_calls; (4) feeds tool responses back as additional history; (5) breaks when the LLM returns a final answer with no tool calls. Available tools include SearchTool (internal and web search), OpenURLTool, ImageGenerationTool, PythonTool (code interpreter), FileReaderTool, and user-defined custom tools.

Prompt assembly: construct_message_history (llm_loop.py, lines 355–637) assembles the message list in this order: [system prompt] [truncated chat history] [custom agent prompt] [context files / project docs] [last user message] [prior tool rounds] [oversized-file metadata] [reminder]. A token budget is computed from the LLM's context window minus tool definitions, then history is truncated from the top to fit. Prompt caching markers are set on the static prefix (system prompt, context files, cacheable history).

Citations and source attribution: The DynamicCitationProcessor (chat/citation_processor.py) processes the LLM stream in real-time: in HYPERLINK mode, it converts citation markers like [1] into hyperlinks formatted as [[1]](url), emitting CitationInfo objects. In REMOVE mode it strips citations from the answer. Citation mappings come from search tool responses (SearchDoc → URL). The run_llm_loop tracks gathered_documents from search tools and the citation_processor is updated with update_citation_processor_from_tool_response after each tool cycle.

Streaming: The entire answer is streamed as Packet objects via an Emitter. Packets include answer tokens (AgentResponseDelta), citations (CitationInfo), tool call metadata (ToolCallDebug), and stop signals (OverallStop). Multiple models can be compared side-by-side via handle_multi_model_stream.

Query rewriting: Before retrieval/generation, semantic_query_rephrase and keyword_query_expansion use an LLM to rewrite the user's query for better search recall, incorporating chat history, user memories, and user information.

Agentic / multi-step answering: The LLM can autonomously decide tool call ordering across multiple cycles. The system supports deep research (run_deep_research_llm_loop) for complex multi-step research queries.

deepset-ai/haystack

answered

Prompt assembly – Haystack uses ChatPromptBuilder and PromptBuilder (haystack/components/builders/) with Jinja2 templates to assemble prompts. The OpenAIChatGenerator (haystack/components/generators/chat/openai.py:57-409) works with ChatMessage objects (with user, assistant, system, tool roles). Generation parameters like temperature, max_completion_tokens, response_format, and stop sequences can be set at init or merged per-key at run time.

Streaming – The OpenAIChatGenerator accepts a streaming_callback: StreamingCallbackT parameter both in __init__ and at run() time (line 348). When a callback is provided, the generator streams ChatCompletionChunk objects through the callback function via OpenAI's Stream response, assembling the full message from chunks (_handle_stream_response). The StreamingChunk dataclass carries each delta: content, delta, and tool_call data.

Tool calling / agentic answering – The Agent component (haystack/components/agents/agent.py:831-921) wraps a ChatGenerator with tools and runs a multi-step loop: call LLM, execute requested tools, append results, repeat until the model replies with text (no tool calls) or max_agent_steps is reached. Each step runs the LLM once and executes all tool calls the model requested in that turn. The Agent returns full conversation history, token usage, step count, and an exit_reason ("text", "max_agent_steps", tool name, etc.). State management with user-defined schemas, hooks (BEFORE_RUN, BEFORE_TOOL, AFTER_TOOL, etc.), and Toolset subclassing provide extensible agent orchestration.

Citations / source attribution – There is no built-in citation-formatting mechanism in the generator components themselves. Source attribution must be implemented by the application pipeline — typically by passing retrieved documents alongside the query in the prompt template. The FaithfulnessEvaluator can verify whether generated statements are inferable from provided contexts.

Structured output – The OpenAIChatGenerator supports OpenAI's response_format parameter including Pydantic model or JSON schema validation (OpenAI Structured Outputs).

Editor's note. Correction: citations are partly built in. AnswerBuilder accepts a reference_pattern (e.g. [(\d+)]) that maps numbered references in the reply back to the input documents and flags them as referenced.

Cinnamon/kotaemon

answered

Answers are generated via AnswerWithContextPipeline (or AnswerWithInlineCitation) which assembles the retrieved evidence and user question into a prompt (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 40–80). Four prompt templates exist: text, table, figure, and chatbot. The chosen template is populated with {context}, {question}, and {lang} variables. The conversation history (up to n_last_interactions) is included as message turn pairs (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 234–240). The system prompt is configurable via UI.

Streaming is the default path: the LLM's .stream() method yields Document(channel="chat") tokens in real-time. If streaming is unsupported, it falls back to .invoke() (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 262–272).

Citation/attribution works two ways:

  • Highlight citation: An LLM function-calling pipeline (CitationPipeline) extracts direct quotes from the evidence using CiteEvidence schema, then match_evidence_with_context() uses SequenceMatcher to find spans in original docs for PDF viewer highlighting (libs/kotaemon/kotaemon/indices/qa/citation.py, lines 22–95).
  • Inline citation: AnswerWithInlineCitation prompts the LLM to output 【N】 markers in the answer alongside START_PHRASE/END_PHRASE delimiting exact spans, then parses and links them to the source documents (libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py, lines 70–361).

Multimodal QA is supported: when evidence_mode is figure and use_multimodal is true, image URLs are injected into the HumanMessage (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 242–256). An optional LLMTrulensScoring relevance score is computed per document in a background thread (libs/ktem/ktem/reasoning/simple.py, lines 297–306). Mindmap generation and citation embedding visualizations run as optional background threads.

Agentic/multi-step answering is implemented via ReAct (libs/ktem/ktem/reasoning/react.py) and ReWOO (libs/ktem/ktem/reasoning/rewoo/) agents. The ReAct agent uses the Think-Action-Observation loop with tools (doc search, Wikipedia, Google, LLM, MCP tools).

HKUDS/RAG-Anything

answered

Prompt assembly. Generation is delegated to LightRAG's aquery() method, which takes a system_prompt parameter (query.py:129-131). RAG-Anything passes through any system_prompt the user provides. For VLM-enhanced queries, a hardcoded system prompt is used: "You are a helpful assistant that can analyze both text and image content to provide comprehensive answers." (query.py:817), optionally extended by user-provided prompts. For multimodal aquery_with_multimodal(), the system uses extensive prompt templates defined in prompt.py — separate prompts for IMAGE_ANALYSIS_SYSTEM, TABLE_ANALYSIS_SYSTEM, EQUATION_ANALYSIS_SYSTEM, and GENERIC_ANALYSIS_SYSTEM — to generate structured JSON descriptions (with detailed_description, entity_info fields) that are then prepended to the user query.

Citations / source attribution. The processed chunks and entities are stored by LightRAG with provenance: each text chunk records page_idx and page_idx_end via the page-provenance system (processor.py:251-385). Multimodal chunks are associated with a doc_id and file path. The QueryParam supports document-level source tracking, but RAG-Anything does not implement explicit citation rendering in generated answers.

Streaming. The aquery() method accepts stream in **kwargs (query.py:284), which is passed through to LightRAG's QueryParam. However, RAG-Anything's own callback/response handling does not explicitly support streaming responses — the flag is mainly used to disable caching (use_cache = not kwargs.get("stream", False)).

Agentic or multi-step answering. RAG-Anything does not implement agentic reasoning loops or multi-step answering. It is a single-turn retrieval-and-generate pipeline. The multimodal query pipeline (aquery_with_multimodal) does have a two-step flow (pre-process multimodal content → enhanced query → retrieve → generate) but this is not adaptive or iterative.

SciPhi-AI/R2R

answered

Prompt assembly. System prompts are stored in the database via PostgresPromptsHandler. The Agent base class (py/core/base/agent/agent.py:81-97) loads a static prompt (e.g., static_rag_agent) and a dynamic reasoning prompt at startup. The R2RAgent.arun() (py/core/agent/base.py:100-151) assembles messages from the conversation history, passes them to llm_provider.aget_completion(), and iterates up to max_iterations (default 10). GenerationConfig (py/shared/abstractions/llm.py) controls model, temperature, max_tokens, streaming, and extended thinking.

Citations / source attribution. During streaming (py/core/agent/base.py:351-638), the R2RStreamingAgent tracks a SearchResultsCollector and CitationTracker. As the LLM streams tokens, find_new_citation_spans() detects bracket references like [abc1234] in the text. Each citation emits an SSE citation_event with the short ID, span boundaries, and full source payload on first occurrence. On finish_reason=stop, consolidated citations with all spans and payloads are emitted in a final_answer_event (py/core/agent/base.py:599-633). Citations are stored in Message.metadata for persistence.

Streaming. The R2RStreamingAgent (py/core/agent/base.py:351-638) yields SSE-formatted events: thinking (model chain-of-thought), message (partial text tokens), tool_call (tool name/arguments), citation (source reference), and final_answer (complete answer with structured citations). The endpoint returns text/event-stream content type via FastAPI's StreamingResponse (py/core/main/api/v3/retrieval_router.py:408-427).

Agentic / multi-step. AgentFactory (py/core/main/services/retrieval_service.py:60-245) creates one of 10 agent types based on mode (RAG or research), streaming preference, and XML tool format. R2RRAGAgent and R2RResearchAgent support tool usage via ToolRegistry — tools include web_search, web_scrape, search_file_knowledge, get_file_content, reasoning, critique, python_executor. The research agent adds its own reasoning/critique/Python tools (py/core/agent/research.py:35-80).

← How is retrieval performed? · How is quality evaluated or observed? →