LLMs Technical Reviews
Home / RAG engines / onyx

onyx-dot-app/onyx

Self-hosted AI chat and enterprise search: 40+ connectors index into OpenSearch, and an LLM tool-calling loop answers with citations.

GitHub ↗★ 32kPythonMIT (ee/ under a separate licence)commit a8d6de7 · 2026-10-06homepage ↗

Overview

Onyx (formerly Danswer) is a self-hosted chat and search application for companies. It pulls documents from workplace tools (Google Drive, Confluence, Jira, Slack, GitHub, Salesforce, Zendesk and about 50 other sources, one folder each), chunks and embeds them, and stores them in OpenSearch with their access-control lists. Users then talk to an assistant in a Next.js web UI, Slack bot, desktop app or browser extension. The assistant can search that index, search the web, open URLs, run Python, generate images or call custom and MCP tools, and it answers with numbered citations that link back to the source documents.

It is a full product, not a library. A deployment is a set of services: a FastAPI API server, Celery workers for fetching and processing documents, two model servers for embeddings, PostgreSQL, Redis, OpenSearch, an object store and the web frontend. Most of the codebase is MIT-licensed. Code under ee/ directories is under a separate enterprise licence and adds features such as permission sync and multi-tenancy.

At the pinned commit, retrieval is not a fixed “retrieve then generate” pipeline. Search is one tool in an agent loop. The LLM decides when to call internal_search, and the search tool runs its own small pipeline: query expansion, several hybrid OpenSearch queries, reciprocal rank fusion, an LLM pass that selects relevant sections, and context expansion around the selected chunks. A separate deep-research loop handles longer multi-step questions.

Architecture

flowchart LR
  SRC["Connectors (Drive, Slack, Jira...)"] --> CEL["Celery docfetching / docprocessing"]
  CEL --> IDX["Indexing pipeline: chunk, enrich, embed"]
  IDX --> EMB["Indexing model server"]
  IDX --> OS["OpenSearch hybrid index"]
  IDX --> PG["PostgreSQL (docs, ACLs, settings)"]
  UI["Web UI / Slack / API"] --> API["FastAPI /chat/send-chat-message"]
  API --> LOOP["run_llm_loop (LiteLLM)"]
  LOOP --> TOOLS["Tools: search, web, Python, MCP..."]
  TOOLS --> ST["SearchTool"]
  ST --> OS
  ST --> QEMB["Inference model server"]
  LOOP --> CIT["Citation processor -> streamed packets"]
Component Path Role
Connectors backend/onyx/connectors/ Load, poll and slim connector interfaces plus one folder per source
Background workers backend/onyx/background/celery/ Celery apps for docfetching, docprocessing, pruning, permission sync and beat scheduling
Indexing pipeline backend/onyx/indexing/ run_indexing_pipeline, Chunker, DefaultIndexingEmbedder, vector-DB writes
File processing backend/onyx/file_processing/ Text extraction for PDF, Office, email, HTML; optional Unstructured API
Document index backend/onyx/document_index/ DocumentIndex interface and its only implementation, OpenSearchDocumentIndex
Search backend/onyx/context/search/ search_pipeline, filter building, search_chunks
Chat backend/onyx/chat/ process_message.py (turn setup, streaming), llm_loop.py (agent loop), citation processing
Tools backend/onyx/tools/ Search, web search, open URL, Python, image generation, file reader, custom OpenAPI and MCP tools
Model server backend/model_server/ SentenceTransformers embedding service, run as separate inference and indexing containers
Web web/ Next.js chat and admin UI

How a request flows

Take a chat message that needs internal documents:

  1. API. POST /chat/send-chat-message enters handle_send_chat_message, which checks the WRITE_CHAT permission and token rate limits (chat_backend.py). It streams packets from handle_stream_message_objects, or from handle_multi_model_stream when several models answer side by side (process_message.py).
  2. Turn setup. _stream_chat_turn builds the turn with build_chat_turn (history, persona, files, tools) and then runs _run_models (process_message.py).
  3. Agent loop. run_llm_loop runs up to MAX_LLM_CYCLES (default 6). Each cycle picks tool_choice: REQUIRED for a forced tool, NONE on the last cycle, otherwise AUTO (llm_loop.py). It rebuilds the message history within the token budget, calls the model through LiteLLM, and runs any tool calls.
  4. Query expansion. When the model calls internal_search, SearchTool runs semantic_query_rephrase, keyword_query_expansion, a source-scope decision and a time-filter decision in parallel, all LLM calls (search_tool.py).
  5. Parallel searches. Each semantic, keyword and original query becomes one search_pipeline call, together with an optional Slack federated search. The results are merged with weighted_reciprocal_rank_fusion (search_tool.py).
  6. Retrieval. search_pipeline builds IndexFilters (ACL, document sets, time, project, hierarchy), strips stopwords for the keyword side, calls search_chunks, and applies the EE post-query censoring hook (pipeline.py). search_chunks embeds the query and calls hybrid_retrieval, or runs pure BM25 when hybrid_alpha is exactly 0 (search_runner.py).
  7. Selection and expansion. The fused sections are trimmed to a token budget. select_sections_for_expansion asks the LLM to pick the most relevant ones. Each selected section is expanded with neighbouring chunks, and overlapping sections are merged (search_tool.py, document_filter.py).
  8. Answer. The tool returns numbered JSON documents to the model. The next LLM cycle writes the answer, and the citation processor turns [1] markers into links while packets stream to the client.

Key components

Indexing pipeline

run_indexing_pipeline picks the active search settings (the FUTURE secondary index during a model switch), decides whether contextual RAG and image summarisation are on, builds a Chunker, and hands the batch to index_doc_batch_with_handler (indexing_pipeline.py). Inside, documents are upserted to Postgres, unchanged documents are skipped by a timestamp or content-hash gate, image sections are summarised by a vision LLM, chunks are optionally enriched with LLM document and chunk summaries, and the embeddings are streamed to OpenSearch with backoff.

Chunker

Chunker uses three chonkie SentenceChunkers: a 128-token blurb, the main chunk at DOC_EMBEDDING_CONTEXT_SIZE (512 tokens by default) with CHUNK_OVERLAP = 0, and optional 150-token mini-chunks for multipass indexing (chunker.py). Title and metadata text are added to each chunk. Optional “large chunks” combine several neighbours.

Embeddings

DefaultIndexingEmbedder embeds the chunk texts and, separately, each distinct document title once. The title embedding becomes its own vector field (embedder.py). The default self-hosted model is nomic-ai/nomic-embed-text-v1, served by the model server. Hosted providers can be configured instead.

DocumentIndex is an abstract interface, but get_default_document_index only builds OpenSearch indices, as a primary/secondary pair (factory.py). hybrid_retrieval sends a hybrid query and names a search pipeline (opensearch_document_index.py). The pipeline normalises scores with min-max or z-score and combines them with fixed weights: 0.1 title vector, 0.45 content vector, 0.45 keyword in the default layout (search.py). The query_type argument is ignored here, so hybrid_alpha values between 0 and 1 do not change the blend. Only 0.0 changes behaviour, by switching to keyword-only search. There is no cross-encoder reranker. The model server’s README says rerankers were dropped in favour of LLM selection.

Agent loop and citations

run_llm_loop keeps the system prompt, custom agent prompt, project files and reminders in a fixed order and truncates old history to fit the context window (llm_loop.py). Long tool outputs from earlier turns are replaced with short placeholders. The citation processor runs on the token stream in HYPERLINK or REMOVE mode.

Extending it

  • Connectors. Subclass LoadConnector, PollConnector and/or SlimConnector from connectors/interfaces.py, register the source, and add tests under backend/tests/daily/connectors. connectors/README.md describes the flow.
  • Tools. Admins can add OpenAPI-described custom tools and MCP servers without code. New built-in tools subclass Tool in tools/interface.py.
  • Agents (personas). System prompt, task prompt, document sets, tools and a forced tool are per-persona settings, so most “custom RAG app” work is configuration.
  • Ingestion hook. _apply_document_ingestion_hook in the indexing pipeline lets a configured hook rewrite a document’s sections or drop the document during indexing.
  • Evals. backend/onyx/evals/ runs chat datasets locally or in Braintrust and checks which tools were called. Traces can go to Langfuse or Braintrust.

Running it

  • Docker Compose. deployment/docker_compose/docker-compose.yml defines api_server, background, web_server, inference_model_server, indexing_model_server, relational_db (Postgres), opensearch, cache (Redis), nginx, an object store and a code-interpreter service. install.sh wraps the setup, docker-compose.prod.yml adds Let’s Encrypt, and docker-compose.onyx-lite.yml is a smaller variant.
  • Kubernetes and cloud. A Helm chart, Terraform modules and an AWS ECS Fargate template are under deployment/.
  • Models. Any LiteLLM-supported chat model, configured in the admin UI. Embeddings come from the local model server or a hosted API.
  • Resources. OpenSearch, two embedding servers and several Celery workers make this a heavy stack. Plan for a real server, not a laptop.

Strengths and caveats

  • Strength: connector breadth with permissions. Many sources, incremental polling, pruning, and ACLs stored on every chunk. Search filters by the user’s access at query time.
  • Strength: agentic search done carefully. Parallel query rewrites, RRF fusion, LLM section selection and context expansion are all in one readable SearchTool.
  • Strength: operable. Zero-downtime embedding-model switches through a secondary index, a content-hash gate, chunk batch stores and tracing for every LLM flow.
  • Caveat: OpenSearch only. The index interface is abstract, but there is one backend. You cannot point Onyx at an existing vector database.
  • Caveat: many LLM calls per search. One internal_search call can mean four parallel planning calls, a selection call and expansion calls before the answer. Latency and token cost add up, especially with contextual RAG turned on at indexing time.
  • Caveat: fixed fusion weights. The hybrid_alpha knob is mostly vestigial on the OpenSearch path. Tuning the keyword/vector balance means changing code or configuration constants.
  • Caveat: split licence. Permission sync, multi-tenancy and some admin features live in ee/ under a non-MIT licence.

Sources: code at a8d6de7, verified Q&A.

How it answers the RAG engines questions

Each answer was drafted by a code-reading agent at commit a8d6de7. Its citations were checked mechanically. Compare with the other rag engines →

How are documents parsed and chunked?

answered

Supported formats: The system accepts PDF, Word (docx), PowerPoint (pptx), Excel (xlsx, xlsm), plain text (txt, md, json, xml, yaml, csv, tsv, sql, conf, log), email (eml/message/rfc822), epub, HTML, and images (png, jpg, jpeg, webp). MIME type classification is in backend/onyx/file_processing/file_types.py (OnyxMimeTypes, lines 39–76).

OCR and layout parsing: Text extraction uses extract_file_text.py (line 1+), which delegates PDF to PyMuPDF (via markitdown), Office formats to python-pptx, python-docx, and openpyxl, and optionally routes documents through the Unstructured.io API (backend/onyx/file_processing/unstructured.py, lines 54–70) for richer layout parsing. Images embedded in documents are extracted and optionally summarized by a vision LLM (image_summarization.py / indexing_pipeline.py:process_image_sections, lines 851–959).

Chunking strategy: The Chunker class (backend/onyx/indexing/chunker.py, lines 124–189) uses the chonkie SentenceChunker configured by DOC_EMBEDDING_CONTEXT_SIZE tokens (default from shared_configs/). It splits on sentence boundaries with zero overlap. Each chunk prepends a title prefix and appends a metadata suffix (key-value pairs as natural language). A DocumentChunker orchestrates per-document splitting (backend/onyx/indexing/chunking/document_chunker.py). The pipeline also generates optional mini-chunks (smaller sub-chunks for multipass retrieval) and large chunks (combining LARGE_CHUNK_RATIO small chunks via generate_large_chunks, lines 111–121).

Table handling: Tabular content (CSV, xlsx) is processed by a dedicated tabular section chunker and embedded as text. Tables are extracted as text representations rather than preserving native cell structure.

Contextual RAG enrichment: During indexing, two optional LLM passes enrich chunks: a document summary (USE_DOCUMENT_SUMMARY) and a per-chunk context description (USE_CHUNK_SUMMARY). These are generated in parallel (add_contextual_summaries, indexing_pipeline.py:1104–1143) and stored alongside the chunk embedding.

Editor's note. Correction: PDF text is extracted with pypdfium2 in an isolated process, with pypdf as the fallback, not PyMuPDF via markitdown (backend/onyx/file_processing/extract_file_text.py).

How are embeddings and indexes built and stored?

answered

Embedding models: Embedding is done by the DefaultIndexingEmbedder (backend/onyx/indexing/embedder.py, lines 89–253) which wraps an EmbeddingModel (search_nlp_models.py). Supported providers include OpenAI (text-embedding-3-*), Cohere, VoyageAI, VertexAI, and custom OpenAI-compatible APIs. The embedding server host/port is configurable (INDEXING_MODEL_SERVER_HOST/PORT). Each chunk receives a full embedding, optional mini-chunk embeddings (for multipass scoring), and a title embedding (lines 156–183). Embeddings are normalized or dimension-reduced per SearchSettings.

Vector store: The sole document index backend is OpenSearch (via the OpenSearchDocumentIndex class at backend/onyx/document_index/opensearch/opensearch_document_index.py, line 1+). It implements the DocumentIndex interface (backend/onyx/document_index/interfaces.py, lines 474–497) which requires hybrid (vector + keyword) search, metadata updates, deletion, and random retrieval.

Hybrid index structure: Each chunk is stored as an OpenSearch document with a dense vector field (content_vector), lexical text fields (content for BM25), and structured metadata fields: access_control_list, source_type, document_sets, tenant_id, cc_pair_ids, created_at, last_updated, hidden, boost, personas, user_projects, and title vectors. The schema is defined in opensearch/schema.py.

Indexing pipeline: The end-to-end flow lives in run_indexing_pipeline (indexing_pipeline.py:1641–1724). Phases: (1) upsert documents to PostgreSQL, (2) filter by change-gate (timestamp or content-hash, get_docs_to_update, lines 365–447), (3) process image sections, (4) chunk via Chunker.chunk(), (5) optionally add contextual-RAG LLM enrichment, (6) embed via IndexingEmbedder, (7) write to OpenSearch via write_chunks_to_vector_db_with_backoff (vector_db_insertion.py, line 19+). The secondary (FUTURE) index is supported for zero-downtime model migrations.

Storage: Chunks are persisted to ChunkBatchStore (temporary on-disk store) between embedding and vector-db write to decouple memory from OpenSearch latency. The document index stores chunk counts for efficient tail-truncation (IndexingMetadata.ChunkCounts).

How is retrieval performed?

answered

Search pipeline entry: search_pipeline (backend/onyx/context/search/pipeline.py, lines 260–344) builds IndexFilters from user-provided filters, persona document sets, ACLs, time ranges, and tenant IDs. It then calls search_chunks.

Search runner: search_chunks (backend/onyx/context/search/retrieval/search_runner.py, lines 89–161) runs parallel queries: the normal hybrid/keyword search against OpenSearch, plus federated retrieval functions from external connectors (e.g. Slack). Results are deduplicated and merged by score via combine_retrieval_results (lines 27–48).

Hybrid search: _embed_and_hybrid_search (lines 51–75) embeds the query via get_query_embedding and calls document_index.hybrid_retrieval() on the OpenSearchDocumentIndex. The hybrid search combines k-NN vector search (cosine similarity) with BM25 keyword scoring via OpenSearch's hybrid query with normalization pipelines (min_max or zscore). A hybrid_alpha parameter controls the blend (≤0.2 treats as keyword-only, ≥0.8 as semantic-only). The query type classification (QueryType.KEYWORD vs SEMANTIC) also affects scoring profiles.

Keyword-only search: When hybrid_alpha=0.0, the embedding step is skipped entirely and pure BM25 retrieval runs (_keyword_search, lines 78–86).

Query rewriting: Before retrieval, semantic_query_rephrase (secondary_llm_flows/query_expansion.py, lines 73–153) rewrites the user's query into a self-contained search query using chat history context. keyword_query_expansion (lines 156–234) generates up to 3 keyword-centric search queries from the same context.

Filters: _build_index_filters (pipeline.py, lines 40–139) constructs the IndexFilters dataclass with fields for: source_type, document_set, tags, access_control_list, cc_pair_access, tenant_id, time_range, created_at_range, updated_at_range, attached_document_ids, hierarchy_node_ids, and a forced_document_set (for search UI enforced scope). Filters are enforced as AND clauses in the OpenSearch query DSL.

Reranking: Reranking is invoked separately via the rerank_pipeline at the context/search level — the OpenSearch document index supports a search_type parameter for reranking pipelines.

Post-query censoring: Some connectors (Salesforce) apply field-level permissions via post_query_chunk_censoring (EE implementation, referenced at pipeline.py, lines 334–342).

Editor's note. Correction: on the OpenSearch path hybrid_alpha does not set the vector/keyword blend; query_type is ignored and the normalization pipeline uses fixed weights (0.1 title vector, 0.45 content vector, 0.45 keyword), so only hybrid_alpha == 0.0 (pure BM25) changes behaviour. There is also no reranker: the search tool fuses its parallel queries with weighted reciprocal rank fusion and then has the LLM select sections (select_sections_for_expansion).

How are answers generated and grounded?

answered

Answer generation entry point: _stream_chat_turn (process_message.py:1650–1750) orchestrates the entire streaming answer generation flow. It calls build_chat_turn for setup, then _run_models which invokes the LLM.

LLM loop with multi-turn tool calling: run_llm_loop (chat/llm_loop.py, lines 787–1467) runs up to MAX_LLM_CYCLES (default 6, configurable) iterations. Each cycle: (1) constructs the message history including system prompt, prior conversation, context files, and tool definitions; (2) calls the LLM via run_llm_step; (3) if the LLM returns tool calls, executes them via run_tool_calls; (4) feeds tool responses back as additional history; (5) breaks when the LLM returns a final answer with no tool calls. Available tools include SearchTool (internal and web search), OpenURLTool, ImageGenerationTool, PythonTool (code interpreter), FileReaderTool, and user-defined custom tools.

Prompt assembly: construct_message_history (llm_loop.py, lines 355–637) assembles the message list in this order: [system prompt] [truncated chat history] [custom agent prompt] [context files / project docs] [last user message] [prior tool rounds] [oversized-file metadata] [reminder]. A token budget is computed from the LLM's context window minus tool definitions, then history is truncated from the top to fit. Prompt caching markers are set on the static prefix (system prompt, context files, cacheable history).

Citations and source attribution: The DynamicCitationProcessor (chat/citation_processor.py) processes the LLM stream in real-time: in HYPERLINK mode, it converts citation markers like [1] into hyperlinks formatted as [[1]](url), emitting CitationInfo objects. In REMOVE mode it strips citations from the answer. Citation mappings come from search tool responses (SearchDoc → URL). The run_llm_loop tracks gathered_documents from search tools and the citation_processor is updated with update_citation_processor_from_tool_response after each tool cycle.

Streaming: The entire answer is streamed as Packet objects via an Emitter. Packets include answer tokens (AgentResponseDelta), citations (CitationInfo), tool call metadata (ToolCallDebug), and stop signals (OverallStop). Multiple models can be compared side-by-side via handle_multi_model_stream.

Query rewriting: Before retrieval/generation, semantic_query_rephrase and keyword_query_expansion use an LLM to rewrite the user's query for better search recall, incorporating chat history, user memories, and user information.

Agentic / multi-step answering: The LLM can autonomously decide tool call ordering across multiple cycles. The system supports deep research (run_deep_research_llm_loop) for complex multi-step research queries.

How is quality evaluated or observed?

answered

Built-in eval framework: Onyx has an eval system at backend/onyx/evals/eval.py (line 1+) with a module-level entry point that sends multi-turn or single-turn chat messages and evaluates the responses. The EvalConfigurationOptions (models.py, lines 86–98) configures the LLM, search permissions email, tool types to enable, dataset name, and Braintrust project/experiment name.

Eval providers: The framework supports two eval backends. Braintrust (backend/onyx/evals/providers/braintrust.py): runs evaluation tasks via the Braintrust SDK with a tool_assertion_scorer that scores whether expected tools were called, tracks pass/fail per turn, and reports MultiTurnEvalResult or EvalToolResult. It supports both local data and remote datasets from Braintrust. Local (backend/onyx/evals/providers/local.py): runs the task inline with a simple pass/fail assertion check. No persistent dashboard UI for local results.

Assertions: Per-turn ToolAssertion (models.py, lines 14–18) specifies which tools should be called and whether all must be invoked (require_all). The EvalToolResult records tools called, citations, timing, and assertion pass/fail.

Timing: EvalTimings (models.py, lines 21–29) captures total_ms, llm_first_token_ms, per-tool execution times, and stream processing time.

Tracing/Observability: All LLM calls are tagged with a standardized LLMFlow enum (backend/onyx/tracing/flows.py, lines 16–68). The tracing framework (tracing/framework/create.py) can send spans to Langfuse (backend/onyx/tracing/langfuse_tracing_processor.py) and Braintrust (backend/onyx/tracing/braintrust_tracing_processor.py). Every LLM invocation (chat, embedding, rerank, image generation, voice, intent classification) opens a generation span tagged with its flow. The indexing pipeline also has its own tracing (INDEXING_PIPELINE_TRACE_NAME).

User usage tracking: The user_usage_processor.py (tracing/processors/) tracks per-user token usage for monitoring and quota enforcement.

However, Onyx does not have built-in RAGAS-style evaluation metrics (faithfulness, answer relevancy, context precision) computed automatically. Quality evaluation relies on the Braintrust integration for structured LLM-as-judge scoring and manual assertion checks.

How is it deployed and operated?

answered

Deployment model: Onyx is a self-hosted service deployed via Docker Compose. The deployment configurations live in deployment/docker_compose/ with variants for production (docker-compose.prod.yml), development (docker-compose.dev.yml), multi-tenant (docker-compose.multitenant.yml), air-gapped, and testing setups. The base docker-compose.yml defines the core services.

Required infrastructure: The system requires PostgreSQL (relational store), Redis (caching, Celery broker, locks), OpenSearch (document/vector index), and optionally MinIO/S3 (file/blob storage). The docker-compose.resources.yml file provisions these. For production, a reverse proxy (nginx with Let's Encrypt via init-letsencrypt.sh) is included.

Service architecture: The backend comprises multiple Celery workers (primary, docfetching, docprocessing, light, heavy, monitoring, beat) orchestrated by Docker Compose. A supervisord.conf manages the FastAPI API server (main.py) and the web server inside the API container. The model server (model_server/) runs as a separate Docker image (Dockerfile.model_server).

API: The FastAPI application serves REST endpoints for search, chat, document management, connector administration, and user management. All calls go through the frontend at port 3000 (proxied to the API), as stated in CLAUDE.md.

UI: The Next.js frontend (web/) provides the admin dashboard and chat interface. A separate mobile app (mobile/, React Native + Expo) is also available.

Multi-tenancy: Enterprise Edition supports multi-tenant deployments with separate schemas per tenant. The DynamicTenantScheduler handles per-tenant Celery beat tasks. OpenSearch indices are tenant-scoped. Alembic migrations have a separate schema_private config for tenant schema migrations.

Scaling: OpenSearch can be horizontally scaled. Celery workers use thread pools with configurable concurrency per pod. Redis handles inter-process queuing. The Kubernetes helm chart (deployment/helm/) and AWS ECS Fargate (deployment/aws_ecs_fargate/) deployment options are available for larger deployments.

Lite (demo) deployment: docker-compose.onyx-lite.yml provides a minimal configuration for evaluation purposes.

Third-party dependencies: Embedding and LLM calls go through LiteLLM (multi-provider gateway). Optional Unstructured.io API for enhanced document parsing. Langfuse and Braintrust integrations for observability and evals.