onyx-dot-app/onyx
Self-hosted AI chat and enterprise search: 40+ connectors index into OpenSearch, and an LLM tool-calling loop answers with citations.
Overview
Onyx (formerly Danswer) is a self-hosted chat and search application for companies. It pulls documents from workplace tools (Google Drive, Confluence, Jira, Slack, GitHub, Salesforce, Zendesk and about 50 other sources, one folder each), chunks and embeds them, and stores them in OpenSearch with their access-control lists. Users then talk to an assistant in a Next.js web UI, Slack bot, desktop app or browser extension. The assistant can search that index, search the web, open URLs, run Python, generate images or call custom and MCP tools, and it answers with numbered citations that link back to the source documents.
It is a full product, not a library. A deployment is a set of services: a FastAPI API server, Celery workers for fetching and processing documents, two model servers for embeddings, PostgreSQL, Redis, OpenSearch, an object store and the web frontend. Most of the codebase is MIT-licensed. Code under ee/ directories is under a separate enterprise licence and adds features such as permission sync and multi-tenancy.
At the pinned commit, retrieval is not a fixed “retrieve then generate” pipeline. Search is one tool in an agent loop. The LLM decides when to call internal_search, and the search tool runs its own small pipeline: query expansion, several hybrid OpenSearch queries, reciprocal rank fusion, an LLM pass that selects relevant sections, and context expansion around the selected chunks. A separate deep-research loop handles longer multi-step questions.
Architecture
flowchart LR
SRC["Connectors (Drive, Slack, Jira...)"] --> CEL["Celery docfetching / docprocessing"]
CEL --> IDX["Indexing pipeline: chunk, enrich, embed"]
IDX --> EMB["Indexing model server"]
IDX --> OS["OpenSearch hybrid index"]
IDX --> PG["PostgreSQL (docs, ACLs, settings)"]
UI["Web UI / Slack / API"] --> API["FastAPI /chat/send-chat-message"]
API --> LOOP["run_llm_loop (LiteLLM)"]
LOOP --> TOOLS["Tools: search, web, Python, MCP..."]
TOOLS --> ST["SearchTool"]
ST --> OS
ST --> QEMB["Inference model server"]
LOOP --> CIT["Citation processor -> streamed packets"]
| Component | Path | Role |
|---|---|---|
| Connectors | backend/onyx/connectors/ |
Load, poll and slim connector interfaces plus one folder per source |
| Background workers | backend/onyx/background/celery/ |
Celery apps for docfetching, docprocessing, pruning, permission sync and beat scheduling |
| Indexing pipeline | backend/onyx/indexing/ |
run_indexing_pipeline, Chunker, DefaultIndexingEmbedder, vector-DB writes |
| File processing | backend/onyx/file_processing/ |
Text extraction for PDF, Office, email, HTML; optional Unstructured API |
| Document index | backend/onyx/document_index/ |
DocumentIndex interface and its only implementation, OpenSearchDocumentIndex |
| Search | backend/onyx/context/search/ |
search_pipeline, filter building, search_chunks |
| Chat | backend/onyx/chat/ |
process_message.py (turn setup, streaming), llm_loop.py (agent loop), citation processing |
| Tools | backend/onyx/tools/ |
Search, web search, open URL, Python, image generation, file reader, custom OpenAPI and MCP tools |
| Model server | backend/model_server/ |
SentenceTransformers embedding service, run as separate inference and indexing containers |
| Web | web/ |
Next.js chat and admin UI |
How a request flows
Take a chat message that needs internal documents:
- API.
POST /chat/send-chat-messageentershandle_send_chat_message, which checks theWRITE_CHATpermission and token rate limits (chat_backend.py). It streams packets fromhandle_stream_message_objects, or fromhandle_multi_model_streamwhen several models answer side by side (process_message.py). - Turn setup.
_stream_chat_turnbuilds the turn withbuild_chat_turn(history, persona, files, tools) and then runs_run_models(process_message.py). - Agent loop.
run_llm_loopruns up toMAX_LLM_CYCLES(default 6). Each cycle pickstool_choice:REQUIREDfor a forced tool,NONEon the last cycle, otherwiseAUTO(llm_loop.py). It rebuilds the message history within the token budget, calls the model through LiteLLM, and runs any tool calls. - Query expansion. When the model calls
internal_search,SearchToolrunssemantic_query_rephrase,keyword_query_expansion, a source-scope decision and a time-filter decision in parallel, all LLM calls (search_tool.py). - Parallel searches. Each semantic, keyword and original query becomes one
search_pipelinecall, together with an optional Slack federated search. The results are merged withweighted_reciprocal_rank_fusion(search_tool.py). - Retrieval.
search_pipelinebuildsIndexFilters(ACL, document sets, time, project, hierarchy), strips stopwords for the keyword side, callssearch_chunks, and applies the EE post-query censoring hook (pipeline.py).search_chunksembeds the query and callshybrid_retrieval, or runs pure BM25 whenhybrid_alphais exactly 0 (search_runner.py). - Selection and expansion. The fused sections are trimmed to a token budget.
select_sections_for_expansionasks the LLM to pick the most relevant ones. Each selected section is expanded with neighbouring chunks, and overlapping sections are merged (search_tool.py, document_filter.py). - Answer. The tool returns numbered JSON documents to the model. The next LLM cycle writes the answer, and the citation processor turns
[1]markers into links while packets stream to the client.
Key components
Indexing pipeline
run_indexing_pipeline picks the active search settings (the FUTURE secondary index during a model switch), decides whether contextual RAG and image summarisation are on, builds a Chunker, and hands the batch to index_doc_batch_with_handler (indexing_pipeline.py). Inside, documents are upserted to Postgres, unchanged documents are skipped by a timestamp or content-hash gate, image sections are summarised by a vision LLM, chunks are optionally enriched with LLM document and chunk summaries, and the embeddings are streamed to OpenSearch with backoff.
Chunker
Chunker uses three chonkie SentenceChunkers: a 128-token blurb, the main chunk at DOC_EMBEDDING_CONTEXT_SIZE (512 tokens by default) with CHUNK_OVERLAP = 0, and optional 150-token mini-chunks for multipass indexing (chunker.py). Title and metadata text are added to each chunk. Optional “large chunks” combine several neighbours.
Embeddings
DefaultIndexingEmbedder embeds the chunk texts and, separately, each distinct document title once. The title embedding becomes its own vector field (embedder.py). The default self-hosted model is nomic-ai/nomic-embed-text-v1, served by the model server. Hosted providers can be configured instead.
OpenSearch hybrid search
DocumentIndex is an abstract interface, but get_default_document_index only builds OpenSearch indices, as a primary/secondary pair (factory.py). hybrid_retrieval sends a hybrid query and names a search pipeline (opensearch_document_index.py). The pipeline normalises scores with min-max or z-score and combines them with fixed weights: 0.1 title vector, 0.45 content vector, 0.45 keyword in the default layout (search.py). The query_type argument is ignored here, so hybrid_alpha values between 0 and 1 do not change the blend. Only 0.0 changes behaviour, by switching to keyword-only search. There is no cross-encoder reranker. The model server’s README says rerankers were dropped in favour of LLM selection.
Agent loop and citations
run_llm_loop keeps the system prompt, custom agent prompt, project files and reminders in a fixed order and truncates old history to fit the context window (llm_loop.py). Long tool outputs from earlier turns are replaced with short placeholders. The citation processor runs on the token stream in HYPERLINK or REMOVE mode.
Extending it
- Connectors. Subclass
LoadConnector,PollConnectorand/orSlimConnectorfromconnectors/interfaces.py, register the source, and add tests underbackend/tests/daily/connectors.connectors/README.mddescribes the flow. - Tools. Admins can add OpenAPI-described custom tools and MCP servers without code. New built-in tools subclass
Toolintools/interface.py. - Agents (personas). System prompt, task prompt, document sets, tools and a forced tool are per-persona settings, so most “custom RAG app” work is configuration.
- Ingestion hook.
_apply_document_ingestion_hookin the indexing pipeline lets a configured hook rewrite a document’s sections or drop the document during indexing. - Evals.
backend/onyx/evals/runs chat datasets locally or in Braintrust and checks which tools were called. Traces can go to Langfuse or Braintrust.
Running it
- Docker Compose.
deployment/docker_compose/docker-compose.ymldefinesapi_server,background,web_server,inference_model_server,indexing_model_server,relational_db(Postgres),opensearch,cache(Redis),nginx, an object store and acode-interpreterservice.install.shwraps the setup,docker-compose.prod.ymladds Let’s Encrypt, anddocker-compose.onyx-lite.ymlis a smaller variant. - Kubernetes and cloud. A Helm chart, Terraform modules and an AWS ECS Fargate template are under
deployment/. - Models. Any LiteLLM-supported chat model, configured in the admin UI. Embeddings come from the local model server or a hosted API.
- Resources. OpenSearch, two embedding servers and several Celery workers make this a heavy stack. Plan for a real server, not a laptop.
Strengths and caveats
- Strength: connector breadth with permissions. Many sources, incremental polling, pruning, and ACLs stored on every chunk. Search filters by the user’s access at query time.
- Strength: agentic search done carefully. Parallel query rewrites, RRF fusion, LLM section selection and context expansion are all in one readable
SearchTool. - Strength: operable. Zero-downtime embedding-model switches through a secondary index, a content-hash gate, chunk batch stores and tracing for every LLM flow.
- Caveat: OpenSearch only. The index interface is abstract, but there is one backend. You cannot point Onyx at an existing vector database.
- Caveat: many LLM calls per search. One
internal_searchcall can mean four parallel planning calls, a selection call and expansion calls before the answer. Latency and token cost add up, especially with contextual RAG turned on at indexing time. - Caveat: fixed fusion weights. The
hybrid_alphaknob is mostly vestigial on the OpenSearch path. Tuning the keyword/vector balance means changing code or configuration constants. - Caveat: split licence. Permission sync, multi-tenancy and some admin features live in
ee/under a non-MIT licence.
Sources: code at a8d6de7, verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit a8d6de7. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredSupported formats: The system accepts PDF, Word (docx), PowerPoint (pptx), Excel (xlsx, xlsm), plain text (txt, md, json, xml, yaml, csv, tsv, sql, conf, log), email (eml/message/rfc822), epub, HTML, and images (png, jpg, jpeg, webp). MIME type classification is in backend/onyx/file_processing/file_types.py (OnyxMimeTypes, lines 39–76).
OCR and layout parsing: Text extraction uses extract_file_text.py (line 1+), which delegates PDF to PyMuPDF (via markitdown), Office formats to python-pptx, python-docx, and openpyxl, and optionally routes documents through the Unstructured.io API (backend/onyx/file_processing/unstructured.py, lines 54–70) for richer layout parsing. Images embedded in documents are extracted and optionally summarized by a vision LLM (image_summarization.py / indexing_pipeline.py:process_image_sections, lines 851–959).
Chunking strategy: The Chunker class (backend/onyx/indexing/chunker.py, lines 124–189) uses the chonkie SentenceChunker configured by DOC_EMBEDDING_CONTEXT_SIZE tokens (default from shared_configs/). It splits on sentence boundaries with zero overlap. Each chunk prepends a title prefix and appends a metadata suffix (key-value pairs as natural language). A DocumentChunker orchestrates per-document splitting (backend/onyx/indexing/chunking/document_chunker.py). The pipeline also generates optional mini-chunks (smaller sub-chunks for multipass retrieval) and large chunks (combining LARGE_CHUNK_RATIO small chunks via generate_large_chunks, lines 111–121).
Table handling: Tabular content (CSV, xlsx) is processed by a dedicated tabular section chunker and embedded as text. Tables are extracted as text representations rather than preserving native cell structure.
Contextual RAG enrichment: During indexing, two optional LLM passes enrich chunks: a document summary (USE_DOCUMENT_SUMMARY) and a per-chunk context description (USE_CHUNK_SUMMARY). These are generated in parallel (add_contextual_summaries, indexing_pipeline.py:1104–1143) and stored alongside the chunk embedding.
How are embeddings and indexes built and stored?
answeredEmbedding models: Embedding is done by the DefaultIndexingEmbedder (backend/onyx/indexing/embedder.py, lines 89–253) which wraps an EmbeddingModel (search_nlp_models.py). Supported providers include OpenAI (text-embedding-3-*), Cohere, VoyageAI, VertexAI, and custom OpenAI-compatible APIs. The embedding server host/port is configurable (INDEXING_MODEL_SERVER_HOST/PORT). Each chunk receives a full embedding, optional mini-chunk embeddings (for multipass scoring), and a title embedding (lines 156–183). Embeddings are normalized or dimension-reduced per SearchSettings.
Vector store: The sole document index backend is OpenSearch (via the OpenSearchDocumentIndex class at backend/onyx/document_index/opensearch/opensearch_document_index.py, line 1+). It implements the DocumentIndex interface (backend/onyx/document_index/interfaces.py, lines 474–497) which requires hybrid (vector + keyword) search, metadata updates, deletion, and random retrieval.
Hybrid index structure: Each chunk is stored as an OpenSearch document with a dense vector field (content_vector), lexical text fields (content for BM25), and structured metadata fields: access_control_list, source_type, document_sets, tenant_id, cc_pair_ids, created_at, last_updated, hidden, boost, personas, user_projects, and title vectors. The schema is defined in opensearch/schema.py.
Indexing pipeline: The end-to-end flow lives in run_indexing_pipeline (indexing_pipeline.py:1641–1724). Phases: (1) upsert documents to PostgreSQL, (2) filter by change-gate (timestamp or content-hash, get_docs_to_update, lines 365–447), (3) process image sections, (4) chunk via Chunker.chunk(), (5) optionally add contextual-RAG LLM enrichment, (6) embed via IndexingEmbedder, (7) write to OpenSearch via write_chunks_to_vector_db_with_backoff (vector_db_insertion.py, line 19+). The secondary (FUTURE) index is supported for zero-downtime model migrations.
Storage: Chunks are persisted to ChunkBatchStore (temporary on-disk store) between embedding and vector-db write to decouple memory from OpenSearch latency. The document index stores chunk counts for efficient tail-truncation (IndexingMetadata.ChunkCounts).
How is retrieval performed?
answeredSearch pipeline entry: search_pipeline (backend/onyx/context/search/pipeline.py, lines 260–344) builds IndexFilters from user-provided filters, persona document sets, ACLs, time ranges, and tenant IDs. It then calls search_chunks.
Search runner: search_chunks (backend/onyx/context/search/retrieval/search_runner.py, lines 89–161) runs parallel queries: the normal hybrid/keyword search against OpenSearch, plus federated retrieval functions from external connectors (e.g. Slack). Results are deduplicated and merged by score via combine_retrieval_results (lines 27–48).
Hybrid search: _embed_and_hybrid_search (lines 51–75) embeds the query via get_query_embedding and calls document_index.hybrid_retrieval() on the OpenSearchDocumentIndex. The hybrid search combines k-NN vector search (cosine similarity) with BM25 keyword scoring via OpenSearch's hybrid query with normalization pipelines (min_max or zscore). A hybrid_alpha parameter controls the blend (≤0.2 treats as keyword-only, ≥0.8 as semantic-only). The query type classification (QueryType.KEYWORD vs SEMANTIC) also affects scoring profiles.
Keyword-only search: When hybrid_alpha=0.0, the embedding step is skipped entirely and pure BM25 retrieval runs (_keyword_search, lines 78–86).
Query rewriting: Before retrieval, semantic_query_rephrase (secondary_llm_flows/query_expansion.py, lines 73–153) rewrites the user's query into a self-contained search query using chat history context. keyword_query_expansion (lines 156–234) generates up to 3 keyword-centric search queries from the same context.
Filters: _build_index_filters (pipeline.py, lines 40–139) constructs the IndexFilters dataclass with fields for: source_type, document_set, tags, access_control_list, cc_pair_access, tenant_id, time_range, created_at_range, updated_at_range, attached_document_ids, hierarchy_node_ids, and a forced_document_set (for search UI enforced scope). Filters are enforced as AND clauses in the OpenSearch query DSL.
Reranking: Reranking is invoked separately via the rerank_pipeline at the context/search level — the OpenSearch document index supports a search_type parameter for reranking pipelines.
Post-query censoring: Some connectors (Salesforce) apply field-level permissions via post_query_chunk_censoring (EE implementation, referenced at pipeline.py, lines 334–342).
hybrid_alpha does not set the vector/keyword blend; query_type is ignored and the normalization pipeline uses fixed weights (0.1 title vector, 0.45 content vector, 0.45 keyword), so only hybrid_alpha == 0.0 (pure BM25) changes behaviour. There is also no reranker: the search tool fuses its parallel queries with weighted reciprocal rank fusion and then has the LLM select sections (select_sections_for_expansion).How are answers generated and grounded?
answeredAnswer generation entry point: _stream_chat_turn (process_message.py:1650–1750) orchestrates the entire streaming answer generation flow. It calls build_chat_turn for setup, then _run_models which invokes the LLM.
LLM loop with multi-turn tool calling: run_llm_loop (chat/llm_loop.py, lines 787–1467) runs up to MAX_LLM_CYCLES (default 6, configurable) iterations. Each cycle: (1) constructs the message history including system prompt, prior conversation, context files, and tool definitions; (2) calls the LLM via run_llm_step; (3) if the LLM returns tool calls, executes them via run_tool_calls; (4) feeds tool responses back as additional history; (5) breaks when the LLM returns a final answer with no tool calls. Available tools include SearchTool (internal and web search), OpenURLTool, ImageGenerationTool, PythonTool (code interpreter), FileReaderTool, and user-defined custom tools.
Prompt assembly: construct_message_history (llm_loop.py, lines 355–637) assembles the message list in this order: [system prompt] [truncated chat history] [custom agent prompt] [context files / project docs] [last user message] [prior tool rounds] [oversized-file metadata] [reminder]. A token budget is computed from the LLM's context window minus tool definitions, then history is truncated from the top to fit. Prompt caching markers are set on the static prefix (system prompt, context files, cacheable history).
Citations and source attribution: The DynamicCitationProcessor (chat/citation_processor.py) processes the LLM stream in real-time: in HYPERLINK mode, it converts citation markers like [1] into hyperlinks formatted as [[1]](url), emitting CitationInfo objects. In REMOVE mode it strips citations from the answer. Citation mappings come from search tool responses (SearchDoc → URL). The run_llm_loop tracks gathered_documents from search tools and the citation_processor is updated with update_citation_processor_from_tool_response after each tool cycle.
Streaming: The entire answer is streamed as Packet objects via an Emitter. Packets include answer tokens (AgentResponseDelta), citations (CitationInfo), tool call metadata (ToolCallDebug), and stop signals (OverallStop). Multiple models can be compared side-by-side via handle_multi_model_stream.
Query rewriting: Before retrieval/generation, semantic_query_rephrase and keyword_query_expansion use an LLM to rewrite the user's query for better search recall, incorporating chat history, user memories, and user information.
Agentic / multi-step answering: The LLM can autonomously decide tool call ordering across multiple cycles. The system supports deep research (run_deep_research_llm_loop) for complex multi-step research queries.
How is quality evaluated or observed?
answeredBuilt-in eval framework: Onyx has an eval system at backend/onyx/evals/eval.py (line 1+) with a module-level entry point that sends multi-turn or single-turn chat messages and evaluates the responses. The EvalConfigurationOptions (models.py, lines 86–98) configures the LLM, search permissions email, tool types to enable, dataset name, and Braintrust project/experiment name.
Eval providers: The framework supports two eval backends. Braintrust (backend/onyx/evals/providers/braintrust.py): runs evaluation tasks via the Braintrust SDK with a tool_assertion_scorer that scores whether expected tools were called, tracks pass/fail per turn, and reports MultiTurnEvalResult or EvalToolResult. It supports both local data and remote datasets from Braintrust. Local (backend/onyx/evals/providers/local.py): runs the task inline with a simple pass/fail assertion check. No persistent dashboard UI for local results.
Assertions: Per-turn ToolAssertion (models.py, lines 14–18) specifies which tools should be called and whether all must be invoked (require_all). The EvalToolResult records tools called, citations, timing, and assertion pass/fail.
Timing: EvalTimings (models.py, lines 21–29) captures total_ms, llm_first_token_ms, per-tool execution times, and stream processing time.
Tracing/Observability: All LLM calls are tagged with a standardized LLMFlow enum (backend/onyx/tracing/flows.py, lines 16–68). The tracing framework (tracing/framework/create.py) can send spans to Langfuse (backend/onyx/tracing/langfuse_tracing_processor.py) and Braintrust (backend/onyx/tracing/braintrust_tracing_processor.py). Every LLM invocation (chat, embedding, rerank, image generation, voice, intent classification) opens a generation span tagged with its flow. The indexing pipeline also has its own tracing (INDEXING_PIPELINE_TRACE_NAME).
User usage tracking: The user_usage_processor.py (tracing/processors/) tracks per-user token usage for monitoring and quota enforcement.
However, Onyx does not have built-in RAGAS-style evaluation metrics (faithfulness, answer relevancy, context precision) computed automatically. Quality evaluation relies on the Braintrust integration for structured LLM-as-judge scoring and manual assertion checks.
How is it deployed and operated?
answeredDeployment model: Onyx is a self-hosted service deployed via Docker Compose. The deployment configurations live in deployment/docker_compose/ with variants for production (docker-compose.prod.yml), development (docker-compose.dev.yml), multi-tenant (docker-compose.multitenant.yml), air-gapped, and testing setups. The base docker-compose.yml defines the core services.
Required infrastructure: The system requires PostgreSQL (relational store), Redis (caching, Celery broker, locks), OpenSearch (document/vector index), and optionally MinIO/S3 (file/blob storage). The docker-compose.resources.yml file provisions these. For production, a reverse proxy (nginx with Let's Encrypt via init-letsencrypt.sh) is included.
Service architecture: The backend comprises multiple Celery workers (primary, docfetching, docprocessing, light, heavy, monitoring, beat) orchestrated by Docker Compose. A supervisord.conf manages the FastAPI API server (main.py) and the web server inside the API container. The model server (model_server/) runs as a separate Docker image (Dockerfile.model_server).
API: The FastAPI application serves REST endpoints for search, chat, document management, connector administration, and user management. All calls go through the frontend at port 3000 (proxied to the API), as stated in CLAUDE.md.
UI: The Next.js frontend (web/) provides the admin dashboard and chat interface. A separate mobile app (mobile/, React Native + Expo) is also available.
Multi-tenancy: Enterprise Edition supports multi-tenant deployments with separate schemas per tenant. The DynamicTenantScheduler handles per-tenant Celery beat tasks. OpenSearch indices are tenant-scoped. Alembic migrations have a separate schema_private config for tenant schema migrations.
Scaling: OpenSearch can be horizontally scaled. Celery workers use thread pools with configurable concurrency per pod. Redis handles inter-process queuing. The Kubernetes helm chart (deployment/helm/) and AWS ECS Fargate (deployment/aws_ecs_fargate/) deployment options are available for larger deployments.
Lite (demo) deployment: docker-compose.onyx-lite.yml provides a minimal configuration for evaluation purposes.
Third-party dependencies: Embedding and LLM calls go through LiteLLM (multi-provider gateway). Optional Unstructured.io API for enhanced document parsing. Langfuse and Braintrust integrations for observability and evals.