# Mintplex-Labs/anything-llm

> Self-hosted Node.js chat-with-your-documents app with workspaces, ten vector-store backends, ~40 LLM providers and an agent mode.

- Category: [RAG engines](https://llms-technical-reviews.com/rag/)
- Repository: https://github.com/Mintplex-Labs/anything-llm (reviewed at commit `297808f9451ceb5795ed208588437f3b526a057a`, 2026-10-06)
- Stars: 66755 · Language: JavaScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/anything-llm/

## Overview

AnythingLLM is Mintplex Labs' self-hosted application for chatting with your own documents. It is a product, not a library. You run it as a desktop app or a Docker container, create **workspaces**, drop files into them, and chat with any of about forty LLM providers over the embedded chunks. Around that core it has grown a large feature surface: threads, multi-user roles with daily quotas, an embeddable chat widget, a developer REST API with an OpenAI-compatible endpoint, user memories, a rule-based model router, scheduled jobs, MCP servers and an agent mode (`@agent`) built on its own `aibitat` framework.

The code is a JavaScript monorepo with three runtime pieces: `server/` (Express API, Prisma ORM on SQLite, all RAG and agent logic), `collector/` (a separate Express process that turns files and links into text), and `frontend/` (React and Vite). Every pluggable subsystem follows the same pattern. An environment variable such as `VECTOR_DB`, `EMBEDDING_ENGINE` or `LLM_PROVIDER` is resolved by a `switch` in `server/utils/helpers/index.js` to a provider class with a shared method set. The defaults (LanceDB on disk, a local ONNX embedder, SQLite) mean a fresh install needs no external services except an LLM.

The RAG design is simple and pragmatic: recursive character chunking, dense-only top-N search with a similarity threshold, optional cross-encoder reranking (LanceDB only), and a "cannonball" middle-out truncation to fit the context window. There is no hybrid search, no query rewriting and no evaluation tooling.

## Architecture

```mermaid
flowchart LR
  UI["React frontend"] --> API["server: Express endpoints"]
  EMB["Embed widget / dev API"] --> API
  API --> CHAT["streamChatWithWorkspace"]
  CHAT --> AG["grepAgents -> aibitat (WebSocket)"]
  CHAT --> VDB["VectorDb provider"]
  CHAT --> LLM["LLM provider connector"]
  API --> COLAPI["CollectorApi"]
  COLAPI --> COL["collector: converters + OCR"]
  COL --> DOCS["storage/documents JSON"]
  API --> ADD["Document.addDocuments"]
  ADD --> DOCS
  ADD --> SPLIT["TextSplitter"]
  SPLIT --> EMBE["Embedding engine"]
  EMBE --> VDB
  API --> DB["Prisma (SQLite)"]
```

| Component | Path | Role |
|---|---|---|
| HTTP server | `server/index.js`, `server/endpoints/` | Express app, auth middleware, workspace, chat, document, admin, agent and API routes |
| Chat pipeline | `server/utils/chats/stream.js` | Slash commands, agent hand-off, context assembly, vector search, compression, streaming |
| Provider switches | `server/utils/helpers/index.js` | `getVectorDbClass`, `getLLMProvider`, `getEmbeddingEngineSelection` |
| Vector stores | `server/utils/vectorDbProviders/` | `VectorDatabase` base class plus LanceDB, PGVector, Pinecone, Chroma, ChromaCloud, Weaviate, Qdrant, Milvus, Zilliz, AstraDB |
| Embedders and reranker | `server/utils/EmbeddingEngines/`, `EmbeddingRerankers/native` | 14 embedding backends; local `Xenova/ms-marco-MiniLM-L-6-v2` cross-encoder |
| LLM connectors | `server/utils/AiProviders/` | One class per provider with `compressMessages`, `streamGetChatCompletion`, `handleStream` |
| Chunking | `server/utils/TextSplitter/` | Wrapper over LangChain's `RecursiveCharacterTextSplitter` with metadata header and prefix |
| Collector | `collector/` | File, link and connector ingestion; per-extension converters; Tesseract OCR |
| Agents | `server/utils/agents/aibitat/` | Multi-agent chat runtime, provider adapters, plugins (RAG memory, web browsing, SQL, files, Gmail, Outlook, MCP) |
| Background jobs | `server/jobs/` | Embedding worker, watched-document sync, memory extraction, scheduled agent jobs |

## How a request flows

Uploading a PDF and then asking a question in a workspace:

1. **Upload.** The frontend posts the file to `/workspace/:slug/upload`. The server calls `Collector.processDocument`, which forwards it to the collector process with a signed payload that `verifyPayloadIntegrity` checks.
2. **Convert.** `processSingleFile` checks that the path stays inside the hot directory, picks a converter by extension, falls back to plain text for unknown text types, and rejects everything else ([processSingleFile/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/index.js#L24-L91)). For `.pdf`, `asPdf` extracts text per page and runs OCR only when no text comes back ([asPDF/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/convert/asPDF/index.js#L12-L34)). The result is a JSON document with `pageContent` and metadata in server storage.
3. **Embed.** When the user moves the file into the workspace, `Document.addDocuments` loops over the files and calls `VectorDb.addDocumentToNamespace(workspace.slug, ...)`, emitting progress events as it goes ([documents.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/models/documents.js#L83-L175)). With the native embedder the same loop runs in a child-process worker, so an out-of-memory error kills the worker and not the server.
4. **Chunk and store.** The LanceDB adapter first checks a per-file vector cache. On a miss it builds a `TextSplitter` from system settings (chunk size capped by the embedder's limit, 20-character default overlap, a metadata header and an embedder prefix), embeds the chunks, and upserts them into a table named after the workspace slug ([lance/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L312-L415)).
5. **Ask.** `POST /workspace/:slug/stream-chat` opens an SSE response, enforces the multi-user daily quota, and calls `streamChatWithWorkspace` ([chat.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/endpoints/chat.js#L23-L75)).
6. **Divert or route.** Slash commands are handled first. Then `grepAgents` checks for an `@agent` handle, or for native tool calling in `automatic` mode, and hands off to the agent WebSocket if so. Otherwise the model router may pick a different connector ([stream.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L24-L101)).
7. **Assemble context.** Pinned documents (whole files, token-capped) and parsed attachments go in first. Then `performSimilaritySearch` embeds the message and fetches the top `topN` chunks above `similarityThreshold`, excluding chunks from pinned files ([stream.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L147-L222)). `fillSourceWindow` backfills from sources cited earlier in the thread when search returns fewer than `topN`. In `query` mode an empty context returns a fixed refusal instead of calling the LLM.
8. **Compress and stream.** `chatPrompt` builds the system prompt (with memories), and the connector's `compressMessages` fits system, context, history and prompt into the window. Then `streamGetChatCompletion` and `handleStream` write SSE chunks, and the chat is saved with its `sources` and metrics ([stream.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L274-L291)).

## Key components

### Vector store adapters

`VectorDatabase` in `base.js` defines the contract: `addDocumentToNamespace`, `deleteDocumentFromNamespace`, `performSimilaritySearch`, `namespaceCount`, `reset` and so on. A namespace is a workspace slug. `getVectorDbClass` picks the adapter from `VECTOR_DB` and falls back to LanceDB ([helpers/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L87-L127)). Each adapter does its own chunking and embedding instead of sharing a pipeline, so behaviour can drift between backends. The clearest case is reranking: only the LanceDB adapter implements the `rerank` path. It pulls `max(10, min(50, 10% of the table))` cosine candidates and re-scores them with the local cross-encoder ([lance/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L99-L171)). On the other nine backends the workspace "rerank" setting has no effect.

### Chunking

`TextSplitter` delegates to `RecursiveCharacterTextSplitter` with a 1,000-character default size and a 20-character overlap. After splitting, it prepends a `<document_metadata>` header (title, source, published date) to every chunk, so chunks can exceed the nominal size ([TextSplitter/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/TextSplitter/index.js#L156-L218)). Embedders such as nomic add their own `search_document:` style prefix.

### Embedders

`getEmbeddingEngineSelection` maps `EMBEDDING_ENGINE` to 14 engines and defaults to the native embedder ([helpers/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L308-L360)). The native engine runs `Xenova/all-MiniLM-L6-v2` by default through transformers.js, with nomic-embed and multilingual-e5 as options.

### Context-window compression

Each connector sets limits of 15% of the window for system, 15% for history and 70% for the user prompt. `cannonball` removes tokens from the middle of any oversize part and inserts `--prompt truncated for brevity--` ([helpers/chat/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/index.js#L333-L360)). No summarisation model is used. The authors say plainly in the code comments that they chose this for speed and generality ([helpers/chat/index.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/index.js#L6-L47)).

### Agent mode

When `grepAgents` fires, the HTTP stream tells the client to open `/agent-invocation/:uuid` over WebSocket, and the rest of the turn runs in `aibitat` ([agents.js](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/agents.js#L56-L120)). `aibitat` is an in-house multi-agent chat runtime with its own provider adapters and a plugin system. The `rag-memory` plugin gives the agent `search` and `store` over the same workspace vectors. Other plugins cover web browsing and scraping, SQL, file creation, charts, Gmail, Google Calendar, Outlook, scheduled jobs and MCP tools.

## Extending it

- **New LLM, embedder or vector store.** Add a class in the matching folder that implements the shared methods, then add a `case` to the switch in `helpers/index.js` and an option in the frontend settings. There is no plugin loader. Providers are compiled into the app.
- **Agent skills.** Write an `aibitat` plugin, install community skills through the hub endpoints, connect MCP servers through the `MCP` hypervisor, or build no-code "agent flows".
- **Data connectors.** The collector's `extensions` endpoints load GitHub and GitLab repos, Confluence, DrupalWiki, Obsidian vaults, Paperless-ngx, YouTube transcripts and website crawls into the same document format.
- **Integration.** The developer API (`/api/v1/...`, API-key protected) exposes workspaces, documents and chat. `/v1/openai/chat/completions` lets OpenAI-compatible clients talk to a workspace. The embed widget puts a workspace chat on any site.

## Running it

- **Docker.** One image runs both processes. The entrypoint runs `prisma migrate deploy`, starts `server/index.js` (port 3001 by default) and `collector/index.js` (port 8888), and exits if either one dies. Set `STORAGE_DIR` to a mounted volume, or documents, vectors and the SQLite database are lost on restart.
- **Desktop.** A desktop app for macOS, Windows and Linux is distributed separately from this repository.
- **Bare metal.** Install Node.js and Yarn, run `yarn setup`, then build the frontend and start the server and collector yourself (`BARE_METAL.md`).
- **Services.** None are required beyond an LLM endpoint, local (Ollama, LM Studio, LocalAI) or cloud. External vector databases are optional.
- **Telemetry.** Anonymous usage events are sent unless `DISABLE_TELEMETRY=true`.

## Strengths and caveats

- **Strength: works out of the box.** SQLite, embedded LanceDB and a local ONNX embedder mean a private, offline-capable RAG setup with a single container.
- **Strength: very wide provider coverage.** About 40 LLM connectors, 14 embedders and 10 vector stores, all switchable from the UI.
- **Strength: a product, not a kit.** Users, roles, quotas, threads, an embed widget, an API and agents are already built, which is the main reason to choose it over a framework.
- **Caveat: basic retrieval.** Dense-only search with fixed top-N, no BM25 or hybrid, no query rewriting, and no metadata filters at query time. Reranking exists only on LanceDB.
- **Caveat: lossy compression.** Middle-out truncation can cut the part of a document or prompt that mattered, and nothing reports it beyond the inline marker.
- **Caveat: per-adapter duplication.** Chunking and embedding are re-implemented in every vector adapter, so features and fixes land unevenly.
- **Caveat: single-node design.** State lives in SQLite and local files, and progress tracking uses process globals. There is no story for horizontal scaling or real multi-tenancy.
- **Caveat: no evaluation.** Apart from per-message token and speed metrics, there is nothing to measure retrieval or answer quality.

*Sources: code at 297808f, deepwiki-open wiki (10 pages), verified Q&A.*

## How Mintplex-Labs/anything-llm answers the RAG engines questions

### How are documents parsed and chunked? (answered)

**Supported formats.** A separate `collector` service (Express.js on port 8888) handles all document ingestion, invoked by the main server via `CollectorApi`. The file-type registry at `collector/utils/constants.js` maps extensions to converter scripts. Text types (`.txt`, `.md`, `.org`, `.adoc`, `.rst`, `.csv`, `.json`) use a simple text extractor. Binary documents: `.pdf` uses PDFLoader with per-page splitting; `.docx`, `.pptx`, `.odt`, `.odp` use office-mime converters (mammoth/docx2js); `.xlsx` is parsed via asXlsx; `.epub` via asEPub; `.mbox` email archives are supported. Audio (`.mp3`, `.wav`, `.m4a`, `.ogg`, `.opus`, `.webm`, `.mp4`) is transcribed to text via whisper (local or cloud). Images (`.png`, `.jpg`, `.webp`) are captioned. YouTube transcripts, website-depth scraping, Confluence/DrupalWiki/Obsidian/Paperless-NGX integrations exist as extension endpoints.

**OCR and layout.** PDF parsing tries text-extraction first via `PDFLoader` (`splitPages: true`). If no text content results, it falls back to `OCRLoader` (Tesseract-based with configurable target languages via `TARGET_OCR_LANG`). Tables are not explicitly detected or extracted — PDF text is concatenated page-wise and no table-structure preservation logic was found. Markdown or HTML tables arrive as raw text.

**Chunking.** The `TextSplitter` class wraps LangChain's `RecursiveCharacterTextSplitter`. Default chunk size is 1,000 characters (configurable via `text_splitter_chunk_size` in SystemSettings, capped at the embedder's `embeddingMaxChunkLength`). Default overlap is 20 characters (configurable via `text_splitter_chunk_overlap`). A `chunkPrefix` can be prepended per embedder requirement (e.g., `search_document: ` for nomic), and a `chunkHeaderMeta` block is prepended with document title, published date, and source URL — wrapped in `<document_metadata>...</document_metadata>` tags so the LLM can cite sources.

> **Editor's note.** Correction: images are not captioned. asImage.js runs Tesseract OCR via OCRLoader.ocrImage and stores the recognised text.

Citations: [collector/utils/constants.js:45-86](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/utils/constants.js#L45-L86) · [collector/processSingleFile/index.js:24-91](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/index.js#L24-L91) · [collector/processSingleFile/convert/asPDF/index.js:12-84](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/convert/asPDF/index.js#L12-L84) · [server/utils/TextSplitter/index.js:1-218](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/TextSplitter/index.js#L1-L218) · [server/utils/helpers/index.js:621-630](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L621-L630)

### How are embeddings and indexes built and stored? (answered)

**Embedding models.** AnythingLLM supports 14 embedding engines, selected via `EMBEDDING_ENGINE` env: `native` (local ONNX via @xenova/transformers), `openai`, `azure`, `localai`, `ollama`, `lmstudio`, `cohere`, `voyageai`, `litellm`, `mistral`, `generic-openai`, `gemini`, `openrouter`, `lemonade`. The native embedder defaults to `Xenova/all-MiniLM-L6-v2` (23MB, 1000-char chunks, 25 concurrent chunks) with two alternatives: `Xenova/nomic-embed-text-v1` (139MB, 16000-char chunks, 5 concurrent, uses `search_document:`/`search_query:` prefixes) and `MintplexLabs/multilingual-e5-small` (487MB, 100+ languages, passages prefixed with `passage:`/`query:`). For local on-device embedding, models are downloaded on first use from HuggingFace Hub with a fallback CDN at `cdn.anythingllm.com`.

**Vector stores.** Ten backends are supported: LanceDB (default, embedded, file-based), Pinecone, Chroma, ChromaCloud, Weaviate, QDrant, Milvus, Zilliz, AstraDB, and PGVector. Each implements the `VectorDatabase` base class with `addDocumentToNamespace`, `performSimilaritySearch`, `deleteDocumentFromNamespace`. Workspaces map 1:1 to vector-db namespaces (using `workspace.slug`). Chunking and embedding happen inside the vector-store adapter's `addDocumentToNamespace()` — every provider reads `text_splitter_chunk_size`/`text_splitter_chunk_overlap` from SystemSettings, splits via `TextSplitter`, embeds via the selected `EmbedderEngine`, and upserts.

**Hybrid / BM25.** There is no built-in BM25 or hybrid search. Retrieval is dense-only via cosine similarity on embedding vectors. Full-text search for workspace/thread *names* uses fast-levenshtein matching, but document-level keyword search is not implemented.

**Metadata.** Document metadata stored per chunk includes: `text` (the chunk content), `title`, `description`, `docAuthor`, `docSource`, `published`, `chunkSource`, `token_count_estimate` and scoring fields. Pinned documents bypass the vector store entirely and are injected directly as context.


Citations: [server/utils/EmbeddingEngines/native/index.js:1-310](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/EmbeddingEngines/native/index.js#L1-L310) · [server/utils/EmbeddingEngines/native/constants.js:1-60](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/EmbeddingEngines/native/constants.js#L1-L60) · [server/utils/helpers/index.js:308-360](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L308-L360) · [server/utils/vectorDbProviders/lance/index.js:312-415](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L312-L415) · [server/utils/vectorDbProviders/pinecone/index.js:156-210](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/pinecone/index.js#L156-L210)

### How is retrieval performed? (answered)

**Dense search.** Retrieval is purely dense (embedding-based). The `performSimilaritySearch()` method on every vector-db adapter takes the user's query text, embeds it via `LLMConnector.embedTextInput()`, then runs a cosine-similarity vector search against the workspace namespace. Parameters: `similarityThreshold` (default 0.25) and `topN` (default 4, configurable per workspace). Results below the threshold are discarded. Pinned documents are deduplicated from search results via `sourceIdentifier()` (a composite of `title+published`).

**Reranking.** Each vector-db provider has an optional reranking path, toggled per workspace by `vectorSearchMode === "rerank"`. The LanceDB implementation (`rerankedSimilarityResponse`) first fetches up to 50 results (capped at `max(10, min(50, ceil(totalEmbeddings * 0.1)))`), then passes them to a `NativeEmbeddingReranker` built on `Xenova/ms-marco-MiniLM-L-6-v2` (a cross-encoder from @xenova/transformers). The reranker scores query-document pairs in batches of 10, sorts globally, returns the top-K. Pinecone uses a simpler single-pass vector search (Pinecone server-side doesn't support reranking — only native reranker is available).

**Query rewriting/decomposition.** No query rewriting or decomposition is implemented. The raw user message is embedded as-is.

**Filters.** The only filter is `filterIdentifiers` — a list of pinned-document identifiers to exclude from results. This prevents the vector store from returning chunks that are already injected via pinned documents. No metadata-field filters (date range, author, source type) are passed to the vector DB at query time.

**Context backfill.** After vector search, `fillSourceWindow()` backfills from recent chat history when fewer than `topN` sources are found. This ensures follow-up questions retain relevant context even when the vector search returns 0 results — earlier sources from the same thread are reused.

> **Editor's note.** Correction: reranking is implemented only in the LanceDB adapter (rerankedSimilarityResponse); the other nine vector-store adapters ignore the workspace's vectorSearchMode=rerank setting and always do plain similarity search.

Citations: [server/utils/vectorDbProviders/lance/index.js:99-170](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L99-L170) · [server/utils/vectorDbProviders/lance/index.js:417-467](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L417-L467) · [server/utils/EmbeddingRerankers/native/index.js:229-276](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/EmbeddingRerankers/native/index.js#L229-L276) · [server/utils/helpers/chat/index.js:403-463](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/index.js#L403-L463) · [server/utils/chats/stream.js:194-240](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L194-L240)

### How are answers generated and grounded? (answered)

**Prompt assembly.** The `streamChatWithWorkspace()` function builds a system prompt by calling `chatPrompt()` which takes the workspace's configured prompt template, expands system-prompt variables, and appends user memories (if enabled) as a `## Things I Remember About You` section (scoped per user, reranked via the native cross-encoder when there are many memories). Context from RAG is appended as numbered `[CONTEXT N]`/`[END CONTEXT N]` blocks via `#appendContext()`. The chat history is formatted as alternating user/assistant messages. The final message array is built by each provider's `constructPrompt()` method: a system role message, the history messages, and a user role message.

**Compression.** The `messageArrayCompressor` (or `messageStringCompressor` for string-mode models) enforces token-budget proportions: system 15%, history 15%, user 70%. It uses a "cannonball" strategy — when a component exceeds its limit, it truncates from the middle outwards, inserting a `--prompt truncated for brevity--` marker. History is the most aggressively compressed component, with a preference for keeping the 3 most recent exchanges even if they must be cannonballed individually.

**Citation / source attribution.** Vector search results carry their chunk metadata (title, source URL, published date, score) which are returned alongside the text response as `sources`. The chunk metadata is included in the context block sent to the LLM, and the UI renders these sources as clickable citations. The `fillSourceWindow` mechanism ensures the context window has consistently `topN` sources for follow-up questions by backfilling from previous turns.

**Streaming.** Every LLM provider implements `streamGetChatCompletion()` and `handleStream()`. The response is delivered as SSE (Server-Sent Events) with chunk types: `textResponseChunk`, `textResponse`, `abort`, `finalizeResponseStream`. OpenAI's Responses API streams reasoning summary as thinking blocks. Cost metrics (prompt tokens, completion tokens, duration, tokens/sec) are computed per turn.

**Multi-step / agentic.** When the workspace has agent skills enabled, `grepAgents()` diverts to the agent system (aibitat) which supports multi-turn tool-using agents with MCP server integration, Gmail/Google Calendar/Outlook skills, SQL agents, file creation, web browsing, and scheduled-job creation. These run as independent subagents with their own provider connections.


Citations: [server/utils/chats/stream.js:273-340](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L273-L340) · [server/utils/chats/index.js:122-144](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/index.js#L122-L144) · [server/utils/memories/index.js:1-139](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/memories/index.js#L1-L139) · [server/utils/helpers/chat/index.js:49-192](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/index.js#L49-L192) · [server/utils/AiProviders/openAi/index.js:49-135](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/AiProviders/openAi/index.js#L49-L135)

### How is quality evaluated or observed? (answered)

There is no built-in RAG evaluation framework — no LLM-as-judge metrics, no answer relevancy/faithfulness scoring, no ground-truth datasets, and no A/B comparison tool. The evaluation story is limited to operational observability:

**Telemetry.** An anonymous telemetry system (`Telemetry.sendTelemetry()`) reports high-level usage events (`server_boot`, `sent_chat`, `embed_sent_chat`, `workspace_created`) when `DISABLE_TELEMETRY` is not set. These include commit hash, LLM provider, embedder, vector-db type, and multi-user mode. In development mode telemetry is stubbed.

**Event logging.** The `EventLogs` model logs events (`sent_chat`, `embed_created`, `embed_updated`, etc.) to the Prisma SQLite database with user ID, timestamp, and metadata JSON. Events can be queried by type or user. These serve as an audit trail but provide no quality metrics.

**Per-request metrics.** The `LLMPerformanceMonitor` captures per-request metrics: `prompt_tokens`, `completion_tokens`, `total_tokens`, `duration` (seconds), `outputTps` (tokens per second). These are stored alongside each chat message and exposed via the API as `metrics`. Cost tracking (`addChatCostToMetrics`) enriches metrics with per-model pricing when pricing data is configured.

**Observability hooks.** The system uses `EmbeddingWorkerManager.emitProgress()` for real-time chunk-embedding progress and `abortConnectorOnClientDisconnect()` for signaling. HTTP logging is optional via middleware. There are no OpenTelemetry, Langfuse, or LangSmith integrations — nor any built-in dashboard for RAG quality metrics.


Citations: [server/utils/telemetry/index.js:1-33](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/telemetry/index.js#L1-L33) · [server/models/eventLogs.js:1-129](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/models/eventLogs.js#L1-L129) · [server/utils/helpers/chat/LLMPerformanceMonitor.js:1-20](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/LLMPerformanceMonitor.js#L1-L20) · [server/utils/chats/stream.js:340-365](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L340-L365) · [server/endpoints/chat.js:76-93](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/endpoints/chat.js#L76-L93)

### How is it deployed and operated? (answered)

**Library vs service.** AnythingLLM is a self-hosted web application (not a library or SaaS). It runs as three Node.js processes: the main server (Express.js on port 3001), the document collector (Express.js on port 8888), and the frontend (Vite/React dev server in development, prebuilt static files in production). The server and collector communicate over the loopback interface with a shared communication key for integrity verification.

**UI.** The frontend is a React application built with Vite, using i18next for 30+ language translations. It provides workspaces, threads, chat interface with streaming, document management, settings, model configuration, agent skill configuration, and embeddable chat widgets. An embed widget can be embedded on external websites as a floating iframe/button with configurable branding, prompt, model, and temperature (if overrides are allowed).

**API.** The server exposes REST endpoints for workspace management, document CRUD, streaming chat (`/workspace/:slug/stream-chat`), thread management, and an OpenAI-compatible API (`/api` endpoints). Multi-user mode supports admin/manager/default roles with daily message quotas.

**Required infrastructure.** Minimum: Node.js 18+, SQLite (Prisma ORM). No external database is required — everything runs on SQLite with file-based storage for documents, models, and vector data (LanceDB default). PostgreSQL is supported (Prisma datasource can be swapped). For production, a Docker container bundles all three services. The `docker-compose.yml` mounts storage volumes and `.env` configuration. The `docker-healthcheck.sh` provides container monitoring.

**Scaling and multi-tenancy.** The system is single-tenant by architecture. Multiple users are supported via the multi-user mode with scoped workspaces, threads, and memories. Each workspace has its own vector-db namespace, prompt configuration, and model assignment. The model router supports rule-based (regex and LLM-classified) model assignment per workspace. Document synchronization (`sync-watched-documents`) and memory extraction (`extract-memories`) run as background jobs via the `BackgroundService` worker. There are no sharding, load-balancing, or horizontal-scaling primitives — multi-tenancy means multiple users on one instance, not multi-instance clustering.


Citations: [server/index.js:1-50](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/index.js#L1-L50) · [collector/index.js:1-228](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/index.js#L1-L228) · [server/utils/helpers/index.js:87-127](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L87-L127) · [server/utils/BackgroundWorkers/index.js:1-75](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/BackgroundWorkers/index.js#L1-L75)
