Mintplex-Labs/anything-llm
Self-hosted Node.js chat-with-your-documents app with workspaces, ten vector-store backends, ~40 LLM providers and an agent mode.
Overview
AnythingLLM is Mintplex Labs’ self-hosted application for chatting with your own documents. It is a product, not a library. You run it as a desktop app or a Docker container, create workspaces, drop files into them, and chat with any of about forty LLM providers over the embedded chunks. Around that core it has grown a large feature surface: threads, multi-user roles with daily quotas, an embeddable chat widget, a developer REST API with an OpenAI-compatible endpoint, user memories, a rule-based model router, scheduled jobs, MCP servers and an agent mode (@agent) built on its own aibitat framework.
The code is a JavaScript monorepo with three runtime pieces: server/ (Express API, Prisma ORM on SQLite, all RAG and agent logic), collector/ (a separate Express process that turns files and links into text), and frontend/ (React and Vite). Every pluggable subsystem follows the same pattern. An environment variable such as VECTOR_DB, EMBEDDING_ENGINE or LLM_PROVIDER is resolved by a switch in server/utils/helpers/index.js to a provider class with a shared method set. The defaults (LanceDB on disk, a local ONNX embedder, SQLite) mean a fresh install needs no external services except an LLM.
The RAG design is simple and pragmatic: recursive character chunking, dense-only top-N search with a similarity threshold, optional cross-encoder reranking (LanceDB only), and a “cannonball” middle-out truncation to fit the context window. There is no hybrid search, no query rewriting and no evaluation tooling.
Architecture
flowchart LR
UI["React frontend"] --> API["server: Express endpoints"]
EMB["Embed widget / dev API"] --> API
API --> CHAT["streamChatWithWorkspace"]
CHAT --> AG["grepAgents -> aibitat (WebSocket)"]
CHAT --> VDB["VectorDb provider"]
CHAT --> LLM["LLM provider connector"]
API --> COLAPI["CollectorApi"]
COLAPI --> COL["collector: converters + OCR"]
COL --> DOCS["storage/documents JSON"]
API --> ADD["Document.addDocuments"]
ADD --> DOCS
ADD --> SPLIT["TextSplitter"]
SPLIT --> EMBE["Embedding engine"]
EMBE --> VDB
API --> DB["Prisma (SQLite)"]
| Component | Path | Role |
|---|---|---|
| HTTP server | server/index.js, server/endpoints/ |
Express app, auth middleware, workspace, chat, document, admin, agent and API routes |
| Chat pipeline | server/utils/chats/stream.js |
Slash commands, agent hand-off, context assembly, vector search, compression, streaming |
| Provider switches | server/utils/helpers/index.js |
getVectorDbClass, getLLMProvider, getEmbeddingEngineSelection |
| Vector stores | server/utils/vectorDbProviders/ |
VectorDatabase base class plus LanceDB, PGVector, Pinecone, Chroma, ChromaCloud, Weaviate, Qdrant, Milvus, Zilliz, AstraDB |
| Embedders and reranker | server/utils/EmbeddingEngines/, EmbeddingRerankers/native |
14 embedding backends; local Xenova/ms-marco-MiniLM-L-6-v2 cross-encoder |
| LLM connectors | server/utils/AiProviders/ |
One class per provider with compressMessages, streamGetChatCompletion, handleStream |
| Chunking | server/utils/TextSplitter/ |
Wrapper over LangChain’s RecursiveCharacterTextSplitter with metadata header and prefix |
| Collector | collector/ |
File, link and connector ingestion; per-extension converters; Tesseract OCR |
| Agents | server/utils/agents/aibitat/ |
Multi-agent chat runtime, provider adapters, plugins (RAG memory, web browsing, SQL, files, Gmail, Outlook, MCP) |
| Background jobs | server/jobs/ |
Embedding worker, watched-document sync, memory extraction, scheduled agent jobs |
How a request flows
Uploading a PDF and then asking a question in a workspace:
- Upload. The frontend posts the file to
/workspace/:slug/upload. The server callsCollector.processDocument, which forwards it to the collector process with a signed payload thatverifyPayloadIntegritychecks. - Convert.
processSingleFilechecks that the path stays inside the hot directory, picks a converter by extension, falls back to plain text for unknown text types, and rejects everything else (processSingleFile/index.js). For.pdf,asPdfextracts text per page and runs OCR only when no text comes back (asPDF/index.js). The result is a JSON document withpageContentand metadata in server storage. - Embed. When the user moves the file into the workspace,
Document.addDocumentsloops over the files and callsVectorDb.addDocumentToNamespace(workspace.slug, ...), emitting progress events as it goes (documents.js). With the native embedder the same loop runs in a child-process worker, so an out-of-memory error kills the worker and not the server. - Chunk and store. The LanceDB adapter first checks a per-file vector cache. On a miss it builds a
TextSplitterfrom system settings (chunk size capped by the embedder’s limit, 20-character default overlap, a metadata header and an embedder prefix), embeds the chunks, and upserts them into a table named after the workspace slug (lance/index.js). - Ask.
POST /workspace/:slug/stream-chatopens an SSE response, enforces the multi-user daily quota, and callsstreamChatWithWorkspace(chat.js). - Divert or route. Slash commands are handled first. Then
grepAgentschecks for an@agenthandle, or for native tool calling inautomaticmode, and hands off to the agent WebSocket if so. Otherwise the model router may pick a different connector (stream.js). - Assemble context. Pinned documents (whole files, token-capped) and parsed attachments go in first. Then
performSimilaritySearchembeds the message and fetches the toptopNchunks abovesimilarityThreshold, excluding chunks from pinned files (stream.js).fillSourceWindowbackfills from sources cited earlier in the thread when search returns fewer thantopN. Inquerymode an empty context returns a fixed refusal instead of calling the LLM. - Compress and stream.
chatPromptbuilds the system prompt (with memories), and the connector’scompressMessagesfits system, context, history and prompt into the window. ThenstreamGetChatCompletionandhandleStreamwrite SSE chunks, and the chat is saved with itssourcesand metrics (stream.js).
Key components
Vector store adapters
VectorDatabase in base.js defines the contract: addDocumentToNamespace, deleteDocumentFromNamespace, performSimilaritySearch, namespaceCount, reset and so on. A namespace is a workspace slug. getVectorDbClass picks the adapter from VECTOR_DB and falls back to LanceDB (helpers/index.js). Each adapter does its own chunking and embedding instead of sharing a pipeline, so behaviour can drift between backends. The clearest case is reranking: only the LanceDB adapter implements the rerank path. It pulls max(10, min(50, 10% of the table)) cosine candidates and re-scores them with the local cross-encoder (lance/index.js). On the other nine backends the workspace “rerank” setting has no effect.
Chunking
TextSplitter delegates to RecursiveCharacterTextSplitter with a 1,000-character default size and a 20-character overlap. After splitting, it prepends a <document_metadata> header (title, source, published date) to every chunk, so chunks can exceed the nominal size (TextSplitter/index.js). Embedders such as nomic add their own search_document: style prefix.
Embedders
getEmbeddingEngineSelection maps EMBEDDING_ENGINE to 14 engines and defaults to the native embedder (helpers/index.js). The native engine runs Xenova/all-MiniLM-L6-v2 by default through transformers.js, with nomic-embed and multilingual-e5 as options.
Context-window compression
Each connector sets limits of 15% of the window for system, 15% for history and 70% for the user prompt. cannonball removes tokens from the middle of any oversize part and inserts --prompt truncated for brevity-- (helpers/chat/index.js). No summarisation model is used. The authors say plainly in the code comments that they chose this for speed and generality (helpers/chat/index.js).
Agent mode
When grepAgents fires, the HTTP stream tells the client to open /agent-invocation/:uuid over WebSocket, and the rest of the turn runs in aibitat (agents.js). aibitat is an in-house multi-agent chat runtime with its own provider adapters and a plugin system. The rag-memory plugin gives the agent search and store over the same workspace vectors. Other plugins cover web browsing and scraping, SQL, file creation, charts, Gmail, Google Calendar, Outlook, scheduled jobs and MCP tools.
Extending it
- New LLM, embedder or vector store. Add a class in the matching folder that implements the shared methods, then add a
caseto the switch inhelpers/index.jsand an option in the frontend settings. There is no plugin loader. Providers are compiled into the app. - Agent skills. Write an
aibitatplugin, install community skills through the hub endpoints, connect MCP servers through theMCPhypervisor, or build no-code “agent flows”. - Data connectors. The collector’s
extensionsendpoints load GitHub and GitLab repos, Confluence, DrupalWiki, Obsidian vaults, Paperless-ngx, YouTube transcripts and website crawls into the same document format. - Integration. The developer API (
/api/v1/..., API-key protected) exposes workspaces, documents and chat./v1/openai/chat/completionslets OpenAI-compatible clients talk to a workspace. The embed widget puts a workspace chat on any site.
Running it
- Docker. One image runs both processes. The entrypoint runs
prisma migrate deploy, startsserver/index.js(port 3001 by default) andcollector/index.js(port 8888), and exits if either one dies. SetSTORAGE_DIRto a mounted volume, or documents, vectors and the SQLite database are lost on restart. - Desktop. A desktop app for macOS, Windows and Linux is distributed separately from this repository.
- Bare metal. Install Node.js and Yarn, run
yarn setup, then build the frontend and start the server and collector yourself (BARE_METAL.md). - Services. None are required beyond an LLM endpoint, local (Ollama, LM Studio, LocalAI) or cloud. External vector databases are optional.
- Telemetry. Anonymous usage events are sent unless
DISABLE_TELEMETRY=true.
Strengths and caveats
- Strength: works out of the box. SQLite, embedded LanceDB and a local ONNX embedder mean a private, offline-capable RAG setup with a single container.
- Strength: very wide provider coverage. About 40 LLM connectors, 14 embedders and 10 vector stores, all switchable from the UI.
- Strength: a product, not a kit. Users, roles, quotas, threads, an embed widget, an API and agents are already built, which is the main reason to choose it over a framework.
- Caveat: basic retrieval. Dense-only search with fixed top-N, no BM25 or hybrid, no query rewriting, and no metadata filters at query time. Reranking exists only on LanceDB.
- Caveat: lossy compression. Middle-out truncation can cut the part of a document or prompt that mattered, and nothing reports it beyond the inline marker.
- Caveat: per-adapter duplication. Chunking and embedding are re-implemented in every vector adapter, so features and fixes land unevenly.
- Caveat: single-node design. State lives in SQLite and local files, and progress tracking uses process globals. There is no story for horizontal scaling or real multi-tenancy.
- Caveat: no evaluation. Apart from per-message token and speed metrics, there is nothing to measure retrieval or answer quality.
Sources: code at 297808f, deepwiki-open wiki (10 pages), verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit 297808f. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredSupported formats. A separate collector service (Express.js on port 8888) handles all document ingestion, invoked by the main server via CollectorApi. The file-type registry at collector/utils/constants.js maps extensions to converter scripts. Text types (.txt, .md, .org, .adoc, .rst, .csv, .json) use a simple text extractor. Binary documents: .pdf uses PDFLoader with per-page splitting; .docx, .pptx, .odt, .odp use office-mime converters (mammoth/docx2js); .xlsx is parsed via asXlsx; .epub via asEPub; .mbox email archives are supported. Audio (.mp3, .wav, .m4a, .ogg, .opus, .webm, .mp4) is transcribed to text via whisper (local or cloud). Images (.png, .jpg, .webp) are captioned. YouTube transcripts, website-depth scraping, Confluence/DrupalWiki/Obsidian/Paperless-NGX integrations exist as extension endpoints.
OCR and layout. PDF parsing tries text-extraction first via PDFLoader (splitPages: true). If no text content results, it falls back to OCRLoader (Tesseract-based with configurable target languages via TARGET_OCR_LANG). Tables are not explicitly detected or extracted — PDF text is concatenated page-wise and no table-structure preservation logic was found. Markdown or HTML tables arrive as raw text.
Chunking. The TextSplitter class wraps LangChain's RecursiveCharacterTextSplitter. Default chunk size is 1,000 characters (configurable via text_splitter_chunk_size in SystemSettings, capped at the embedder's embeddingMaxChunkLength). Default overlap is 20 characters (configurable via text_splitter_chunk_overlap). A chunkPrefix can be prepended per embedder requirement (e.g., search_document: for nomic), and a chunkHeaderMeta block is prepended with document title, published date, and source URL — wrapped in <document_metadata>...</document_metadata> tags so the LLM can cite sources.
How are embeddings and indexes built and stored?
answeredEmbedding models. AnythingLLM supports 14 embedding engines, selected via EMBEDDING_ENGINE env: native (local ONNX via @xenova/transformers), openai, azure, localai, ollama, lmstudio, cohere, voyageai, litellm, mistral, generic-openai, gemini, openrouter, lemonade. The native embedder defaults to Xenova/all-MiniLM-L6-v2 (23MB, 1000-char chunks, 25 concurrent chunks) with two alternatives: Xenova/nomic-embed-text-v1 (139MB, 16000-char chunks, 5 concurrent, uses search_document:/search_query: prefixes) and MintplexLabs/multilingual-e5-small (487MB, 100+ languages, passages prefixed with passage:/query:). For local on-device embedding, models are downloaded on first use from HuggingFace Hub with a fallback CDN at cdn.anythingllm.com.
Vector stores. Ten backends are supported: LanceDB (default, embedded, file-based), Pinecone, Chroma, ChromaCloud, Weaviate, QDrant, Milvus, Zilliz, AstraDB, and PGVector. Each implements the VectorDatabase base class with addDocumentToNamespace, performSimilaritySearch, deleteDocumentFromNamespace. Workspaces map 1:1 to vector-db namespaces (using workspace.slug). Chunking and embedding happen inside the vector-store adapter's addDocumentToNamespace() — every provider reads text_splitter_chunk_size/text_splitter_chunk_overlap from SystemSettings, splits via TextSplitter, embeds via the selected EmbedderEngine, and upserts.
Hybrid / BM25. There is no built-in BM25 or hybrid search. Retrieval is dense-only via cosine similarity on embedding vectors. Full-text search for workspace/thread names uses fast-levenshtein matching, but document-level keyword search is not implemented.
Metadata. Document metadata stored per chunk includes: text (the chunk content), title, description, docAuthor, docSource, published, chunkSource, token_count_estimate and scoring fields. Pinned documents bypass the vector store entirely and are injected directly as context.
How is retrieval performed?
answeredDense search. Retrieval is purely dense (embedding-based). The performSimilaritySearch() method on every vector-db adapter takes the user's query text, embeds it via LLMConnector.embedTextInput(), then runs a cosine-similarity vector search against the workspace namespace. Parameters: similarityThreshold (default 0.25) and topN (default 4, configurable per workspace). Results below the threshold are discarded. Pinned documents are deduplicated from search results via sourceIdentifier() (a composite of title+published).
Reranking. Each vector-db provider has an optional reranking path, toggled per workspace by vectorSearchMode === "rerank". The LanceDB implementation (rerankedSimilarityResponse) first fetches up to 50 results (capped at max(10, min(50, ceil(totalEmbeddings * 0.1)))), then passes them to a NativeEmbeddingReranker built on Xenova/ms-marco-MiniLM-L-6-v2 (a cross-encoder from @xenova/transformers). The reranker scores query-document pairs in batches of 10, sorts globally, returns the top-K. Pinecone uses a simpler single-pass vector search (Pinecone server-side doesn't support reranking — only native reranker is available).
Query rewriting/decomposition. No query rewriting or decomposition is implemented. The raw user message is embedded as-is.
Filters. The only filter is filterIdentifiers — a list of pinned-document identifiers to exclude from results. This prevents the vector store from returning chunks that are already injected via pinned documents. No metadata-field filters (date range, author, source type) are passed to the vector DB at query time.
Context backfill. After vector search, fillSourceWindow() backfills from recent chat history when fewer than topN sources are found. This ensures follow-up questions retain relevant context even when the vector search returns 0 results — earlier sources from the same thread are reused.
How are answers generated and grounded?
answeredPrompt assembly. The streamChatWithWorkspace() function builds a system prompt by calling chatPrompt() which takes the workspace's configured prompt template, expands system-prompt variables, and appends user memories (if enabled) as a ## Things I Remember About You section (scoped per user, reranked via the native cross-encoder when there are many memories). Context from RAG is appended as numbered [CONTEXT N]/[END CONTEXT N] blocks via #appendContext(). The chat history is formatted as alternating user/assistant messages. The final message array is built by each provider's constructPrompt() method: a system role message, the history messages, and a user role message.
Compression. The messageArrayCompressor (or messageStringCompressor for string-mode models) enforces token-budget proportions: system 15%, history 15%, user 70%. It uses a "cannonball" strategy — when a component exceeds its limit, it truncates from the middle outwards, inserting a --prompt truncated for brevity-- marker. History is the most aggressively compressed component, with a preference for keeping the 3 most recent exchanges even if they must be cannonballed individually.
Citation / source attribution. Vector search results carry their chunk metadata (title, source URL, published date, score) which are returned alongside the text response as sources. The chunk metadata is included in the context block sent to the LLM, and the UI renders these sources as clickable citations. The fillSourceWindow mechanism ensures the context window has consistently topN sources for follow-up questions by backfilling from previous turns.
Streaming. Every LLM provider implements streamGetChatCompletion() and handleStream(). The response is delivered as SSE (Server-Sent Events) with chunk types: textResponseChunk, textResponse, abort, finalizeResponseStream. OpenAI's Responses API streams reasoning summary as thinking blocks. Cost metrics (prompt tokens, completion tokens, duration, tokens/sec) are computed per turn.
Multi-step / agentic. When the workspace has agent skills enabled, grepAgents() diverts to the agent system (aibitat) which supports multi-turn tool-using agents with MCP server integration, Gmail/Google Calendar/Outlook skills, SQL agents, file creation, web browsing, and scheduled-job creation. These run as independent subagents with their own provider connections.
How is quality evaluated or observed?
answeredThere is no built-in RAG evaluation framework — no LLM-as-judge metrics, no answer relevancy/faithfulness scoring, no ground-truth datasets, and no A/B comparison tool. The evaluation story is limited to operational observability:
Telemetry. An anonymous telemetry system (Telemetry.sendTelemetry()) reports high-level usage events (server_boot, sent_chat, embed_sent_chat, workspace_created) when DISABLE_TELEMETRY is not set. These include commit hash, LLM provider, embedder, vector-db type, and multi-user mode. In development mode telemetry is stubbed.
Event logging. The EventLogs model logs events (sent_chat, embed_created, embed_updated, etc.) to the Prisma SQLite database with user ID, timestamp, and metadata JSON. Events can be queried by type or user. These serve as an audit trail but provide no quality metrics.
Per-request metrics. The LLMPerformanceMonitor captures per-request metrics: prompt_tokens, completion_tokens, total_tokens, duration (seconds), outputTps (tokens per second). These are stored alongside each chat message and exposed via the API as metrics. Cost tracking (addChatCostToMetrics) enriches metrics with per-model pricing when pricing data is configured.
Observability hooks. The system uses EmbeddingWorkerManager.emitProgress() for real-time chunk-embedding progress and abortConnectorOnClientDisconnect() for signaling. HTTP logging is optional via middleware. There are no OpenTelemetry, Langfuse, or LangSmith integrations — nor any built-in dashboard for RAG quality metrics.
How is it deployed and operated?
answeredLibrary vs service. AnythingLLM is a self-hosted web application (not a library or SaaS). It runs as three Node.js processes: the main server (Express.js on port 3001), the document collector (Express.js on port 8888), and the frontend (Vite/React dev server in development, prebuilt static files in production). The server and collector communicate over the loopback interface with a shared communication key for integrity verification.
UI. The frontend is a React application built with Vite, using i18next for 30+ language translations. It provides workspaces, threads, chat interface with streaming, document management, settings, model configuration, agent skill configuration, and embeddable chat widgets. An embed widget can be embedded on external websites as a floating iframe/button with configurable branding, prompt, model, and temperature (if overrides are allowed).
API. The server exposes REST endpoints for workspace management, document CRUD, streaming chat (/workspace/:slug/stream-chat), thread management, and an OpenAI-compatible API (/api endpoints). Multi-user mode supports admin/manager/default roles with daily message quotas.
Required infrastructure. Minimum: Node.js 18+, SQLite (Prisma ORM). No external database is required — everything runs on SQLite with file-based storage for documents, models, and vector data (LanceDB default). PostgreSQL is supported (Prisma datasource can be swapped). For production, a Docker container bundles all three services. The docker-compose.yml mounts storage volumes and .env configuration. The docker-healthcheck.sh provides container monitoring.
Scaling and multi-tenancy. The system is single-tenant by architecture. Multiple users are supported via the multi-user mode with scoped workspaces, threads, and memories. Each workspace has its own vector-db namespace, prompt configuration, and model assignment. The model router supports rule-based (regex and LLM-classified) model assignment per workspace. Document synchronization (sync-watched-documents) and memory extraction (extract-memories) run as background jobs via the BackgroundService worker. There are no sharding, load-balancing, or horizontal-scaling primitives — multi-tenancy means multiple users on one instance, not multi-instance clustering.