LLMs Technical Reviews

How is it deployed and operated?

Library vs service; UI; API; required infrastructure; scaling and multi-tenancy.

Verdict

AnythingLLM is the easiest to run: one container, or a desktop app. Onyx and RAGFlow are the most complete multi-user services. LlamaIndex and Haystack suit teams that want to embed RAG in their own code.

Multi-service platforms. RAGFlow is one Go binary that starts as separate API, admin, ingestor, syncer and DeepDoc processes. It runs beside MySQL, MinIO, Kvrocks, NATS JetStream, ClickHouse and a search engine, with tenant-scoped models and datasets. Onyx runs a FastAPI server, Celery workers, two embedding model servers, Postgres, Redis and OpenSearch. It ships a Helm chart and an ECS template. Multi-tenancy and permission sync are under the separate ee/ licence. Quivr needs Postgres, Temporal, Weaviate, TEI and SeaweedFS plus plugin sidecars. Its API is v0 and the README marks it as evaluation stage. R2R needs only Postgres with pgvector. Its default simple orchestration runs ingestion inside the upload request, and Hatchet workers come only with the full config. Authentication is off by default, and the default admin password is well known.

Single-node applications. AnythingLLM bundles SQLite, LanceDB and a local embedder. Several users share one instance, and there is no horizontal scaling. Kotaemon is a Gradio app with SQLite and no REST API. It starts with an admin/admin login.

Libraries. LlamaIndex persists JSON stores through fsspec and has no server. Haystack has no server either, so serving is left to the separate Hayhooks project. RAG-Anything keeps one knowledge base per working_dir and needs the MinerU CLI and LibreOffice on the host. PageIndex stores JSON files in ./.pageindex. When given an API key it hands indexing and storage to the vendor’s hosted service.

Pick: AnythingLLM or Kotaemon for a private RAG tool on one machine. Pick: Onyx or RAGFlow for a self-hosted, multi-user company deployment. Pick: LlamaIndex, Haystack or R2R’s API to build RAG into your own product.

Per-project answers

infiniflow/ragflow

answered

Service architecture: RAGFlow is a Docker-based service with multiple profiles. The main API server runs on port 9380 (Go, via cmd/ragflow_server.go), an optional admin server on 9381, and nginx serves the web UI on ports 80/443. Task executors are pluggable workers that consume ingestion/parsing tasks from a message queue.

Infrastructure requirements: The minimum stack requires MySQL (metadata/settings), MinIO (document storage), Kvrocks (cache/task queue), and a vector store (Elasticsearch, Infinity, OceanBase, or SereneDB). Optional components include NATS (message queue), ClickHouse (analytics), and OpenSearch (hybrid search). The docker-compose (docker/docker-compose.yml, docker-compose-base.yml) configures all services including ES, Infinity, OceanBase, etc., gated behind profiles (elasticsearch, infinity, oceanbase, opensearch). Memory limits, ulimits (nofile=65535), and Docker healthchecks are preconfigured.

UI: The web frontend (web/) is a React/TypeScript application built with Vite, using Ant Design icons, AntV G2/G6 for graphs, and Lexical for rich text. It connects to the Go API server over HTTP. The UI includes knowledge base management, chat/dialog management, search app configuration, document upload, evaluation dashboards, and agent canvas workflows.

API: The system exposes REST APIs under /api/v1/ and /api/v1/openai/<chat_id>/chat/completions for OpenAI-compatible endpoints. The Go router (internal/router/router.go) handles all routes. There's also an MCP (Model Context Protocol) server endpoint on port 9382, supporting SSE and Streamable HTTP transports.

Multi-tenancy: Multi-tenancy is built in at the model layer — all entities carry TenantID/CreatedBy, and the evaluation/chat/knowledgebase tables all filter by tenant. Models are per-tenant via tenant_model_instance. The tenant_model_provider.go and tenant_llm.go tables store provider credentials and LLM configurations per tenant. The Search/Share system (service/search.go) supports access control with "me" vs "team" permission levels.

Scaling: Task executors can be horizontally scaled using --workers=N or range-based --consumer-no-beg/--consumer-no-end flags. The entrypoint.sh (docker/entrypoint.sh) allows selectively disabling the web server, task executor, or data sync for separate container roles. Redis/Kvrocks serves as the task queue backend, and NATS can be used for high-throughput messaging, both naturally supporting distributed worker pools.

Editor's note. Correction: the Go ingestor's queue is NATS JetStream (ingestor.mq_type: 'nats'), and NATS is in the default ragflow-go Compose profile, so it is not optional. Kvrocks is the cache. The MCP listener on 9382 is commented out by default. In the entrypoint, the --consumer-no-beg/--consumer-no-end range starts a single ingestor, and only --workers=N starts N of them.

Mintplex-Labs/anything-llm

answered

Library vs service. AnythingLLM is a self-hosted web application (not a library or SaaS). It runs as three Node.js processes: the main server (Express.js on port 3001), the document collector (Express.js on port 8888), and the frontend (Vite/React dev server in development, prebuilt static files in production). The server and collector communicate over the loopback interface with a shared communication key for integrity verification.

UI. The frontend is a React application built with Vite, using i18next for 30+ language translations. It provides workspaces, threads, chat interface with streaming, document management, settings, model configuration, agent skill configuration, and embeddable chat widgets. An embed widget can be embedded on external websites as a floating iframe/button with configurable branding, prompt, model, and temperature (if overrides are allowed).

API. The server exposes REST endpoints for workspace management, document CRUD, streaming chat (/workspace/:slug/stream-chat), thread management, and an OpenAI-compatible API (/api endpoints). Multi-user mode supports admin/manager/default roles with daily message quotas.

Required infrastructure. Minimum: Node.js 18+, SQLite (Prisma ORM). No external database is required — everything runs on SQLite with file-based storage for documents, models, and vector data (LanceDB default). PostgreSQL is supported (Prisma datasource can be swapped). For production, a Docker container bundles all three services. The docker-compose.yml mounts storage volumes and .env configuration. The docker-healthcheck.sh provides container monitoring.

Scaling and multi-tenancy. The system is single-tenant by architecture. Multiple users are supported via the multi-user mode with scoped workspaces, threads, and memories. Each workspace has its own vector-db namespace, prompt configuration, and model assignment. The model router supports rule-based (regex and LLM-classified) model assignment per workspace. Document synchronization (sync-watched-documents) and memory extraction (extract-memories) run as background jobs via the BackgroundService worker. There are no sharding, load-balancing, or horizontal-scaling primitives — multi-tenancy means multiple users on one instance, not multi-instance clustering.

run-llama/llama_index

answered

LlamaIndex is distributed as a Python library (PyPI package llama-index-core), not as a managed service. There is no built-in UI or API server in core — the CLI (llama_index/core/command_line/) has been deprecated and moved to its own llama-index-cli package. The chat_ui directory exists but is experimental. Usage is purely programmatic: users write Python scripts or integrate into web frameworks (FastAPI, Flask, etc.).

Infrastructure: All state is handled through pluggable storage backends. StorageContext (llama_index/core/storage/storage_context.py:52-72) combines a BaseDocumentStore (default SimpleDocumentStore — in-memory dicts persisted to JSON), BaseIndexStore (index metadata), vector stores (in-memory or external), and GraphStore. Persistence uses fsspec — files can be saved to local disk or cloud storage (S3, GCS). Each store has persist()/load() methods for round-trip serialization.

Production vector stores: Users plug in external vector database integrations (Chroma, Pinecone, Qdrant, Weaviate, Elasticsearch, FAISS, etc.) via pip install llama-index-vector-stores-<name>. These handle indexing, sharding, and replication outside of LlamaIndex. The SimpleVectorStore supports only single-node in-memory use.

Scaling: The ingestion pipeline (llama_index/core/ingestion/pipeline.py:72-112) supports multiprocessing and async parallelism via ProcessPoolExecutor. The library supports async throughout (parallel embedding, parallel retrieval, parallel sub-question execution via run_jobs).

Multi-tenancy: Not built-in. Each StorageContext instance represents one index. Users manage isolation at the application layer. There is no built-in authentication, rate limiting, or user session management.

LLM configuration: Global Settings dataclass (llama_index/core/settings.py:18-30) provides lazy-initialized defaults for llm, embed_model, callback_manager, tokenizer, node_parser, and transformations. These can be overridden per-index or per-query engine instance.

The-Vibe-Company/quivr

answered

Service, not library: Deployed as a distributed service with multiple roles (api, worker, migrate) from a single Go binary (cmd/quivr/main.go). Infrastructure (deploy/compose/compose.yaml:1-90, deploy/railway/services.json): PostgreSQL 17 (metadata, routing, activity), Temporal (workflow orchestration via go.temporal.io/sdk), Weaviate standalone (vector+BM25 search index), TEI container (HuggingFace text-embeddings-inference, CPU-only), SeaweedFS (S3-compatible blob storage for artifacts). Two Go processes — api serves HTTP (REST/JSON on port 8080, admin UI, also serves the web frontend) and worker runs Temporal activities (ingestion pipeline, enrichment, backfill, rebuilds, evaluation). Web UI: Node.js/React app at quivr-search/. API: Documented via OpenAPI at contracts/http/v0/openapi.yaml. Also exposes an MCP interface (quivr mcp) for AI agents. Multi-tenancy: Via API keys mapped to corpus.Scope (Organization + action permissions + corpus list, app/run.go:70). Scalability: Concurrent search limited to 64 per API process (transport/httpapi/search.go:16). Worker capacity is elastic via Temporal task queue quivr-content-v0. The current Railway demo is single-node with 8 containers and no HA for Temporal dev server. Configuration: Single QUIVR_CONFIG JSON file with database URL, S3 credentials, TEI/Weaviate/Temporal addresses, TLS, telemetry, API keys/scopes, plugin pins, delivery/webhook settings, and observability toggles (app/run.go:50-90). Plugins: Run as sidecar processes (Go or Python) communicating over HTTP via Plugin API v0, pinned by deployment configuration. First-party plugins are built into the core Docker image (deploy/images/quivr.Dockerfile); third-party plugins use plugin.Dockerfile.

VectifyAI/PageIndex

answered

Dual-mode architecture: library and cloud service. PageIndex is both a Python library (pip install -U pageindex) and a cloud API service (api.pageindex.ai). The same PageIndexClient switches modes based on whether an API key is provided (client.py:593-602,623-668).

Local mode deployment. In local mode, everything runs on the user's machine. Dependencies: openai-agents, litellm (for LLM routing), pypdfium2 (for PDF parsing), PyPDF2 (fallback extraction), Pillow (imaging). No external infrastructure — no vector DB, no database server. Documents are stored on disk via DocStore (local_store.py). LLM calls route through LiteLLM to OpenAI, Anthropic, or any OpenAI-compatible endpoint. The user provides their own API keys (OPENAI_API_KEY, etc.) (README.md:87-96).

Cloud mode deployment. With a PAGEINDEX_API_KEY, documents are uploaded to PageIndex Cloud (cloud_api.py:38-86). The cloud handles all parsing, OCR, indexing, and storage. The SDK client talks REST to https://api.pageindex.ai for document management and /chat/completions for the managed chat endpoint. For own-model chat over cloud documents, the agent runs in-process using the same tools proxied through an MCP bridge to the cloud (agent_tools.py:1490-1510, mcp_bridge.py).

Supported models. The indexing model (for summaries and expand) and chat model (for answering) are independently configurable. Model names follow LiteLLM's convention: bare names → OpenAI, anthropic/claude-* → Anthropic, bedrock/* → AWS Bedrock, vertex_ai/* → GCP Vertex, openai/* → explicit OpenAI routing (client.py:92-98,112-135). The default indexing model is gpt-5.6-luna; default chat model is gpt-5.6-sol (from config.yaml).

APIs. The client provides several protocol options: default answer lane (simplified Chat Completions), protocol="chat_completions" (full Chat Completions envelope), protocol="responses" (native OpenAI Responses), and protocol="messages" (native Anthropic Messages). It also offers an MCP server for Claude integration (mcp_bridge.py) and plain Python function tools for the OpenAI Agents SDK and Claude Agent SDK (integrations/openai_agents.py, integrations/claude_agent_sdk.py, integrations/anthropic_sdk.py).

No UI. The open-source SDK has no bundled UI. The PageIndex App is a separate hosted product at app.pageindex.ai (README.md:36).

Scaling. Local scaling is limited to single-machine disk storage. Cloud scaling supports "millions of documents" through the PageIndex File System (pages, not this repo's code; README.md:35). Multi-tenancy is cloud-only via API keys.

Packaging. Published via PyPI as pageindex. Requires Python ≥3.10. Optional extras: pageindex[claude] for Claude Agent SDK; pageindex[anthropic] for Anthropic SDK; pageindex[openai] (empty, keeps the install flag valid) (pyproject.toml:50-54).

Editor's note. Correction: mcp_bridge.py is an MCP client for the hosted PageIndex MCP server, not an MCP server. In local mode as_claude_mcp() returns an in-process Claude Agent SDK MCP server, which needs pageindex[claude]. The local store defaults to ./.pageindex.

onyx-dot-app/onyx

answered

Deployment model: Onyx is a self-hosted service deployed via Docker Compose. The deployment configurations live in deployment/docker_compose/ with variants for production (docker-compose.prod.yml), development (docker-compose.dev.yml), multi-tenant (docker-compose.multitenant.yml), air-gapped, and testing setups. The base docker-compose.yml defines the core services.

Required infrastructure: The system requires PostgreSQL (relational store), Redis (caching, Celery broker, locks), OpenSearch (document/vector index), and optionally MinIO/S3 (file/blob storage). The docker-compose.resources.yml file provisions these. For production, a reverse proxy (nginx with Let's Encrypt via init-letsencrypt.sh) is included.

Service architecture: The backend comprises multiple Celery workers (primary, docfetching, docprocessing, light, heavy, monitoring, beat) orchestrated by Docker Compose. A supervisord.conf manages the FastAPI API server (main.py) and the web server inside the API container. The model server (model_server/) runs as a separate Docker image (Dockerfile.model_server).

API: The FastAPI application serves REST endpoints for search, chat, document management, connector administration, and user management. All calls go through the frontend at port 3000 (proxied to the API), as stated in CLAUDE.md.

UI: The Next.js frontend (web/) provides the admin dashboard and chat interface. A separate mobile app (mobile/, React Native + Expo) is also available.

Multi-tenancy: Enterprise Edition supports multi-tenant deployments with separate schemas per tenant. The DynamicTenantScheduler handles per-tenant Celery beat tasks. OpenSearch indices are tenant-scoped. Alembic migrations have a separate schema_private config for tenant schema migrations.

Scaling: OpenSearch can be horizontally scaled. Celery workers use thread pools with configurable concurrency per pod. Redis handles inter-process queuing. The Kubernetes helm chart (deployment/helm/) and AWS ECS Fargate (deployment/aws_ecs_fargate/) deployment options are available for larger deployments.

Lite (demo) deployment: docker-compose.onyx-lite.yml provides a minimal configuration for evaluation purposes.

Third-party dependencies: Embedding and LLM calls go through LiteLLM (multi-provider gateway). Optional Unstructured.io API for enhanced document parsing. Langfuse and Braintrust integrations for observability and evals.

deepset-ai/haystack

answered

Library, not a service. Haystack is a Python framework published as the haystack-ai package on PyPI and conda-forge (pyproject.toml:5). It is installed via pip install haystack-ai and embedded into a user's own application — it does not ship a built-in HTTP server, CLI daemon, or web UI. There are no server classes, no FastAPI/Flask endpoints, and no __main__ entry point in the repository.

Deployment via Hayhooks (separate project). The README (README.md:76) and dedicated documentation page (docs-website/docs/development/hayhooks.mdx) point to Hayhooks, a separate open-source project (deepset-ai/hayhooks) that wraps Haystack pipelines as REST APIs, MCP servers, or OpenAI-compatible chat-completion endpoints. A user writes a PipelineWrapper subclass with a run_api method, deploys a YAML pipeline definition, and Hayhooks serves it via uvicorn on port 1416. Hayhooks also supports file uploads, streaming, CORS, SSL, and a Chainlit UI.

API Surface. The core API is the Pipeline class (haystack/core/pipeline/pipeline.py:118) with synchronous run() and asynchronous run_async() / run_async_generator() methods. Pipelines are directed graphs of @component-decorated classes (haystack/core/component/component.py:8). Components connect via input/output sockets, and the pipeline orchestrates execution in topological order with support for branches, loops, and conditional routing. The Pipeline.loads() / Pipeline.dumps() methods (haystack/core/pipeline/base.py:294-370) serialize pipelines to YAML or JSON for portability.

Required Infrastructure. Haystack itself requires only Python 3.10+ and the haystack-ai package with its dependencies (networkx, pydantic, httpx, openai SDK, etc. — pyproject.toml:44-63). However, a useful Haystack application typically also needs external LLM provider APIs (OpenAI, Anthropic, Mistral, etc.), a vector database (external integrations like Weaviate, Pinecone, Qdrant, Milvus, or the bundled InMemoryDocumentStore), and optionally embedding services, file converters, and rankers. API credentials are handled via the Secret.from_env_var() pattern (haystack/utils/auth.py:182-215) which loads tokens from environment variables.

Scaling and Multi-tenancy. Haystack provides no built-in load balancing, horizontal scaling, or multi-tenant isolation. The async pipeline support (Pipeline.run_async(), haystack/core/pipeline/pipeline.py:1096) enables non-blocking execution suitable for concurrent requests in an async web server (e.g., FastAPI + uvicorn), but scaling out, queuing, tenant-aware routing, and resource isolation are the responsibility of the deployment layer. The official Docker image (docker/Dockerfile.base) packages Haystack on python:3.12-slim and is built via Docker Bake for linux/amd64 and linux/arm64 (docker/docker-bake.hcl:38-39), intended as a base image for derived containers rather than a standalone service.

Telemetry. Haystack includes anonymous usage telemetry sent to PostHog (haystack/telemetry/_telemetry.py:6), opt-out via the HAYSTACK_TELEMETRY_ENABLED environment variable, reporting pipeline runs and system specs — a concern for deployment environments that require strict data sovereignty.

Cinnamon/kotaemon

answered

Kotaemon is a self-hosted web application, not a library or API service. It runs on Gradio (app.py, line 18–26) and is served as a full UI. The entry point is python app.py which launches ktem.main.App with .queue().launch(). A Gradio server listens on configurable GRADIO_SERVER_NAME and GRADIO_SERVER_PORT (default 7860). The UI includes Chat, File management (indices), Settings, Resources, and Help tabs (libs/ktem/ktem/main.py, lines 46–82).

Docker deployment is the recommended path. The Dockerfile (Dockerfile) provides four variants: lite (Python 3.11-slim with basic dependencies + pdf.js), full (adds Tesseract OCR, LibreOffice, Unstructured, torch), paddle (adds PaddleOCR GPU support with CUDA 11.3), and ollama (bundles Ollama with nomic-embed-text). Images are published on GHCR (ghcr.io/cinnamon/kotaemon).

Infrastructure requirements: SQLite database (for users, indices, model configs), a filestorage directory for uploaded files, and persistent directories for vectorstore and docstore — all under ./ktem_app_data by default. Multi-user support is built in (SQLAlchemy-backed user management with login page, private/public collections, admin default user) (flowsettings.py, lines 76–84). SSO is optionally supported via KH_SSO_ENABLED. The first-run setup wizard can be enabled.

Scaling and multi-tenancy: The app uses SQLite by default (single-server). It supports private (per-user) and public file collections via the private index config and user_id column in Source/Index tables (libs/ktem/ktem/index/file/index.py, lines 62–126). The index_manager manages multiple indices (File Collection, GraphRAG collections, etc.) with separate tables, vector stores, and doc stores per index (libs/ktem/ktem/index/file/index.py, lines 153–163). Embedding and LLM models are managed via per-user managers (SQL-persisted) with a pool pattern (libs/ktem/ktem/embeddings/manager.py, libs/ktem/ktem/llms/manager.py). Elasticsearch is listed as an alternative docstore for deployments that need to scale beyond SQLite's full-text search capabilities. The app also provides a settings.yaml config file and fly.toml for Fly.io deployment.

HKUDS/RAG-Anything

answered

Library, not a service. RAG-Anything is a Python library (PyPI package raganything) with no built-in server, REST API, or web UI. It does not include FastAPI endpoints, a Gradio interface, or a CLI server. It is designed to be imported and used programmatically within a Python application.

Infrastructure and scaling. Required infrastructure includes: a document parser (MinerU, installed as a system-level CLI tool; Docling as Python API; or PaddleOCR optionally), plus an LLM and embedding model accessible via user-provided callables (any OpenAI-compatible API, Ollama local, or custom functions as shown in examples/). The RAGAnythingConfig exposes only max_concurrent_files (config.py:68-71) for limited parallelism. There is no built-in horizontal scaling, sharding, or multi-tenant isolation — those would need to be handled by the LightRAG storage backends (e.g., external MongoDB/Neo4j).

Multi-tenancy. Not supported natively. Each RAGAnything instance owns one working_dir (config.py:18-19), which maps to one LightRAG workspace. Separate workspaces would require separate instances or manually swapped working_dir paths.

UI and API. None in the repository. The examples/ directory shows usage patterns: raganything_example.py demonstrates end-to-end processing and queries, ollama_integration_example.py shows local deployment with Ollama, and lmstudio_integration_example.py / vllm_integration_example.py show alternative LLM backends. These are scripts, not a web service.

Operating the parser. MinerU requires system-level installation (mineru[core] PyPI package, which installs the mineru CLI) and may need model data downloaded on first run. Office file support requires LibreOffice installed on the host (parser.py:296-460). These are significant operational dependencies for a production deployment. The MineruExecutionError class (parser.py:111-123) includes known-failure diagnosis with actionable remediation hints for common issues like version mismatches.

Environment configuration. All parser and processing options are configurable via environment variables or the RAGAnythingConfig dataclass (config.py:13-171), including WORKING_DIR, PARSE_METHOD, PARSER, ENABLE_IMAGE_PROCESSING, MAX_CONCURRENT_FILES, and more.

SciPhi-AI/R2R

answered

Library vs service. R2R is both a Python library (pip install r2r) and a FastAPI web service. The library exposes R2RClient and R2RAsyncClient (py/r2r/__init__.py:1-19) for programmatic access. The service runs via uvicorn on the core.main.app_entry:app FastAPI application (py/core/main/app_entry.py:109-139).

API. Versioned REST API under /v3/ with routers for search, RAG, agent conversations, ingestion, documents, collections, users, conversations, prompts, graphs, chunks, indices, and system health (py/core/main/app.py:14-23). The /retrieval/search, /retrieval/rag, and /retrieval/agent endpoints are the primary retrieval/RAG surface. Authentication is pluggable via AuthProvider (Supabase, JWT, Clerk). CORS is wide open (allow_origins=["*"]).

Configuration. Multiple toml config files (py/core/main/config.py:29-43) with environment variable overrides. Key env vars: R2R_POSTGRES_HOST/PORT/USER/PASSWORD/DBNAME, OPENAI_API_KEY, ANTHROPIC_API_KEY, MISTRAL_API_KEY, R2R_SENTRY_DSN, R2R_CONFIG_NAME, R2R_PROJECT_NAME.

Required infrastructure. PostgreSQL 15+ with pgvector extension is mandatory. Optional: Unstructured API (or self-hosted Unstructured service) for advanced document parsing; Mistral API for OCR; Redis for APScheduler jobs; Sentry DSN for error tracking.

Deployment. Docker image (py/Dockerfile) runs uvicorn on port 8000 (configurable via R2R_PORT). The Docker build uses multi-stage with Rust compilation for tokenizer dependencies. A CMD of uvicorn core.main.app_entry:app --host $R2R_HOST --port $R2R_PORT launches the service.

Scaling and multi-tenancy. No horizontal scaling infrastructure — the SimpleOrchestrationProvider (py/core/providers/orchestration/simple.py) is a stub with no background worker. A HatchetOrchestrationProvider exists for orchestration via an external Hatchet server. Multi-tenancy is handled via collection-based scoping: users and documents belong to collections, and all queries filter by collection_ids (py/core/providers/database/chunks.py:178-179 creates a GIN index on collection_ids). The database schema uses PostgreSQL schemas (one per project via R2R_PROJECT_NAME) for logical isolation.

Editor's note. Correction: the default simple orchestration is not a stub; it runs ingestion and graph workflows inline inside the HTTP request (Hatchet is used only by the full configs). Redis is not used: the scheduler is in-process APScheduler.

← How is quality evaluated or observed?