# RAG engines: comparison

> End-to-end retrieval-augmented generation engines and frameworks — document parsing, chunking, indexing, retrieval, reranking and grounded answers.

Canonical page: https://llms-technical-reviews.com/compare/rag/

## How are documents parsed and chunked?

[RAGFlow](/p/ragflow/) parses scanned, mixed-format and table-heavy files best. [RAG-Anything](/p/rag-anything/) is the one to use when figures and equations carry the meaning. [Onyx](/p/onyx/) is strongest at pulling content out of workplace apps rather than parsing it.

**Layout-aware parsing.** RAGFlow runs its own ONNX OCR, layout and table-structure models (DeepDoc) and offers twelve chunking templates. The General template cuts at 512 tokens on sentence delimiters. [Kotaemon](/p/kotaemon/) lets each user switch the PDF loader in settings (Adobe, Azure Document Intelligence, Docling, two PaddleOCR modes). It splits text at 1,024 tokens with 256 overlap and keeps tables and figures whole. RAG-Anything parses with MinerU by default. A vision model describes each image, table and equation, and that description becomes a graph entity. LightRAG chunks the plain text.

**Connector-fed apps with plain text extraction.** Onyx has about 50 source connectors. It reads PDFs with pypdfium2 and chunks them into 512-token sentence chunks with no overlap. [AnythingLLM](/p/anything-llm/) splits on characters (1,000 per chunk, 20 overlap). It runs Tesseract OCR only on PDFs that yield no text, and on images, which it OCRs rather than captions. [R2R](/p/r2r/) maps about 35 file types and defaults to 1,024-character chunks with 512 overlap. Only the `recursive` and `character` splitters work. `basic` and `by_title` raise `NotImplementedError`.

**Framework building blocks.** In [LlamaIndex](/p/llama_index/) the default `SentenceSplitter` makes 1,024-token chunks with 200 overlap, and semantic and code splitters are also available. OCR comes only from integrations. [Haystack](/p/haystack/)'s `DocumentSplitter` defaults to 200 words with no overlap, and it ships no OCR engine.

**No classic chunking.** [PageIndex](/p/pageindex/) builds a section tree from pdfium font statistics. Local mode accepts only PDFs and has no OCR. [Quivr](/p/quivr/) cuts 384-token windows with 48 overlap. Its first-party PDF normalizer extracts text one page at a time with no OCR, and it has no Office support.

Pick: RAGFlow for scans, tables and mixed formats.
Pick: RAG-Anything when charts, figures and formulas must be answerable.
Pick: PageIndex for long, well-structured digital PDFs where section boundaries matter more than chunks.

Per-project answers: https://llms-technical-reviews.com/rag/q/ingestion/index.md

## How are embeddings and indexes built and stored?

[R2R](/p/r2r/) has the simplest strong index: vectors and keyword search in one Postgres. [RAGFlow](/p/ragflow/) has the most complete hybrid index at scale. [LlamaIndex](/p/llama_index/) gives the widest choice of backend.

**Hybrid index in one search engine.** RAGFlow stores a BM25 token field (`content_ltks`), a dense vector and position metadata on every chunk. It runs on one of five engines, with Elasticsearch as the default. [Onyx](/p/onyx/) supports OpenSearch only. It embeds each document title as its own vector, stores ACLs on every chunk, and builds a secondary index so the embedding model can be switched without downtime. R2R uses pgvector plus a generated English `tsvector`. HNSW or IVFFlat indexes have to be created explicitly through the API. [Quivr](/p/quivr/) treats Weaviate as a rebuildable copy of canonical Postgres and S3 data. Each embedding model owns its own "vector space", which can be evaluated beside the live one and then promoted.

**Pluggable dense stores.** [AnythingLLM](/p/anything-llm/) has ten vector adapters (LanceDB by default) and a local MiniLM embedder. It has no keyword index, and each adapter re-implements chunking and embedding. [Kotaemon](/p/kotaemon/) pairs a local Chroma vector store with a LanceDB document store that provides full-text search. LlamaIndex has about 80 vector-store integrations. Its built-in `SimpleVectorStore` scores every vector in Python and has no hybrid mode, and BM25 is a separate package. [Haystack](/p/haystack/)'s `DocumentStore` protocol requires only four methods. Its in-memory store does brute-force vector search and implements BM25L/Okapi/Plus itself.

**Graph plus vectors.** [RAG-Anything](/p/rag-anything/) writes chunks, entities and relations into LightRAG (<1.5) stores. It builds no BM25 or keyword index, so there is no lexical retrieval.

**No vectors.** [PageIndex](/p/pageindex/) stores one JSON tree per document. Each node has a page range and an LLM summary of up to 150 words.

Pick: R2R if you want one Postgres for vectors, keywords and users.
Pick: RAGFlow or Onyx for large corpora that need lexical plus semantic recall.
Pick: LlamaIndex or Haystack to keep the vector database you already run.

Per-project answers: https://llms-technical-reviews.com/rag/q/indexing/index.md

## How is retrieval performed?

[RAGFlow](/p/ragflow/) and [R2R](/p/r2r/) give the most tunable hybrid retrieval. [Onyx](/p/onyx/) does the most per query, with LLM query expansion and section selection.

**Hybrid with fusion.** RAGFlow sends BM25 and KNN queries together. It then blends token and vector similarity with `vector_similarity_weight` (default 0.3), unless a rerank model is set. Empty results trigger looser retries and a dense-only fallback. R2R runs separate vector and full-text SQL queries and fuses them with weighted reciprocal rank fusion. HyDE and RAG-Fusion can be switched on per request. Its only reranker is a TEI endpoint, which is off by default. Onyx fuses several LLM-rewritten queries with RRF and then asks the LLM to choose sections. It has no reranker. Its OpenSearch blend uses fixed weights, so `hybrid_alpha` only matters when it is 0 (pure BM25). [Quivr](/p/quivr/) runs Weaviate hybrid with alpha 0.5. Its `deep` profile ranks the same as `default` unless the optional rerank plugin is installed, and that plugin calls an external paid service.

**Hybrid that is weaker than it looks.** [Kotaemon](/p/kotaemon/) puts keyword hits in front of vector hits and does not fuse them. Its dedupe check never matches, and the default Cohere reranker returns its input unchanged when no key is set. Follow-up questions are searched exactly as typed.

**Dense only.** [AnythingLLM](/p/anything-llm/) returns the top-N chunks above a similarity threshold, with no metadata filters. Reranking works only on its LanceDB backend.

**Composable frameworks.** [LlamaIndex](/p/llama_index/) offers `QueryFusionRetriever` (four LLM rewrites, then fusion) plus rerankers as postprocessors. Whether hybrid works depends on the vector backend. [Haystack](/p/haystack/)'s `MultiRetriever` merges named retrievers with RRF, and rankers and query expansion are separate components.

**Graph or structure navigation.** [RAG-Anything](/p/rag-anything/) uses LightRAG's `mix` mode by default, which adds vector chunks to graph retrieval. [PageIndex](/p/pageindex/) has no similarity search: an agent reads the section tree and fetches page ranges.

Pick: RAGFlow or R2R for tunable hybrid search over large corpora.
Pick: Onyx for permission-filtered company search.
Pick: PageIndex for structure-dependent questions over a few long documents.

Per-project answers: https://llms-technical-reviews.com/rag/q/retrieval/index.md

## How are answers generated and grounded?

[RAGFlow](/p/ragflow/) does the most to ground answers, because it attaches citations even when the model leaves them out. [Kotaemon](/p/kotaemon/) shows evidence most clearly. [R2R](/p/r2r/) has the cleanest citation stream for building your own product.

**The engine enforces citations.** If the model writes no `[ID:n]` markers, RAGFlow embeds each answer sentence and attaches up to four chunks by cosine similarity. Its reasoning levels add a planner graph or a tool-using agent. R2R turns bracketed short ids into `Citation` objects. When streaming, it emits a `citation` event the first time each id appears. [Onyx](/p/onyx/) gives the model numbered documents inside an agent loop of up to six cycles. A stream processor turns `[1]` markers into links. In highlight mode, Kotaemon runs a separate function-calling pass that extracts verbatim quotes and fuzzy-matches them to chunk spans in its PDF viewer. [PageIndex](/p/pageindex/) has the model write `<cite doc page/>` tags. These are page-level in local mode.

**Sources attached, citation left to you.** [AnythingLLM](/p/anything-llm/) returns retrieved chunks as `sources`. To fit the context window, it cuts text out of the middle of the system prompt, history or user message. [LlamaIndex](/p/llama_index/) returns `source_nodes` with every response and uses compact-and-refine synthesis by default. `CitationQueryEngine` re-splits sources into numbered 512-token chunks. [Haystack](/p/haystack/)'s `AnswerBuilder` can map `[n]` references back to documents through `reference_pattern`. Its `Agent` handles multi-step answering, with hooks.

**Delegated or absent.** [RAG-Anything](/p/rag-anything/) leaves the answer to LightRAG and does not render citations. Its VLM-enhanced query puts retrieved images into the prompt. [Quivr](/p/quivr/) does not generate answers at all. It returns segments with Unicode code-point offsets, and its MCP tools tell an external agent how to cite them.

Pick: RAGFlow or Onyx when every answer must be traceable without depending on the model.
Pick: Kotaemon when users need to see the highlighted evidence.
Pick: R2R or Quivr when you build your own front end on top of an API.

Per-project answers: https://llms-technical-reviews.com/rag/q/generation/index.md

## How is quality evaluated or observed?

[LlamaIndex](/p/llama_index/) and [Haystack](/p/haystack/) are the only projects with answer-quality evaluators you can run. Among the applications, [Onyx](/p/onyx/) has the most usable evaluation and tracing.

**Evaluator libraries.** LlamaIndex ships faithfulness, relevancy, correctness and pairwise LLM judges. It also has `RetrieverEvaluator` (hit rate, MRR) and `BatchEvalRunner`, which runs evaluations concurrently. Haystack's evaluators are pipeline components: faithfulness, context relevance, MAP, MRR, NDCG, recall and semantic answer similarity. Its tracer interface is pluggable, and anonymous usage telemetry is on by default.

**Product-level evals and tracing.** Onyx runs chat datasets locally or in Braintrust and checks which tools were called. Every LLM flow is traced to Langfuse or Braintrust. It has no RAGAS-style metrics. [RAGFlow](/p/ragflow/) traces each chat to Langfuse (per-tenant keys) and has OpenTelemetry spans. We found its evaluation tables to be schema only: no service, route or UI runs them. [Quivr](/p/quivr/) exposes metrics, OpenTelemetry and per-organisation rollups. It also lets you index a new embedding model as an evaluation vector space next to the live one, which compares models but does not score answers.

**Runtime signals only.** [Kotaemon](/p/kotaemon/) has an LLM grade each retrieved chunk and computes a `qa_score` from logprobs. It warns when relevance falls below 0.3. We found no evaluation harness, but the evidence was incomplete. [AnythingLLM](/p/anything-llm/) records tokens, speed and cost for each message. [R2R](/p/r2r/) has Sentry and JSON logs, plus a Ragas cookbook that runs outside the server. [RAG-Anything](/p/rag-anything/)'s `MetricsCallback` counts documents and times each stage. [PageIndex](/p/pageindex/) turns Agents SDK tracing off. Its only quality checks run during indexing (`verify_toc`, tree cost before and after optimisation).

Pick: LlamaIndex or Haystack to measure retrieval and faithfulness in code.
Pick: Onyx or RAGFlow for production traces in Langfuse.
Pick: for every other project, plan to add an external evaluation framework such as Ragas.

Per-project answers: https://llms-technical-reviews.com/rag/q/evaluation/index.md

## How is it deployed and operated?

[AnythingLLM](/p/anything-llm/) is the easiest to run: one container, or a desktop app. [Onyx](/p/onyx/) and [RAGFlow](/p/ragflow/) are the most complete multi-user services. [LlamaIndex](/p/llama_index/) and [Haystack](/p/haystack/) suit teams that want to embed RAG in their own code.

**Multi-service platforms.** RAGFlow is one Go binary that starts as separate API, admin, ingestor, syncer and DeepDoc processes. It runs beside MySQL, MinIO, Kvrocks, NATS JetStream, ClickHouse and a search engine, with tenant-scoped models and datasets. Onyx runs a FastAPI server, Celery workers, two embedding model servers, Postgres, Redis and OpenSearch. It ships a Helm chart and an ECS template. Multi-tenancy and permission sync are under the separate `ee/` licence. [Quivr](/p/quivr/) needs Postgres, Temporal, Weaviate, TEI and SeaweedFS plus plugin sidecars. Its API is `v0` and the README marks it as evaluation stage. [R2R](/p/r2r/) needs only Postgres with pgvector. Its default `simple` orchestration runs ingestion inside the upload request, and Hatchet workers come only with the `full` config. Authentication is off by default, and the default admin password is well known.

**Single-node applications.** AnythingLLM bundles SQLite, LanceDB and a local embedder. Several users share one instance, and there is no horizontal scaling. [Kotaemon](/p/kotaemon/) is a Gradio app with SQLite and no REST API. It starts with an `admin`/`admin` login.

**Libraries.** LlamaIndex persists JSON stores through fsspec and has no server. Haystack has no server either, so serving is left to the separate Hayhooks project. [RAG-Anything](/p/rag-anything/) keeps one knowledge base per `working_dir` and needs the MinerU CLI and LibreOffice on the host. [PageIndex](/p/pageindex/) stores JSON files in `./.pageindex`. When given an API key it hands indexing and storage to the vendor's hosted service.

Pick: AnythingLLM or Kotaemon for a private RAG tool on one machine.
Pick: Onyx or RAGFlow for a self-hosted, multi-user company deployment.
Pick: LlamaIndex, Haystack or R2R's API to build RAG into your own product.

Per-project answers: https://llms-technical-reviews.com/rag/q/deploy/index.md

## Projects

- [infiniflow/ragflow](https://llms-technical-reviews.com/p/ragflow/index.md) — Go RAG server with DeepDoc layout parsing, hybrid BM25 + vector search over pluggable engines, and agentic chat modes.
- [Mintplex-Labs/anything-llm](https://llms-technical-reviews.com/p/anything-llm/index.md) — Self-hosted Node.js chat-with-your-documents app with workspaces, ten vector-store backends, ~40 LLM providers and an agent mode.
- [run-llama/llama_index](https://llms-technical-reviews.com/p/llama_index/index.md) — Python RAG and agent framework with a core of index, retriever, synthesizer and agent abstractions plus ~550 integration packages.
- [The-Vibe-Company/quivr](https://llms-technical-reviews.com/p/quivr/index.md) — Go engine that ingests content streams durably and serves hybrid search and alerts; parsing, embedding and ranking live in plugins.
- [VectifyAI/PageIndex](https://llms-technical-reviews.com/p/pageindex/index.md) — Python SDK that builds a page-ranged section tree per PDF and lets an LLM agent read it with tools instead of vector search.
- [onyx-dot-app/onyx](https://llms-technical-reviews.com/p/onyx/index.md) — Self-hosted AI chat and enterprise search: 40+ connectors index into OpenSearch, and an LLM tool-calling loop answers with citations.
- [deepset-ai/haystack](https://llms-technical-reviews.com/p/haystack/index.md) — Python framework for typed component pipelines and a tool-calling Agent, with an in-memory store and integrations for the rest.
- [Cinnamon/kotaemon](https://llms-technical-reviews.com/p/kotaemon/index.md) — Gradio app for chatting with documents: keyword plus vector retrieval, rerankers, highlighted citations, GraphRAG/LightRAG indexes and ReAct agents.
- [HKUDS/RAG-Anything](https://llms-technical-reviews.com/p/rag-anything/index.md) — Multimodal document RAG library on LightRAG: parses PDFs, Office, images, audio and video, and adds VLM-captioned entities to its graph.
- [SciPhi-AI/R2R](https://llms-technical-reviews.com/p/r2r/index.md) — Python RAG server on Postgres/pgvector: hybrid and graph search, cited streaming answers, and RAG and research agents behind a REST API.