LLMs Technical Reviews
Home / RAG engines / Comparison

RAG engines: ragflow vs anything-llm vs llama_index vs quivr vs PageIndex vs onyx vs haystack vs kotaemon vs RAG-Anything vs R2R

End-to-end retrieval-augmented generation engines and frameworks — document parsing, chunking, indexing, retrieval, reranking and grounded answers. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.

At a glance

● answered from code · — not applicable (the project does not do this) · ? insufficient evidence

How are documents parsed and chunked?

RAGFlow parses scanned, mixed-format and table-heavy files best. RAG-Anything is the one to use when figures and equations carry the meaning. Onyx is strongest at pulling content out of workplace apps rather than parsing it.

Layout-aware parsing. RAGFlow runs its own ONNX OCR, layout and table-structure models (DeepDoc) and offers twelve chunking templates. The General template cuts at 512 tokens on sentence delimiters. Kotaemon lets each user switch the PDF loader in settings (Adobe, Azure Document Intelligence, Docling, two PaddleOCR modes). It splits text at 1,024 tokens with 256 overlap and keeps tables and figures whole. RAG-Anything parses with MinerU by default. A vision model describes each image, table and equation, and that description becomes a graph entity. LightRAG chunks the plain text.

Connector-fed apps with plain text extraction. Onyx has about 50 source connectors. It reads PDFs with pypdfium2 and chunks them into 512-token sentence chunks with no overlap. AnythingLLM splits on characters (1,000 per chunk, 20 overlap). It runs Tesseract OCR only on PDFs that yield no text, and on images, which it OCRs rather than captions. R2R maps about 35 file types and defaults to 1,024-character chunks with 512 overlap. Only the recursive and character splitters work. basic and by_title raise NotImplementedError.

Framework building blocks. In LlamaIndex the default SentenceSplitter makes 1,024-token chunks with 200 overlap, and semantic and code splitters are also available. OCR comes only from integrations. Haystack's DocumentSplitter defaults to 200 words with no overlap, and it ships no OCR engine.

No classic chunking. PageIndex builds a section tree from pdfium font statistics. Local mode accepts only PDFs and has no OCR. Quivr cuts 384-token windows with 48 overlap. Its first-party PDF normalizer extracts text one page at a time with no OCR, and it has no Office support.

Pick: RAGFlow for scans, tables and mixed formats. Pick: RAG-Anything when charts, figures and formulas must be answerable. Pick: PageIndex for long, well-structured digital PDFs where section boundaries matter more than chunks.

How are embeddings and indexes built and stored?

R2R has the simplest strong index: vectors and keyword search in one Postgres. RAGFlow has the most complete hybrid index at scale. LlamaIndex gives the widest choice of backend.

Hybrid index in one search engine. RAGFlow stores a BM25 token field (content_ltks), a dense vector and position metadata on every chunk. It runs on one of five engines, with Elasticsearch as the default. Onyx supports OpenSearch only. It embeds each document title as its own vector, stores ACLs on every chunk, and builds a secondary index so the embedding model can be switched without downtime. R2R uses pgvector plus a generated English tsvector. HNSW or IVFFlat indexes have to be created explicitly through the API. Quivr treats Weaviate as a rebuildable copy of canonical Postgres and S3 data. Each embedding model owns its own "vector space", which can be evaluated beside the live one and then promoted.

Pluggable dense stores. AnythingLLM has ten vector adapters (LanceDB by default) and a local MiniLM embedder. It has no keyword index, and each adapter re-implements chunking and embedding. Kotaemon pairs a local Chroma vector store with a LanceDB document store that provides full-text search. LlamaIndex has about 80 vector-store integrations. Its built-in SimpleVectorStore scores every vector in Python and has no hybrid mode, and BM25 is a separate package. Haystack's DocumentStore protocol requires only four methods. Its in-memory store does brute-force vector search and implements BM25L/Okapi/Plus itself.

Graph plus vectors. RAG-Anything writes chunks, entities and relations into LightRAG (<1.5) stores. It builds no BM25 or keyword index, so there is no lexical retrieval.

No vectors. PageIndex stores one JSON tree per document. Each node has a page range and an LLM summary of up to 150 words.

Pick: R2R if you want one Postgres for vectors, keywords and users. Pick: RAGFlow or Onyx for large corpora that need lexical plus semantic recall. Pick: LlamaIndex or Haystack to keep the vector database you already run.

How is retrieval performed?

RAGFlow and R2R give the most tunable hybrid retrieval. Onyx does the most per query, with LLM query expansion and section selection.

Hybrid with fusion. RAGFlow sends BM25 and KNN queries together. It then blends token and vector similarity with vector_similarity_weight (default 0.3), unless a rerank model is set. Empty results trigger looser retries and a dense-only fallback. R2R runs separate vector and full-text SQL queries and fuses them with weighted reciprocal rank fusion. HyDE and RAG-Fusion can be switched on per request. Its only reranker is a TEI endpoint, which is off by default. Onyx fuses several LLM-rewritten queries with RRF and then asks the LLM to choose sections. It has no reranker. Its OpenSearch blend uses fixed weights, so hybrid_alpha only matters when it is 0 (pure BM25). Quivr runs Weaviate hybrid with alpha 0.5. Its deep profile ranks the same as default unless the optional rerank plugin is installed, and that plugin calls an external paid service.

Hybrid that is weaker than it looks. Kotaemon puts keyword hits in front of vector hits and does not fuse them. Its dedupe check never matches, and the default Cohere reranker returns its input unchanged when no key is set. Follow-up questions are searched exactly as typed.

Dense only. AnythingLLM returns the top-N chunks above a similarity threshold, with no metadata filters. Reranking works only on its LanceDB backend.

Composable frameworks. LlamaIndex offers QueryFusionRetriever (four LLM rewrites, then fusion) plus rerankers as postprocessors. Whether hybrid works depends on the vector backend. Haystack's MultiRetriever merges named retrievers with RRF, and rankers and query expansion are separate components.

Graph or structure navigation. RAG-Anything uses LightRAG's mix mode by default, which adds vector chunks to graph retrieval. PageIndex has no similarity search: an agent reads the section tree and fetches page ranges.

Pick: RAGFlow or R2R for tunable hybrid search over large corpora. Pick: Onyx for permission-filtered company search. Pick: PageIndex for structure-dependent questions over a few long documents.

How are answers generated and grounded?

RAGFlow does the most to ground answers, because it attaches citations even when the model leaves them out. Kotaemon shows evidence most clearly. R2R has the cleanest citation stream for building your own product.

The engine enforces citations. If the model writes no [ID:n] markers, RAGFlow embeds each answer sentence and attaches up to four chunks by cosine similarity. Its reasoning levels add a planner graph or a tool-using agent. R2R turns bracketed short ids into Citation objects. When streaming, it emits a citation event the first time each id appears. Onyx gives the model numbered documents inside an agent loop of up to six cycles. A stream processor turns [1] markers into links. In highlight mode, Kotaemon runs a separate function-calling pass that extracts verbatim quotes and fuzzy-matches them to chunk spans in its PDF viewer. PageIndex has the model write <cite doc page/> tags. These are page-level in local mode.

Sources attached, citation left to you. AnythingLLM returns retrieved chunks as sources. To fit the context window, it cuts text out of the middle of the system prompt, history or user message. LlamaIndex returns source_nodes with every response and uses compact-and-refine synthesis by default. CitationQueryEngine re-splits sources into numbered 512-token chunks. Haystack's AnswerBuilder can map [n] references back to documents through reference_pattern. Its Agent handles multi-step answering, with hooks.

Delegated or absent. RAG-Anything leaves the answer to LightRAG and does not render citations. Its VLM-enhanced query puts retrieved images into the prompt. Quivr does not generate answers at all. It returns segments with Unicode code-point offsets, and its MCP tools tell an external agent how to cite them.

Pick: RAGFlow or Onyx when every answer must be traceable without depending on the model. Pick: Kotaemon when users need to see the highlighted evidence. Pick: R2R or Quivr when you build your own front end on top of an API.

How is quality evaluated or observed?

LlamaIndex and Haystack are the only projects with answer-quality evaluators you can run. Among the applications, Onyx has the most usable evaluation and tracing.

Evaluator libraries. LlamaIndex ships faithfulness, relevancy, correctness and pairwise LLM judges. It also has RetrieverEvaluator (hit rate, MRR) and BatchEvalRunner, which runs evaluations concurrently. Haystack's evaluators are pipeline components: faithfulness, context relevance, MAP, MRR, NDCG, recall and semantic answer similarity. Its tracer interface is pluggable, and anonymous usage telemetry is on by default.

Product-level evals and tracing. Onyx runs chat datasets locally or in Braintrust and checks which tools were called. Every LLM flow is traced to Langfuse or Braintrust. It has no RAGAS-style metrics. RAGFlow traces each chat to Langfuse (per-tenant keys) and has OpenTelemetry spans. We found its evaluation tables to be schema only: no service, route or UI runs them. Quivr exposes metrics, OpenTelemetry and per-organisation rollups. It also lets you index a new embedding model as an evaluation vector space next to the live one, which compares models but does not score answers.

Runtime signals only. Kotaemon has an LLM grade each retrieved chunk and computes a qa_score from logprobs. It warns when relevance falls below 0.3. We found no evaluation harness, but the evidence was incomplete. AnythingLLM records tokens, speed and cost for each message. R2R has Sentry and JSON logs, plus a Ragas cookbook that runs outside the server. RAG-Anything's MetricsCallback counts documents and times each stage. PageIndex turns Agents SDK tracing off. Its only quality checks run during indexing (verify_toc, tree cost before and after optimisation).

Pick: LlamaIndex or Haystack to measure retrieval and faithfulness in code. Pick: Onyx or RAGFlow for production traces in Langfuse. Pick: for every other project, plan to add an external evaluation framework such as Ragas.

How is it deployed and operated?

AnythingLLM is the easiest to run: one container, or a desktop app. Onyx and RAGFlow are the most complete multi-user services. LlamaIndex and Haystack suit teams that want to embed RAG in their own code.

Multi-service platforms. RAGFlow is one Go binary that starts as separate API, admin, ingestor, syncer and DeepDoc processes. It runs beside MySQL, MinIO, Kvrocks, NATS JetStream, ClickHouse and a search engine, with tenant-scoped models and datasets. Onyx runs a FastAPI server, Celery workers, two embedding model servers, Postgres, Redis and OpenSearch. It ships a Helm chart and an ECS template. Multi-tenancy and permission sync are under the separate ee/ licence. Quivr needs Postgres, Temporal, Weaviate, TEI and SeaweedFS plus plugin sidecars. Its API is v0 and the README marks it as evaluation stage. R2R needs only Postgres with pgvector. Its default simple orchestration runs ingestion inside the upload request, and Hatchet workers come only with the full config. Authentication is off by default, and the default admin password is well known.

Single-node applications. AnythingLLM bundles SQLite, LanceDB and a local embedder. Several users share one instance, and there is no horizontal scaling. Kotaemon is a Gradio app with SQLite and no REST API. It starts with an admin/admin login.

Libraries. LlamaIndex persists JSON stores through fsspec and has no server. Haystack has no server either, so serving is left to the separate Hayhooks project. RAG-Anything keeps one knowledge base per working_dir and needs the MinerU CLI and LibreOffice on the host. PageIndex stores JSON files in ./.pageindex. When given an API key it hands indexing and storage to the vendor's hosted service.

Pick: AnythingLLM or Kotaemon for a private RAG tool on one machine. Pick: Onyx or RAGFlow for a self-hosted, multi-user company deployment. Pick: LlamaIndex, Haystack or R2R's API to build RAG into your own product.

The projects