# RAG engines

> End-to-end retrieval-augmented generation engines and frameworks — document parsing, chunking, indexing, retrieval, reranking and grounded answers.

RAG engines answer questions from your own documents. They parse files, index the content, retrieve the passages a question needs and have an LLM answer from them, usually with citations. The ten projects here split along four axes. The first is form: full self-hosted products with a UI and workers, API servers, or Python libraries you call from your own code. The second is parsing depth, from OCR and table-structure models to plain text extraction. The third is retrieval: hybrid BM25 plus vector search, dense-only search, a knowledge graph, or an agent that navigates a document's table of contents. The fourth is grounding: whether citations are enforced, only attached as sources, or left to you. When choosing, check that the parser handles your formats and scans. Check that keyword and vector results are really fused, and whether reranking needs a paid key. Check how citations are produced, whether you can measure answer quality, and how much infrastructure you must operate.


## Projects

- [infiniflow/ragflow](https://llms-technical-reviews.com/p/ragflow/) — Go RAG server with DeepDoc layout parsing, hybrid BM25 + vector search over pluggable engines, and agentic chat modes. (★91730, Go)
- [Mintplex-Labs/anything-llm](https://llms-technical-reviews.com/p/anything-llm/) — Self-hosted Node.js chat-with-your-documents app with workspaces, ten vector-store backends, ~40 LLM providers and an agent mode. (★66755, JavaScript)
- [run-llama/llama_index](https://llms-technical-reviews.com/p/llama_index/) — Python RAG and agent framework with a core of index, retriever, synthesizer and agent abstractions plus ~550 integration packages. (★52422, Python)
- [The-Vibe-Company/quivr](https://llms-technical-reviews.com/p/quivr/) — Go engine that ingests content streams durably and serves hybrid search and alerts; parsing, embedding and ranking live in plugins. (★39580, Go)
- [VectifyAI/PageIndex](https://llms-technical-reviews.com/p/pageindex/) — Python SDK that builds a page-ranged section tree per PDF and lets an LLM agent read it with tools instead of vector search. (★38760, Python)
- [onyx-dot-app/onyx](https://llms-technical-reviews.com/p/onyx/) — Self-hosted AI chat and enterprise search: 40+ connectors index into OpenSearch, and an LLM tool-calling loop answers with citations. (★32336, Python)
- [deepset-ai/haystack](https://llms-technical-reviews.com/p/haystack/) — Python framework for typed component pipelines and a tool-calling Agent, with an in-memory store and integrations for the rest. (★26683, Python)
- [Cinnamon/kotaemon](https://llms-technical-reviews.com/p/kotaemon/) — Gradio app for chatting with documents: keyword plus vector retrieval, rerankers, highlighted citations, GraphRAG/LightRAG indexes and ReAct agents. (★25796, Python)
- [HKUDS/RAG-Anything](https://llms-technical-reviews.com/p/rag-anything/) — Multimodal document RAG library on LightRAG: parses PDFs, Office, images, audio and video, and adds VLM-captioned entities to its graph. (★23483, Python)
- [SciPhi-AI/R2R](https://llms-technical-reviews.com/p/r2r/) — Python RAG server on Postgres/pgvector: hybrid and graph search, cited streaming answers, and RAG and research agents behind a REST API. (★8010, Python)


## Comparison questions

- [How are documents parsed and chunked?](https://llms-technical-reviews.com/rag/q/ingestion/) — Supported formats; OCR and layout parsing; table handling; chunking strategy and sizes.
- [How are embeddings and indexes built and stored?](https://llms-technical-reviews.com/rag/q/indexing/) — Embedding models; vector stores supported; hybrid / keyword (BM25) indexes; metadata.
- [How is retrieval performed?](https://llms-technical-reviews.com/rag/q/retrieval/) — Dense / sparse / hybrid search; reranking; query rewriting or decomposition; filters.
- [How are answers generated and grounded?](https://llms-technical-reviews.com/rag/q/generation/) — Prompt assembly; citations / source attribution; streaming; agentic or multi-step answering.
- [How is quality evaluated or observed?](https://llms-technical-reviews.com/rag/q/evaluation/) — Built-in evals, metrics, tracing / observability hooks; if absent, say so.
- [How is it deployed and operated?](https://llms-technical-reviews.com/rag/q/deploy/) — Library vs service; UI; API; required infrastructure; scaling and multi-tenancy.

Full comparison: https://llms-technical-reviews.com/compare/rag/