RAG engines
End-to-end retrieval-augmented generation engines and frameworks — document parsing, chunking, indexing, retrieval, reranking and grounded answers.
RAG engines take your documents and answer questions from them. They parse files, build an index, find the passages a question needs and ask an LLM to answer with citations. The projects here vary a lot. Some are full self-hosted products with a UI, workers and a search cluster. Others are libraries you call from your own code. Some rely on vector and keyword search. Others drop embeddings and let an LLM agent navigate a document's structure. When you choose one, check six things. Does the parser handle your formats, and scanned pages or tables? Which search backends does it support? Does retrieval mix lexical and semantic signals, and does it rerank? How are citations attached? Is there any built-in way to measure answer quality? How much infrastructure must you run?
Projects (2)
| Project | Stars |
|---|---|
| infiniflow/ragflowGo RAG server with DeepDoc layout parsing, hybrid BM25 + vector search over pluggable engines, and agentic chat modes. | ★ 92k |
| VectifyAI/PageIndexPython SDK that builds a page-ranged section tree per PDF and lets an LLM agent read it with tools instead of vector search. | ★ 39k |
In the research queue: run-llama/llama_index, onyx-dot-app/onyx, deepset-ai/haystack, Mintplex-Labs/anything-llm, HKUDS/RAG-Anything, The-Vibe-Company/quivr, SciPhi-AI/R2R, Cinnamon/kotaemon.
Comparison questions
Each question is answered separately for every project in this category, from that project's source code.
- How are documents parsed and chunked?Supported formats; OCR and layout parsing; table handling; chunking strategy and sizes.
- How are embeddings and indexes built and stored?Embedding models; vector stores supported; hybrid / keyword (BM25) indexes; metadata.
- How is retrieval performed?Dense / sparse / hybrid search; reranking; query rewriting or decomposition; filters.
- How are answers generated and grounded?Prompt assembly; citations / source attribution; streaming; agentic or multi-step answering.
- How is quality evaluated or observed?Built-in evals, metrics, tracing / observability hooks; if absent, say so.
- How is it deployed and operated?Library vs service; UI; API; required infrastructure; scaling and multi-tenancy.