How is retrieval performed?
Dense / sparse / hybrid search; reranking; query rewriting or decomposition; filters.
Verdict
RAGFlow and R2R give the most tunable hybrid retrieval. Onyx does the most per query, with LLM query expansion and section selection.
Hybrid with fusion. RAGFlow sends BM25 and KNN queries together. It then blends token and vector similarity with vector_similarity_weight (default 0.3), unless a rerank model is set. Empty results trigger looser retries and a dense-only fallback. R2R runs separate vector and full-text SQL queries and fuses them with weighted reciprocal rank fusion. HyDE and RAG-Fusion can be switched on per request. Its only reranker is a TEI endpoint, which is off by default. Onyx fuses several LLM-rewritten queries with RRF and then asks the LLM to choose sections. It has no reranker. Its OpenSearch blend uses fixed weights, so hybrid_alpha only matters when it is 0 (pure BM25). Quivr runs Weaviate hybrid with alpha 0.5. Its deep profile ranks the same as default unless the optional rerank plugin is installed, and that plugin calls an external paid service.
Hybrid that is weaker than it looks. Kotaemon puts keyword hits in front of vector hits and does not fuse them. Its dedupe check never matches, and the default Cohere reranker returns its input unchanged when no key is set. Follow-up questions are searched exactly as typed.
Dense only. AnythingLLM returns the top-N chunks above a similarity threshold, with no metadata filters. Reranking works only on its LanceDB backend.
Composable frameworks. LlamaIndex offers QueryFusionRetriever (four LLM rewrites, then fusion) plus rerankers as postprocessors. Whether hybrid works depends on the vector backend. Haystack’s MultiRetriever merges named retrievers with RRF, and rankers and query expansion are separate components.
Graph or structure navigation. RAG-Anything uses LightRAG’s mix mode by default, which adds vector chunks to graph retrieval. PageIndex has no similarity search: an agent reads the section tree and fetches page ranges.
Pick: RAGFlow or R2R for tunable hybrid search over large corpora. Pick: Onyx for permission-filtered company search. Pick: PageIndex for structure-dependent questions over a few long documents.
Per-project answers
infiniflow/ragflow
answeredHybrid search: The core retrieval is implemented in service/nlp/retrieval.go. The RetrievalService.Search() method (line 671) builds parallel BM25 (matchText) and dense vector (matchDense) query expressions and optionally a fusionExpr that combines them. Defaults: KNN topK=1024, numCandidates=2048, similarityThreshold=0.2, vectorSimilarityWeight=0.3 (1.0 for vector-only), rankFeature with pagerank_fea=10.0, rerankCandidatesCount=64. When no embedding model is available, it falls back to keyword-only search. The fusion expression weight is controlled by VectorSimilarityWeight, and if empty results are returned the system retries with progressively relaxed thresholds (min_match=0.1, similarity=0.17) and optionally pure dense-only fallback.
Dense / sparse / fusion: The types/SearchRequest struct (engine/types/types.go:34-56) carries MatchExprs — a list of text match, dense vector match, and fusion expressions. The vector expression is built via GetVector() which calls the embedding model. For hybrid, all three are passed to the engine: BM25 text match + vector kNN + fusion (weighted sum).
Reranking: Reranking (service/nlp/reranker.go) supports model-based reranking (via a RerankModel provider, e.g., Jina, Cohere, BGE-reranker) and engine-dependent fallback — Infinity uses pre-normalized scores, Elasticsearch recomputes via RerankStandard() which combines token cosine similarity tkSim and vector cosine similarity vtSim weighted by tkWeight (0.7) and vtWeight (0.3).
Query rewriting: The chat pipeline (chat_pipeline.go phase 8, lines 188-189) applies multi-turn refinement (refine_multiturn), cross-language translation (cross_languages), metadata filtering (meta_data_filter), and keyword extraction via LLM (remaining service/generator.go:51-95 uses a configured ChatModel with a keyword_prompt template).
DeepResearcher (agentic retrieval): When reasoning=true and agentic mode is enabled (chat_pipeline.go:267-268), the system runs a recursive DeepResearcher (up to depth=3) that iterates between KB search, web search and optional Knowledge Graph retrieval, with sufficiency checking and multi-query generation at each layer.
Filters: The RetrievalRequest carries Filter (arbitrary key-value metadata conditions) and DocIDs for document-scoped search. The available_int=1 filter defaults to exclude unavailable chunks. Metadata pushdown filtering is supported via FilterDocIdsByMetaPushdown() on the engine interface.
DeepResearcher (service/deep_researcher.go) has no production caller. Reasoning levels 1-4 run the agentic graph in internal/rag/agentic-rag through retrievalbridge.NewHarnessRetriever, and level 5 runs the smart-reasoning eino agent (internal/agentic_rag). On non-Infinity engines the first search uses fixed fusion weights 0.001,1. vector_similarity_weight is applied afterwards (tkWeight = 1 - weight), so 0.7/0.3 is only the default.Mintplex-Labs/anything-llm
answeredDense search. Retrieval is purely dense (embedding-based). The performSimilaritySearch() method on every vector-db adapter takes the user's query text, embeds it via LLMConnector.embedTextInput(), then runs a cosine-similarity vector search against the workspace namespace. Parameters: similarityThreshold (default 0.25) and topN (default 4, configurable per workspace). Results below the threshold are discarded. Pinned documents are deduplicated from search results via sourceIdentifier() (a composite of title+published).
Reranking. Each vector-db provider has an optional reranking path, toggled per workspace by vectorSearchMode === "rerank". The LanceDB implementation (rerankedSimilarityResponse) first fetches up to 50 results (capped at max(10, min(50, ceil(totalEmbeddings * 0.1)))), then passes them to a NativeEmbeddingReranker built on Xenova/ms-marco-MiniLM-L-6-v2 (a cross-encoder from @xenova/transformers). The reranker scores query-document pairs in batches of 10, sorts globally, returns the top-K. Pinecone uses a simpler single-pass vector search (Pinecone server-side doesn't support reranking — only native reranker is available).
Query rewriting/decomposition. No query rewriting or decomposition is implemented. The raw user message is embedded as-is.
Filters. The only filter is filterIdentifiers — a list of pinned-document identifiers to exclude from results. This prevents the vector store from returning chunks that are already injected via pinned documents. No metadata-field filters (date range, author, source type) are passed to the vector DB at query time.
Context backfill. After vector search, fillSourceWindow() backfills from recent chat history when fewer than topN sources are found. This ensures follow-up questions retain relevant context even when the vector search returns 0 results — earlier sources from the same thread are reused.
run-llama/llama_index
answeredRetrieval is orchestrated by the RetrieverQueryEngine (llama_index/core/query_engine/retriever_query_engine.py:25-56), which calls a retriever, applies node postprocessors, then synthesizes a response.
Dense retrieval: VectorIndexRetriever (llama_index/core/indices/vector_store/retrievers/retriever.py:24-80) embeds the query using the configured embed model, builds a VectorStoreQuery with the embedding, similarity_top_k, filters, and mode, then delegates to the vector store's query() method. The vector store returns VectorStoreQueryResult with nodes, similarities, and IDs.
Sparse and hybrid: VectorStoreQueryMode.SPARSE skips the embedding step and sends the raw query string. VectorStoreQueryMode.HYBRID sends both the embedding and query string, with alpha controlling the dense/sparse blend. The KeywordTableIndex provides an alternative LLM-keyword-based sparse retrieval.
Query rewriting/decomposition: QueryFusionRetriever (llama_index/core/retrievers/fusion_retriever.py:33-70) uses an LLM to generate multiple query variants from the original, retrieves from each variant across a set of sub-retrievers, then fuses results via reciprocal rank fusion, relative score, or simple re-ranking. SubQuestionQueryEngine (llama_index/core/query_engine/sub_question_query_engine.py:37-60) breaks a complex query into sub-questions, dispatches each to a different query engine tool, then synthesizes the final answer.
Reranking: Postprocessors run after retrieval. LLMRerank (llama_index/core/postprocessor/llm_rerank.py:23-65) uses an LLM to select the top-N most relevant nodes from candidate batches. SentenceTransformerRerank (llama_index/core/postprocessor/sbert_rerank.py:12-56) uses a cross-encoder model. Other postprocessors include SimilarityPostprocessor, KeywordNodePostprocessor, and MetadataReplacementPostProcessor.
Filters: MetadataFilters (llama_index/core/vector_stores/types.py:142-200) supports AND/OR/NOT conditions with operators including ==, !=, >, <, IN, ANY, ALL, TEXT_MATCH.
The-Vibe-Company/quivr
answeredThe engine itself does no ranking — it delegates to a retrieval plugin via a multi-round protocol. Three modes (retrieval/retrieval.go:209-211): lexical (BM25-only), semantic (vector-only), hybrid (both, the default). The Request struct carries Query, CorpusIDs, Mode, Profile, Limit (max 50), SourceNamespaces (max 50), Metadata filters, and an optional pre-encoded Vector (retrieval/retrieval.go:73-98). Search proceeds in rounds (retrieval/plugin.go:189-256): (1) the plugin is called with a SearchRequest containing the current session state; (2) the plugin responds with either requests for candidates (primitive+space combinations) or a final ranking; (3) the engine serves candidates by querying Weaviate with authorization, generation routing, source namespace, and metadata filters applied (weaviate/projection.go:470-564); (4) the plugin receives the candidates and can request more or return the ranking. The core.retrieve plugin (plugins/core-retrieve/main.go) runs in at most two rounds: round 1 requests candidates (single BM25 for lexical, near_vector per served space for semantic, hybrid per space with configurable alpha 0.5 and fusion type for hybrid), then final ranking merges by descending score, deduplicates per segment, and tie-breaks by segment ID. Query encoding is done by the ingestion plugin that owns each vector space (TEI for legacy E5, plugin's EmbedQuery for plugin-owned spaces). Reranking: Not built in; the deep profile exists for future reranking. Query rewriting/decomposition: Not implemented. Filters: Source namespace filter via SourceNamespaceProjected check (retrieval/retrieval.go:290), metadata filter via MetadataProjected check (line:284) and corpus.ResolveFilters validation.
deep profile ranks like default), but the optional first-party jev.rerank plugin serves deep and reranks candidates through the external TypeSafe API.VectifyAI/PageIndex
answeredThere is no traditional retrieval. PageIndex does not perform dense/sparse/hybrid search, BM25, query rewriting, decomposition, or reranking. Instead, retrieval is an agentic LLM reasoning process over the tree index.
The agent-based retrieval mechanism. The chat model receives a system prompt that includes the tree structure and instructions for navigating it (agent_tools.py:1587-1627). It uses a set of MCP-style tools to explore the document: browse_documents (list documents), get_document (check status/metadata), get_document_structure (fetch the hierarchical tree outline), get_page_content (retrieve page text by page numbers), and remove_document (agent_tools.py:68-303). The get_document_structure tool returns the tree outline paginated across multiple parts when large. The agent decides which tool to call, what parameters to use, and what pages to read based on the question — effectively performing its own query routing.
Document targeting. For scoped retrieval, the system prepends a targeting block as the first user message: the document's metadata and a directive to work within it (local_chat.py:1695-1777). For folder-scoped searches (cloud only), a folder_targeting_block is added. The targeting block replaces what would be query filtering in a vector system.
Multi-document search. Multiple doc_id values can be passed to chat(). The targeting block lists all documents, and the agent can call get_document_structure and get_page_content on each. The conversation's scope is enforced at the tool layer via _allowed_ids (agent_tools.py:1238-1239).
Filters. Local mode supports folder_id/sort/query on browse_documents but returns an error — those are cloud-only (agent_tools.py:727-749). Local metadata is stored and returned but not searchable as a filter. The tree structure's start_index/end_index page ranges serve as an implicit filter: the agent can target specific page ranges via get_page_content("5-10").
Cloud differences. With a cloud API key, the client connects to the PageIndex MCP server (mcp_bridge.py) which serves the live tool set including folders. The cloud also provides a managed chat endpoint at /chat/completions/ (cloud_api.py:272-340) that selects its own model and can enable_citations. In own-model chat, retrieval runs the same agent loop but over cloud-hosted documents, proxied through the same tool interface.
onyx-dot-app/onyx
answeredSearch pipeline entry: search_pipeline (backend/onyx/context/search/pipeline.py, lines 260–344) builds IndexFilters from user-provided filters, persona document sets, ACLs, time ranges, and tenant IDs. It then calls search_chunks.
Search runner: search_chunks (backend/onyx/context/search/retrieval/search_runner.py, lines 89–161) runs parallel queries: the normal hybrid/keyword search against OpenSearch, plus federated retrieval functions from external connectors (e.g. Slack). Results are deduplicated and merged by score via combine_retrieval_results (lines 27–48).
Hybrid search: _embed_and_hybrid_search (lines 51–75) embeds the query via get_query_embedding and calls document_index.hybrid_retrieval() on the OpenSearchDocumentIndex. The hybrid search combines k-NN vector search (cosine similarity) with BM25 keyword scoring via OpenSearch's hybrid query with normalization pipelines (min_max or zscore). A hybrid_alpha parameter controls the blend (≤0.2 treats as keyword-only, ≥0.8 as semantic-only). The query type classification (QueryType.KEYWORD vs SEMANTIC) also affects scoring profiles.
Keyword-only search: When hybrid_alpha=0.0, the embedding step is skipped entirely and pure BM25 retrieval runs (_keyword_search, lines 78–86).
Query rewriting: Before retrieval, semantic_query_rephrase (secondary_llm_flows/query_expansion.py, lines 73–153) rewrites the user's query into a self-contained search query using chat history context. keyword_query_expansion (lines 156–234) generates up to 3 keyword-centric search queries from the same context.
Filters: _build_index_filters (pipeline.py, lines 40–139) constructs the IndexFilters dataclass with fields for: source_type, document_set, tags, access_control_list, cc_pair_access, tenant_id, time_range, created_at_range, updated_at_range, attached_document_ids, hierarchy_node_ids, and a forced_document_set (for search UI enforced scope). Filters are enforced as AND clauses in the OpenSearch query DSL.
Reranking: Reranking is invoked separately via the rerank_pipeline at the context/search level — the OpenSearch document index supports a search_type parameter for reranking pipelines.
Post-query censoring: Some connectors (Salesforce) apply field-level permissions via post_query_chunk_censoring (EE implementation, referenced at pipeline.py, lines 334–342).
hybrid_alpha does not set the vector/keyword blend; query_type is ignored and the normalization pipeline uses fixed weights (0.1 title vector, 0.45 content vector, 0.45 keyword), so only hybrid_alpha == 0.0 (pure BM25) changes behaviour. There is also no reranker: the search tool fuses its parallel queries with weighted reciprocal rank fusion and then has the LLM select sections (select_sections_for_expansion).deepset-ai/haystack
answeredDense/embedding retrieval – InMemoryEmbeddingRetriever (haystack/components/retrievers/in_memory/embedding_retriever.py:12-237) accepts a query_embedding (list of floats, pre-computed by a TextEmbedder), passes it to document_store.embedding_retrieval(), which computes cosine or dot-product similarity scores via np.dot (haystack/document_stores/in_memory/document_store.py:853-918). Scores can be scaled to [0,1] via expit for dot-product or (score+1)/2 for cosine. The TextEmbeddingRetriever component (haystack/components/retrievers/text_embedding_retriever.py:14-70) composes a TextEmbedder and an EmbeddingRetriever into a single run(query: str) component.
Sparse/BM25 retrieval – InMemoryBM25Retriever (haystack/components/retrievers/in_memory/bm25_retriever.py:13-197) accepts a raw query: str and calls document_store.bm25_retrieval(), which scores documents using the rank_bm25 library and returns top-k results.
Hybrid search – The MultiRetriever (haystack/components/retrievers/multi_retriever.py:19-120) runs multiple named retrievers concurrently via a thread pool, then merges results using concatenate + deduplication or reciprocal_rank_fusion (RRF). A top_k parameter applies consistent global ranking after RRF.
Reranking – The LLMRanker (haystack/components/rankers/llm_ranker.py:52-307) takes query + candidate documents, builds a prompt with instructions and JSON-schema response format, calls a ChatGenerator (default: gpt-4.1-mini with temperature=0.0 and structured output), and re-orders documents by the LLM's relevance assessment. It deduplicates input documents before ranking and falls back to input order on failure. The LostInTheMiddleRanker, MetaFieldRanker, and MetaFieldGroupingRanker are also available.
Query rewriting – QueryExpander (haystack/components/query/query_expander.py:56-60) uses an LLM to generate semantically similar query variants, which can then be fed to MultiQueryTextRetriever (haystack/components/retrievers/multi_query_text_retriever.py:16-82) for parallel retrieval across all query variants with deduplicated results.
Filters – All retrievers accept filters dictionaries that support comparison operators (==, !=, <, <=, >, >=, in, not in) and logic operators (AND, OR, NOT). The FilterPolicy (REPLACE/MERGE) controls how init-time and run-time filters combine.
Cinnamon/kotaemon
answeredRetrieval is performed by VectorRetrieval.run() (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 134–303), which supports three modes via retrieval_mode:
- vector: embeds the query, queries the vector store, fetches full text from the doc store.
- text: performs BM25 full-text search on the doc store (LanceDB FTS or Elasticsearch BM25).
- hybrid: runs both in parallel threads, de-duplicates results (giving priority to vector hits with scores), and merges them (
libs/kotaemon/kotaemon/indices/vectorindex.py, lines 186–237). This is the default mode.
Results are re-ranked by configurable rerankers: CohereReranking (via Cohere's rerank API), LLMReranking (LLM-based YES/NO relevance filter), LLMTrulensScoring (0–10 relevance grader that normalizes to 0–1 and stores as llm_trulens_score), and VoyageAIReranking (libs/kotaemon/kotaemon/indices/rankings/). The reranker list is applied in order; if a reranker is LLMReranking, results are first truncated to top_k (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 240–245).
Filters are supported via LlamaIndex MetadataFilters with file_id in selected doc IDs (libs/ktem/ktem/index/file/pipelines.py, lines 158–167). MMR (Maximum Marginal Relevance) can be enabled for diversity (libs/ktem/ktem/index/file/pipelines.py, lines 169–172).
Query rewriting is optional: AddQueryContextPipeline uses the LLM to generate a search query from conversation history (libs/ktem/ktem/reasoning/simple.py, lines 42–83). Question decomposition is available via FullDecomposeQAPipeline which splits complex questions into sub-questions and retrieves per sub-question (libs/ktem/ktem/reasoning/simple.py, lines 487–609). Agentic retrieval is also supported: ReactAgentPipeline integrates a ReAct agent with tools including doc search, Wikipedia, Google, and LLM (libs/ktem/ktem/reasoning/react.py, lines 181–306). Web search retrievers (Tavily, Jina) are available (libs/kotaemon/kotaemon/indices/retrievers/).
HKUDS/RAG-Anything
answeredDense / sparse / hybrid search. All retrieval is performed by LightRAG through its aquery() method. The user selects the mode: local (graph-traversal only), global (entity-level graph search), hybrid (vector + graph), naive (vector-only chunk retrieval), mix (combines local and global graph results), or bypass (response generation without retrieval). These are passed as QueryParam(mode=...) (query.py:188-197). RAG-Anything's aquery() and aquery_vlm_enhanced() methods are thin wrappers around LightRAG's aquery().
Reranking. LightRAG exposes rerank_model_func in its configuration, which can be passed through lightrag_kwargs (raganything.py:95). RAG-Anything does not add its own reranking layer.
Query rewriting or decomposition. RAG-Anything does not implement query rewriting or multi-step decomposition. The query is sent as-is to LightRAG, which performs the configured retrieval mode.
Filters. No explicit filter mechanism exists in RAG-Anything for filtering by metadata at retrieval time. LightRAG's own graph traversal effectively acts as a structural filter (local/global modes constrain the search space). For multimodal queries, aquery_with_multimodal() (query.py:221-382) pre-processes multimodal content (images, tables, equations) using LLM/VLM captioning to produce an enhanced query string before calling aquery(). For VLM-enhanced queries (aquery_vlm_enhanced, query.py:384-463), the system retrieves context from LightRAG, scans for image paths, validates them against safe directories, encodes valid images to base64, and builds a multimodal message payload for the vision model.
Source attribution. The QueryParam object supports passing file_paths for citation; LightRAG tracks which chunks belong to which document via the doc_status system, and chunks maintain their full_doc_id reference (processor.py:303).
hybrid combines the local (entity) and global (relation) graph retrievals, and mix, the default in RAG-Anything's aquery, combines the graph retrieval with vector chunk retrieval.SciPhi-AI/R2R
answeredDense / sparse / hybrid search. The RetrievalService (py/core/main/services/retrieval_service.py:248-280) dispatches to three strategies: basic (vanilla semantic + optional graph), HyDE (LLM-generated hypothetical documents), and RAG Fusion (multi-query with RRF). Internally, _vector_search_logic (py/core/main/services/retrieval_service.py:651-729) routes to semantic_search, full_text_search, or hybrid_search on PostgresChunksHandler depending on SearchSettings flags. Semantic search (py/core/providers/database/chunks.py:327-482) computes cosine distance via vec <=> $1::vector(N) with a two-stage binary re-ranking path for INT1 quantization. Full-text search uses ts_rank(fts, websearch_to_tsquery('english', $1)). Hybrid search (py/core/providers/database/chunks.py:538-640) fuses both via weighted RRF.
Search strategies. HyDE (py/core/main/services/retrieval_service.py:554-614) generates N hypothetical documents via LLM, embeds each, runs parallel searches, and re-ranks results against the original query. RAG Fusion (py/core/main/services/retrieval_service.py:326-403) generates num_sub_queries alternative phrasings via LLM, runs each independently, then fuses via RRF and optionally re-ranks.
Reranking. The embedding provider's arerank() method is called after initial search. OpenAIEmbeddingProvider passes through (py/core/providers/embeddings/openai.py:227-243). LiteLLMEmbeddingProvider optionally uses a HuggingFace TEI reranking endpoint (py/core/providers/embeddings/litellm.py:47-59).
Graph search. _graph_search_logic (py/core/main/services/retrieval_service.py:732-794) searches entities, relationships, and communities via embedding similarity on their description vectors. Graph entities have their own vector tables with pgvector.
Filters. Rich filter system (py/core/providers/database/filters.py) supporting $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $like, $ilike, $overlap, $contains, $and, $or. Applied to the metadata JSONB column and top-level columns like document_id, owner_id, collection_ids.
← How are embeddings and indexes built and stored? · How are answers generated and grounded? →