# How are embeddings and indexes built and stored?

> RAG engines — a good answer covers: Embedding models; vector stores supported; hybrid / keyword (BM25) indexes; metadata.

Canonical page: https://llms-technical-reviews.com/rag/q/indexing/

## Verdict

[R2R](/p/r2r/) has the simplest strong index: vectors and keyword search in one Postgres. [RAGFlow](/p/ragflow/) has the most complete hybrid index at scale. [LlamaIndex](/p/llama_index/) gives the widest choice of backend.

**Hybrid index in one search engine.** RAGFlow stores a BM25 token field (`content_ltks`), a dense vector and position metadata on every chunk. It runs on one of five engines, with Elasticsearch as the default. [Onyx](/p/onyx/) supports OpenSearch only. It embeds each document title as its own vector, stores ACLs on every chunk, and builds a secondary index so the embedding model can be switched without downtime. R2R uses pgvector plus a generated English `tsvector`. HNSW or IVFFlat indexes have to be created explicitly through the API. [Quivr](/p/quivr/) treats Weaviate as a rebuildable copy of canonical Postgres and S3 data. Each embedding model owns its own "vector space", which can be evaluated beside the live one and then promoted.

**Pluggable dense stores.** [AnythingLLM](/p/anything-llm/) has ten vector adapters (LanceDB by default) and a local MiniLM embedder. It has no keyword index, and each adapter re-implements chunking and embedding. [Kotaemon](/p/kotaemon/) pairs a local Chroma vector store with a LanceDB document store that provides full-text search. LlamaIndex has about 80 vector-store integrations. Its built-in `SimpleVectorStore` scores every vector in Python and has no hybrid mode, and BM25 is a separate package. [Haystack](/p/haystack/)'s `DocumentStore` protocol requires only four methods. Its in-memory store does brute-force vector search and implements BM25L/Okapi/Plus itself.

**Graph plus vectors.** [RAG-Anything](/p/rag-anything/) writes chunks, entities and relations into LightRAG (<1.5) stores. It builds no BM25 or keyword index, so there is no lexical retrieval.

**No vectors.** [PageIndex](/p/pageindex/) stores one JSON tree per document. Each node has a page range and an LLM summary of up to 150 words.

Pick: R2R if you want one Postgres for vectors, keywords and users.
Pick: RAGFlow or Onyx for large corpora that need lexical plus semantic recall.
Pick: LlamaIndex or Haystack to keep the vector database you already run.

## Per-project answers

### infiniflow/ragflow (answered)

**Embedding models:** RAGFlow supports a broad range of embedding providers through a plugin-style model architecture (`internal/entity/models/`). The factory (`factory.go`) resolves model requests to provider-specific drivers — there are 80+ model provider files for OpenAI, Anthropic, Google/Gemini, DeepSeek, Ollama, Cohere, Jina, Voyage, vLLM, Xinference, HuggingFace, and many Chinese providers (e.g. Baichuan, ZhipuAI, Qwen, Hunyuan). Each supports chat, text-embedding, reranking, or TTS modalities. Models are registered per-tenant in the `tenant_model_instance` table and resolved at runtime by `ModelSolver` (`service/model_solver.go`).

**Vector stores:** Four backends are supported (`engine/engine.go:33-38`): Elasticsearch, Infinity, OceanBase, and SereneDB, plus OpenSearch (via the OS config block in `service_conf.yaml.template`). Each implements the `DocEngine` interface (`engine/engine.go:41-93`) with chunk CRUD, search, metadata operations, and SQL execution. The Go code has Elasticsearch (`engine/elasticsearch/chunk.go`) and Infinity (`engine/infinity/`) implementations; OceanBase and SereneDB share a PostgreSQL-compatible layer. Chunks are created via `CreateChunkStore()` which sets up an ES index with `number_of_shards=1`, `number_of_replicas=0` and appropriate mappings including `tag_feas` as `rank_features` for tag-based ranking. Bulk writes are batched at 500 chunks or 20MB.

**Hybrid / keyword (BM25) indexes:** Chunks carry both `content_ltks` (tokenized text for BM25 matching) and embedding vectors. BM25 is implemented via the QueryBuilder (`service/nlp/query_builder.go`) that produces `matchText` expressions for the engine. OpenSearch has dedicated hybrid search pipeline configuration (`service_conf.yaml.template:59`).

**Metadata:** The `DocEngine` interface defines `CreateMetadataStore`, `InsertMetadata`, `UpdateMetadata`, `SearchMetadata` operations (`engine/engine.go:54-61`). Metadata fields per chunk include `docnm_kwd` (document name), `kb_id`, `img_id`, `title_tks`, `position_int`, `page_num_int`, `chunk_order_int`, `content_with_weight`, `doc_type_kwd` (text/image/table), `mom_id` (parent chunk for child fragments), and user-defined metadata values set via filter conditions.

> **Editor's note.** Correction: `DocEngine` has five engine types: Elasticsearch, Infinity, OceanBase, SeekDB and SereneDB. OceanBase and SeekDB share one implementation (`IsOceanBaseFamily`), and SereneDB is a separate package. There is no OpenSearch engine under `internal/engine/`, only a Compose profile and a mapping file.

Citations: [internal/engine/engine.go:30-93](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/engine.go#L30-L93) · [internal/engine/elasticsearch/chunk.go:54-151](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/engine/elasticsearch/chunk.go#L54-L151) · [internal/service/nlp/retrieval.go:697-710](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/nlp/retrieval.go#L697-L710)

### Mintplex-Labs/anything-llm (answered)

**Embedding models.** AnythingLLM supports 14 embedding engines, selected via `EMBEDDING_ENGINE` env: `native` (local ONNX via @xenova/transformers), `openai`, `azure`, `localai`, `ollama`, `lmstudio`, `cohere`, `voyageai`, `litellm`, `mistral`, `generic-openai`, `gemini`, `openrouter`, `lemonade`. The native embedder defaults to `Xenova/all-MiniLM-L6-v2` (23MB, 1000-char chunks, 25 concurrent chunks) with two alternatives: `Xenova/nomic-embed-text-v1` (139MB, 16000-char chunks, 5 concurrent, uses `search_document:`/`search_query:` prefixes) and `MintplexLabs/multilingual-e5-small` (487MB, 100+ languages, passages prefixed with `passage:`/`query:`). For local on-device embedding, models are downloaded on first use from HuggingFace Hub with a fallback CDN at `cdn.anythingllm.com`.

**Vector stores.** Ten backends are supported: LanceDB (default, embedded, file-based), Pinecone, Chroma, ChromaCloud, Weaviate, QDrant, Milvus, Zilliz, AstraDB, and PGVector. Each implements the `VectorDatabase` base class with `addDocumentToNamespace`, `performSimilaritySearch`, `deleteDocumentFromNamespace`. Workspaces map 1:1 to vector-db namespaces (using `workspace.slug`). Chunking and embedding happen inside the vector-store adapter's `addDocumentToNamespace()` — every provider reads `text_splitter_chunk_size`/`text_splitter_chunk_overlap` from SystemSettings, splits via `TextSplitter`, embeds via the selected `EmbedderEngine`, and upserts.

**Hybrid / BM25.** There is no built-in BM25 or hybrid search. Retrieval is dense-only via cosine similarity on embedding vectors. Full-text search for workspace/thread *names* uses fast-levenshtein matching, but document-level keyword search is not implemented.

**Metadata.** Document metadata stored per chunk includes: `text` (the chunk content), `title`, `description`, `docAuthor`, `docSource`, `published`, `chunkSource`, `token_count_estimate` and scoring fields. Pinned documents bypass the vector store entirely and are injected directly as context.


Citations: [server/utils/EmbeddingEngines/native/index.js:1-310](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/EmbeddingEngines/native/index.js#L1-L310) · [server/utils/EmbeddingEngines/native/constants.js:1-60](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/EmbeddingEngines/native/constants.js#L1-L60) · [server/utils/helpers/index.js:308-360](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L308-L360) · [server/utils/vectorDbProviders/lance/index.js:312-415](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/lance/index.js#L312-L415) · [server/utils/vectorDbProviders/pinecone/index.js:156-210](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/vectorDbProviders/pinecone/index.js#L156-L210)

### run-llama/llama_index (answered)

**Embedding models** are accessed through the `BaseEmbedding` abstract class (`llama_index/core/base/embeddings/base.py:72-80`). The `resolve_embed_model()` function (`llama_index/core/embeddings/utils.py:30-50`) resolves a string name (e.g., `"default"` → OpenAI's `text-embedding-ada-002`) or a LangChain embedding into a `BaseEmbedding` instance. The `VectorStoreIndex._get_node_with_embedding()` embeds nodes in batches (`llama_index/core/indices/vector_store/base.py:126-148`).

**Vector stores** implement `BasePydanticVectorStore` (`llama_index/core/vector_stores/types.py`). The built-in `SimpleVectorStore` (`llama_index/core/vector_stores/simple.py:64-80`) keeps an in-memory dict of `node_id → embedding` and persists to JSON via fsspec. Dozens of production vector store integrations exist as separate packages (Chroma, Pinecone, Qdrant, Weaviate, FAISS, Elasticsearch, Milvus, etc.). The `StorageContext` (`llama_index/core/storage/storage_context.py:52-72`) holds the vector store, document store, index store, and graph store.

**Hybrid/keyword indexing**: The `VectorStoreQueryMode` enum (`llama_index/core/vector_stores/types.py:45-60`) defines `HYBRID`, `SPARSE`, and `SEMANTIC_HYBRID` modes. The `VectorIndexRetriever` passes an `alpha` parameter to the vector store to weight dense vs sparse scores (`llama_index/core/indices/vector_store/retrievers/retriever.py:42-73`). Support for hybrid search depends on the underlying vector store. Separately, `BaseKeywordTableIndex` (`llama_index/core/indices/keyword_table/base.py:43-97`) uses an LLM to extract keywords from each chunk and builds an inverted mapping of keyword → node IDs — a form of keyword-based sparse retrieval.

**Metadata** is stored in each node's `metadata` dict and persisted to the vector store. `MetadataFilters` (`llama_index/core/vector_stores/types.py:142-200`) with operators (`==`, `>`, `<`, `IN`, `ANY`, `TEXT_MATCH`, etc.) support filtering at query time.


Citations: [llama-index-core/llama_index/core/base/embeddings/base.py:72-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/embeddings/base.py#L72-L80) · [llama-index-core/llama_index/core/embeddings/utils.py:30-50](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/embeddings/utils.py#L30-L50) · [llama-index-core/llama_index/core/indices/vector_store/base.py:126-148](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/vector_store/base.py#L126-L148) · [llama-index-core/llama_index/core/vector_stores/types.py:45-60](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/types.py#L45-L60) · [llama-index-core/llama_index/core/vector_stores/simple.py:64-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/simple.py#L64-L80) · [llama-index-core/llama_index/core/indices/keyword_table/base.py:43-97](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/keyword_table/base.py#L43-L97)

### The-Vibe-Company/quivr (answered)

**Embedding model**: `intfloat/multilingual-e5-small` (384 dimensions, cosine distance), served through a TEI (Text Embeddings Inference) container (deploy/compose/compose.yaml:65-79). The TEI encoder (adapters/tei/encoder.go:26-65) validates 384-dim unit-norm vectors, prefixes documents with `passage: ` and queries with `query: ` (plugins/core-ingest/main.go:131). The engine's legacy E5 space is an older built-in encoder for pre-THE-777 Corpora (app/run.go:55). **Vector store**: Only **Weaviate** is supported (adapters/weaviate/projection.go). Vectors are stored as named vectors per generation: the legacy `semantic_text_v1` for pre-named-space generations, or `s_<hash>` for plugin-owned spaces (projection.go:54-59). Distance metrics include cosine (default), dot, and L2-squared (projection.go:62-70). The index uses HNSW (projection.go:73). **Hybrid/BM25**: Weaviate provides built-in BM25 alongside vector search. The engine constructs GraphQL queries: `bm25:{query}` for lexical, `nearVector:{vector}` for semantic, and `hybrid:{query, vector, alpha, fusionType}` for hybrid (projection.go:552-558). Keyword fields default to `title^2` and `body`, with an optional `lexicalText` property from plugins. **Metadata**: Projected alongside vectors into Weaviate as individual properties (adapters/weaviate/metadata.go). The `Generation.MetadataProjected` flag gates filter availability (content/searchable.go:58-68); filters apply as Weaviate `where` clauses (projection.go:480-502). Embeddings are stored as artifacts in SeaweedFS (S3) with SHA-256 content hashes (content/embeddings.go:43-61), carrying full provenance (version, segment, space, producer info).

> **Editor's note.** Correction: TEI with multilingual-e5-small is the default via core.ingest, but the optional hosted.embed plugin embeds through OpenAI-compatible or Cohere v2 APIs in its own vector space, which can be evaluated and promoted.

Citations: [internal/adapters/tei/encoder.go:26-80](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/adapters/tei/encoder.go#L26-L80) · [internal/adapters/weaviate/projection.go:46-75](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/adapters/weaviate/projection.go#L46-L75) · [internal/adapters/weaviate/projection.go:470-564](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/adapters/weaviate/projection.go#L470-L564) · [internal/content/embeddings.go:43-61](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/content/embeddings.go#L43-L61) · [internal/content/searchable.go:1-68](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/content/searchable.go#L1-L68)

### VectifyAI/PageIndex (answered)

**No embeddings, no vector stores.** PageIndex does not compute vector embeddings or use vector databases. The index is a **hierarchical tree of sections** with page ranges stored as JSON. This "tree index" replaces what a vector index would do in traditional RAG.

**Tree structure.** Each node has `{title, node_id, start_index, end_index, (summary,) (text,) nodes}` (`local_api.py:396-414`, `types.py`). The tree is built bottom-up: leaves from layout headings, parent nodes spanning the range of their children. Intro nodes cover pages a parent opens with before its first child (`tree_optimize.py`). Node summaries are LLM-generated (the `summary_model`) using a bottom-up scheduler — leaves from their own page text, parents composed from child summaries plus residual pages (`utils.py:1140-1157`). Small nodes (under ~300 tokens) keep raw text as their summary.

**Standard mode indexing.** The classic pipeline (`page_index_classic.py:1200-1283`): (1) detect TOC pages via LLM, (2) extract the TOC structure with page indices, (3) verify correctness by spot-checking titles against page text (`verify_toc`, line 1066-1120), (4) fix incorrect items with retries (`fix_incorrect_toc_with_retries`, line 1044-1060), (5) post-process into a nested tree, (6) recursively split large nodes (>10 pages and >20k tokens; `process_large_node_recursively`, line 1168-1198), (7) add node text and summaries, (8) generate an optional one-sentence doc description (`generate_doc_description`, `utils.py:1183-1198`). Configurable via `config.yaml`: defaults `toc_check_page_num=20`, `max_page_num_each_node=10`, `max_token_num_each_node=20000`.

**Flash mode indexing.** `flash/api.py:180-248`: (1) LLM-free layout-based TOC extraction (`extract_toc`), (2) optional embedded bookmark integration, (3) deterministic `merge` pass (collapses nodes where navigating to children is more expensive than scanning the parent; `tree_optimize.py:18-24`), (4) optional LLM `expand` pass (proposes subsections where a node's span exceeds cost of routing; `tree_optimize.py:12-16`), (5) bottom-up summaries. The entire Flash pipeline is LLM-free when `summary=False, optimize=False`.

**Local storage.** The local store (`local_store.py`) writes JSON files to `~/.pageindex/`: each document as a directory containing `doc.json` (metadata), `tree.json` (the tree structure), `pages.json` (page text). A `manifest.json` indexes all documents. Atomic writes with `_write_json_atomic` use temp-file + `os.replace`. No external database is required.

**Cloud indexing.** Cloud mode delegates indexing to PageIndex's managed pipeline at `api.pageindex.ai` (`cloud_api.py:38-86`). The client uploads the PDF and receives a `doc_id`; the cloud handles parsing, OCR, image understanding, tree construction, and stores the result. Metadata, folders, and block-level citations are cloud-only features (`README.md:199-208`).

> **Editor's note.** Correction: the local store defaults to `./.pageindex` relative to the working directory (`storage_path`, `client.py`), not `~/.pageindex`.

Citations: [pageindex/local_api.py:396-414](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_api.py#L396-L414) · [pageindex/utils.py:1140-1157](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/utils.py#L1140-L1157) · [pageindex/page_index_classic.py:1200-1283](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1200-L1283) · [pageindex/tree_optimize.py:1-54](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/tree_optimize.py#L1-L54)

### onyx-dot-app/onyx (answered)

**Embedding models:** Embedding is done by the `DefaultIndexingEmbedder` (`backend/onyx/indexing/embedder.py`, lines 89–253) which wraps an `EmbeddingModel` (`search_nlp_models.py`). Supported providers include OpenAI (`text-embedding-3-*`), Cohere, VoyageAI, VertexAI, and custom OpenAI-compatible APIs. The embedding server host/port is configurable (`INDEXING_MODEL_SERVER_HOST/PORT`). Each chunk receives a **full embedding**, optional **mini-chunk embeddings** (for multipass scoring), and a **title embedding** (lines 156–183). Embeddings are normalized or dimension-reduced per SearchSettings.

**Vector store:** The sole document index backend is **OpenSearch** (via the `OpenSearchDocumentIndex` class at `backend/onyx/document_index/opensearch/opensearch_document_index.py`, line 1+). It implements the `DocumentIndex` interface (`backend/onyx/document_index/interfaces.py`, lines 474–497) which requires hybrid (vector + keyword) search, metadata updates, deletion, and random retrieval.

**Hybrid index structure:** Each chunk is stored as an OpenSearch document with a dense vector field (`content_vector`), lexical text fields (`content` for BM25), and structured metadata fields: `access_control_list`, `source_type`, `document_sets`, `tenant_id`, `cc_pair_ids`, `created_at`, `last_updated`, `hidden`, `boost`, `personas`, `user_projects`, and title vectors. The schema is defined in `opensearch/schema.py`.

**Indexing pipeline:** The end-to-end flow lives in `run_indexing_pipeline` (`indexing_pipeline.py:1641–1724`). Phases: (1) upsert documents to PostgreSQL, (2) filter by change-gate (timestamp or content-hash, `get_docs_to_update`, lines 365–447), (3) process image sections, (4) chunk via `Chunker.chunk()`, (5) optionally add contextual-RAG LLM enrichment, (6) embed via `IndexingEmbedder`, (7) write to OpenSearch via `write_chunks_to_vector_db_with_backoff` (`vector_db_insertion.py`, line 19+). The secondary (FUTURE) index is supported for zero-downtime model migrations.

**Storage:** Chunks are persisted to `ChunkBatchStore` (temporary on-disk store) between embedding and vector-db write to decouple memory from OpenSearch latency. The document index stores chunk counts for efficient tail-truncation (`IndexingMetadata.ChunkCounts`).


Citations: [backend/onyx/indexing/embedder.py:89-253](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/embedder.py#L89-L253) · [backend/onyx/indexing/indexing_pipeline.py:1641-1724](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/indexing_pipeline.py#L1641-L1724) · [backend/onyx/indexing/indexing_pipeline.py:365-447](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/indexing_pipeline.py#L365-L447) · [backend/onyx/indexing/vector_db_insertion.py:19-107](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/vector_db_insertion.py#L19-L107)

### deepset-ai/haystack (answered)

**Embedding models** – Embedders are Haystack components that call external APIs or local models. The `OpenAITextEmbedder` (`haystack/components/embedders/openai_text_embedder.py:16-254`) and `OpenAIDocumentEmbedder` embed text and documents using OpenAI's `embeddings.create` endpoint (default: `text-embedding-3-small`). They accept `dimensions`, `prefix`/`suffix`, and full OpenAI client configuration. API clients are lazily initialized in `warm_up()` (line 116), not `__init__`, keeping components serializable without credentials. The embedders directory also includes `AzureTextEmbedder`/`AzureDocumentEmbedder` variants and mock embedders for testing.

**Document stores** – The only built-in store is `InMemoryDocumentStore` (`haystack/document_stores/in_memory/document_store.py:98`), which stores documents in process-global dictionaries. All other stores — Elasticsearch, OpenSearch, Pinecone, Qdrant, Weaviate, Chroma, Astra DB, Pgvector — are provided as separate integration packages (e.g., `elasticsearch-haystack`, `pinecone-haystack`). The `DocumentStore` protocol (`haystack/document_stores/types/protocol.py:11-136`) defines the interface all stores implement: `write_documents`, `filter_documents`, `delete_documents`, `count_documents`. The protocol also defines methods for `embedding_retrieval` and `bm25_retrieval` that stores implement with their backend's native similarity search.

**Hybrid/keyword (BM25) indexes** – The in-memory store maintains a BM25 index incrementally during `write_documents` (lines 543-553). Each document is tokenized via `_tokenize_bm25` (custom CJK-aware regex, line 58), term frequencies are stored in `BM25DocumentStats` dataclasses, and an IDF vocabulary counter and average document length are updated incrementally. The store supports multiple BM25 variants via the rank_bm25 library (configurable via `bm25_algorithm`). Plugging the BM25 retriever and embedding retriever into the `MultiRetriever` component (`haystack/components/retrievers/multi_retriever.py:19-120`) enables hybrid search with `reciprocal_rank_fusion` joining mode.

**Metadata** – All document stores support structured metadata filtering via a nested comparison/logic filter syntax. The `DuplicatePolicy` enum (`haystack/document_stores/types/policy.py`) controls write behavior: `NONE`/`FAIL`, `SKIP`, `OVERWRITE`.

> **Editor's note.** Correction: the DocumentStore protocol only requires count/filter/write/delete (plus to_dict/from_dict); embedding_retrieval and bm25_retrieval are store-specific methods, not protocol members. The in-memory store implements BM25L, BM25Okapi and BM25Plus itself rather than calling the rank_bm25 library.

Citations: [haystack/document_stores/in_memory/document_store.py:98-553](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/document_stores/in_memory/document_store.py#L98-L553) · [haystack/components/retrievers/multi_retriever.py:19-120](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/retrievers/multi_retriever.py#L19-L120)

### Cinnamon/kotaemon (answered)

Embeddings are produced by pluggable models: OpenAI, Azure OpenAI, Cohere, VoyageAI, Google, Mistral, HuggingFace (via LangChain wrappers), TEI (Text-Embedding-Inference) endpoints, and local FastEmbed models (e.g. `BAAI/bge-small-en-v1.5`) (`libs/kotaemon/kotaemon/embeddings/` directory). The `VectorIndexing` class calls `self.embedding(docs)` to generate embeddings, then adds them to both a vector store and a doc store (`libs/kotaemon/kotaemon/indices/vectorindex.py`, lines 84–93). Vector stores available: Chroma (persistent client), LanceDB, Qdrant, Milvus (with lazy init by dimension), and SimpleFileVectorStore (`libs/kotaemon/kotaemon/storages/vectorstores/`). Default in `flowsettings.py` is Chroma at `ktem_app_data/user_data/vectorstore` (line 100–103). Document stores support full-text search (BM25) via Elasticsearch (with custom BM25 similarity, line 38–46 of `elasticsearch.py`) or LanceDB (with FTS index using `en_stem` tokenizer, line 56–60 of `lancedb.py`). The default docstore in `flowsettings.py` is `LanceDBDocumentStore` (line 95–97). Hybrid retrieval is supported by pairing a vector store with a BM25-capable doc store. Metadata stored includes `file_id`, `page_label`, `file_name`, `type`, `thumbnail_doc_id`, and `window` for sentence-window splitters. The indexing also writes chunk content to a markdown cache directory and records source-to-chunk relationships in an SQL `Index` table (`libs/ktem/ktem/index/file/pipelines.py`, lines 426–464).


Citations: [libs/kotaemon/kotaemon/indices/vectorindex.py:84-93](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/vectorindex.py#L84-L93) · [libs/kotaemon/kotaemon/storages/vectorstores/chroma.py:8-58](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/storages/vectorstores/chroma.py#L8-L58) · [libs/kotaemon/kotaemon/storages/docstores/elasticsearch.py:10-61](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/storages/docstores/elasticsearch.py#L10-L61) · [libs/kotaemon/kotaemon/storages/docstores/lancedb.py:11-61](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/storages/docstores/lancedb.py#L11-L61)

### HKUDS/RAG-Anything (answered)

RAG-Anything does not implement its own indexing — it delegates entirely to **LightRAG**, which is a dependency (`lightrag-hku<1.5` in `requirements.txt:6`). The `RAGAnything` class wraps a `LightRAG` instance (`raganything.py:69`) and passes through `lightrag_kwargs` for storage backend selection.

**Embedding models.** The user provides an `embedding_func` callable (e.g. `openai_embed` from `lightrag.llm.openai`) when constructing `RAGAnything`. This is passed directly to `LightRAG(...)` (`raganything.py:425`). No embedding model is bundled; any embedding function compatible with LightRAG's `EmbeddingFunc` signature works.

**Vector stores and graph storage.** LightRAG supports pluggable backends for `kv_storage`, `vector_storage`, `graph_storage`, and `doc_status_storage`, all exposed through `lightrag_kwargs` (`raganything.py:86-96`). These default to LightRAG's built-in options (typically file-based JSON/JSONL for development, with production backends like MongoDB, Neo4j, or Milvus available by configuring `lightrag_kwargs`). The RAG-Anything code additionally creates two KV namespaces: `parse_cache` for caching parsed content lists and `multimodal_status` for multimodal completion tracking (`raganything.py:447-465`).

**Hybrid / keyword (BM25) indexes.** This is handled entirely by LightRAG's query modes. RAG-Anything surfaces them as the `mode` parameter (`local`, `global`, `hybrid`, `naive`, `mix`, `bypass`) in `aquery()` (`query.py:129-131`). LightRAG internally maintains BM25 indexes alongside vector and graph indexes for hybrid search.

**Metadata.** Each chunk stores `page_idx` and `page_idx_end` via the page-provenance system in `processor.py:251-385`. File-level metadata (path, content hash, parser config) is tracked in `doc_status` records (`processor.py:128-159`). The `use_full_path` config option (`config.py:126-128`) controls whether file references use the absolute path or just the basename.

> **Editor's note.** Correction: neither RAG-Anything nor the pinned LightRAG (<1.5) builds a BM25 or keyword index; retrieval runs over LightRAG's vector stores and knowledge graph only. The requirements pin also notes that RAG-Anything was merged into LightRAG 1.5+, and this package stays on the 1.4.x API.

Citations: [raganything/raganything.py:86-96](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/raganything.py#L86-L96) · [raganything/raganything.py:420-465](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/raganything.py#L420-L465) · [raganything/query.py:129-131](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L129-L131) · [raganything/processor.py:251-295](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/processor.py#L251-L295) · [raganything/config.py:126-128](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/config.py#L126-L128)

### SciPhi-AI/R2R (answered)

**Embedding models.** Three embedding providers: `OpenAIEmbeddingProvider` (`py/core/providers/embeddings/openai.py:21-60`) supporting text-embedding-ada-002, text-embedding-3-small, and text-embedding-3-large with configurable dimensions; `LiteLLMEmbeddingProvider` (`py/core/providers/embeddings/litellm.py:25-65`) routing through LiteLLM for any provider (Amazon, HuggingFace, etc.); and `OllamaEmbeddingProvider` for local models. EmbeddingConfig (`py/core/base/providers/embedding.py:21-44`) configures provider, base_model, base_dimension, batch_size, concurrent_request_limit, and vector quantization settings. Retry with exponential backoff is built in.

**Vector store.** All vectors are stored in PostgreSQL using pgvector. `PostgresChunksHandler` (`py/core/providers/database/chunks.py:75-182`) creates tables with a `vec` column (configurable dimension), optional `vec_binary bit(N)` column for INT1 quantization, `text TEXT`, and `metadata JSONB`. A full-text search column (`fts tsvector`) is auto-generated as `to_tsvector('english', text)`. Index methods supported: IVFFlat (`IndexArgsIVFFlat` with `n_lists`) and HNSW (`IndexArgsHNSW` with `m` and `ef_construction`) via `py/shared/abstractions/vector.py:27-107`. Distance measures: cosine, L2, max-inner-product, L1, hamming, jaccard. Vector quantization types: FP32, FP16, INT1 (binary), SPARSE.

**Hybrid / keyword (BM25).** Full-text search uses PostgreSQL's `websearch_to_tsquery('english', ...)` with `ts_rank` ordering (`py/core/providers/database/chunks.py:484-536`). Hybrid search (`py/core/providers/database/chunks.py:538-640`) combines semantic and full-text via weighted Reciprocal Rank Fusion with configurable `semantic_weight`, `full_text_weight`, and `rrf_k`.

**Metadata.** Stored as JSONB alongside each vector entry. All document metadata (title, version, chunk_order, page_number, parser_generated) is carried through to the metadata column. Filtered at query time via the filter system in `py/core/providers/database/filters.py`.


Citations: [py/core/providers/embeddings/openai.py:21-60](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/embeddings/openai.py#L21-L60) · [py/core/providers/embeddings/litellm.py:25-65](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/embeddings/litellm.py#L25-L65) · [py/core/base/providers/embedding.py:21-44](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/base/providers/embedding.py#L21-L44) · [py/core/providers/database/chunks.py:75-182](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/database/chunks.py#L75-L182) · [py/core/providers/database/chunks.py:484-640](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/database/chunks.py#L484-L640) · [py/shared/abstractions/vector.py:27-107](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/shared/abstractions/vector.py#L27-L107)
