How are embeddings and indexes built and stored?
Embedding models; vector stores supported; hybrid / keyword (BM25) indexes; metadata.
Verdict
R2R has the simplest strong index: vectors and keyword search in one Postgres. RAGFlow has the most complete hybrid index at scale. LlamaIndex gives the widest choice of backend.
Hybrid index in one search engine. RAGFlow stores a BM25 token field (content_ltks), a dense vector and position metadata on every chunk. It runs on one of five engines, with Elasticsearch as the default. Onyx supports OpenSearch only. It embeds each document title as its own vector, stores ACLs on every chunk, and builds a secondary index so the embedding model can be switched without downtime. R2R uses pgvector plus a generated English tsvector. HNSW or IVFFlat indexes have to be created explicitly through the API. Quivr treats Weaviate as a rebuildable copy of canonical Postgres and S3 data. Each embedding model owns its own “vector space”, which can be evaluated beside the live one and then promoted.
Pluggable dense stores. AnythingLLM has ten vector adapters (LanceDB by default) and a local MiniLM embedder. It has no keyword index, and each adapter re-implements chunking and embedding. Kotaemon pairs a local Chroma vector store with a LanceDB document store that provides full-text search. LlamaIndex has about 80 vector-store integrations. Its built-in SimpleVectorStore scores every vector in Python and has no hybrid mode, and BM25 is a separate package. Haystack’s DocumentStore protocol requires only four methods. Its in-memory store does brute-force vector search and implements BM25L/Okapi/Plus itself.
Graph plus vectors. RAG-Anything writes chunks, entities and relations into LightRAG (<1.5) stores. It builds no BM25 or keyword index, so there is no lexical retrieval.
No vectors. PageIndex stores one JSON tree per document. Each node has a page range and an LLM summary of up to 150 words.
Pick: R2R if you want one Postgres for vectors, keywords and users. Pick: RAGFlow or Onyx for large corpora that need lexical plus semantic recall. Pick: LlamaIndex or Haystack to keep the vector database you already run.
Per-project answers
infiniflow/ragflow
answeredEmbedding models: RAGFlow supports a broad range of embedding providers through a plugin-style model architecture (internal/entity/models/). The factory (factory.go) resolves model requests to provider-specific drivers — there are 80+ model provider files for OpenAI, Anthropic, Google/Gemini, DeepSeek, Ollama, Cohere, Jina, Voyage, vLLM, Xinference, HuggingFace, and many Chinese providers (e.g. Baichuan, ZhipuAI, Qwen, Hunyuan). Each supports chat, text-embedding, reranking, or TTS modalities. Models are registered per-tenant in the tenant_model_instance table and resolved at runtime by ModelSolver (service/model_solver.go).
Vector stores: Four backends are supported (engine/engine.go:33-38): Elasticsearch, Infinity, OceanBase, and SereneDB, plus OpenSearch (via the OS config block in service_conf.yaml.template). Each implements the DocEngine interface (engine/engine.go:41-93) with chunk CRUD, search, metadata operations, and SQL execution. The Go code has Elasticsearch (engine/elasticsearch/chunk.go) and Infinity (engine/infinity/) implementations; OceanBase and SereneDB share a PostgreSQL-compatible layer. Chunks are created via CreateChunkStore() which sets up an ES index with number_of_shards=1, number_of_replicas=0 and appropriate mappings including tag_feas as rank_features for tag-based ranking. Bulk writes are batched at 500 chunks or 20MB.
Hybrid / keyword (BM25) indexes: Chunks carry both content_ltks (tokenized text for BM25 matching) and embedding vectors. BM25 is implemented via the QueryBuilder (service/nlp/query_builder.go) that produces matchText expressions for the engine. OpenSearch has dedicated hybrid search pipeline configuration (service_conf.yaml.template:59).
Metadata: The DocEngine interface defines CreateMetadataStore, InsertMetadata, UpdateMetadata, SearchMetadata operations (engine/engine.go:54-61). Metadata fields per chunk include docnm_kwd (document name), kb_id, img_id, title_tks, position_int, page_num_int, chunk_order_int, content_with_weight, doc_type_kwd (text/image/table), mom_id (parent chunk for child fragments), and user-defined metadata values set via filter conditions.
DocEngine has five engine types: Elasticsearch, Infinity, OceanBase, SeekDB and SereneDB. OceanBase and SeekDB share one implementation (IsOceanBaseFamily), and SereneDB is a separate package. There is no OpenSearch engine under internal/engine/, only a Compose profile and a mapping file.Mintplex-Labs/anything-llm
answeredEmbedding models. AnythingLLM supports 14 embedding engines, selected via EMBEDDING_ENGINE env: native (local ONNX via @xenova/transformers), openai, azure, localai, ollama, lmstudio, cohere, voyageai, litellm, mistral, generic-openai, gemini, openrouter, lemonade. The native embedder defaults to Xenova/all-MiniLM-L6-v2 (23MB, 1000-char chunks, 25 concurrent chunks) with two alternatives: Xenova/nomic-embed-text-v1 (139MB, 16000-char chunks, 5 concurrent, uses search_document:/search_query: prefixes) and MintplexLabs/multilingual-e5-small (487MB, 100+ languages, passages prefixed with passage:/query:). For local on-device embedding, models are downloaded on first use from HuggingFace Hub with a fallback CDN at cdn.anythingllm.com.
Vector stores. Ten backends are supported: LanceDB (default, embedded, file-based), Pinecone, Chroma, ChromaCloud, Weaviate, QDrant, Milvus, Zilliz, AstraDB, and PGVector. Each implements the VectorDatabase base class with addDocumentToNamespace, performSimilaritySearch, deleteDocumentFromNamespace. Workspaces map 1:1 to vector-db namespaces (using workspace.slug). Chunking and embedding happen inside the vector-store adapter's addDocumentToNamespace() — every provider reads text_splitter_chunk_size/text_splitter_chunk_overlap from SystemSettings, splits via TextSplitter, embeds via the selected EmbedderEngine, and upserts.
Hybrid / BM25. There is no built-in BM25 or hybrid search. Retrieval is dense-only via cosine similarity on embedding vectors. Full-text search for workspace/thread names uses fast-levenshtein matching, but document-level keyword search is not implemented.
Metadata. Document metadata stored per chunk includes: text (the chunk content), title, description, docAuthor, docSource, published, chunkSource, token_count_estimate and scoring fields. Pinned documents bypass the vector store entirely and are injected directly as context.
run-llama/llama_index
answeredEmbedding models are accessed through the BaseEmbedding abstract class (llama_index/core/base/embeddings/base.py:72-80). The resolve_embed_model() function (llama_index/core/embeddings/utils.py:30-50) resolves a string name (e.g., "default" → OpenAI's text-embedding-ada-002) or a LangChain embedding into a BaseEmbedding instance. The VectorStoreIndex._get_node_with_embedding() embeds nodes in batches (llama_index/core/indices/vector_store/base.py:126-148).
Vector stores implement BasePydanticVectorStore (llama_index/core/vector_stores/types.py). The built-in SimpleVectorStore (llama_index/core/vector_stores/simple.py:64-80) keeps an in-memory dict of node_id → embedding and persists to JSON via fsspec. Dozens of production vector store integrations exist as separate packages (Chroma, Pinecone, Qdrant, Weaviate, FAISS, Elasticsearch, Milvus, etc.). The StorageContext (llama_index/core/storage/storage_context.py:52-72) holds the vector store, document store, index store, and graph store.
Hybrid/keyword indexing: The VectorStoreQueryMode enum (llama_index/core/vector_stores/types.py:45-60) defines HYBRID, SPARSE, and SEMANTIC_HYBRID modes. The VectorIndexRetriever passes an alpha parameter to the vector store to weight dense vs sparse scores (llama_index/core/indices/vector_store/retrievers/retriever.py:42-73). Support for hybrid search depends on the underlying vector store. Separately, BaseKeywordTableIndex (llama_index/core/indices/keyword_table/base.py:43-97) uses an LLM to extract keywords from each chunk and builds an inverted mapping of keyword → node IDs — a form of keyword-based sparse retrieval.
Metadata is stored in each node's metadata dict and persisted to the vector store. MetadataFilters (llama_index/core/vector_stores/types.py:142-200) with operators (==, >, <, IN, ANY, TEXT_MATCH, etc.) support filtering at query time.
The-Vibe-Company/quivr
answeredEmbedding model: intfloat/multilingual-e5-small (384 dimensions, cosine distance), served through a TEI (Text Embeddings Inference) container (deploy/compose/compose.yaml:65-79). The TEI encoder (adapters/tei/encoder.go:26-65) validates 384-dim unit-norm vectors, prefixes documents with passage: and queries with query: (plugins/core-ingest/main.go:131). The engine's legacy E5 space is an older built-in encoder for pre-THE-777 Corpora (app/run.go:55). Vector store: Only Weaviate is supported (adapters/weaviate/projection.go). Vectors are stored as named vectors per generation: the legacy semantic_text_v1 for pre-named-space generations, or s_<hash> for plugin-owned spaces (projection.go:54-59). Distance metrics include cosine (default), dot, and L2-squared (projection.go:62-70). The index uses HNSW (projection.go:73). Hybrid/BM25: Weaviate provides built-in BM25 alongside vector search. The engine constructs GraphQL queries: bm25:{query} for lexical, nearVector:{vector} for semantic, and hybrid:{query, vector, alpha, fusionType} for hybrid (projection.go:552-558). Keyword fields default to title^2 and body, with an optional lexicalText property from plugins. Metadata: Projected alongside vectors into Weaviate as individual properties (adapters/weaviate/metadata.go). The Generation.MetadataProjected flag gates filter availability (content/searchable.go:58-68); filters apply as Weaviate where clauses (projection.go:480-502). Embeddings are stored as artifacts in SeaweedFS (S3) with SHA-256 content hashes (content/embeddings.go:43-61), carrying full provenance (version, segment, space, producer info).
VectifyAI/PageIndex
answeredNo embeddings, no vector stores. PageIndex does not compute vector embeddings or use vector databases. The index is a hierarchical tree of sections with page ranges stored as JSON. This "tree index" replaces what a vector index would do in traditional RAG.
Tree structure. Each node has {title, node_id, start_index, end_index, (summary,) (text,) nodes} (local_api.py:396-414, types.py). The tree is built bottom-up: leaves from layout headings, parent nodes spanning the range of their children. Intro nodes cover pages a parent opens with before its first child (tree_optimize.py). Node summaries are LLM-generated (the summary_model) using a bottom-up scheduler — leaves from their own page text, parents composed from child summaries plus residual pages (utils.py:1140-1157). Small nodes (under ~300 tokens) keep raw text as their summary.
Standard mode indexing. The classic pipeline (page_index_classic.py:1200-1283): (1) detect TOC pages via LLM, (2) extract the TOC structure with page indices, (3) verify correctness by spot-checking titles against page text (verify_toc, line 1066-1120), (4) fix incorrect items with retries (fix_incorrect_toc_with_retries, line 1044-1060), (5) post-process into a nested tree, (6) recursively split large nodes (>10 pages and >20k tokens; process_large_node_recursively, line 1168-1198), (7) add node text and summaries, (8) generate an optional one-sentence doc description (generate_doc_description, utils.py:1183-1198). Configurable via config.yaml: defaults toc_check_page_num=20, max_page_num_each_node=10, max_token_num_each_node=20000.
Flash mode indexing. flash/api.py:180-248: (1) LLM-free layout-based TOC extraction (extract_toc), (2) optional embedded bookmark integration, (3) deterministic merge pass (collapses nodes where navigating to children is more expensive than scanning the parent; tree_optimize.py:18-24), (4) optional LLM expand pass (proposes subsections where a node's span exceeds cost of routing; tree_optimize.py:12-16), (5) bottom-up summaries. The entire Flash pipeline is LLM-free when summary=False, optimize=False.
Local storage. The local store (local_store.py) writes JSON files to ~/.pageindex/: each document as a directory containing doc.json (metadata), tree.json (the tree structure), pages.json (page text). A manifest.json indexes all documents. Atomic writes with _write_json_atomic use temp-file + os.replace. No external database is required.
Cloud indexing. Cloud mode delegates indexing to PageIndex's managed pipeline at api.pageindex.ai (cloud_api.py:38-86). The client uploads the PDF and receives a doc_id; the cloud handles parsing, OCR, image understanding, tree construction, and stores the result. Metadata, folders, and block-level citations are cloud-only features (README.md:199-208).
./.pageindex relative to the working directory (storage_path, client.py), not ~/.pageindex.onyx-dot-app/onyx
answeredEmbedding models: Embedding is done by the DefaultIndexingEmbedder (backend/onyx/indexing/embedder.py, lines 89–253) which wraps an EmbeddingModel (search_nlp_models.py). Supported providers include OpenAI (text-embedding-3-*), Cohere, VoyageAI, VertexAI, and custom OpenAI-compatible APIs. The embedding server host/port is configurable (INDEXING_MODEL_SERVER_HOST/PORT). Each chunk receives a full embedding, optional mini-chunk embeddings (for multipass scoring), and a title embedding (lines 156–183). Embeddings are normalized or dimension-reduced per SearchSettings.
Vector store: The sole document index backend is OpenSearch (via the OpenSearchDocumentIndex class at backend/onyx/document_index/opensearch/opensearch_document_index.py, line 1+). It implements the DocumentIndex interface (backend/onyx/document_index/interfaces.py, lines 474–497) which requires hybrid (vector + keyword) search, metadata updates, deletion, and random retrieval.
Hybrid index structure: Each chunk is stored as an OpenSearch document with a dense vector field (content_vector), lexical text fields (content for BM25), and structured metadata fields: access_control_list, source_type, document_sets, tenant_id, cc_pair_ids, created_at, last_updated, hidden, boost, personas, user_projects, and title vectors. The schema is defined in opensearch/schema.py.
Indexing pipeline: The end-to-end flow lives in run_indexing_pipeline (indexing_pipeline.py:1641–1724). Phases: (1) upsert documents to PostgreSQL, (2) filter by change-gate (timestamp or content-hash, get_docs_to_update, lines 365–447), (3) process image sections, (4) chunk via Chunker.chunk(), (5) optionally add contextual-RAG LLM enrichment, (6) embed via IndexingEmbedder, (7) write to OpenSearch via write_chunks_to_vector_db_with_backoff (vector_db_insertion.py, line 19+). The secondary (FUTURE) index is supported for zero-downtime model migrations.
Storage: Chunks are persisted to ChunkBatchStore (temporary on-disk store) between embedding and vector-db write to decouple memory from OpenSearch latency. The document index stores chunk counts for efficient tail-truncation (IndexingMetadata.ChunkCounts).
deepset-ai/haystack
answeredEmbedding models – Embedders are Haystack components that call external APIs or local models. The OpenAITextEmbedder (haystack/components/embedders/openai_text_embedder.py:16-254) and OpenAIDocumentEmbedder embed text and documents using OpenAI's embeddings.create endpoint (default: text-embedding-3-small). They accept dimensions, prefix/suffix, and full OpenAI client configuration. API clients are lazily initialized in warm_up() (line 116), not __init__, keeping components serializable without credentials. The embedders directory also includes AzureTextEmbedder/AzureDocumentEmbedder variants and mock embedders for testing.
Document stores – The only built-in store is InMemoryDocumentStore (haystack/document_stores/in_memory/document_store.py:98), which stores documents in process-global dictionaries. All other stores — Elasticsearch, OpenSearch, Pinecone, Qdrant, Weaviate, Chroma, Astra DB, Pgvector — are provided as separate integration packages (e.g., elasticsearch-haystack, pinecone-haystack). The DocumentStore protocol (haystack/document_stores/types/protocol.py:11-136) defines the interface all stores implement: write_documents, filter_documents, delete_documents, count_documents. The protocol also defines methods for embedding_retrieval and bm25_retrieval that stores implement with their backend's native similarity search.
Hybrid/keyword (BM25) indexes – The in-memory store maintains a BM25 index incrementally during write_documents (lines 543-553). Each document is tokenized via _tokenize_bm25 (custom CJK-aware regex, line 58), term frequencies are stored in BM25DocumentStats dataclasses, and an IDF vocabulary counter and average document length are updated incrementally. The store supports multiple BM25 variants via the rank_bm25 library (configurable via bm25_algorithm). Plugging the BM25 retriever and embedding retriever into the MultiRetriever component (haystack/components/retrievers/multi_retriever.py:19-120) enables hybrid search with reciprocal_rank_fusion joining mode.
Metadata – All document stores support structured metadata filtering via a nested comparison/logic filter syntax. The DuplicatePolicy enum (haystack/document_stores/types/policy.py) controls write behavior: NONE/FAIL, SKIP, OVERWRITE.
Cinnamon/kotaemon
answeredEmbeddings are produced by pluggable models: OpenAI, Azure OpenAI, Cohere, VoyageAI, Google, Mistral, HuggingFace (via LangChain wrappers), TEI (Text-Embedding-Inference) endpoints, and local FastEmbed models (e.g. BAAI/bge-small-en-v1.5) (libs/kotaemon/kotaemon/embeddings/ directory). The VectorIndexing class calls self.embedding(docs) to generate embeddings, then adds them to both a vector store and a doc store (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 84–93). Vector stores available: Chroma (persistent client), LanceDB, Qdrant, Milvus (with lazy init by dimension), and SimpleFileVectorStore (libs/kotaemon/kotaemon/storages/vectorstores/). Default in flowsettings.py is Chroma at ktem_app_data/user_data/vectorstore (line 100–103). Document stores support full-text search (BM25) via Elasticsearch (with custom BM25 similarity, line 38–46 of elasticsearch.py) or LanceDB (with FTS index using en_stem tokenizer, line 56–60 of lancedb.py). The default docstore in flowsettings.py is LanceDBDocumentStore (line 95–97). Hybrid retrieval is supported by pairing a vector store with a BM25-capable doc store. Metadata stored includes file_id, page_label, file_name, type, thumbnail_doc_id, and window for sentence-window splitters. The indexing also writes chunk content to a markdown cache directory and records source-to-chunk relationships in an SQL Index table (libs/ktem/ktem/index/file/pipelines.py, lines 426–464).
HKUDS/RAG-Anything
answeredRAG-Anything does not implement its own indexing — it delegates entirely to LightRAG, which is a dependency (lightrag-hku<1.5 in requirements.txt:6). The RAGAnything class wraps a LightRAG instance (raganything.py:69) and passes through lightrag_kwargs for storage backend selection.
Embedding models. The user provides an embedding_func callable (e.g. openai_embed from lightrag.llm.openai) when constructing RAGAnything. This is passed directly to LightRAG(...) (raganything.py:425). No embedding model is bundled; any embedding function compatible with LightRAG's EmbeddingFunc signature works.
Vector stores and graph storage. LightRAG supports pluggable backends for kv_storage, vector_storage, graph_storage, and doc_status_storage, all exposed through lightrag_kwargs (raganything.py:86-96). These default to LightRAG's built-in options (typically file-based JSON/JSONL for development, with production backends like MongoDB, Neo4j, or Milvus available by configuring lightrag_kwargs). The RAG-Anything code additionally creates two KV namespaces: parse_cache for caching parsed content lists and multimodal_status for multimodal completion tracking (raganything.py:447-465).
Hybrid / keyword (BM25) indexes. This is handled entirely by LightRAG's query modes. RAG-Anything surfaces them as the mode parameter (local, global, hybrid, naive, mix, bypass) in aquery() (query.py:129-131). LightRAG internally maintains BM25 indexes alongside vector and graph indexes for hybrid search.
Metadata. Each chunk stores page_idx and page_idx_end via the page-provenance system in processor.py:251-385. File-level metadata (path, content hash, parser config) is tracked in doc_status records (processor.py:128-159). The use_full_path config option (config.py:126-128) controls whether file references use the absolute path or just the basename.
SciPhi-AI/R2R
answeredEmbedding models. Three embedding providers: OpenAIEmbeddingProvider (py/core/providers/embeddings/openai.py:21-60) supporting text-embedding-ada-002, text-embedding-3-small, and text-embedding-3-large with configurable dimensions; LiteLLMEmbeddingProvider (py/core/providers/embeddings/litellm.py:25-65) routing through LiteLLM for any provider (Amazon, HuggingFace, etc.); and OllamaEmbeddingProvider for local models. EmbeddingConfig (py/core/base/providers/embedding.py:21-44) configures provider, base_model, base_dimension, batch_size, concurrent_request_limit, and vector quantization settings. Retry with exponential backoff is built in.
Vector store. All vectors are stored in PostgreSQL using pgvector. PostgresChunksHandler (py/core/providers/database/chunks.py:75-182) creates tables with a vec column (configurable dimension), optional vec_binary bit(N) column for INT1 quantization, text TEXT, and metadata JSONB. A full-text search column (fts tsvector) is auto-generated as to_tsvector('english', text). Index methods supported: IVFFlat (IndexArgsIVFFlat with n_lists) and HNSW (IndexArgsHNSW with m and ef_construction) via py/shared/abstractions/vector.py:27-107. Distance measures: cosine, L2, max-inner-product, L1, hamming, jaccard. Vector quantization types: FP32, FP16, INT1 (binary), SPARSE.
Hybrid / keyword (BM25). Full-text search uses PostgreSQL's websearch_to_tsquery('english', ...) with ts_rank ordering (py/core/providers/database/chunks.py:484-536). Hybrid search (py/core/providers/database/chunks.py:538-640) combines semantic and full-text via weighted Reciprocal Rank Fusion with configurable semantic_weight, full_text_weight, and rrf_k.
Metadata. Stored as JSONB alongside each vector entry. All document metadata (title, version, chunk_order, page_number, parser_generated) is carried through to the metadata column. Filtered at query time via the filter system in py/core/providers/database/filters.py.
← How are documents parsed and chunked? · How is retrieval performed? →