# run-llama/llama_index

> Python RAG and agent framework with a core of index, retriever, synthesizer and agent abstractions plus ~550 integration packages.

- Category: [RAG engines](https://llms-technical-reviews.com/rag/)
- Repository: https://github.com/run-llama/llama_index (reviewed at commit `81f0e06f6e61a5659bd409cac4a199e28c2fd85e`, 2026-10-05)
- Stars: 52422 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/llama_index/

## Overview

LlamaIndex is a Python framework for connecting LLMs to your own data. Its vocabulary has become standard in RAG: `Document`s are split into `Node`s, nodes are embedded into an index, a retriever fetches `NodeWithScore`s, postprocessors filter and rerank them, and a response synthesizer turns them into an answer with `source_nodes` attached. On top of that sit chat engines, sub-question and router query engines, structured-output programs, evaluators, and event-driven agents.

The repository is a monorepo of separately published packages. `llama-index-core` holds the abstractions and a minimal default implementation of each: an in-memory `SimpleVectorStore`, JSON-backed doc and index stores, `SentenceSplitter`, and the synthesizers. Nearly everything that talks to the outside world lives under `llama-index-integrations/`, which has about 550 packages: roughly 100 LLMs, 60 embedding providers, 80 vector stores, 150 readers, plus retrievers, postprocessors, tools and observability hooks. Each one is installed as `llama-index-<kind>-<name>` and imported under a namespace package such as `llama_index.vector_stores.chroma`. Even the default OpenAI LLM and embedding model are integrations that core imports lazily.

Two recent splits matter when reading the code. The workflow engine that powers agents (`Workflow`, `@step`, `Context`) is now the separate `llama-index-workflows` package, and `llama_index.core.workflow` mostly re-exports it. Instrumentation (dispatchers, spans, events) has moved into `llama-index-instrumentation`. The older `CallbackManager` still runs alongside it.

## Architecture

```mermaid
flowchart LR
  R["Readers"] --> D["Documents"]
  D --> IP["IngestionPipeline / run_transformations"]
  IP --> NP["Node parsers + extractors"]
  NP --> EM["Embed model"]
  EM --> VS["Vector store"]
  IP --> DS["Docstore (dedup)"]
  VS --> IDX["VectorStoreIndex"]
  IDX --> RET["VectorIndexRetriever"]
  RET --> PP["Node postprocessors"]
  PP --> SYN["Response synthesizer"]
  SYN --> LLM["LLM"]
  QE["RetrieverQueryEngine"] --> RET
  QE --> SYN
  AG["Agents (Workflow)"] --> TOOLS["QueryEngineTool / FunctionTool"]
  TOOLS --> QE
  S["Settings (lazy defaults)"] -.-> EM
  S -.-> LLM
```

| Component | Path | Role |
|---|---|---|
| Schema | `llama-index-core/llama_index/core/schema.py` | `Document`, `TextNode`, `ImageNode`, `NodeWithScore`, `TransformComponent`, `QueryBundle` |
| Settings | `llama-index-core/llama_index/core/settings.py` | Global lazy defaults for `llm`, `embed_model`, `node_parser`, `transformations`, `callback_manager` |
| Ingestion | `.../core/ingestion/pipeline.py` | `run_transformations`, `IngestionPipeline` with cache, docstore dedup and multiprocessing |
| Node parsers | `.../core/node_parser/` | `SentenceSplitter`, `TokenTextSplitter`, `SemanticSplitterNodeParser`, `CodeSplitter`, markdown/HTML/JSON parsers |
| Indices | `.../core/indices/` | `VectorStoreIndex`, `SummaryIndex`, keyword-table, tree, property-graph and document-summary indices |
| Vector stores | `.../core/vector_stores/` | `BasePydanticVectorStore` contract, `VectorStoreQuery`, `MetadataFilters`, `SimpleVectorStore` |
| Storage | `.../core/storage/` | `StorageContext` bundling docstore, index store, vector stores and graph store; fsspec persistence |
| Query side | `.../core/retrievers/`, `postprocessor/`, `response_synthesizers/`, `query_engine/` | Retrieval, fusion, reranking, synthesis and engine composition |
| Agents | `.../core/agent/workflow/` | `FunctionAgent`, `ReActAgent`, `CodeActAgent`, multi-agent `AgentWorkflow` |
| Evaluation | `.../core/evaluation/` | LLM-judge evaluators, retrieval metrics, `BatchEvalRunner` |
| Integrations | `llama-index-integrations/` | Provider packages (LLMs, embeddings, vector stores, readers and more) |

## How a request flows

Take the canonical five lines: `index = VectorStoreIndex.from_documents(docs)` then `index.as_query_engine().query("...")`.

1. **Defaults.** Nothing is configured, so `Settings` resolves lazily. `embed_model` becomes `resolve_embed_model("default")`, which imports `llama-index-embeddings-openai` and validates `OPENAI_API_KEY` ([utils.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/embeddings/utils.py#L30-L75)). The node parser defaults to a `SentenceSplitter` ([settings.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/settings.py#L138-L147)).
2. **Transform.** `from_documents` records each document's hash in the docstore and calls `run_transformations`, which applies each `TransformComponent` in order. If a cache is given, each step is looked up by a hash of the nodes and the transform ([base.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/base.py#L89-L129), [pipeline.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/ingestion/pipeline.py#L72-L112)).
3. **Chunk.** `SentenceSplitter` (1,024 tokens, 200 overlap) splits recursively by paragraph, then by NLTK sentence, then by a punctuation regex, then by spaces, and merges the pieces back up to the chunk size ([sentence.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/sentence.py#L198-L246)).
4. **Embed and store.** `_add_nodes_to_index` embeds nodes in batches and calls `vector_store.add`. If the store does not keep text (`stores_text=False`, as with `SimpleVectorStore`), the nodes also go into the docstore and index struct without their embeddings ([vector_store/base.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/vector_store/base.py#L219-L258)).
5. **Build the engine.** `as_query_engine` calls `as_retriever` and wraps it with `RetrieverQueryEngine.from_args`, which defaults to `ResponseMode.COMPACT` ([base.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/base.py#L491-L516), [retriever_query_engine.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/query_engine/retriever_query_engine.py#L63-L80)).
6. **Retrieve.** `VectorIndexRetriever._retrieve` embeds the query unless the mode is `SPARSE` or `TEXT_SEARCH`. It then builds a `VectorStoreQuery` carrying `similarity_top_k`, `filters`, `mode`, `alpha` and the hybrid/sparse top-k values ([retriever.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/vector_store/retrievers/retriever.py#L92-L144)). `SimpleVectorStore.query` applies metadata filters in Python and scores every remaining embedding. It supports default, MMR and learner modes and raises on anything else, including hybrid ([simple.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/simple.py#L244-L314)).
7. **Postprocess and synthesize.** `_query` runs the retriever, applies each node postprocessor in order, and calls the synthesizer inside a `QUERY` callback event ([retriever_query_engine.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/query_engine/retriever_query_engine.py#L142-L217)). `CompactAndRefine` repacks the chunks to fill the context window, then runs the refine loop: answer with the first packed chunk, refine with each next one ([compact_and_refine.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/compact_and_refine.py#L33-L59)).
8. **Return.** The `Response` (or `StreamingResponse` when `streaming=True`) carries the answer text plus `source_nodes`, the retrieved nodes with scores.

## Key components

### Nodes, transformations and ingestion

Every ingestion step is a `TransformComponent` that maps a list of nodes to a list of nodes: splitters, metadata extractors (title, summary, keywords, questions answered) and embedding models alike. `IngestionPipeline.run` adds three production features on top of `run_transformations`: a transformation cache, docstore-based deduplication with `UPSERTS`, `DUPLICATES_ONLY` or `UPSERTS_AND_DELETE` strategies, and a `spawn` multiprocessing pool when `num_workers > 1`. It writes embedded nodes straight to the attached vector store ([pipeline.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/ingestion/pipeline.py#L606-L661)).

### Vector store contract

`BasePydanticVectorStore` asks for `add`, `delete` and `query` (plus async versions), and declares `stores_text` and `is_embedding_query`. All retrieval options travel in a single `VectorStoreQuery`. Query modes include `DEFAULT`, `SPARSE`, `HYBRID`, `TEXT_SEARCH`, `SEMANTIC_HYBRID`, `MMR` and learner modes ([types.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/types.py#L45-L61)). Filters support `AND`, `OR` and `NOT`. Each integration decides which modes and filter operators it honours, so "hybrid search in LlamaIndex" really means hybrid search in a particular backend.

### Retrieval composition

Beyond the vector retriever, core provides `QueryFusionRetriever`. It generates `num_queries=4` rewrites with an LLM, runs every sub-retriever on every query, and fuses the results with reciprocal rank, relative-score, distance-based or simple fusion ([fusion_retriever.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/retrievers/fusion_retriever.py#L24-L41)). Core also has router and recursive retrievers and auto-merging retrieval. BM25 is not in core; it is the `llama-index-retrievers-bm25` integration. Rerankers are node postprocessors: `LLMRerank`, `RankGPT`, `SentenceTransformerRerank`, plus similarity cutoffs, recency and metadata replacement.

### Response synthesizers

`BaseSynthesizer.synthesize` emits instrumentation events, short-circuits on empty input, and dispatches to `get_response` with node content rendered at `MetadataMode.LLM` ([base.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/base.py#L232-L290)). The modes are `refine`, `compact` (the default), `tree_summarize`, `simple_summarize`, `accumulate`, `generation`, `no_text` and `context_only`. They trade LLM calls against context use. `compact` is a sensible default; `refine` costs one call per chunk.

### Agents on workflows

`BaseWorkflowAgent` is a `Workflow` subclass whose `@step` methods form the loop. `init_run` loads memory, `setup_agent` adds the system prompt and state, `run_agent_step` calls the subclass's `take_step`, `parse_agent_output` enforces `max_iterations` and routes to tool calls, `call_tool` runs them, and `aggregate_tool_results` loops back ([base_agent.py](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/agent/workflow/base_agent.py#L393-L491)). `FunctionAgent` uses native tool calling, `ReActAgent` parses text, and `CodeActAgent` executes generated code. `AgentWorkflow` hands control between named agents. RAG is plugged in by wrapping a query engine as a `QueryEngineTool`.

## Extending it

- **New provider.** Subclass `LLM`/`CustomLLM`, `BaseEmbedding`, `BasePydanticVectorStore`, `BaseReader` or `BaseRetriever`, then publish it as a `llama-index-<kind>-<name>` package with its own `pyproject.toml`. The `llama-dev` tool in the repo manages and tests these packages.
- **Custom transformation.** Implement `TransformComponent.__call__(nodes)` and drop it into `transformations=[...]` or an `IngestionPipeline`.
- **Postprocessors and synthesizers.** Subclass `BaseNodePostprocessor._postprocess_nodes` or pass custom `PromptTemplate`s (`text_qa_template`, `refine_template`) to the query engine.
- **Workflows.** Write your own `Workflow` with typed events and `@step` methods for pipelines that do not fit the query-engine mould.
- **Observability.** Register span and event handlers on the instrumentation dispatcher, or install one of the callback/observability integrations (OpenTelemetry, Arize Phoenix, Langfuse and others).

## Running it

- **Install.** `pip install llama-index` pulls core plus the OpenAI LLM and embedding integrations; file readers, vector stores and other providers are installed one package at a time. `pip install llama-index-core` gives the bare abstractions. Core requires Pydantic v2, NLTK, tiktoken, SQLAlchemy, networkx, numpy and fsspec, and the workflows package.
- **Persistence.** `storage_context.persist(persist_dir=...)` writes the docstore, index store and `SimpleVectorStore` as JSON through fsspec, so local disk, S3 and GCS all work. Reload with `load_index_from_storage`. Anything larger belongs in an external vector store integration.
- **Services.** None are required beyond model access. The defaults call OpenAI, and setting `Settings.llm` and `Settings.embed_model` to local integrations (Ollama, HuggingFace) makes it fully offline.
- **No server.** Core has no HTTP server or UI. The old CLI is now the separate `llama-index-cli` package.

## Strengths and caveats

- **Strength: coverage.** No other framework has as many maintained connectors. Swapping a vector store or LLM is usually a one-line change because the contracts are narrow.
- **Strength: composable query side.** Retrievers, postprocessors and synthesizers are independent pieces, and fusion, routing, sub-questions and auto-merging are built in.
- **Strength: ingestion is production-aware.** Caching, docstore upserts and multiprocessing come with the pipeline.
- **Caveat: the in-memory store is a toy.** `SimpleVectorStore` scores every embedding in a Python loop and has no hybrid mode. Use it for prototypes only.
- **Caveat: implicit globals.** `Settings` silently falls back to OpenAI for both LLM and embeddings, so a missing configuration shows up as an API-key error deep in a call stack.
- **Caveat: capability depends on the backend.** Query modes, filter operators and async support vary between vector store integrations, and the core types do not tell you which ones a given store supports.
- **Caveat: many layers and some legacy.** Callbacks and instrumentation coexist, the workflow engine lives in another package, and old `ServiceContext`-era names still appear, so tracing behaviour through the stack takes effort.

*Sources: code at 81f0e06, deepwiki-open wiki (12 pages), verified Q&A.*

## How run-llama/llama_index answers the RAG engines questions

### How are documents parsed and chunked? (answered)

Documents enter through **readers** which implement `BaseReader` or `BasePydanticReader` (`llama_index/core/readers/base.py:19-48`). Core file readers (in `llama-index-integrations/readers/file/`) support PDF, DOCX, EPUB, HTML, Markdown, CSV/Excel, IPYNB, PPTX, RTF, XML, images, and video/audio. Table-specific readers (`PandasCSVReader`, `PandasExcelReader`) parse tabular data. The `UnstructuredReader` wraps the unstructured.io library for OCR and layout parsing but is an integration, not built into core. No built-in OCR or table-extraction logic exists in the core library — those require external integrations.

Chunking is done by **node parsers** (subclasses of `NodeParser` at `llama_index/core/node_parser/interface.py:50-68`). The primary chunkers are:
- **`SentenceSplitter`** (`llama_index/core/node_parser/text/sentence.py:34-100`): default chunk_size=1024 tokens, chunk_overlap=200 tokens, prefers sentence/paragraph boundaries, uses regex `[^,.;。？！]+[,.;。？！]?|[,.;。？！]` as secondary splitter.
- **`TokenTextSplitter`** (`llama_index/core/node_parser/text/token.py:22-84`): splits by token count (using a configurable tokenizer), falls back from separator to `
` to character-level splitting.
- **`SemanticSplitterNodeParser`** (`llama_index/core/node_parser/text/semantic_splitter.py:35-78`): groups sentences by embedding similarity — embeds sentence windows and breaks at a configurable percentile threshold of cosine dissimilarity.
- **`CodeSplitter`** (`llama_index/core/node_parser/text/code.py:19-30`): language-aware AST-based splitting for code files.

Chunking runs in the **ingestion pipeline** (`llama_index/core/ingestion/pipeline.py:72-112`) where `run_transformations()` applies a sequence of `TransformComponent` instances with caching via `IngestionCache`. The pipeline supports parallelism via `ProcessPoolExecutor`.


Citations: [llama-index-core/llama_index/core/readers/base.py:19-48](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/readers/base.py#L19-L48) · [llama-index-core/llama_index/core/node_parser/interface.py:50-68](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/interface.py#L50-L68) · [llama-index-core/llama_index/core/node_parser/text/sentence.py:34-100](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/sentence.py#L34-L100) · [llama-index-core/llama_index/core/node_parser/text/token.py:22-84](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/token.py#L22-L84) · [llama-index-core/llama_index/core/node_parser/text/semantic_splitter.py:35-78](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/semantic_splitter.py#L35-L78) · [llama-index-core/llama_index/core/ingestion/pipeline.py:72-112](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/ingestion/pipeline.py#L72-L112)

### How are embeddings and indexes built and stored? (answered)

**Embedding models** are accessed through the `BaseEmbedding` abstract class (`llama_index/core/base/embeddings/base.py:72-80`). The `resolve_embed_model()` function (`llama_index/core/embeddings/utils.py:30-50`) resolves a string name (e.g., `"default"` → OpenAI's `text-embedding-ada-002`) or a LangChain embedding into a `BaseEmbedding` instance. The `VectorStoreIndex._get_node_with_embedding()` embeds nodes in batches (`llama_index/core/indices/vector_store/base.py:126-148`).

**Vector stores** implement `BasePydanticVectorStore` (`llama_index/core/vector_stores/types.py`). The built-in `SimpleVectorStore` (`llama_index/core/vector_stores/simple.py:64-80`) keeps an in-memory dict of `node_id → embedding` and persists to JSON via fsspec. Dozens of production vector store integrations exist as separate packages (Chroma, Pinecone, Qdrant, Weaviate, FAISS, Elasticsearch, Milvus, etc.). The `StorageContext` (`llama_index/core/storage/storage_context.py:52-72`) holds the vector store, document store, index store, and graph store.

**Hybrid/keyword indexing**: The `VectorStoreQueryMode` enum (`llama_index/core/vector_stores/types.py:45-60`) defines `HYBRID`, `SPARSE`, and `SEMANTIC_HYBRID` modes. The `VectorIndexRetriever` passes an `alpha` parameter to the vector store to weight dense vs sparse scores (`llama_index/core/indices/vector_store/retrievers/retriever.py:42-73`). Support for hybrid search depends on the underlying vector store. Separately, `BaseKeywordTableIndex` (`llama_index/core/indices/keyword_table/base.py:43-97`) uses an LLM to extract keywords from each chunk and builds an inverted mapping of keyword → node IDs — a form of keyword-based sparse retrieval.

**Metadata** is stored in each node's `metadata` dict and persisted to the vector store. `MetadataFilters` (`llama_index/core/vector_stores/types.py:142-200`) with operators (`==`, `>`, `<`, `IN`, `ANY`, `TEXT_MATCH`, etc.) support filtering at query time.


Citations: [llama-index-core/llama_index/core/base/embeddings/base.py:72-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/embeddings/base.py#L72-L80) · [llama-index-core/llama_index/core/embeddings/utils.py:30-50](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/embeddings/utils.py#L30-L50) · [llama-index-core/llama_index/core/indices/vector_store/base.py:126-148](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/vector_store/base.py#L126-L148) · [llama-index-core/llama_index/core/vector_stores/types.py:45-60](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/types.py#L45-L60) · [llama-index-core/llama_index/core/vector_stores/simple.py:64-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/simple.py#L64-L80) · [llama-index-core/llama_index/core/indices/keyword_table/base.py:43-97](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/keyword_table/base.py#L43-L97)

### How is retrieval performed? (answered)

Retrieval is orchestrated by the **`RetrieverQueryEngine`** (`llama_index/core/query_engine/retriever_query_engine.py:25-56`), which calls a retriever, applies node postprocessors, then synthesizes a response.

**Dense retrieval**: `VectorIndexRetriever` (`llama_index/core/indices/vector_store/retrievers/retriever.py:24-80`) embeds the query using the configured embed model, builds a `VectorStoreQuery` with the embedding, `similarity_top_k`, `filters`, and `mode`, then delegates to the vector store's `query()` method. The vector store returns `VectorStoreQueryResult` with nodes, similarities, and IDs.

**Sparse and hybrid**: `VectorStoreQueryMode.SPARSE` skips the embedding step and sends the raw query string. `VectorStoreQueryMode.HYBRID` sends both the embedding and query string, with `alpha` controlling the dense/sparse blend. The `KeywordTableIndex` provides an alternative LLM-keyword-based sparse retrieval.

**Query rewriting/decomposition**: `QueryFusionRetriever` (`llama_index/core/retrievers/fusion_retriever.py:33-70`) uses an LLM to generate multiple query variants from the original, retrieves from each variant across a set of sub-retrievers, then fuses results via reciprocal rank fusion, relative score, or simple re-ranking. `SubQuestionQueryEngine` (`llama_index/core/query_engine/sub_question_query_engine.py:37-60`) breaks a complex query into sub-questions, dispatches each to a different query engine tool, then synthesizes the final answer.

**Reranking**: Postprocessors run after retrieval. `LLMRerank` (`llama_index/core/postprocessor/llm_rerank.py:23-65`) uses an LLM to select the top-N most relevant nodes from candidate batches. `SentenceTransformerRerank` (`llama_index/core/postprocessor/sbert_rerank.py:12-56`) uses a cross-encoder model. Other postprocessors include `SimilarityPostprocessor`, `KeywordNodePostprocessor`, and `MetadataReplacementPostProcessor`.

**Filters**: `MetadataFilters` (`llama_index/core/vector_stores/types.py:142-200`) supports `AND`/`OR`/`NOT` conditions with operators including `==`, `!=`, `>`, `<`, `IN`, `ANY`, `ALL`, `TEXT_MATCH`.


Citations: [llama-index-core/llama_index/core/query_engine/retriever_query_engine.py:25-56](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/query_engine/retriever_query_engine.py#L25-L56) · [llama-index-core/llama_index/core/indices/vector_store/retrievers/retriever.py:24-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/indices/vector_store/retrievers/retriever.py#L24-L80) · [llama-index-core/llama_index/core/retrievers/fusion_retriever.py:33-70](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/retrievers/fusion_retriever.py#L33-L70) · [llama-index-core/llama_index/core/postprocessor/llm_rerank.py:23-65](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/postprocessor/llm_rerank.py#L23-L65) · [llama-index-core/llama_index/core/postprocessor/sbert_rerank.py:12-56](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/postprocessor/sbert_rerank.py#L12-L56) · [llama-index-core/llama_index/core/vector_stores/types.py:142-200](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/types.py#L142-L200)

### How are answers generated and grounded? (answered)

**Prompt assembly**: The `BaseSynthesizer` (`llama_index/core/response_synthesizers/base.py:64-107`) holds a `_text_qa_template` (default `DEFAULT_TEXT_QA_PROMPT`) and `_refine_template`. The QA prompt receives `context_str` (concatenated node text at `MetadataMode.LLM`) and `query_str`. For chat models, parallel `_chat_content_qa_template` and `_chat_content_refine_template` variants use `ChatMessage` blocks.

**Response modes** (`llama_index/core/response_synthesizers/type.py:4-58`): `REFINE` iterates through nodes, building an initial answer then refining with each subsequent chunk. `COMPACT` (via `CompactAndRefine` at `response_synthesizers/compact_and_refine.py:13-60`) packs multiple chunks into the context window before refining. `TREE_SUMMARIZE` builds a bottom-up summary tree. `GENERATION` (`llama_index/core/response_synthesizers/generation.py:32-50`) ignores context entirely. `SIMPLE_SUMMARIZE` stuffs all text into one prompt. `NO_TEXT` and `CONTEXT_ONLY` return nodes as-is.

**Source attribution**: Every `Response` object (`llama_index/core/base/response/schema.py:14-42`) carries `source_nodes: List[NodeWithScore]` alongside the response text. `Response.get_formatted_sources()` renders truncated node content with node IDs. Source nodes flow from retrieval through synthesis unchanged.

**Streaming**: When `streaming=True` is passed, calls return `StreamingResponse` or `AsyncStreamingResponse` (`llama_index/core/base/response/schema.py:108-200`), which wrap a generator/async generator yielding tokens. The response also carries `source_nodes` for display.

**Agentic/multi-step answering**: The `AgentWorkflow` and `FunctionAgent` (`llama_index/core/agent/`) support tool-calling agents that can iterate, reflect, and call retrieval tools. `SubQuestionQueryEngine` decomposes queries into parallel sub-questions. `ReActAgent` implements the ReAct pattern with chat formatters and output parsers.


Citations: [llama-index-core/llama_index/core/response_synthesizers/base.py:64-107](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/base.py#L64-L107) · [llama-index-core/llama_index/core/response_synthesizers/compact_and_refine.py:13-60](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/compact_and_refine.py#L13-L60) · [llama-index-core/llama_index/core/response_synthesizers/generation.py:32-50](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/response_synthesizers/generation.py#L32-L50) · [llama-index-core/llama_index/core/base/response/schema.py:14-42](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/response/schema.py#L14-L42) · [llama-index-core/llama_index/core/base/response/schema.py:108-200](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/base/response/schema.py#L108-L200)

### How is quality evaluated or observed? (answered)

LlamaIndex has a substantial **built-in evaluation suite** in `llama_index.core.evaluation`. Evaluators are subclasses of `BaseEvaluator` and include:
- **`FaithfulnessEvaluator`** (`llama_index/core/evaluation/faithfulness.py:16-44`): LLM-judge that checks if each claim in the response is supported by the context, answering YES/NO with example few-shot prompts.
- **`RelevancyEvaluator`** / **`AnswerRelevancyEvaluator`**: measure how relevant the response is to the query and context.
- **`CorrectnessEvaluator`**: compares response against a reference answer.
- **`ContextRelevancyEvaluator`**: evaluates whether retrieved context is relevant to the query.
- **`SemanticSimilarityEvaluator`**: embedding-based similarity between response and reference.
- **`PairwiseComparisonEvaluator`**: A/B comparison of two responses.
- **`GuidelineEvaluator`**: checks responses against custom guidelines.

**Retrieval-specific evaluation**: `RetrieverEvaluator` evaluates retrieval quality with `HitRate` and `MRR` from `llama_index/core/evaluation/retrieval/metrics.py`.

**Batch evaluation**: `BatchEvalRunner` (`llama_index/core/evaluation/batch_runner.py:75-80`) runs multiple evaluators across queries in parallel with semaphore concurrency control and exponential-backoff retries.

**Observability/tracing**: Two-layer system — (1) `CallbackManager` (`llama_index/core/callbacks/base.py:28-80`) with `CBEventType` events (`CHUNKING`, `NODE_PARSING`, `EMBEDDING`, `LLM`, `QUERY`, `RETRIEVE`, `SYNTHESIZE`, etc.) and contextvar-based trace stacking. (2) The `instrumentation` package (`llama_index/core/instrumentation/__init__.py`) provides typed events (`QueryStartEvent`, `RetrievalEndEvent`, `SynthesizeStartEvent`, etc.) with `@dispatcher.span` decorators on key methods. `NullEventHandler` and `NullSpanHandler` are defaults; users register custom handlers for production monitoring.


Citations: [llama-index-core/llama_index/core/evaluation/faithfulness.py:16-44](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/faithfulness.py#L16-L44) · [llama-index-core/llama_index/core/evaluation/batch_runner.py:75-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/batch_runner.py#L75-L80) · [llama-index-core/llama_index/core/evaluation/__init__.py:1-44](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/__init__.py#L1-L44) · [llama-index-core/llama_index/core/callbacks/base.py:28-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/callbacks/base.py#L28-L80) · [llama-index-core/llama_index/core/callbacks/schema.py:16-47](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/callbacks/schema.py#L16-L47)

### How is it deployed and operated? (answered)

LlamaIndex is distributed as a **Python library** (PyPI package `llama-index-core`), not as a managed service. There is **no built-in UI or API server** in core — the CLI (`llama_index/core/command_line/`) has been deprecated and moved to its own `llama-index-cli` package. The `chat_ui` directory exists but is experimental. Usage is purely programmatic: users write Python scripts or integrate into web frameworks (FastAPI, Flask, etc.).

**Infrastructure**: All state is handled through pluggable storage backends. `StorageContext` (`llama_index/core/storage/storage_context.py:52-72`) combines a `BaseDocumentStore` (default `SimpleDocumentStore` — in-memory dicts persisted to JSON), `BaseIndexStore` (index metadata), vector stores (in-memory or external), and `GraphStore`. Persistence uses `fsspec` — files can be saved to local disk or cloud storage (S3, GCS). Each store has `persist()`/`load()` methods for round-trip serialization.

**Production vector stores**: Users plug in external vector database integrations (Chroma, Pinecone, Qdrant, Weaviate, Elasticsearch, FAISS, etc.) via `pip install llama-index-vector-stores-<name>`. These handle indexing, sharding, and replication outside of LlamaIndex. The `SimpleVectorStore` supports only single-node in-memory use.

**Scaling**: The ingestion pipeline (`llama_index/core/ingestion/pipeline.py:72-112`) supports multiprocessing and async parallelism via `ProcessPoolExecutor`. The library supports async throughout (parallel embedding, parallel retrieval, parallel sub-question execution via `run_jobs`).

**Multi-tenancy**: Not built-in. Each `StorageContext` instance represents one index. Users manage isolation at the application layer. There is no built-in authentication, rate limiting, or user session management.

**LLM configuration**: Global `Settings` dataclass (`llama_index/core/settings.py:18-30`) provides lazy-initialized defaults for `llm`, `embed_model`, `callback_manager`, `tokenizer`, `node_parser`, and `transformations`. These can be overridden per-index or per-query engine instance.


Citations: [llama-index-core/llama_index/core/storage/storage_context.py:52-72](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/storage/storage_context.py#L52-L72) · [llama-index-core/llama_index/core/ingestion/pipeline.py:72-112](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/ingestion/pipeline.py#L72-L112) · [llama-index-core/llama_index/core/settings.py:18-30](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/settings.py#L18-L30) · [llama-index-core/llama_index/core/vector_stores/simple.py:64-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/vector_stores/simple.py#L64-L80)
