LLMs Technical Reviews
Home / RAG engines / llama_index

run-llama/llama_index

Python RAG and agent framework with a core of index, retriever, synthesizer and agent abstractions plus ~550 integration packages.

GitHub ↗★ 52kPythonMITcommit 81f0e06 · 2026-10-05homepage ↗

Overview

LlamaIndex is a Python framework for connecting LLMs to your own data. Its vocabulary has become standard in RAG: Documents are split into Nodes, nodes are embedded into an index, a retriever fetches NodeWithScores, postprocessors filter and rerank them, and a response synthesizer turns them into an answer with source_nodes attached. On top of that sit chat engines, sub-question and router query engines, structured-output programs, evaluators, and event-driven agents.

The repository is a monorepo of separately published packages. llama-index-core holds the abstractions and a minimal default implementation of each: an in-memory SimpleVectorStore, JSON-backed doc and index stores, SentenceSplitter, and the synthesizers. Nearly everything that talks to the outside world lives under llama-index-integrations/, which has about 550 packages: roughly 100 LLMs, 60 embedding providers, 80 vector stores, 150 readers, plus retrievers, postprocessors, tools and observability hooks. Each one is installed as llama-index-<kind>-<name> and imported under a namespace package such as llama_index.vector_stores.chroma. Even the default OpenAI LLM and embedding model are integrations that core imports lazily.

Two recent splits matter when reading the code. The workflow engine that powers agents (Workflow, @step, Context) is now the separate llama-index-workflows package, and llama_index.core.workflow mostly re-exports it. Instrumentation (dispatchers, spans, events) has moved into llama-index-instrumentation. The older CallbackManager still runs alongside it.

Architecture

flowchart LR
  R["Readers"] --> D["Documents"]
  D --> IP["IngestionPipeline / run_transformations"]
  IP --> NP["Node parsers + extractors"]
  NP --> EM["Embed model"]
  EM --> VS["Vector store"]
  IP --> DS["Docstore (dedup)"]
  VS --> IDX["VectorStoreIndex"]
  IDX --> RET["VectorIndexRetriever"]
  RET --> PP["Node postprocessors"]
  PP --> SYN["Response synthesizer"]
  SYN --> LLM["LLM"]
  QE["RetrieverQueryEngine"] --> RET
  QE --> SYN
  AG["Agents (Workflow)"] --> TOOLS["QueryEngineTool / FunctionTool"]
  TOOLS --> QE
  S["Settings (lazy defaults)"] -.-> EM
  S -.-> LLM
Component Path Role
Schema llama-index-core/llama_index/core/schema.py Document, TextNode, ImageNode, NodeWithScore, TransformComponent, QueryBundle
Settings llama-index-core/llama_index/core/settings.py Global lazy defaults for llm, embed_model, node_parser, transformations, callback_manager
Ingestion .../core/ingestion/pipeline.py run_transformations, IngestionPipeline with cache, docstore dedup and multiprocessing
Node parsers .../core/node_parser/ SentenceSplitter, TokenTextSplitter, SemanticSplitterNodeParser, CodeSplitter, markdown/HTML/JSON parsers
Indices .../core/indices/ VectorStoreIndex, SummaryIndex, keyword-table, tree, property-graph and document-summary indices
Vector stores .../core/vector_stores/ BasePydanticVectorStore contract, VectorStoreQuery, MetadataFilters, SimpleVectorStore
Storage .../core/storage/ StorageContext bundling docstore, index store, vector stores and graph store; fsspec persistence
Query side .../core/retrievers/, postprocessor/, response_synthesizers/, query_engine/ Retrieval, fusion, reranking, synthesis and engine composition
Agents .../core/agent/workflow/ FunctionAgent, ReActAgent, CodeActAgent, multi-agent AgentWorkflow
Evaluation .../core/evaluation/ LLM-judge evaluators, retrieval metrics, BatchEvalRunner
Integrations llama-index-integrations/ Provider packages (LLMs, embeddings, vector stores, readers and more)

How a request flows

Take the canonical five lines: index = VectorStoreIndex.from_documents(docs) then index.as_query_engine().query("...").

  1. Defaults. Nothing is configured, so Settings resolves lazily. embed_model becomes resolve_embed_model("default"), which imports llama-index-embeddings-openai and validates OPENAI_API_KEY (utils.py). The node parser defaults to a SentenceSplitter (settings.py).
  2. Transform. from_documents records each document’s hash in the docstore and calls run_transformations, which applies each TransformComponent in order. If a cache is given, each step is looked up by a hash of the nodes and the transform (base.py, pipeline.py).
  3. Chunk. SentenceSplitter (1,024 tokens, 200 overlap) splits recursively by paragraph, then by NLTK sentence, then by a punctuation regex, then by spaces, and merges the pieces back up to the chunk size (sentence.py).
  4. Embed and store. _add_nodes_to_index embeds nodes in batches and calls vector_store.add. If the store does not keep text (stores_text=False, as with SimpleVectorStore), the nodes also go into the docstore and index struct without their embeddings (vector_store/base.py).
  5. Build the engine. as_query_engine calls as_retriever and wraps it with RetrieverQueryEngine.from_args, which defaults to ResponseMode.COMPACT (base.py, retriever_query_engine.py).
  6. Retrieve. VectorIndexRetriever._retrieve embeds the query unless the mode is SPARSE or TEXT_SEARCH. It then builds a VectorStoreQuery carrying similarity_top_k, filters, mode, alpha and the hybrid/sparse top-k values (retriever.py). SimpleVectorStore.query applies metadata filters in Python and scores every remaining embedding. It supports default, MMR and learner modes and raises on anything else, including hybrid (simple.py).
  7. Postprocess and synthesize. _query runs the retriever, applies each node postprocessor in order, and calls the synthesizer inside a QUERY callback event (retriever_query_engine.py). CompactAndRefine repacks the chunks to fill the context window, then runs the refine loop: answer with the first packed chunk, refine with each next one (compact_and_refine.py).
  8. Return. The Response (or StreamingResponse when streaming=True) carries the answer text plus source_nodes, the retrieved nodes with scores.

Key components

Nodes, transformations and ingestion

Every ingestion step is a TransformComponent that maps a list of nodes to a list of nodes: splitters, metadata extractors (title, summary, keywords, questions answered) and embedding models alike. IngestionPipeline.run adds three production features on top of run_transformations: a transformation cache, docstore-based deduplication with UPSERTS, DUPLICATES_ONLY or UPSERTS_AND_DELETE strategies, and a spawn multiprocessing pool when num_workers > 1. It writes embedded nodes straight to the attached vector store (pipeline.py).

Vector store contract

BasePydanticVectorStore asks for add, delete and query (plus async versions), and declares stores_text and is_embedding_query. All retrieval options travel in a single VectorStoreQuery. Query modes include DEFAULT, SPARSE, HYBRID, TEXT_SEARCH, SEMANTIC_HYBRID, MMR and learner modes (types.py). Filters support AND, OR and NOT. Each integration decides which modes and filter operators it honours, so “hybrid search in LlamaIndex” really means hybrid search in a particular backend.

Retrieval composition

Beyond the vector retriever, core provides QueryFusionRetriever. It generates num_queries=4 rewrites with an LLM, runs every sub-retriever on every query, and fuses the results with reciprocal rank, relative-score, distance-based or simple fusion (fusion_retriever.py). Core also has router and recursive retrievers and auto-merging retrieval. BM25 is not in core; it is the llama-index-retrievers-bm25 integration. Rerankers are node postprocessors: LLMRerank, RankGPT, SentenceTransformerRerank, plus similarity cutoffs, recency and metadata replacement.

Response synthesizers

BaseSynthesizer.synthesize emits instrumentation events, short-circuits on empty input, and dispatches to get_response with node content rendered at MetadataMode.LLM (base.py). The modes are refine, compact (the default), tree_summarize, simple_summarize, accumulate, generation, no_text and context_only. They trade LLM calls against context use. compact is a sensible default; refine costs one call per chunk.

Agents on workflows

BaseWorkflowAgent is a Workflow subclass whose @step methods form the loop. init_run loads memory, setup_agent adds the system prompt and state, run_agent_step calls the subclass’s take_step, parse_agent_output enforces max_iterations and routes to tool calls, call_tool runs them, and aggregate_tool_results loops back (base_agent.py). FunctionAgent uses native tool calling, ReActAgent parses text, and CodeActAgent executes generated code. AgentWorkflow hands control between named agents. RAG is plugged in by wrapping a query engine as a QueryEngineTool.

Extending it

  • New provider. Subclass LLM/CustomLLM, BaseEmbedding, BasePydanticVectorStore, BaseReader or BaseRetriever, then publish it as a llama-index-<kind>-<name> package with its own pyproject.toml. The llama-dev tool in the repo manages and tests these packages.
  • Custom transformation. Implement TransformComponent.__call__(nodes) and drop it into transformations=[...] or an IngestionPipeline.
  • Postprocessors and synthesizers. Subclass BaseNodePostprocessor._postprocess_nodes or pass custom PromptTemplates (text_qa_template, refine_template) to the query engine.
  • Workflows. Write your own Workflow with typed events and @step methods for pipelines that do not fit the query-engine mould.
  • Observability. Register span and event handlers on the instrumentation dispatcher, or install one of the callback/observability integrations (OpenTelemetry, Arize Phoenix, Langfuse and others).

Running it

  • Install. pip install llama-index pulls core plus the OpenAI LLM and embedding integrations; file readers, vector stores and other providers are installed one package at a time. pip install llama-index-core gives the bare abstractions. Core requires Pydantic v2, NLTK, tiktoken, SQLAlchemy, networkx, numpy and fsspec, and the workflows package.
  • Persistence. storage_context.persist(persist_dir=...) writes the docstore, index store and SimpleVectorStore as JSON through fsspec, so local disk, S3 and GCS all work. Reload with load_index_from_storage. Anything larger belongs in an external vector store integration.
  • Services. None are required beyond model access. The defaults call OpenAI, and setting Settings.llm and Settings.embed_model to local integrations (Ollama, HuggingFace) makes it fully offline.
  • No server. Core has no HTTP server or UI. The old CLI is now the separate llama-index-cli package.

Strengths and caveats

  • Strength: coverage. No other framework has as many maintained connectors. Swapping a vector store or LLM is usually a one-line change because the contracts are narrow.
  • Strength: composable query side. Retrievers, postprocessors and synthesizers are independent pieces, and fusion, routing, sub-questions and auto-merging are built in.
  • Strength: ingestion is production-aware. Caching, docstore upserts and multiprocessing come with the pipeline.
  • Caveat: the in-memory store is a toy. SimpleVectorStore scores every embedding in a Python loop and has no hybrid mode. Use it for prototypes only.
  • Caveat: implicit globals. Settings silently falls back to OpenAI for both LLM and embeddings, so a missing configuration shows up as an API-key error deep in a call stack.
  • Caveat: capability depends on the backend. Query modes, filter operators and async support vary between vector store integrations, and the core types do not tell you which ones a given store supports.
  • Caveat: many layers and some legacy. Callbacks and instrumentation coexist, the workflow engine lives in another package, and old ServiceContext-era names still appear, so tracing behaviour through the stack takes effort.

Sources: code at 81f0e06, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the RAG engines questions

Each answer was drafted by a code-reading agent at commit 81f0e06. Its citations were checked mechanically. Compare with the other rag engines →

How are documents parsed and chunked?

answered

Documents enter through readers which implement BaseReader or BasePydanticReader (llama_index/core/readers/base.py:19-48). Core file readers (in llama-index-integrations/readers/file/) support PDF, DOCX, EPUB, HTML, Markdown, CSV/Excel, IPYNB, PPTX, RTF, XML, images, and video/audio. Table-specific readers (PandasCSVReader, PandasExcelReader) parse tabular data. The UnstructuredReader wraps the unstructured.io library for OCR and layout parsing but is an integration, not built into core. No built-in OCR or table-extraction logic exists in the core library — those require external integrations.

Chunking is done by node parsers (subclasses of NodeParser at llama_index/core/node_parser/interface.py:50-68). The primary chunkers are:

  • SentenceSplitter (llama_index/core/node_parser/text/sentence.py:34-100): default chunk_size=1024 tokens, chunk_overlap=200 tokens, prefers sentence/paragraph boundaries, uses regex [^,.;。?!]+[,.;。?!]?|[,.;。?!] as secondary splitter.
  • TokenTextSplitter (llama_index/core/node_parser/text/token.py:22-84): splits by token count (using a configurable tokenizer), falls back from separator to to character-level splitting.
  • SemanticSplitterNodeParser (llama_index/core/node_parser/text/semantic_splitter.py:35-78): groups sentences by embedding similarity — embeds sentence windows and breaks at a configurable percentile threshold of cosine dissimilarity.
  • CodeSplitter (llama_index/core/node_parser/text/code.py:19-30): language-aware AST-based splitting for code files.

Chunking runs in the ingestion pipeline (llama_index/core/ingestion/pipeline.py:72-112) where run_transformations() applies a sequence of TransformComponent instances with caching via IngestionCache. The pipeline supports parallelism via ProcessPoolExecutor.

How are embeddings and indexes built and stored?

answered

Embedding models are accessed through the BaseEmbedding abstract class (llama_index/core/base/embeddings/base.py:72-80). The resolve_embed_model() function (llama_index/core/embeddings/utils.py:30-50) resolves a string name (e.g., "default" → OpenAI's text-embedding-ada-002) or a LangChain embedding into a BaseEmbedding instance. The VectorStoreIndex._get_node_with_embedding() embeds nodes in batches (llama_index/core/indices/vector_store/base.py:126-148).

Vector stores implement BasePydanticVectorStore (llama_index/core/vector_stores/types.py). The built-in SimpleVectorStore (llama_index/core/vector_stores/simple.py:64-80) keeps an in-memory dict of node_id → embedding and persists to JSON via fsspec. Dozens of production vector store integrations exist as separate packages (Chroma, Pinecone, Qdrant, Weaviate, FAISS, Elasticsearch, Milvus, etc.). The StorageContext (llama_index/core/storage/storage_context.py:52-72) holds the vector store, document store, index store, and graph store.

Hybrid/keyword indexing: The VectorStoreQueryMode enum (llama_index/core/vector_stores/types.py:45-60) defines HYBRID, SPARSE, and SEMANTIC_HYBRID modes. The VectorIndexRetriever passes an alpha parameter to the vector store to weight dense vs sparse scores (llama_index/core/indices/vector_store/retrievers/retriever.py:42-73). Support for hybrid search depends on the underlying vector store. Separately, BaseKeywordTableIndex (llama_index/core/indices/keyword_table/base.py:43-97) uses an LLM to extract keywords from each chunk and builds an inverted mapping of keyword → node IDs — a form of keyword-based sparse retrieval.

Metadata is stored in each node's metadata dict and persisted to the vector store. MetadataFilters (llama_index/core/vector_stores/types.py:142-200) with operators (==, >, <, IN, ANY, TEXT_MATCH, etc.) support filtering at query time.

How is retrieval performed?

answered

Retrieval is orchestrated by the RetrieverQueryEngine (llama_index/core/query_engine/retriever_query_engine.py:25-56), which calls a retriever, applies node postprocessors, then synthesizes a response.

Dense retrieval: VectorIndexRetriever (llama_index/core/indices/vector_store/retrievers/retriever.py:24-80) embeds the query using the configured embed model, builds a VectorStoreQuery with the embedding, similarity_top_k, filters, and mode, then delegates to the vector store's query() method. The vector store returns VectorStoreQueryResult with nodes, similarities, and IDs.

Sparse and hybrid: VectorStoreQueryMode.SPARSE skips the embedding step and sends the raw query string. VectorStoreQueryMode.HYBRID sends both the embedding and query string, with alpha controlling the dense/sparse blend. The KeywordTableIndex provides an alternative LLM-keyword-based sparse retrieval.

Query rewriting/decomposition: QueryFusionRetriever (llama_index/core/retrievers/fusion_retriever.py:33-70) uses an LLM to generate multiple query variants from the original, retrieves from each variant across a set of sub-retrievers, then fuses results via reciprocal rank fusion, relative score, or simple re-ranking. SubQuestionQueryEngine (llama_index/core/query_engine/sub_question_query_engine.py:37-60) breaks a complex query into sub-questions, dispatches each to a different query engine tool, then synthesizes the final answer.

Reranking: Postprocessors run after retrieval. LLMRerank (llama_index/core/postprocessor/llm_rerank.py:23-65) uses an LLM to select the top-N most relevant nodes from candidate batches. SentenceTransformerRerank (llama_index/core/postprocessor/sbert_rerank.py:12-56) uses a cross-encoder model. Other postprocessors include SimilarityPostprocessor, KeywordNodePostprocessor, and MetadataReplacementPostProcessor.

Filters: MetadataFilters (llama_index/core/vector_stores/types.py:142-200) supports AND/OR/NOT conditions with operators including ==, !=, >, <, IN, ANY, ALL, TEXT_MATCH.

How are answers generated and grounded?

answered

Prompt assembly: The BaseSynthesizer (llama_index/core/response_synthesizers/base.py:64-107) holds a _text_qa_template (default DEFAULT_TEXT_QA_PROMPT) and _refine_template. The QA prompt receives context_str (concatenated node text at MetadataMode.LLM) and query_str. For chat models, parallel _chat_content_qa_template and _chat_content_refine_template variants use ChatMessage blocks.

Response modes (llama_index/core/response_synthesizers/type.py:4-58): REFINE iterates through nodes, building an initial answer then refining with each subsequent chunk. COMPACT (via CompactAndRefine at response_synthesizers/compact_and_refine.py:13-60) packs multiple chunks into the context window before refining. TREE_SUMMARIZE builds a bottom-up summary tree. GENERATION (llama_index/core/response_synthesizers/generation.py:32-50) ignores context entirely. SIMPLE_SUMMARIZE stuffs all text into one prompt. NO_TEXT and CONTEXT_ONLY return nodes as-is.

Source attribution: Every Response object (llama_index/core/base/response/schema.py:14-42) carries source_nodes: List[NodeWithScore] alongside the response text. Response.get_formatted_sources() renders truncated node content with node IDs. Source nodes flow from retrieval through synthesis unchanged.

Streaming: When streaming=True is passed, calls return StreamingResponse or AsyncStreamingResponse (llama_index/core/base/response/schema.py:108-200), which wrap a generator/async generator yielding tokens. The response also carries source_nodes for display.

Agentic/multi-step answering: The AgentWorkflow and FunctionAgent (llama_index/core/agent/) support tool-calling agents that can iterate, reflect, and call retrieval tools. SubQuestionQueryEngine decomposes queries into parallel sub-questions. ReActAgent implements the ReAct pattern with chat formatters and output parsers.

How is quality evaluated or observed?

answered

LlamaIndex has a substantial built-in evaluation suite in llama_index.core.evaluation. Evaluators are subclasses of BaseEvaluator and include:

  • FaithfulnessEvaluator (llama_index/core/evaluation/faithfulness.py:16-44): LLM-judge that checks if each claim in the response is supported by the context, answering YES/NO with example few-shot prompts.
  • RelevancyEvaluator / AnswerRelevancyEvaluator: measure how relevant the response is to the query and context.
  • CorrectnessEvaluator: compares response against a reference answer.
  • ContextRelevancyEvaluator: evaluates whether retrieved context is relevant to the query.
  • SemanticSimilarityEvaluator: embedding-based similarity between response and reference.
  • PairwiseComparisonEvaluator: A/B comparison of two responses.
  • GuidelineEvaluator: checks responses against custom guidelines.

Retrieval-specific evaluation: RetrieverEvaluator evaluates retrieval quality with HitRate and MRR from llama_index/core/evaluation/retrieval/metrics.py.

Batch evaluation: BatchEvalRunner (llama_index/core/evaluation/batch_runner.py:75-80) runs multiple evaluators across queries in parallel with semaphore concurrency control and exponential-backoff retries.

Observability/tracing: Two-layer system — (1) CallbackManager (llama_index/core/callbacks/base.py:28-80) with CBEventType events (CHUNKING, NODE_PARSING, EMBEDDING, LLM, QUERY, RETRIEVE, SYNTHESIZE, etc.) and contextvar-based trace stacking. (2) The instrumentation package (llama_index/core/instrumentation/__init__.py) provides typed events (QueryStartEvent, RetrievalEndEvent, SynthesizeStartEvent, etc.) with @dispatcher.span decorators on key methods. NullEventHandler and NullSpanHandler are defaults; users register custom handlers for production monitoring.

How is it deployed and operated?

answered

LlamaIndex is distributed as a Python library (PyPI package llama-index-core), not as a managed service. There is no built-in UI or API server in core — the CLI (llama_index/core/command_line/) has been deprecated and moved to its own llama-index-cli package. The chat_ui directory exists but is experimental. Usage is purely programmatic: users write Python scripts or integrate into web frameworks (FastAPI, Flask, etc.).

Infrastructure: All state is handled through pluggable storage backends. StorageContext (llama_index/core/storage/storage_context.py:52-72) combines a BaseDocumentStore (default SimpleDocumentStore — in-memory dicts persisted to JSON), BaseIndexStore (index metadata), vector stores (in-memory or external), and GraphStore. Persistence uses fsspec — files can be saved to local disk or cloud storage (S3, GCS). Each store has persist()/load() methods for round-trip serialization.

Production vector stores: Users plug in external vector database integrations (Chroma, Pinecone, Qdrant, Weaviate, Elasticsearch, FAISS, etc.) via pip install llama-index-vector-stores-<name>. These handle indexing, sharding, and replication outside of LlamaIndex. The SimpleVectorStore supports only single-node in-memory use.

Scaling: The ingestion pipeline (llama_index/core/ingestion/pipeline.py:72-112) supports multiprocessing and async parallelism via ProcessPoolExecutor. The library supports async throughout (parallel embedding, parallel retrieval, parallel sub-question execution via run_jobs).

Multi-tenancy: Not built-in. Each StorageContext instance represents one index. Users manage isolation at the application layer. There is no built-in authentication, rate limiting, or user session management.

LLM configuration: Global Settings dataclass (llama_index/core/settings.py:18-30) provides lazy-initialized defaults for llm, embed_model, callback_manager, tokenizer, node_parser, and transformations. These can be overridden per-index or per-query engine instance.