deepset-ai/haystack
Python framework for typed component pipelines and a tool-calling Agent, with an in-memory store and integrations for the rest.
Overview
Haystack is deepset’s Python framework for building LLM applications out of small, typed building blocks. Everything is a component: a class with a run() method whose parameters and declared outputs become typed input and output sockets. A Pipeline wires components into a graph, which may branch and loop, and runs them when their inputs are ready. Retrieval-augmented generation is the canonical use, but the same machinery runs routing, extraction, evaluation and agent workflows.
The core package (haystack-ai) is deliberately narrow. It ships the pipeline engine, the Document/ChatMessage data classes, converters and splitters, OpenAI and Azure generators and embedders, an InMemoryDocumentStore with brute-force vector search and in-house BM25, retrievers for that store, rankers, routers, joiners and evaluators. Vector databases, other model providers and local embedding models live in separate integration packages. The core only defines the DocumentStore protocol they implement.
At this commit the centre of gravity has moved towards agents. The Agent component owns tool calling (the old standalone ToolInvoker and the separate AsyncPipeline class are gone). It comes with a hook system (haystack/hooks/) for budgets, context compaction, human-in-the-loop confirmation and tool-result offloading. Pipelines themselves can be paused and resumed through breakpoints and snapshots.
Architecture
flowchart LR
U["Your code"] --> P["Pipeline.run / run_async"]
P --> Q["Priority queue scheduler"]
Q --> C["Components (@component)"]
C --> CONV["Converters + splitters"]
C --> EMB["Embedders"]
C --> RET["Retrievers"]
RET --> DS["DocumentStore protocol"]
DS --> MEM["InMemoryDocumentStore"]
DS --> EXT["Integration stores"]
C --> PB["ChatPromptBuilder (Jinja2)"]
PB --> GEN["Chat generators"]
C --> AG["Agent"]
AG --> GEN
AG --> TOOLS["Tools / Toolsets"]
AG --> HOOKS["Hooks"]
P --> TR["Tracer spans"]
| Component | Path | Role |
|---|---|---|
| Component model | haystack/core/component/ |
@component decorator, ComponentMeta socket parsing, output_types |
| Pipeline engine | haystack/core/pipeline/pipeline.py, base.py |
Graph building (add_component, connect), sync and async schedulers, YAML/JSON (de)serialisation |
| Breakpoints | haystack/core/pipeline/breakpoint.py |
Pipeline snapshots for pause, resume and post-mortem |
| SuperComponent | haystack/core/super_component/ |
Wraps a whole pipeline as one component with mapped inputs and outputs |
| Data classes | haystack/dataclasses/ |
Document, ChatMessage, StreamingChunk, ToolCall and more |
| Document stores | haystack/document_stores/ |
DocumentStore protocol, DuplicatePolicy, InMemoryDocumentStore |
| Components | haystack/components/ |
Converters, preprocessors, embedders, retrievers, rankers, builders, generators, routers, joiners, evaluators, agents |
| Tools | haystack/tools/ |
Tool, @tool, ComponentTool, PipelineTool, Toolset, SearchableToolset |
| Hooks | haystack/hooks/ |
Agent lifecycle hooks: budget, compaction, human-in-the-loop, tool-result offloading |
| Tracing and telemetry | haystack/tracing/, haystack/telemetry/ |
Span API with pluggable tracers; anonymous PostHog usage telemetry |
How a request flows
Take a basic RAG pipeline: TextEmbedder -> InMemoryEmbeddingRetriever -> ChatPromptBuilder -> OpenAIChatGenerator, run with pipe.run({"embedder": {"text": q}, "prompt_builder": {"query": q}}).
- Build. Each component class went through
@component, which re-creates it withComponentMetaand registers it for deserialisation (component.py). On instantiation the metaclass reads therunsignature into input sockets and theoutput_typesinto output sockets, and records whether a realrun_asynccoroutine exists (component.py).Pipeline.connect("retriever.documents", "prompt_builder.documents")type-checks the two sockets and adds an edge to anetworkxmultigraph. - Prepare.
run()callswarm_up()on every component (this is where API clients are created), normalises the input data, validates it, and sorts component names so that scheduling is deterministic (pipeline.py). - Schedule.
_fill_queuegives each component a priority:HIGHEST,READY,DEFERorBLOCKED, depending on which sockets have values and whether all predecessors have run (base.py). The main loop pops the next runnable component and breaks ties between deferred components by topological order. It stops when only blocked components remain (pipeline.py). A per-component visit cap (max_runs_per_component) guards loops (base.py). - Run a component.
_run_componentopens a tracing span, deep-copies the inputs, callsinstance.run(**inputs), wraps any exception inPipelineRuntimeError, and checks that the output is a mapping with the declared keys (pipeline.py). - Retrieve. The retriever calls
InMemoryDocumentStore.embedding_retrieval. It filters documents in Python, stacks all embeddings into a NumPy matrix, computes a dot product (or cosine after normalisation), sorts, and returns the topkwith scores (document_store.py). There is no ANN index; that is what the integration stores are for. - Prompt and generate.
ChatPromptBuilderrenders a Jinja2 template in a sandboxed environment intoChatMessages.OpenAIChatGenerator.runmergesgeneration_kwargs, calls the chat completions API, assembles the stream throughstreaming_callbackif one is set, and returnsreplies(openai.py). - Route outputs.
_write_component_outputspushes each output to its receivers’ sockets. Unconsumed outputs become pipeline outputs. The queue is refilled when it goes stale, and the loop continues until nothing can run (pipeline.py). - On failure. If a component raises (or a
Breakpointfires), the loop builds a pipeline snapshot of inputs, visit counts and partial outputs, saves it, and attaches it to the exception. You can then pass it back torun(pipeline_snapshot=...)to resume (pipeline.py).
Key components
The scheduler
Haystack does not do a single topological pass. It re-evaluates readiness after every component, which is what makes cycles (for example a generator, validator and retry loop) and conditional branches possible. Multiple senders into one list-typed socket are promoted to a lazy variadic socket. In the sync path their values arrive ordered by sender name, not by connect() order (base.py). The async path (run_async, run_async_generator, stream) uses the same priorities but schedules ready components as asyncio tasks under a semaphore (concurrency_limit=4 by default). It consumes inputs before creating each task so that siblings do not race (pipeline.py). Sync-only components are offloaded with asyncio.to_thread.
Agent
Agent wraps any chat generator plus tools. Each step runs the BEFORE_LLM hooks, calls the generator, and stops if there are no tools or the reply has no tool calls. Otherwise it runs the BEFORE_TOOL hooks, executes the pending calls (up to tool_concurrency_limit=4 in parallel), runs the AFTER_TOOL hooks, and checks the tool-based exit_conditions (agent.py). The outer loop caps at max_agent_steps=100 and records an exit_reason (agent.py). State is a typed dict (state_schema) that tools can read from and write to through inputs_from_state/outputs_to_state. The hook points are before_run, before_llm, before_tool, after_tool and after_run. Because a hook can rewrite pending tool calls or set stop_run, approval flows and token budgets live outside the loop instead of inside it.
Document store and retrieval
The DocumentStore protocol has only four required operations: count_documents, filter_documents, write_documents and delete_documents (protocol.py). Retrieval methods are store-specific and are reached through matching retriever components. InMemoryDocumentStore keeps documents in process-global dicts keyed by index name (or per-instance with shared=False). It updates BM25 statistics incrementally on write and scores with its own BM25L/Okapi/Plus implementations, with no third-party BM25 dependency (document_store.py). Hybrid search is a composition: MultiRetriever runs named retrievers in a thread pool and merges with reciprocal rank fusion, which is always used when a global top_k is set (multi_retriever.py). Rankers (LLMRanker, LostInTheMiddleRanker, meta-field rankers) and query expansion (QueryExpander plus the multi-query retrievers) are separate components.
Generation and answers
OpenAIChatGenerator creates its client lazily in warm_up(), so a component can be serialised without live credentials (openai.py). Sibling generators cover the OpenAI Responses API, Azure, and a FallbackChatGenerator. Citations are not automatic, but AnswerBuilder takes a reference_pattern such as \[(\d+)\], maps the numbers back to the input documents, and marks them referenced (answer_builder.py).
Serialisation
Pipeline.dumps()/loads() round-trips a pipeline through YAML. Every component provides to_dict/from_dict, and secrets are stored as Secret.from_env_var references, never as values. Loading is guarded: haystack/core/serialization_security.py restricts which modules may be imported during deserialisation unless you pass allowed_modules= or set HAYSTACK_DESERIALIZATION_ALLOWLIST.
Extending it
- Custom component. Decorate a class with
@component, declare@component.output_types(...)onrun, and optionally add a realasync def run_async. Dynamic sockets are set withcomponent.set_input_types/set_output_types. - Custom document store. Implement the four protocol methods plus your own retrieval methods, and ship a retriever component that calls them. This is how every external store integration is built.
- Tools. Use
@tool/create_tool_from_functionfor plain functions (the schema comes from type hints and docstrings),ComponentToolfor any component,PipelineToolfor a whole pipeline, andToolset/SearchableToolsetfor groups that can grow during a run. - Hooks. Implement the
Hookprotocol (run(state), optionalrun_async) and register it under a hook point on theAgent. - Sub-pipelines.
SuperComponentexposes a pipeline as one component with explicit input and output mappings. - Tracing. Subclass
Tracerand callenable_tracing(). Content tags such as prompts and documents are only recorded when content tracing is turned on.
Running it
- Install.
pip install haystack-ai(Python 3.10 or newer). Core dependencies are small:openai,networkx,pydantic,Jinja2,numpy,jsonschema,posthogand a few utilities. Converters and splitters pull optional extras (pypdf,tiktoken, NLTK and others) lazily and tell you what to install. - No server. The repository has no HTTP server or CLI. You embed pipelines in your own app. deepset’s separate Hayhooks project serves them as REST, MCP or OpenAI-compatible endpoints.
docker/only builds a base image with Haystack installed. - Services. For anything beyond a prototype you need an LLM provider key and an external document store integration. The in-memory store keeps everything in RAM and searches linearly.
- Telemetry. Anonymous PostHog telemetry is on by default. Set
HAYSTACK_TELEMETRY_ENABLED=Falseto turn it off.
Strengths and caveats
- Strength: explicit, typed wiring. Socket types are checked at
connect()time, and the whole graph serialises to YAML, so pipelines are reviewable artefacts rather than opaque chains. - Strength: loops and branches are first-class. The priority-queue scheduler handles cycles, conditional routers and variadic joins, with a visit cap as a safety net.
- Strength: debuggable failures. Breakpoints and automatic snapshots on error let you resume a long pipeline or agent run from the failing component.
- Strength: modern agent plumbing. Hooks for budgets, compaction and human approval are part of the core, not left to the user.
- Caveat: the core is thin on backends. Only OpenAI/Azure models and an in-memory, brute-force store ship in the box. Real deployments depend on integration packages that are versioned separately.
- Caveat: scheduling subtleties. Variadic input order differs between sync (alphabetical by sender) and async (no guarantee). Components that wait on optional inputs can end up deferred or blocked in ways that take reading
base.pyto understand. - Caveat: telemetry on by default. Usage events go to PostHog unless you opt out.
- Caveat: no serving layer. Scaling, auth and multi-tenancy are entirely the host application’s or Hayhooks’ job.
Sources: code at e03f7c6, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit e03f7c6. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredParsing – Haystack provides a converter component for each supported format. The PyPDFToDocument component (haystack/components/converters/pypdf.py:50-287) wraps PyPDF2/PyPDF and offers two extraction modes: PLAIN mode extracts text as it appears in the PDF stream, while LAYOUT mode (experimental, lines 22–23) preserves the rendered layout by considering font height, spacing weight, and vertical whitespace. For images embedded in PDFs, PDFToImageContent (haystack/components/converters/image/pdf_to_image.py:18-80) renders PDF pages to ImageContent objects that can be sent directly to multimodal LLMs — this is the closest Haystack comes to OCR, though no dedicated OCR engine (e.g., Tesseract) is built in. Tables from spreadsheets are handled by XLSXToDocument (haystack/components/converters/xlsx.py:26-59), which uses pandas/openpyxl to read Excel sheets and outputs each sheet as a Document in CSV or Markdown format. Other converters include CSVToDocument, DOCXToDocument, HTMLToDocument, MarkdownToDocument, PPTXToDocument, TextFileToDocument, and JSONConverter — all declared in haystack/components/converters/__init__.py:5-25. A MultiFileConverter (haystack/components/converters/multi_file_converter.py:37-58) dispatches files by MIME type to the appropriate converter via FileTypeRouter.
Chunking – The DocumentSplitter (haystack/components/preprocessors/document_splitter.py:28-58) is the primary chunking component. It supports eight split_by modes: word (default, 200 tokens), passage (double newline), page (form-feed character), period, line, sentence (NLTK-based), token (tiktoken), and function (user-supplied callable). Parameters include split_length, split_overlap (up to split_length-1), and split_threshold (minimum units to avoid orphan chunks). When split_by="token", the encoder defaults to o200k_base (the tokenizer for current OpenAI models). The respect_sentence_boundary option, when combined with split_by="word", uses NLTK to keep whole sentences together. The HierarchicalDocumentSplitter, PythonCodeSplitter, and RecursiveSplitter are also available in the preprocessors directory.
How are embeddings and indexes built and stored?
answeredEmbedding models – Embedders are Haystack components that call external APIs or local models. The OpenAITextEmbedder (haystack/components/embedders/openai_text_embedder.py:16-254) and OpenAIDocumentEmbedder embed text and documents using OpenAI's embeddings.create endpoint (default: text-embedding-3-small). They accept dimensions, prefix/suffix, and full OpenAI client configuration. API clients are lazily initialized in warm_up() (line 116), not __init__, keeping components serializable without credentials. The embedders directory also includes AzureTextEmbedder/AzureDocumentEmbedder variants and mock embedders for testing.
Document stores – The only built-in store is InMemoryDocumentStore (haystack/document_stores/in_memory/document_store.py:98), which stores documents in process-global dictionaries. All other stores — Elasticsearch, OpenSearch, Pinecone, Qdrant, Weaviate, Chroma, Astra DB, Pgvector — are provided as separate integration packages (e.g., elasticsearch-haystack, pinecone-haystack). The DocumentStore protocol (haystack/document_stores/types/protocol.py:11-136) defines the interface all stores implement: write_documents, filter_documents, delete_documents, count_documents. The protocol also defines methods for embedding_retrieval and bm25_retrieval that stores implement with their backend's native similarity search.
Hybrid/keyword (BM25) indexes – The in-memory store maintains a BM25 index incrementally during write_documents (lines 543-553). Each document is tokenized via _tokenize_bm25 (custom CJK-aware regex, line 58), term frequencies are stored in BM25DocumentStats dataclasses, and an IDF vocabulary counter and average document length are updated incrementally. The store supports multiple BM25 variants via the rank_bm25 library (configurable via bm25_algorithm). Plugging the BM25 retriever and embedding retriever into the MultiRetriever component (haystack/components/retrievers/multi_retriever.py:19-120) enables hybrid search with reciprocal_rank_fusion joining mode.
Metadata – All document stores support structured metadata filtering via a nested comparison/logic filter syntax. The DuplicatePolicy enum (haystack/document_stores/types/policy.py) controls write behavior: NONE/FAIL, SKIP, OVERWRITE.
How is retrieval performed?
answeredDense/embedding retrieval – InMemoryEmbeddingRetriever (haystack/components/retrievers/in_memory/embedding_retriever.py:12-237) accepts a query_embedding (list of floats, pre-computed by a TextEmbedder), passes it to document_store.embedding_retrieval(), which computes cosine or dot-product similarity scores via np.dot (haystack/document_stores/in_memory/document_store.py:853-918). Scores can be scaled to [0,1] via expit for dot-product or (score+1)/2 for cosine. The TextEmbeddingRetriever component (haystack/components/retrievers/text_embedding_retriever.py:14-70) composes a TextEmbedder and an EmbeddingRetriever into a single run(query: str) component.
Sparse/BM25 retrieval – InMemoryBM25Retriever (haystack/components/retrievers/in_memory/bm25_retriever.py:13-197) accepts a raw query: str and calls document_store.bm25_retrieval(), which scores documents using the rank_bm25 library and returns top-k results.
Hybrid search – The MultiRetriever (haystack/components/retrievers/multi_retriever.py:19-120) runs multiple named retrievers concurrently via a thread pool, then merges results using concatenate + deduplication or reciprocal_rank_fusion (RRF). A top_k parameter applies consistent global ranking after RRF.
Reranking – The LLMRanker (haystack/components/rankers/llm_ranker.py:52-307) takes query + candidate documents, builds a prompt with instructions and JSON-schema response format, calls a ChatGenerator (default: gpt-4.1-mini with temperature=0.0 and structured output), and re-orders documents by the LLM's relevance assessment. It deduplicates input documents before ranking and falls back to input order on failure. The LostInTheMiddleRanker, MetaFieldRanker, and MetaFieldGroupingRanker are also available.
Query rewriting – QueryExpander (haystack/components/query/query_expander.py:56-60) uses an LLM to generate semantically similar query variants, which can then be fed to MultiQueryTextRetriever (haystack/components/retrievers/multi_query_text_retriever.py:16-82) for parallel retrieval across all query variants with deduplicated results.
Filters – All retrievers accept filters dictionaries that support comparison operators (==, !=, <, <=, >, >=, in, not in) and logic operators (AND, OR, NOT). The FilterPolicy (REPLACE/MERGE) controls how init-time and run-time filters combine.
How are answers generated and grounded?
answeredPrompt assembly – Haystack uses ChatPromptBuilder and PromptBuilder (haystack/components/builders/) with Jinja2 templates to assemble prompts. The OpenAIChatGenerator (haystack/components/generators/chat/openai.py:57-409) works with ChatMessage objects (with user, assistant, system, tool roles). Generation parameters like temperature, max_completion_tokens, response_format, and stop sequences can be set at init or merged per-key at run time.
Streaming – The OpenAIChatGenerator accepts a streaming_callback: StreamingCallbackT parameter both in __init__ and at run() time (line 348). When a callback is provided, the generator streams ChatCompletionChunk objects through the callback function via OpenAI's Stream response, assembling the full message from chunks (_handle_stream_response). The StreamingChunk dataclass carries each delta: content, delta, and tool_call data.
Tool calling / agentic answering – The Agent component (haystack/components/agents/agent.py:831-921) wraps a ChatGenerator with tools and runs a multi-step loop: call LLM, execute requested tools, append results, repeat until the model replies with text (no tool calls) or max_agent_steps is reached. Each step runs the LLM once and executes all tool calls the model requested in that turn. The Agent returns full conversation history, token usage, step count, and an exit_reason ("text", "max_agent_steps", tool name, etc.). State management with user-defined schemas, hooks (BEFORE_RUN, BEFORE_TOOL, AFTER_TOOL, etc.), and Toolset subclassing provide extensible agent orchestration.
Citations / source attribution – There is no built-in citation-formatting mechanism in the generator components themselves. Source attribution must be implemented by the application pipeline — typically by passing retrieved documents alongside the query in the prompt template. The FaithfulnessEvaluator can verify whether generated statements are inferable from provided contexts.
Structured output – The OpenAIChatGenerator supports OpenAI's response_format parameter including Pydantic model or JSON schema validation (OpenAI Structured Outputs).
How is quality evaluated or observed?
answeredBuilt-in evaluators – Haystack ships a set of component-based evaluators under haystack/components/evaluators/. The FaithfulnessEvaluator (haystack/components/evaluators/faithfulness.py:53-269) uses an LLM to split a generated answer into statements and check each against the provided contexts, producing a score of 0–1 representing the proportion of supported statements. The ContextRelevanceEvaluator similarly scores how relevant retrieved contexts are to a question. Other built-in metrics include AnswerExactMatchEvaluator, DocumentMAPEvaluator (Mean Average Precision), DocumentMRREvaluator (Mean Reciprocal Rank), DocumentNDCGEvaluator (Normalized Discounted Cumulative Gain), DocumentRecallEvaluator, and SASEvaluator (Semantic Answer Similarity). The LLMEvaluator (haystack/components/evaluators/llm_evaluator.py:25-80) is the base class for LLM-as-judge evaluations: user-defined instructions, input/output schema, and few-shot examples guide an LLM (default: OpenAI in JSON mode) to produce structured scores. All evaluators follow the Haystack @component pattern and can be plugged into Pipelines.
Tracing / observability – Haystack provides an abstract Tracer interface (haystack/tracing/tracer.py:77-94) with trace() context managers and Span objects for instrumenting operations. A built-in LoggingTracer (haystack/tracing/logging_tracer.py:35-50) logs operation names and tags with customizable ANSI coloring. Third-party tracers (e.g., OpenTelemetry) can be integrated by subclassing Tracer and calling enable_tracing(). Content-level tracing (queries, documents, answers) is opt-in via the HAYSTACK_CONTENT_TRACING_ENABLED environment variable.
Telemetry – Haystack uses PostHog for anonymous usage analytics (haystack/telemetry/__init__.py), tracking when pipelines are executed.
What is absent – There is no built-in evaluation result database or dashboard; the EvalRunResult dataclass (haystack/evaluation/eval_run_result.py) provides a lightweight container, but users are expected to aggregate and log results themselves.
How is it deployed and operated?
answeredLibrary, not a service. Haystack is a Python framework published as the haystack-ai package on PyPI and conda-forge (pyproject.toml:5). It is installed via pip install haystack-ai and embedded into a user's own application — it does not ship a built-in HTTP server, CLI daemon, or web UI. There are no server classes, no FastAPI/Flask endpoints, and no __main__ entry point in the repository.
Deployment via Hayhooks (separate project). The README (README.md:76) and dedicated documentation page (docs-website/docs/development/hayhooks.mdx) point to Hayhooks, a separate open-source project (deepset-ai/hayhooks) that wraps Haystack pipelines as REST APIs, MCP servers, or OpenAI-compatible chat-completion endpoints. A user writes a PipelineWrapper subclass with a run_api method, deploys a YAML pipeline definition, and Hayhooks serves it via uvicorn on port 1416. Hayhooks also supports file uploads, streaming, CORS, SSL, and a Chainlit UI.
API Surface. The core API is the Pipeline class (haystack/core/pipeline/pipeline.py:118) with synchronous run() and asynchronous run_async() / run_async_generator() methods. Pipelines are directed graphs of @component-decorated classes (haystack/core/component/component.py:8). Components connect via input/output sockets, and the pipeline orchestrates execution in topological order with support for branches, loops, and conditional routing. The Pipeline.loads() / Pipeline.dumps() methods (haystack/core/pipeline/base.py:294-370) serialize pipelines to YAML or JSON for portability.
Required Infrastructure. Haystack itself requires only Python 3.10+ and the haystack-ai package with its dependencies (networkx, pydantic, httpx, openai SDK, etc. — pyproject.toml:44-63). However, a useful Haystack application typically also needs external LLM provider APIs (OpenAI, Anthropic, Mistral, etc.), a vector database (external integrations like Weaviate, Pinecone, Qdrant, Milvus, or the bundled InMemoryDocumentStore), and optionally embedding services, file converters, and rankers. API credentials are handled via the Secret.from_env_var() pattern (haystack/utils/auth.py:182-215) which loads tokens from environment variables.
Scaling and Multi-tenancy. Haystack provides no built-in load balancing, horizontal scaling, or multi-tenant isolation. The async pipeline support (Pipeline.run_async(), haystack/core/pipeline/pipeline.py:1096) enables non-blocking execution suitable for concurrent requests in an async web server (e.g., FastAPI + uvicorn), but scaling out, queuing, tenant-aware routing, and resource isolation are the responsibility of the deployment layer. The official Docker image (docker/Dockerfile.base) packages Haystack on python:3.12-slim and is built via Docker Bake for linux/amd64 and linux/arm64 (docker/docker-bake.hcl:38-39), intended as a base image for derived containers rather than a standalone service.
Telemetry. Haystack includes anonymous usage telemetry sent to PostHog (haystack/telemetry/_telemetry.py:6), opt-out via the HAYSTACK_TELEMETRY_ENABLED environment variable, reporting pipeline runs and system specs — a concern for deployment environments that require strict data sovereignty.