UKGovernmentBEIS/inspect_ai
Python eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.
Overview
Inspect is the evaluation framework from the UK AI Security Institute. An eval is plain Python. A @task function returns a Task that binds a dataset of Samples, a solver (a prompt pipeline or a tool-using agent) and one or more scorers. inspect eval runs every sample concurrently and writes a detailed log. Each sample’s log holds every model call, tool call, sandbox command and score as a typed event, and inspect view browses the logs.
Of the tools in this category, it is the general one. It handles classic benchmarks (many are packaged in the separate inspect_evals repository), and the same primitives cover agentic evals with Docker sandboxes, approval policies and human-in-the-loop solvers. For LLM app teams, the solver slot is the key. It doesn’t have to be a prompt: it can be a react() agent with tools, or your own agent code. Through the agent bridge, an agent written against the OpenAI or Anthropic SDK can be run under Inspect unchanged, by pointing its model name at inspect. So you can build a regression suite for your actual app logic with reference answers, model-graded rubrics and code checks.
Inspect does not observe production. It has no OpenTelemetry ingest and no trace store. It runs evals offline over datasets you define. Compared with lm-evaluation-harness, it is chat- and agent-first rather than loglikelihood-first. Compared with garak, it gives you a framework to write adversarial tasks, not a ready catalogue of attacks.
Architecture
flowchart LR
CLI["inspect eval / eval()"] --> EA["eval_async: resolve tasks, models"]
EA --> RUN["eval_run -> task_run"]
RUN --> SCH["Sample scheduler (anyio)"]
SCH --> SAMP["task_run_sample"]
SAMP --> SBX["Sandbox (docker, local)"]
SAMP --> PLAN["Plan / Solver / Agent"]
PLAN --> MODEL["Model.generate -> ModelAPI provider"]
PLAN --> TOOLS["Tools (bash, python, web...)"]
SAMP --> SCR["Scorers"]
SCR --> MET["Metrics + epoch reducers"]
SAMP --> EVT["Events transcript"]
EVT --> LOG["Recorder: .eval zip / .json"]
LOG --> VIEW["inspect view"]
| Component | Path | Role |
|---|---|---|
| Entry points | src/inspect_ai/_eval/eval.py, evalset.py |
eval(), eval_async, eval_retry, and eval_set with resume and retries |
| Task runner | src/inspect_ai/_eval/task/run.py |
task_run, task_run_sample: sandbox setup, plan execution, limits, scoring, logging |
| Task | src/inspect_ai/_eval/task/task.py |
Task: dataset, setup, solver, scorer, metrics, limits, epochs, sandbox, approval |
| Datasets | src/inspect_ai/dataset/ |
Sample, MemoryDataset, CSV/JSON/HF loaders, field mapping |
| Solvers | src/inspect_ai/solver/ |
generate, chain_of_thought, multiple_choice, self_critique, use_tools, Plan |
| Agents | src/inspect_ai/agent/ |
react(), handoffs, agent_bridge, human agent |
| Models | src/inspect_ai/model/ |
Model/ModelAPI, 30+ providers, retries, caching, connection pools |
| Tools | src/inspect_ai/tool/ |
Bash, Python, text editor, web search, browser, computer use, MCP |
| Sandboxes | src/inspect_ai/util/_sandbox/ |
Docker/Compose and local sandboxes, plus a registry for others |
| Scorers | src/inspect_ai/scorer/ |
match, includes, choice, f1, math, model_graded_qa, metrics, reducers |
| Logs | src/inspect_ai/log/ |
EvalLog schema, recorders, read/edit APIs |
How a request flows
Take inspect eval my_evals.py@qa --model openai/gpt-4o:
- Resolve.
eval()(eval.py) wrapseval_async, which loads the@task, resolves the model andmodel_roles(for example agrader), sets up the recorder, and callseval_runfor each batch of tasks (eval.py). - Build the task. The
Taskconstructor takes a dataset (or a dynamicSampleSource),setup,solver(defaultgenerate()),scorer,metrics,epochs,sandbox, approval and review policies, and message, token, time and cost limits (task.py). - Schedule samples.
task_runstarts the sandboxes and fans samples out under a concurrency limit. Model calls share an adaptive connection pool per provider. - Solve. For each sample attempt,
_task_run_sample_attemptbuilds aTaskState, starts the working-limit monitor, and runsstate = await plan(state, generate)inside asolversspan (run.py, L2829-L2832). A solver is justasync (TaskState, Generate) -> TaskState(_solver.py). Tool calls, sandbox exec and model calls are each recorded as events. - Score. Each scorer is awaited as
scorer(state, Target(sample.target))in its own span. Scorers are not allowed to mutatestate.scores(run.py, _scorer.py). - Reduce. With multiple epochs, reducers (
mean,mode,pass_at…) combine per-epoch scores. Metrics such asaccuracy,stderrand bootstrap or Wilson CIs then aggregate across samples. - Log. Samples stream into the recorder. The default
.evalformat is a zip archive of JSON entries, which can be read incrementally (eval.py recorder).read_eval_logloads it back in Python, andinspect viewserves the web viewer.
Key components
Solvers and agents
Solvers compose. chain() and Plan stack prompt engineering steps before generate(). react() is a configurable ReAct loop with tools, a submit tool, retry on refusal, context compaction, and approval or review hooks (_react.py). Agents can hand off to each other and can be used as tools. The human_agent solver lets a person do the task inside the sandbox, which gives you human baselines.
Agent bridge
agent_bridge() patches the OpenAI and Anthropic client libraries so requests for a model named inspect (or inspect/...) go through Inspect’s model layer. The eval’s config wins over the agent’s request settings (bridge.py). That puts a third-party or in-house agent under Inspect’s logging, limits and scorers with no rewrite. It is the most direct way to eval an existing LLM app here.
Model-graded scoring
model_graded_qa and model_graded_fact prompt a grader for GRADE: C/P/I. The grader comes from model, or else the grader role, or else the model under test. Pass a list of models and each grades independently, combined by majority or mode. Off-menu verdicts count as parse failures, not wrong answers (_model.py). The answer text is sanitised so dataset-controlled [BEGIN DATA] markers can’t inject into the grading prompt.
Sandboxes and safety rails
Task(sandbox="docker") gives each sample its own container or Compose project, and tools like bash() and python() execute there. Per-sample message, token, time, working and cost limits plus approval policies keep runaway agents bounded.
Logs and events
Every sample stores a full event transcript: model, tool, sandbox, score, state and span events. eval_retry and eval_set reuse completed samples from logs, so crashed or partial runs resume rather than restart. Hooks (@hooks) fire on run, task and sample lifecycle and before each model call, which is where you would forward data to an external tracker.
Extending it
- Tasks, solvers, scorers, metrics, tools, agents: decorate with
@task,@solver,@scorer,@metric,@tool,@agent. Registered objects can be referenced by name from the CLI. - Model providers and sandboxes:
@modelapiand@sandboxenvregister new backends. Packages expose them through theinspect_aientry-point group, loaded on demand (entrypoints.py). - Hooks: subclass
Hooksand register with@hooksfor telemetry or custom persistence. - Post-hoc:
inspect scorere-scores an existing log with new scorers without re-running the model.
Running it
pip install inspect-aiplus the provider SDK you need (openai,anthropic…). Python 3.10+. Docker is needed only for sandboxed tasks.inspect eval file.py@task --model provider/model, with flags for--limit,--epochs,--max-connectionsand--model-role grader=....inspect eval-setruns suites with retries into one log directory.inspect viewstarts the log viewer locally. Logs can live on local disk or S3 through fsspec.- Model responses can be cached with an expiry policy, and
inspect cachemanages the cache.
Strengths and caveats
- Strength: one model for benchmarks, app checks and agents. The same Task/Solver/Scorer split covers multiple choice, rubric grading and long agent runs in sandboxes.
- Strength: transcripts. Per-sample event logs make it practical to debug why a score happened, not just what it was.
- Strength: evaluate your own agent. The agent bridge and
react()cover both bring-your-own and built-in agent setups. - Strength: careful scoring. Grader panels, parse-failure handling, epochs with reducers and several CI estimators.
- Caveat: offline only. No production trace ingest or OTel. Online monitoring needs another tool, with Inspect hooks as the glue.
- Caveat: no attack library. Red-teaming means writing your own tasks, or pulling them from
inspect_evals. - Caveat: large surface. The runner (
task/run.pyalone is about 4,000 lines) has many knobs, and behaviour around retries, limits and checkpoints takes reading to understand.
Sources: code at aa20052, verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit aa20052. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredInspect provides three tiers of scoring: per-sample scorers, per-task metrics, and epoch reducers.
Built-in scorers (src/inspect_ai/scorer/) each return a Score object with a value (str/int/float/bool/list/dict). They include:
- match — string matching at begin/end/any/exact locations, with case-insensitive and numeric options
- includes — substring containment check
- exact — normalized exact-match for QA
- f1 — SQuAD-style F1 token overlap
- choice — multiple-choice letter grading with unshuffle support
- math — symbolic math via SymPy/LaTeX parsing with sandboxed expression validation, timeout, and complexity limits
- answer — extracts ANSWER:-prefixed answers by letter/word/line pattern
- model_graded_qa / model_graded_fact — LLM-as-a-judge using configurable templates and grader models (covered under llm-judge)
- perplexity — scores via prompt logprobs NLL
- cascade — chains scorers cheapest-first, short-circuiting when threshold met
- multi_scorer — runs multiple scorers in parallel and reduces via majority/mode/mean
- precomputed_scores — loads externally computed scores from JSON/JSONL by sample ID
Metrics (src/inspect_ai/scorer/_metrics/) aggregate per-sample scores into eval-level values:
- accuracy — proportion correct, with CORRECT/INCORRECT/PARTIAL/NOANSWER sentinels and a pluggable ValueToFloat converter
- mean / std / var — arithmetic mean, sample standard deviation, variance
- stderr / bootstrap_stderr / ci / ci_wilson — standard error (plain or clustered by sample metadata), confidence intervals via t-distribution or percentile bootstrap, Wilson score for binary proportions
- frequency / categorical — categorical score distribution (counts or proportions), with StrEnum integration and zero-fill for unobserved categories
- aggregate — extracts one key from dict-valued scores and runs another metric on it
- grouped — partitions scores by metadata key and applies a metric per group
- krippendorff_alpha — inter-rater agreement across multiple judges
Custom scorers are created with the @scorer decorator. The Scorer protocol requires an async callable (state: TaskState, target: Target) -> Score | None. Custom metrics use @metric and accept list[SampleScore] -> Value. Metrics can declare a scores mode: "auto" (reduced), "reduced", or "unreduced" (per-epoch).
Score reducers (src/inspect_ai/scorer/_reducer/) aggregate multi-epoch samples: mean_score, median_score, mode_score, max_score, majority_score, pass_at, pass_k, at_least, collect_score.
How is LLM-as-a-judge implemented?
answeredLLM-as-a-judge is implemented via model_graded_qa and model_graded_fact scorers (src/inspect_ai/scorer/_model.py).
Judge prompts and rubrics. Both scorers accept a template with {question}, {answer}, {criterion}, and {instructions} variables plus any sample metadata keys. The default instructions ask the grader to emit a GRADE: C / GRADE: I (or GRADE: P for partial credit) verdict after step-by-step reasoning. The model_scoring_prompt() function formats the template, neutralizes structural delimiters ([BEGIN DATA]/[END DATA]) in dataset-controlled inputs to prevent prompt injection, and preserves media attachments.
Structured output via regex. The grade is extracted by regex with a leading greedy .* (DOTALL) that ensures the last GRADE: X wins — crucial for injection robustness. A permissive variant captures any word when the default instructions are used, then validates against the grades actually offered. Off-menu verdicts are treated as parse failures (grader_failed), not incorrect answers.
Multi-sample / consensus. When model is a list of models (or model_role binds to a list), each model grades independently via multi_scorer() which runs them concurrently with tg_collect. By default (reducer="majority"), a grade needs >50% of graders to agree. reducer="mode" picks the most common grade with tie-breaking by model order.
Judge model choice. Controlled via the model parameter (takes precedence) or model_role (default "grader", resolved from eval's model_roles). Falls back to the model under evaluation when no judge model is specified.
Calibration / bias controls. No explicit bias-calibration is provided — the approach relies on prompt engineering and majority-vote reducers. Anti-injection measures include structural delimiter neutralization and grade-pattern validation.
How are test datasets and cases defined, generated and versioned?
answeredDataset definition. A Dataset is a Sequence[Sample] (src/inspect_ai/dataset/_dataset.py:144-221). Each Sample has input, target, choices, id, metadata, sandbox, files, setup, and checkpoint fields. The input can be a string or a list of ChatMessage objects.
File formats and sources (src/inspect_ai/dataset/_sources/):
- hf_dataset — loads from Hugging Face datasets with retry logic for transient errors, field mapping via
FieldSpec, and support for splits/shuffling - json_dataset — JSON array of sample dicts
- csv_dataset — CSV file with configurable field mapping
- file_dataset — generic file (auto-detects JSON/CSV/JSONL by extension)
- example_dataset — inline sample generation via a function
Synthetic data generation is done in user code by writing Python functions that return MemoryDataset(list[Sample]) or by implementing a SampleSource protocol. There is no built-in synthetic data DSL beyond the Sample constructor.
Versioning. Inspect does not version datasets natively — versioning is the data provider's responsibility (e.g. Hugging Face dataset revisions pinned by revision= in hf_dataset()). The eval_set system provides log-based versioning through eval set IDs and manifest files.
Benchmark task registry. Tasks are registered with @task decorator and can be discovered via inspect list or list_tasks(). The Task object binds a Dataset, a Solver/Plan, and optional Scorers, metrics, and config.
Golden sets. There is no built-in "golden set" concept — tasks specify their own dataset at construction time. The SampleSource protocol allows dynamic sample feeding. The precomputed_scores scorer attaches externally computed labels to existing logs by sample ID.
How are evals executed and reported?
answeredRunners. Evals are launched via eval() / eval_async() (src/inspect_ai/_eval/eval.py). The top-level orchestrator eval_set() (src/inspect_ai/_eval/evalset.py) manages multiple tasks with shared configuration, sample dispatch, retry logic, and logging. Inside a task, task_run() sets up sandbox lifecycle and prepares task options. Samples are fanned out by SampleScheduler, a live fanout loop that accepts new samples mid-run.
Parallelism. parallel= controls concurrent samples, running in an anyio TaskGroup. Model-level parallelism uses adaptive connection pools (DEFAULT_MAX_CONNECTIONS / DEFAULT_MAX_CONNECTIONS_BATCH). max_tasks controls parallel task execution within eval sets.
Caching. Model outputs are cached with configurable TTL. The inspect cache CLI command (src/inspect_ai/_cli/cache.py) provides clear, prune, list, and size operations.
CI integration. Standard Python library approach — pip install inspect-ai && python your_eval.py. The CLI supports --detach for long-running evals, ACP for cross-machine dispatch, and JSON output modes.
Result storage. Logs use a compact binary .eval format (or JSON) via the Recorder hierarchy (src/inspect_ai/log/_recorders/). FileRecorder writes to local/remote filesystems via fsspec. BufferSampleStore provides crash-recovery through a SQLite buffer.
Comparison / regression. Eval results are structured as EvalResults with EvalScore objects per scorer and EvalMetric dicts per metric (src/inspect_ai/log/_log.py:799-904). The inspect view command starts a web dashboard for browsing logs and comparing runs. There is no built-in regression test framework — users compare metrics programmatically from returned EvalLog objects.
How are traces or production data captured and linked to evaluations?
answeredSDK instrumentation is provided through the events module and the hooks system, not through OpenTelemetry.
Event system (src/inspect_ai/event/). During sample execution, a timeline of typed events is recorded: ModelEvent, ToolEvent, ScoreEvent, SandboxEvent, StateEvent, StepEvent, SubtaskEvent, InputEvent, ApprovalEvent, ErrorEvent, StoreEvent, SpanEvent, and more. These are organized into an EventTree modeling the hierarchical sample run. The Timeline builds a linearized, filterable view.
Hooks (src/inspect_ai/hooks/). Lifecycle hooks fire at eval_set_start/end, run_start/end, task_start/end, sample_start/init/attempt_start/attempt_end/end/scoring, and before_model_generate. Hooks are passed to eval().
Trace logging (src/inspect_ai/_util/trace.py). A custom TRACE log level captures HTTP requests, model calls, and long-running actions. The inspect trace CLI (src/inspect_ai/_cli/trace.py) lists, dumps, filters traces by action type, and surfaces anomalies.
No OpenTelemetry. There is no opentelemetry dependency or exporter. The tracing system is file-based JSONL.
Online vs offline evals. "Online" evals run via eval() and write to log files in real-time. "Offline" analysis reads logged EvalLog objects via read_eval_log(), recomputes metrics via recompute_metrics(), or edits scores via edit_score().
Feedback and annotation. The precomputed_scores scorer attaches externally computed scores by sample ID. edit_score() modifies logged scores with provenance tracking. There is no built-in annotation UI.
Does it support red-teaming or safety testing, and how?
answeredInspect provides infrastructure useful for red-teaming but has no built-in adversarial probes, attack plugins, jailbreak harnesses, or dedicated red-teaming tooling.
What exists:
- Scanner support (
src/inspect_ai/_eval/task/scan.py). Aneval_setcan attachScannerConfigobjects that run per-sample analysis via the optionalinspect_scoutpackage. Designed for post-hoc scanning of completed transcripts, not adversarial generation. - Review system (
src/inspect_ai/review/). TheReviewerprotocol intercepts tool call results mid-execution and cancontinue,terminate, orescalate. - Sandbox environments (
src/inspect_ai/util/_sandbox/). Docker-based sandboxes with diagnostics and egress controls provide isolation for running untrusted model outputs. - Model-level safety settings. Several model provider modules expose safety/abuse-detection parameters (Anthropic, OpenAI, Google, Grok) but as passthroughs, not a managed red-teaming feature.
What is absent:
- No built-in adversarial attack library (no prompt injection generators, no jailbreak test suites, no fuzzing tools)
- No dedicated red-team evaluation harness or report format
- No vulnerability reporting workflow or CVE tracking
- No built-in harmfulness classifiers or refusal detectors
- No automated red-teaming loop that generates increasingly adversarial inputs
How users do red-teaming today: By writing custom @scorer functions that test for specific failure modes, using the Tool system to build adversarial tool environments, and leveraging the Plan/Solver system to construct multi-turn probe sequences. The framework is flexible enough for ad-hoc red-teaming through these primitives, but provides no turnkey solution.