LLMs Technical Reviews

UKGovernmentBEIS/inspect_ai

Python eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.

GitHub ↗★ 2.9kPythonMITcommit aa20052 · 2026-10-06homepage ↗

Overview

Inspect is the evaluation framework from the UK AI Security Institute. An eval is plain Python. A @task function returns a Task that binds a dataset of Samples, a solver (a prompt pipeline or a tool-using agent) and one or more scorers. inspect eval runs every sample concurrently and writes a detailed log. Each sample’s log holds every model call, tool call, sandbox command and score as a typed event, and inspect view browses the logs.

Of the tools in this category, it is the general one. It handles classic benchmarks (many are packaged in the separate inspect_evals repository), and the same primitives cover agentic evals with Docker sandboxes, approval policies and human-in-the-loop solvers. For LLM app teams, the solver slot is the key. It doesn’t have to be a prompt: it can be a react() agent with tools, or your own agent code. Through the agent bridge, an agent written against the OpenAI or Anthropic SDK can be run under Inspect unchanged, by pointing its model name at inspect. So you can build a regression suite for your actual app logic with reference answers, model-graded rubrics and code checks.

Inspect does not observe production. It has no OpenTelemetry ingest and no trace store. It runs evals offline over datasets you define. Compared with lm-evaluation-harness, it is chat- and agent-first rather than loglikelihood-first. Compared with garak, it gives you a framework to write adversarial tasks, not a ready catalogue of attacks.

Architecture

flowchart LR
  CLI["inspect eval / eval()"] --> EA["eval_async: resolve tasks, models"]
  EA --> RUN["eval_run -> task_run"]
  RUN --> SCH["Sample scheduler (anyio)"]
  SCH --> SAMP["task_run_sample"]
  SAMP --> SBX["Sandbox (docker, local)"]
  SAMP --> PLAN["Plan / Solver / Agent"]
  PLAN --> MODEL["Model.generate -> ModelAPI provider"]
  PLAN --> TOOLS["Tools (bash, python, web...)"]
  SAMP --> SCR["Scorers"]
  SCR --> MET["Metrics + epoch reducers"]
  SAMP --> EVT["Events transcript"]
  EVT --> LOG["Recorder: .eval zip / .json"]
  LOG --> VIEW["inspect view"]
Component Path Role
Entry points src/inspect_ai/_eval/eval.py, evalset.py eval(), eval_async, eval_retry, and eval_set with resume and retries
Task runner src/inspect_ai/_eval/task/run.py task_run, task_run_sample: sandbox setup, plan execution, limits, scoring, logging
Task src/inspect_ai/_eval/task/task.py Task: dataset, setup, solver, scorer, metrics, limits, epochs, sandbox, approval
Datasets src/inspect_ai/dataset/ Sample, MemoryDataset, CSV/JSON/HF loaders, field mapping
Solvers src/inspect_ai/solver/ generate, chain_of_thought, multiple_choice, self_critique, use_tools, Plan
Agents src/inspect_ai/agent/ react(), handoffs, agent_bridge, human agent
Models src/inspect_ai/model/ Model/ModelAPI, 30+ providers, retries, caching, connection pools
Tools src/inspect_ai/tool/ Bash, Python, text editor, web search, browser, computer use, MCP
Sandboxes src/inspect_ai/util/_sandbox/ Docker/Compose and local sandboxes, plus a registry for others
Scorers src/inspect_ai/scorer/ match, includes, choice, f1, math, model_graded_qa, metrics, reducers
Logs src/inspect_ai/log/ EvalLog schema, recorders, read/edit APIs

How a request flows

Take inspect eval my_evals.py@qa --model openai/gpt-4o:

  1. Resolve. eval() (eval.py) wraps eval_async, which loads the @task, resolves the model and model_roles (for example a grader), sets up the recorder, and calls eval_run for each batch of tasks (eval.py).
  2. Build the task. The Task constructor takes a dataset (or a dynamic SampleSource), setup, solver (default generate()), scorer, metrics, epochs, sandbox, approval and review policies, and message, token, time and cost limits (task.py).
  3. Schedule samples. task_run starts the sandboxes and fans samples out under a concurrency limit. Model calls share an adaptive connection pool per provider.
  4. Solve. For each sample attempt, _task_run_sample_attempt builds a TaskState, starts the working-limit monitor, and runs state = await plan(state, generate) inside a solvers span (run.py, L2829-L2832). A solver is just async (TaskState, Generate) -> TaskState (_solver.py). Tool calls, sandbox exec and model calls are each recorded as events.
  5. Score. Each scorer is awaited as scorer(state, Target(sample.target)) in its own span. Scorers are not allowed to mutate state.scores (run.py, _scorer.py).
  6. Reduce. With multiple epochs, reducers (mean, mode, pass_at…) combine per-epoch scores. Metrics such as accuracy, stderr and bootstrap or Wilson CIs then aggregate across samples.
  7. Log. Samples stream into the recorder. The default .eval format is a zip archive of JSON entries, which can be read incrementally (eval.py recorder). read_eval_log loads it back in Python, and inspect view serves the web viewer.

Key components

Solvers and agents

Solvers compose. chain() and Plan stack prompt engineering steps before generate(). react() is a configurable ReAct loop with tools, a submit tool, retry on refusal, context compaction, and approval or review hooks (_react.py). Agents can hand off to each other and can be used as tools. The human_agent solver lets a person do the task inside the sandbox, which gives you human baselines.

Agent bridge

agent_bridge() patches the OpenAI and Anthropic client libraries so requests for a model named inspect (or inspect/...) go through Inspect’s model layer. The eval’s config wins over the agent’s request settings (bridge.py). That puts a third-party or in-house agent under Inspect’s logging, limits and scorers with no rewrite. It is the most direct way to eval an existing LLM app here.

Model-graded scoring

model_graded_qa and model_graded_fact prompt a grader for GRADE: C/P/I. The grader comes from model, or else the grader role, or else the model under test. Pass a list of models and each grades independently, combined by majority or mode. Off-menu verdicts count as parse failures, not wrong answers (_model.py). The answer text is sanitised so dataset-controlled [BEGIN DATA] markers can’t inject into the grading prompt.

Sandboxes and safety rails

Task(sandbox="docker") gives each sample its own container or Compose project, and tools like bash() and python() execute there. Per-sample message, token, time, working and cost limits plus approval policies keep runaway agents bounded.

Logs and events

Every sample stores a full event transcript: model, tool, sandbox, score, state and span events. eval_retry and eval_set reuse completed samples from logs, so crashed or partial runs resume rather than restart. Hooks (@hooks) fire on run, task and sample lifecycle and before each model call, which is where you would forward data to an external tracker.

Extending it

  • Tasks, solvers, scorers, metrics, tools, agents: decorate with @task, @solver, @scorer, @metric, @tool, @agent. Registered objects can be referenced by name from the CLI.
  • Model providers and sandboxes: @modelapi and @sandboxenv register new backends. Packages expose them through the inspect_ai entry-point group, loaded on demand (entrypoints.py).
  • Hooks: subclass Hooks and register with @hooks for telemetry or custom persistence.
  • Post-hoc: inspect score re-scores an existing log with new scorers without re-running the model.

Running it

  • pip install inspect-ai plus the provider SDK you need (openai, anthropic…). Python 3.10+. Docker is needed only for sandboxed tasks.
  • inspect eval file.py@task --model provider/model, with flags for --limit, --epochs, --max-connections and --model-role grader=.... inspect eval-set runs suites with retries into one log directory.
  • inspect view starts the log viewer locally. Logs can live on local disk or S3 through fsspec.
  • Model responses can be cached with an expiry policy, and inspect cache manages the cache.

Strengths and caveats

  • Strength: one model for benchmarks, app checks and agents. The same Task/Solver/Scorer split covers multiple choice, rubric grading and long agent runs in sandboxes.
  • Strength: transcripts. Per-sample event logs make it practical to debug why a score happened, not just what it was.
  • Strength: evaluate your own agent. The agent bridge and react() cover both bring-your-own and built-in agent setups.
  • Strength: careful scoring. Grader panels, parse-failure handling, epochs with reducers and several CI estimators.
  • Caveat: offline only. No production trace ingest or OTel. Online monitoring needs another tool, with Inspect hooks as the glue.
  • Caveat: no attack library. Red-teaming means writing your own tasks, or pulling them from inspect_evals.
  • Caveat: large surface. The runner (task/run.py alone is about 4,000 lines) has many knobs, and behaviour around retries, limits and checkpoints takes reading to understand.

Sources: code at aa20052, verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit aa20052. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Inspect provides three tiers of scoring: per-sample scorers, per-task metrics, and epoch reducers.

Built-in scorers (src/inspect_ai/scorer/) each return a Score object with a value (str/int/float/bool/list/dict). They include:

  • match — string matching at begin/end/any/exact locations, with case-insensitive and numeric options
  • includes — substring containment check
  • exact — normalized exact-match for QA
  • f1 — SQuAD-style F1 token overlap
  • choice — multiple-choice letter grading with unshuffle support
  • math — symbolic math via SymPy/LaTeX parsing with sandboxed expression validation, timeout, and complexity limits
  • answer — extracts ANSWER:-prefixed answers by letter/word/line pattern
  • model_graded_qa / model_graded_fact — LLM-as-a-judge using configurable templates and grader models (covered under llm-judge)
  • perplexity — scores via prompt logprobs NLL
  • cascade — chains scorers cheapest-first, short-circuiting when threshold met
  • multi_scorer — runs multiple scorers in parallel and reduces via majority/mode/mean
  • precomputed_scores — loads externally computed scores from JSON/JSONL by sample ID

Metrics (src/inspect_ai/scorer/_metrics/) aggregate per-sample scores into eval-level values:

  • accuracy — proportion correct, with CORRECT/INCORRECT/PARTIAL/NOANSWER sentinels and a pluggable ValueToFloat converter
  • mean / std / var — arithmetic mean, sample standard deviation, variance
  • stderr / bootstrap_stderr / ci / ci_wilson — standard error (plain or clustered by sample metadata), confidence intervals via t-distribution or percentile bootstrap, Wilson score for binary proportions
  • frequency / categorical — categorical score distribution (counts or proportions), with StrEnum integration and zero-fill for unobserved categories
  • aggregate — extracts one key from dict-valued scores and runs another metric on it
  • grouped — partitions scores by metadata key and applies a metric per group
  • krippendorff_alpha — inter-rater agreement across multiple judges

Custom scorers are created with the @scorer decorator. The Scorer protocol requires an async callable (state: TaskState, target: Target) -> Score | None. Custom metrics use @metric and accept list[SampleScore] -> Value. Metrics can declare a scores mode: "auto" (reduced), "reduced", or "unreduced" (per-epoch).

Score reducers (src/inspect_ai/scorer/_reducer/) aggregate multi-epoch samples: mean_score, median_score, mode_score, max_score, majority_score, pass_at, pass_k, at_least, collect_score.

How is LLM-as-a-judge implemented?

answered

LLM-as-a-judge is implemented via model_graded_qa and model_graded_fact scorers (src/inspect_ai/scorer/_model.py).

Judge prompts and rubrics. Both scorers accept a template with {question}, {answer}, {criterion}, and {instructions} variables plus any sample metadata keys. The default instructions ask the grader to emit a GRADE: C / GRADE: I (or GRADE: P for partial credit) verdict after step-by-step reasoning. The model_scoring_prompt() function formats the template, neutralizes structural delimiters ([BEGIN DATA]/[END DATA]) in dataset-controlled inputs to prevent prompt injection, and preserves media attachments.

Structured output via regex. The grade is extracted by regex with a leading greedy .* (DOTALL) that ensures the last GRADE: X wins — crucial for injection robustness. A permissive variant captures any word when the default instructions are used, then validates against the grades actually offered. Off-menu verdicts are treated as parse failures (grader_failed), not incorrect answers.

Multi-sample / consensus. When model is a list of models (or model_role binds to a list), each model grades independently via multi_scorer() which runs them concurrently with tg_collect. By default (reducer="majority"), a grade needs >50% of graders to agree. reducer="mode" picks the most common grade with tie-breaking by model order.

Judge model choice. Controlled via the model parameter (takes precedence) or model_role (default "grader", resolved from eval's model_roles). Falls back to the model under evaluation when no judge model is specified.

Calibration / bias controls. No explicit bias-calibration is provided — the approach relies on prompt engineering and majority-vote reducers. Anti-injection measures include structural delimiter neutralization and grade-pattern validation.

How are test datasets and cases defined, generated and versioned?

answered

Dataset definition. A Dataset is a Sequence[Sample] (src/inspect_ai/dataset/_dataset.py:144-221). Each Sample has input, target, choices, id, metadata, sandbox, files, setup, and checkpoint fields. The input can be a string or a list of ChatMessage objects.

File formats and sources (src/inspect_ai/dataset/_sources/):

  • hf_dataset — loads from Hugging Face datasets with retry logic for transient errors, field mapping via FieldSpec, and support for splits/shuffling
  • json_dataset — JSON array of sample dicts
  • csv_dataset — CSV file with configurable field mapping
  • file_dataset — generic file (auto-detects JSON/CSV/JSONL by extension)
  • example_dataset — inline sample generation via a function

Synthetic data generation is done in user code by writing Python functions that return MemoryDataset(list[Sample]) or by implementing a SampleSource protocol. There is no built-in synthetic data DSL beyond the Sample constructor.

Versioning. Inspect does not version datasets natively — versioning is the data provider's responsibility (e.g. Hugging Face dataset revisions pinned by revision= in hf_dataset()). The eval_set system provides log-based versioning through eval set IDs and manifest files.

Benchmark task registry. Tasks are registered with @task decorator and can be discovered via inspect list or list_tasks(). The Task object binds a Dataset, a Solver/Plan, and optional Scorers, metrics, and config.

Golden sets. There is no built-in "golden set" concept — tasks specify their own dataset at construction time. The SampleSource protocol allows dynamic sample feeding. The precomputed_scores scorer attaches externally computed labels to existing logs by sample ID.

How are evals executed and reported?

answered

Runners. Evals are launched via eval() / eval_async() (src/inspect_ai/_eval/eval.py). The top-level orchestrator eval_set() (src/inspect_ai/_eval/evalset.py) manages multiple tasks with shared configuration, sample dispatch, retry logic, and logging. Inside a task, task_run() sets up sandbox lifecycle and prepares task options. Samples are fanned out by SampleScheduler, a live fanout loop that accepts new samples mid-run.

Parallelism. parallel= controls concurrent samples, running in an anyio TaskGroup. Model-level parallelism uses adaptive connection pools (DEFAULT_MAX_CONNECTIONS / DEFAULT_MAX_CONNECTIONS_BATCH). max_tasks controls parallel task execution within eval sets.

Caching. Model outputs are cached with configurable TTL. The inspect cache CLI command (src/inspect_ai/_cli/cache.py) provides clear, prune, list, and size operations.

CI integration. Standard Python library approach — pip install inspect-ai && python your_eval.py. The CLI supports --detach for long-running evals, ACP for cross-machine dispatch, and JSON output modes.

Result storage. Logs use a compact binary .eval format (or JSON) via the Recorder hierarchy (src/inspect_ai/log/_recorders/). FileRecorder writes to local/remote filesystems via fsspec. BufferSampleStore provides crash-recovery through a SQLite buffer.

Comparison / regression. Eval results are structured as EvalResults with EvalScore objects per scorer and EvalMetric dicts per metric (src/inspect_ai/log/_log.py:799-904). The inspect view command starts a web dashboard for browsing logs and comparing runs. There is no built-in regression test framework — users compare metrics programmatically from returned EvalLog objects.

Editor's note. Correction: --acp-server is not cross-machine dispatch. It exposes a running eval over the Agent Client Protocol so clients (inspect acp, editors) can attach for human-in-the-loop approvals. The .eval log format is a zip archive of JSON entries, not a custom binary format.

How are traces or production data captured and linked to evaluations?

answered

SDK instrumentation is provided through the events module and the hooks system, not through OpenTelemetry.

Event system (src/inspect_ai/event/). During sample execution, a timeline of typed events is recorded: ModelEvent, ToolEvent, ScoreEvent, SandboxEvent, StateEvent, StepEvent, SubtaskEvent, InputEvent, ApprovalEvent, ErrorEvent, StoreEvent, SpanEvent, and more. These are organized into an EventTree modeling the hierarchical sample run. The Timeline builds a linearized, filterable view.

Hooks (src/inspect_ai/hooks/). Lifecycle hooks fire at eval_set_start/end, run_start/end, task_start/end, sample_start/init/attempt_start/attempt_end/end/scoring, and before_model_generate. Hooks are passed to eval().

Trace logging (src/inspect_ai/_util/trace.py). A custom TRACE log level captures HTTP requests, model calls, and long-running actions. The inspect trace CLI (src/inspect_ai/_cli/trace.py) lists, dumps, filters traces by action type, and surfaces anomalies.

No OpenTelemetry. There is no opentelemetry dependency or exporter. The tracing system is file-based JSONL.

Online vs offline evals. "Online" evals run via eval() and write to log files in real-time. "Offline" analysis reads logged EvalLog objects via read_eval_log(), recomputes metrics via recompute_metrics(), or edits scores via edit_score().

Feedback and annotation. The precomputed_scores scorer attaches externally computed scores by sample ID. edit_score() modifies logged scores with provenance tracking. There is no built-in annotation UI.

Does it support red-teaming or safety testing, and how?

answered

Inspect provides infrastructure useful for red-teaming but has no built-in adversarial probes, attack plugins, jailbreak harnesses, or dedicated red-teaming tooling.

What exists:

  • Scanner support (src/inspect_ai/_eval/task/scan.py). An eval_set can attach ScannerConfig objects that run per-sample analysis via the optional inspect_scout package. Designed for post-hoc scanning of completed transcripts, not adversarial generation.
  • Review system (src/inspect_ai/review/). The Reviewer protocol intercepts tool call results mid-execution and can continue, terminate, or escalate.
  • Sandbox environments (src/inspect_ai/util/_sandbox/). Docker-based sandboxes with diagnostics and egress controls provide isolation for running untrusted model outputs.
  • Model-level safety settings. Several model provider modules expose safety/abuse-detection parameters (Anthropic, OpenAI, Google, Grok) but as passthroughs, not a managed red-teaming feature.

What is absent:

  • No built-in adversarial attack library (no prompt injection generators, no jailbreak test suites, no fuzzing tools)
  • No dedicated red-team evaluation harness or report format
  • No vulnerability reporting workflow or CVE tracking
  • No built-in harmfulness classifiers or refusal detectors
  • No automated red-teaming loop that generates increasingly adversarial inputs

How users do red-teaming today: By writing custom @scorer functions that test for specific failure modes, using the Tool system to build adversarial tool environments, and leveraging the Plan/Solver system to construct multi-turn probe sequences. The framework is flexible enough for ad-hoc red-teaming through these primitives, but provides no turnkey solution.