# UKGovernmentBEIS/inspect_ai

> Python eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/UKGovernmentBEIS/inspect_ai (reviewed at commit `aa20052a65b13516f1ee79d10ccceda00c205cc6`, 2026-10-06)
- Stars: 2946 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/inspect_ai/

## Overview

Inspect is the evaluation framework from the UK AI Security Institute. An eval is plain Python. A `@task` function returns a `Task` that binds a dataset of `Sample`s, a **solver** (a prompt pipeline or a tool-using agent) and one or more **scorers**. `inspect eval` runs every sample concurrently and writes a detailed log. Each sample's log holds every model call, tool call, sandbox command and score as a typed event, and `inspect view` browses the logs.

Of the tools in this category, it is the general one. It handles classic benchmarks (many are packaged in the separate `inspect_evals` repository), and the same primitives cover agentic evals with Docker sandboxes, approval policies and human-in-the-loop solvers. For LLM app teams, the solver slot is the key. It doesn't have to be a prompt: it can be a `react()` agent with tools, or your own agent code. Through the **agent bridge**, an agent written against the OpenAI or Anthropic SDK can be run under Inspect unchanged, by pointing its model name at `inspect`. So you can build a regression suite for your actual app logic with reference answers, model-graded rubrics and code checks.

Inspect does not observe production. It has no OpenTelemetry ingest and no trace store. It runs evals offline over datasets you define. Compared with lm-evaluation-harness, it is chat- and agent-first rather than loglikelihood-first. Compared with garak, it gives you a framework to write adversarial tasks, not a ready catalogue of attacks.

## Architecture

```mermaid
flowchart LR
  CLI["inspect eval / eval()"] --> EA["eval_async: resolve tasks, models"]
  EA --> RUN["eval_run -> task_run"]
  RUN --> SCH["Sample scheduler (anyio)"]
  SCH --> SAMP["task_run_sample"]
  SAMP --> SBX["Sandbox (docker, local)"]
  SAMP --> PLAN["Plan / Solver / Agent"]
  PLAN --> MODEL["Model.generate -> ModelAPI provider"]
  PLAN --> TOOLS["Tools (bash, python, web...)"]
  SAMP --> SCR["Scorers"]
  SCR --> MET["Metrics + epoch reducers"]
  SAMP --> EVT["Events transcript"]
  EVT --> LOG["Recorder: .eval zip / .json"]
  LOG --> VIEW["inspect view"]
```

| Component | Path | Role |
|---|---|---|
| Entry points | `src/inspect_ai/_eval/eval.py`, `evalset.py` | `eval()`, `eval_async`, `eval_retry`, and `eval_set` with resume and retries |
| Task runner | `src/inspect_ai/_eval/task/run.py` | `task_run`, `task_run_sample`: sandbox setup, plan execution, limits, scoring, logging |
| Task | `src/inspect_ai/_eval/task/task.py` | `Task`: dataset, setup, solver, scorer, metrics, limits, epochs, sandbox, approval |
| Datasets | `src/inspect_ai/dataset/` | `Sample`, `MemoryDataset`, CSV/JSON/HF loaders, field mapping |
| Solvers | `src/inspect_ai/solver/` | `generate`, `chain_of_thought`, `multiple_choice`, `self_critique`, `use_tools`, `Plan` |
| Agents | `src/inspect_ai/agent/` | `react()`, handoffs, `agent_bridge`, human agent |
| Models | `src/inspect_ai/model/` | `Model`/`ModelAPI`, 30+ providers, retries, caching, connection pools |
| Tools | `src/inspect_ai/tool/` | Bash, Python, text editor, web search, browser, computer use, MCP |
| Sandboxes | `src/inspect_ai/util/_sandbox/` | Docker/Compose and local sandboxes, plus a registry for others |
| Scorers | `src/inspect_ai/scorer/` | `match`, `includes`, `choice`, `f1`, `math`, `model_graded_qa`, metrics, reducers |
| Logs | `src/inspect_ai/log/` | `EvalLog` schema, recorders, read/edit APIs |

## How a request flows

Take `inspect eval my_evals.py@qa --model openai/gpt-4o`:

1. **Resolve.** `eval()` ([eval.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/eval.py#L124-L135)) wraps `eval_async`, which loads the `@task`, resolves the model and `model_roles` (for example a `grader`), sets up the recorder, and calls `eval_run` for each batch of tasks ([eval.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/eval.py#L1080-L1110)).
2. **Build the task.** The `Task` constructor takes a dataset (or a dynamic `SampleSource`), `setup`, `solver` (default `generate()`), `scorer`, `metrics`, `epochs`, `sandbox`, approval and review policies, and message, token, time and cost limits ([task.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/task.py#L81-L130)).
3. **Schedule samples.** `task_run` starts the sandboxes and fans samples out under a concurrency limit. Model calls share an adaptive connection pool per provider.
4. **Solve.** For each sample attempt, `_task_run_sample_attempt` builds a `TaskState`, starts the working-limit monitor, and runs `state = await plan(state, generate)` inside a `solvers` span ([run.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/run.py#L2363-L2380), [L2829-L2832](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/run.py#L2829-L2832)). A solver is just `async (TaskState, Generate) -> TaskState` ([_solver.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/solver/_solver.py#L79-L111)). Tool calls, sandbox exec and model calls are each recorded as events.
5. **Score.** Each scorer is awaited as `scorer(state, Target(sample.target))` in its own span. Scorers are not allowed to mutate `state.scores` ([run.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/run.py#L3070-L3090), [_scorer.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_scorer.py#L35-L64)).
6. **Reduce.** With multiple epochs, reducers (`mean`, `mode`, `pass_at`...) combine per-epoch scores. Metrics such as `accuracy`, `stderr` and bootstrap or Wilson CIs then aggregate across samples.
7. **Log.** Samples stream into the recorder. The default `.eval` format is a zip archive of JSON entries, which can be read incrementally ([eval.py recorder](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/log/_recorders/eval.py#L170-L192)). `read_eval_log` loads it back in Python, and `inspect view` serves the web viewer.

## Key components

### Solvers and agents

Solvers compose. `chain()` and `Plan` stack prompt engineering steps before `generate()`. `react()` is a configurable ReAct loop with tools, a `submit` tool, retry on refusal, context compaction, and approval or review hooks ([_react.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/agent/_react.py#L56-L80)). Agents can hand off to each other and can be used as tools. The `human_agent` solver lets a person do the task inside the sandbox, which gives you human baselines.

### Agent bridge

`agent_bridge()` patches the OpenAI and Anthropic client libraries so requests for a model named `inspect` (or `inspect/...`) go through Inspect's model layer. The eval's config wins over the agent's request settings ([bridge.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/agent/_bridge/bridge.py#L99-L130)). That puts a third-party or in-house agent under Inspect's logging, limits and scorers with no rewrite. It is the most direct way to eval an existing LLM app here.

### Model-graded scoring

`model_graded_qa` and `model_graded_fact` prompt a grader for `GRADE: C/P/I`. The grader comes from `model`, or else the `grader` role, or else the model under test. Pass a list of models and each grades independently, combined by `majority` or `mode`. Off-menu verdicts count as parse failures, not wrong answers ([_model.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L114-L175)). The answer text is sanitised so dataset-controlled `[BEGIN DATA]` markers can't inject into the grading prompt.

### Sandboxes and safety rails

`Task(sandbox="docker")` gives each sample its own container or Compose project, and tools like `bash()` and `python()` execute there. Per-sample message, token, time, working and cost limits plus approval policies keep runaway agents bounded.

### Logs and events

Every sample stores a full event transcript: model, tool, sandbox, score, state and span events. `eval_retry` and `eval_set` reuse completed samples from logs, so crashed or partial runs resume rather than restart. Hooks (`@hooks`) fire on run, task and sample lifecycle and before each model call, which is where you would forward data to an external tracker.

## Extending it

- **Tasks, solvers, scorers, metrics, tools, agents:** decorate with `@task`, `@solver`, `@scorer`, `@metric`, `@tool`, `@agent`. Registered objects can be referenced by name from the CLI.
- **Model providers and sandboxes:** `@modelapi` and `@sandboxenv` register new backends. Packages expose them through the `inspect_ai` entry-point group, loaded on demand ([entrypoints.py](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_util/entrypoints.py#L7-L40)).
- **Hooks:** subclass `Hooks` and register with `@hooks` for telemetry or custom persistence.
- **Post-hoc:** `inspect score` re-scores an existing log with new scorers without re-running the model.

## Running it

- `pip install inspect-ai` plus the provider SDK you need (`openai`, `anthropic`...). Python 3.10+. Docker is needed only for sandboxed tasks.
- `inspect eval file.py@task --model provider/model`, with flags for `--limit`, `--epochs`, `--max-connections` and `--model-role grader=...`. `inspect eval-set` runs suites with retries into one log directory.
- `inspect view` starts the log viewer locally. Logs can live on local disk or S3 through fsspec.
- Model responses can be cached with an expiry policy, and `inspect cache` manages the cache.

## Strengths and caveats

- **Strength: one model for benchmarks, app checks and agents.** The same Task/Solver/Scorer split covers multiple choice, rubric grading and long agent runs in sandboxes.
- **Strength: transcripts.** Per-sample event logs make it practical to debug why a score happened, not just what it was.
- **Strength: evaluate your own agent.** The agent bridge and `react()` cover both bring-your-own and built-in agent setups.
- **Strength: careful scoring.** Grader panels, parse-failure handling, epochs with reducers and several CI estimators.
- **Caveat: offline only.** No production trace ingest or OTel. Online monitoring needs another tool, with Inspect hooks as the glue.
- **Caveat: no attack library.** Red-teaming means writing your own tasks, or pulling them from `inspect_evals`.
- **Caveat: large surface.** The runner (`task/run.py` alone is about 4,000 lines) has many knobs, and behaviour around retries, limits and checkpoints takes reading to understand.

*Sources: code at aa20052, verified Q&A.*

## How UKGovernmentBEIS/inspect_ai answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Inspect provides three tiers of scoring: per-sample scorers, per-task metrics, and epoch reducers.

**Built-in scorers** (`src/inspect_ai/scorer/`) each return a `Score` object with a `value` (str/int/float/bool/list/dict). They include:
- **match** — string matching at begin/end/any/exact locations, with case-insensitive and numeric options
- **includes** — substring containment check
- **exact** — normalized exact-match for QA
- **f1** — SQuAD-style F1 token overlap
- **choice** — multiple-choice letter grading with unshuffle support
- **math** — symbolic math via SymPy/LaTeX parsing with sandboxed expression validation, timeout, and complexity limits
- **answer** — extracts ANSWER:-prefixed answers by letter/word/line pattern
- **model_graded_qa / model_graded_fact** — LLM-as-a-judge using configurable templates and grader models (covered under llm-judge)
- **perplexity** — scores via prompt logprobs NLL
- **cascade** — chains scorers cheapest-first, short-circuiting when threshold met
- **multi_scorer** — runs multiple scorers in parallel and reduces via majority/mode/mean
- **precomputed_scores** — loads externally computed scores from JSON/JSONL by sample ID

**Metrics** (`src/inspect_ai/scorer/_metrics/`) aggregate per-sample scores into eval-level values:
- **accuracy** — proportion correct, with CORRECT/INCORRECT/PARTIAL/NOANSWER sentinels and a pluggable ValueToFloat converter
- **mean / std / var** — arithmetic mean, sample standard deviation, variance
- **stderr / bootstrap_stderr / ci / ci_wilson** — standard error (plain or clustered by sample metadata), confidence intervals via t-distribution or percentile bootstrap, Wilson score for binary proportions
- **frequency / categorical** — categorical score distribution (counts or proportions), with StrEnum integration and zero-fill for unobserved categories
- **aggregate** — extracts one key from dict-valued scores and runs another metric on it
- **grouped** — partitions scores by metadata key and applies a metric per group
- **krippendorff_alpha** — inter-rater agreement across multiple judges

**Custom scorers** are created with the `@scorer` decorator. The `Scorer` protocol requires an async callable `(state: TaskState, target: Target) -> Score | None`. Custom metrics use `@metric` and accept `list[SampleScore] -> Value`. Metrics can declare a `scores` mode: "auto" (reduced), "reduced", or "unreduced" (per-epoch).

**Score reducers** (`src/inspect_ai/scorer/_reducer/`) aggregate multi-epoch samples: `mean_score`, `median_score`, `mode_score`, `max_score`, `majority_score`, `pass_at`, `pass_k`, `at_least`, `collect_score`.


Citations: [src/inspect_ai/scorer/__init__.py:1-123](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/__init__.py#L1-L123) · [src/inspect_ai/scorer/_scorer.py:34-210](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_scorer.py#L34-L210) · [src/inspect_ai/scorer/_metric.py:42-552](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metric.py#L42-L552) · [src/inspect_ai/scorer/_metrics/accuracy.py:14-39](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metrics/accuracy.py#L14-L39) · [src/inspect_ai/scorer/_metrics/std.py:17-604](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metrics/std.py#L17-L604) · [src/inspect_ai/scorer/_cascade.py:13-80](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_cascade.py#L13-L80)

### How is LLM-as-a-judge implemented? (answered)

LLM-as-a-judge is implemented via `model_graded_qa` and `model_graded_fact` scorers (`src/inspect_ai/scorer/_model.py`).

**Judge prompts and rubrics.** Both scorers accept a `template` with `{question}`, `{answer}`, `{criterion}`, and `{instructions}` variables plus any sample metadata keys. The default instructions ask the grader to emit a `GRADE: C` / `GRADE: I` (or `GRADE: P` for partial credit) verdict after step-by-step reasoning. The `model_scoring_prompt()` function formats the template, neutralizes structural delimiters (`[BEGIN DATA]`/`[END DATA]`) in dataset-controlled inputs to prevent prompt injection, and preserves media attachments.

**Structured output via regex.** The grade is extracted by regex with a leading greedy `.*` (DOTALL) that ensures the *last* `GRADE: X` wins — crucial for injection robustness. A permissive variant captures any word when the default instructions are used, then validates against the grades actually offered. Off-menu verdicts are treated as parse failures (`grader_failed`), not incorrect answers.

**Multi-sample / consensus.** When `model` is a list of models (or `model_role` binds to a list), each model grades independently via `multi_scorer()` which runs them concurrently with `tg_collect`. By default (`reducer="majority"`), a grade needs >50% of graders to agree. `reducer="mode"` picks the most common grade with tie-breaking by model order.

**Judge model choice.** Controlled via the `model` parameter (takes precedence) or `model_role` (default `"grader"`, resolved from eval's `model_roles`). Falls back to the model under evaluation when no judge model is specified.

**Calibration / bias controls.** No explicit bias-calibration is provided — the approach relies on prompt engineering and majority-vote reducers. Anti-injection measures include structural delimiter neutralization and grade-pattern validation.


Citations: [src/inspect_ai/scorer/_model.py:34-120](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L34-L120) · [src/inspect_ai/scorer/_model.py:246-362](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L246-L362) · [src/inspect_ai/scorer/_model.py:364-460](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L364-L460) · [src/inspect_ai/scorer/_multi.py:19-57](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_multi.py#L19-L57)

### How are test datasets and cases defined, generated and versioned? (answered)

**Dataset definition.** A `Dataset` is a `Sequence[Sample]` (`src/inspect_ai/dataset/_dataset.py:144-221`). Each `Sample` has `input`, `target`, `choices`, `id`, `metadata`, `sandbox`, `files`, `setup`, and `checkpoint` fields. The input can be a string or a list of `ChatMessage` objects.

**File formats and sources** (`src/inspect_ai/dataset/_sources/`):
- **hf_dataset** — loads from Hugging Face datasets with retry logic for transient errors, field mapping via `FieldSpec`, and support for splits/shuffling
- **json_dataset** — JSON array of sample dicts
- **csv_dataset** — CSV file with configurable field mapping
- **file_dataset** — generic file (auto-detects JSON/CSV/JSONL by extension)
- **example_dataset** — inline sample generation via a function

**Synthetic data generation** is done in user code by writing Python functions that return `MemoryDataset(list[Sample])` or by implementing a `SampleSource` protocol. There is no built-in synthetic data DSL beyond the `Sample` constructor.

**Versioning.** Inspect does not version datasets natively — versioning is the data provider's responsibility (e.g. Hugging Face dataset revisions pinned by `revision=` in `hf_dataset()`). The `eval_set` system provides log-based versioning through eval set IDs and manifest files.

**Benchmark task registry.** Tasks are registered with `@task` decorator and can be discovered via `inspect list` or `list_tasks()`. The `Task` object binds a `Dataset`, a `Solver`/`Plan`, and optional `Scorer`s, metrics, and config.

**Golden sets.** There is no built-in "golden set" concept — tasks specify their own dataset at construction time. The `SampleSource` protocol allows dynamic sample feeding. The `precomputed_scores` scorer attaches externally computed labels to existing logs by sample ID.


Citations: [src/inspect_ai/dataset/__init__.py:3-27](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/__init__.py#L3-L27) · [src/inspect_ai/dataset/_dataset.py:29-121](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_dataset.py#L29-L121) · [src/inspect_ai/dataset/_dataset.py:144-299](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_dataset.py#L144-L299) · [src/inspect_ai/dataset/_sources/hf.py:1-100](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_sources/hf.py#L1-L100)

### How are evals executed and reported? (answered)

**Runners.** Evals are launched via `eval()` / `eval_async()` (`src/inspect_ai/_eval/eval.py`). The top-level orchestrator `eval_set()` (`src/inspect_ai/_eval/evalset.py`) manages multiple tasks with shared configuration, sample dispatch, retry logic, and logging. Inside a task, `task_run()` sets up sandbox lifecycle and prepares task options. Samples are fanned out by `SampleScheduler`, a live fanout loop that accepts new samples mid-run.

**Parallelism.** `parallel=` controls concurrent samples, running in an anyio `TaskGroup`. Model-level parallelism uses adaptive connection pools (`DEFAULT_MAX_CONNECTIONS` / `DEFAULT_MAX_CONNECTIONS_BATCH`). `max_tasks` controls parallel task execution within eval sets.

**Caching.** Model outputs are cached with configurable TTL. The `inspect cache` CLI command (`src/inspect_ai/_cli/cache.py`) provides `clear`, `prune`, `list`, and `size` operations.

**CI integration.** Standard Python library approach — `pip install inspect-ai && python your_eval.py`. The CLI supports `--detach` for long-running evals, ACP for cross-machine dispatch, and JSON output modes.

**Result storage.** Logs use a compact binary `.eval` format (or JSON) via the `Recorder` hierarchy (`src/inspect_ai/log/_recorders/`). `FileRecorder` writes to local/remote filesystems via fsspec. `BufferSampleStore` provides crash-recovery through a SQLite buffer.

**Comparison / regression.** Eval results are structured as `EvalResults` with `EvalScore` objects per scorer and `EvalMetric` dicts per metric (`src/inspect_ai/log/_log.py:799-904`). The `inspect view` command starts a web dashboard for browsing logs and comparing runs. There is no built-in regression test framework — users compare metrics programmatically from returned `EvalLog` objects.

> **Editor's note.** Correction: --acp-server is not cross-machine dispatch. It exposes a running eval over the Agent Client Protocol so clients (inspect acp, editors) can attach for human-in-the-loop approvals. The .eval log format is a zip archive of JSON entries, not a custom binary format.

Citations: [src/inspect_ai/_eval/eval.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/eval.py#L1-L50) · [src/inspect_ai/_eval/task/run.py:139-300](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/run.py#L139-L300) · [src/inspect_ai/_eval/task/scheduler.py:1-100](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/scheduler.py#L1-L100) · [src/inspect_ai/_cli/cache.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/cache.py#L1-L50) · [src/inspect_ai/log/_log.py:799-904](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/log/_log.py#L799-L904) · [src/inspect_ai/_cli/view.py:1-60](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/view.py#L1-L60)

### How are traces or production data captured and linked to evaluations? (answered)

**SDK instrumentation** is provided through the `events` module and the hooks system, not through OpenTelemetry.

**Event system** (`src/inspect_ai/event/`). During sample execution, a timeline of typed events is recorded: `ModelEvent`, `ToolEvent`, `ScoreEvent`, `SandboxEvent`, `StateEvent`, `StepEvent`, `SubtaskEvent`, `InputEvent`, `ApprovalEvent`, `ErrorEvent`, `StoreEvent`, `SpanEvent`, and more. These are organized into an `EventTree` modeling the hierarchical sample run. The `Timeline` builds a linearized, filterable view.

**Hooks** (`src/inspect_ai/hooks/`). Lifecycle hooks fire at `eval_set_start/end`, `run_start/end`, `task_start/end`, `sample_start/init/attempt_start/attempt_end/end/scoring`, and `before_model_generate`. Hooks are passed to `eval()`.

**Trace logging** (`src/inspect_ai/_util/trace.py`). A custom `TRACE` log level captures HTTP requests, model calls, and long-running actions. The `inspect trace` CLI (`src/inspect_ai/_cli/trace.py`) lists, dumps, filters traces by action type, and surfaces anomalies.

**No OpenTelemetry.** There is no `opentelemetry` dependency or exporter. The tracing system is file-based JSONL.

**Online vs offline evals.** "Online" evals run via `eval()` and write to log files in real-time. "Offline" analysis reads logged `EvalLog` objects via `read_eval_log()`, recomputes metrics via `recompute_metrics()`, or edits scores via `edit_score()`.

**Feedback and annotation.** The `precomputed_scores` scorer attaches externally computed scores by sample ID. `edit_score()` modifies logged scores with provenance tracking. There is no built-in annotation UI.


Citations: [src/inspect_ai/event/__init__.py:1-91](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/event/__init__.py#L1-L91) · [src/inspect_ai/hooks/_hooks.py:1-61](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/hooks/_hooks.py#L1-L61) · [src/inspect_ai/_util/trace.py:1-120](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_util/trace.py#L1-L120) · [src/inspect_ai/_cli/trace.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/trace.py#L1-L50) · [src/inspect_ai/log/_log.py:95-180](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/log/_log.py#L95-L180)

### Does it support red-teaming or safety testing, and how? (answered)

Inspect provides infrastructure useful for red-teaming but has no built-in adversarial probes, attack plugins, jailbreak harnesses, or dedicated red-teaming tooling.

**What exists:**
- **Scanner support** (`src/inspect_ai/_eval/task/scan.py`). An `eval_set` can attach `ScannerConfig` objects that run per-sample analysis via the optional `inspect_scout` package. Designed for post-hoc scanning of completed transcripts, not adversarial generation.
- **Review system** (`src/inspect_ai/review/`). The `Reviewer` protocol intercepts tool call results mid-execution and can `continue`, `terminate`, or `escalate`.
- **Sandbox environments** (`src/inspect_ai/util/_sandbox/`). Docker-based sandboxes with diagnostics and egress controls provide isolation for running untrusted model outputs.
- **Model-level safety settings.** Several model provider modules expose safety/abuse-detection parameters (Anthropic, OpenAI, Google, Grok) but as passthroughs, not a managed red-teaming feature.

**What is absent:**
- No built-in adversarial attack library (no prompt injection generators, no jailbreak test suites, no fuzzing tools)
- No dedicated red-team evaluation harness or report format
- No vulnerability reporting workflow or CVE tracking
- No built-in harmfulness classifiers or refusal detectors
- No automated red-teaming loop that generates increasingly adversarial inputs

**How users do red-teaming today:** By writing custom `@scorer` functions that test for specific failure modes, using the `Tool` system to build adversarial tool environments, and leveraging the `Plan`/`Solver` system to construct multi-turn probe sequences. The framework is flexible enough for ad-hoc red-teaming through these primitives, but provides no turnkey solution.


Citations: [src/inspect_ai/_eval/task/scan.py:1-100](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/scan.py#L1-L100) · [src/inspect_ai/review/_review.py:1-25](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/review/_review.py#L1-L25) · [src/inspect_ai/review/_reviewer.py:10-38](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/review/_reviewer.py#L10-L38)
