# openai/evals

> Python CLI and YAML registry of ~460 benchmark evals that run models over JSONL samples and log every match and sample as events.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/openai/evals (reviewed at commit `8eac7a7de5215c907fbddc30efdaf316913eccdd`, 2026-04-14)
- Stars: 19563 · Language: Python · License: n/a
- Canonical page: https://llms-technical-reviews.com/p/openai-evals/

## Overview

OpenAI Evals is a batch benchmark runner plus a large registry of eval definitions. You name a model (or a registered "completion function" or "solver") and an eval, and the `oaieval` CLI loads the eval's YAML spec, reads its JSONL samples, runs the model over every sample in a thread pool, records each comparison as an event, and prints a final metrics dict. The registry at the pinned commit holds 463 eval YAML files under `evals/registry/evals/`, plus model-graded rubrics, solver configs and eval sets.

The design is small and explicit. An eval is a Python class with two methods: `eval_sample` (score one sample, record events) and `run` (load samples, fan out, aggregate). Most of the registry reuses a handful of generic classes (`Match`, `Includes`, `FuzzyMatch`, `ModelBasedClassify`). The larger custom suites under `evals/elsuite/` (sandbagging, steganography, make-me-pay, multistep web tasks and others) are full Python evals built on the newer `Solver` interface.

Contributions are restricted: the README says the project is not accepting evals with custom code, only YAML-plus-data evals on the existing templates ([README.md](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/README.md#L75-L75)). This is a research harness, not an eval platform. There is no tracing, no dashboard, no regression gate and no human-annotation tool. Results go to a JSONL file by default. The README now points users to the hosted Evals product in the OpenAI dashboard, and the most recent commit only pins pre-commit hooks, so treat the repository as a stable reference implementation rather than an actively growing tool.

## Architecture

```mermaid
flowchart LR
  CLI["oaieval CLI"] --> REG["Registry (YAML dirs)"]
  REG --> SPEC["EvalSpec + JSONL samples"]
  REG --> CFN["CompletionFn / Solver"]
  CLI --> EV["Eval class (elsuite)"]
  EV --> POOL["ThreadPool over samples"]
  POOL --> CFN
  CFN --> API["Model API"]
  EV --> MG["Model-graded classify"]
  MG --> CFN
  POOL --> REC["Recorder (events)"]
  REC --> OUT["JSONL / HTTP / Snowflake"]
  EV --> REP["Final report dict"]
```

| Component | Path | Role |
|---|---|---|
| CLI | `evals/cli/oaieval.py`, `evals/cli/oaievalset.py` | Parse args, build completion functions and recorder, run one eval or a resumable set |
| Registry | `evals/registry.py` | Load YAML from `registry/{evals,completion_fns,solvers,modelgraded,eval_sets}`, resolve aliases, map model names to completion classes |
| Eval base | `evals/eval.py` | `Eval` and `SolverEval`: seeded shuffling, thread-pool fan-out, per-sample recorder context |
| Generic evals | `evals/elsuite/basic/`, `evals/elsuite/modelgraded/` | `Match`, `Includes`, `FuzzyMatch`, `JsonValidator`, `ModelBasedClassify` |
| Custom suites | `evals/elsuite/*/` | Bespoke evals (sandbagging, steganography, bluff, theory_of_mind, ...) |
| Completion functions | `evals/api.py`, `evals/completion_fns/` | `CompletionFn` protocol; OpenAI, LangChain, CoT and retrieval wrappers |
| Solvers | `evals/solvers/` | Stateful `Solver` over a `TaskState`; providers for OpenAI, Anthropic, Gemini, Together |
| Recorder | `evals/record.py` | Typed events, batched flushes, Local/HTTP/Snowflake/Dummy backends |
| Metrics | `evals/metrics.py` | Accuracy, bootstrap std, confusion matrix, precision/recall/F1 |
| Data | `evals/registry/data/` | JSONL samples, stored in Git LFS |

## How a request flows

Take `oaieval gpt-3.5-turbo ab`:

1. **Parse and resolve.** `run()` asks the `Registry` for the eval spec, merges `--extra_eval_params`, and builds one completion function per comma-separated name ([oaieval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L118-L171)). The registry loads every YAML file under the default paths, the package's own `registry/` and `~/.evals` ([registry.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry.py#L30-L33), [L287-L310](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry.py#L287-L310)). `ab` is an alias whose `id` points at `ab.dev.v0`, and `_dereference` follows such aliases until it reaches a concrete spec ([registry.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry.py#L156-L191)).
2. **Pick a model class.** `make_completion_fn` returns a dummy, an `OpenAIChatCompletionFn` for names that `is_chat_model` recognises, an `OpenAICompletionFn` for other API model ids, or instantiates a registry entry by its `class` path ([registry.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry.py#L120-L151)).
3. **Build the run.** A `RunSpec` gets a time-plus-random `run_id` ([base.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/base.py#L74-L89)). `build_recorder` picks `DummyRecorder` (dry run), `LocalRecorder` (default, `/tmp/evallogs/<run_id>_....jsonl`), `HttpRecorder` or the Snowflake `Recorder` ([oaieval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L197-L266)).
4. **Fan out.** The eval class is instantiated and `eval.run(recorder)` is called. For `ab.dev.v0` that is `ModelBasedClassify`. `run` loads the JSONL and calls `eval_all_samples`, which shuffles indices with a fixed seed, truncates to `--max_samples`, and maps `eval_sample` over a `ThreadPool` of `EVALS_THREADS` (default 10) workers. Each worker runs inside `recorder.as_default_recorder(sample_id)` with an RNG seeded from the sample id ([eval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L30-L38), [L112-L147](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L112-L147)).
5. **Sample and judge.** `eval_sample` gets the subject model's completion, then calls `classify()` with the judge model, which is the *last* completion function on the command line ([classify.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L28-L102)). Every OpenAI call records a `sampling` event with prompt, output and usage ([openai.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/completion_fns/openai.py#L99-L131)).
6. **Record.** `record_metrics(choice=..., score=...)` appends an event. `record_event` flushes in batches, at most every 100 events or 10 seconds ([record.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L157-L185)).
7. **Aggregate and report.** `run` reads the `metrics` events back, counts each choice and averages scores ([classify.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L104-L127)). The CLI sums token usage from `sampling` events into the result, writes a `final_report` line and logs every key ([oaieval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L226-L294)).

## Key components

### Registry and naming

Everything is a YAML entry keyed by name. Eval names must be `<base>.<split>.<version>` (the constructor rejects names without a split), and a bare base name is an alias with metadata such as `metrics: [accuracy]` and `higher_is_better` ([base.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/base.py#L29-L44)). Duplicate keys across registry paths fail with an assertion, and `key`, `group` and `cls` are reserved. `oaievalset` expands an eval set into one `oaieval` subprocess per eval and keeps a progress file in `/tmp/oaievalset/` so a set can resume ([oaievalset.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaievalset.py#L81-L131)).

### Match-style evals

`Match` sends the sample's `input` (optionally with few-shot examples spliced in before the last message) at temperature 0. `record_and_check_match` then checks whether the output *starts with* any `ideal` string ([match.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/basic/match.py#L30-L65), [api.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/api.py#L55-L105)). The run reports `accuracy` plus a bootstrap spread. That spread is the standard deviation of means over 1,000 random half-size subsamples drawn without replacement, not a classic with-replacement bootstrap ([metrics.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/metrics.py#L12-L23)).

### Model-graded evals

A `ModelGradedSpec` holds a prompt template, the allowed `choice_strings`, a map from sample fields to completions, and optional `choice_scores` ([base.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/base.py#L11-L25)). `classify()` appends an answer instruction for one of four modes (`classify`, `classify_cot`, `cot_classify`, plus a Japanese CoT variant). `get_choice` strips punctuation and scans lines (in reverse for CoT modes) for the first matching choice, returning `__invalid__` otherwise. Invalid answers score as the lowest choice ([classify_utils.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify_utils.py#L13-L128)). This is plain text parsing, with no structured output or schema validation. `multicomp_n` concatenates several subject completions into one judge input for best-of or diversity rubrics.

### Solvers

`Solver` is the newer abstraction for agent-style evals. It receives a `TaskState` (task description plus message history), returns a `SolverResult`, and can chain postprocessors that each emit a `postprocessor` event ([solver.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/solvers/solver.py#L41-L125)). `SolverEval` deep-copies the solver for every sample so that stateful solvers do not leak memory between samples, and supports a "gentle interrupt" that reports partial results on Ctrl-C ([eval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L168-L255)). Plain completion functions and solvers are wrapped into each other as needed ([utils.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/solvers/utils.py#L10-L25)).

### Recorder

Events are `(run_id, event_id, sample_id, type, data, created_by, created_at)` records ([record.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L43-L88)). The sample id comes from a `ContextVar`, which is why helpers like `evals.record.record_metrics` work from anywhere inside `eval_sample`. `LocalRecorder` writes through `blobfile`, so the log path can be local, GCS or Azure, and it can drop `hidden_data_fields` before writing ([record.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L316-L371)).

## Extending it

- **New eval from data only.** Add JSONL under `registry/data/<name>/` and a YAML entry pointing a generic class (`Match`, `Includes`, `FuzzyMatch`, `ModelBasedClassify`) at it ([ab.yaml](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/evals/ab.yaml#L1-L11)).
- **New eval class.** Subclass `Eval` (or `SolverEval`), implement `eval_sample` and `run`, and reference it as `module:Class` in YAML ([eval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L46-L88)).
- **New model or agent.** Implement the `CompletionFn` protocol, a callable that returns an object with `get_completions()` ([api.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/api.py#L16-L40)), or a `Solver`, then register it under `completion_fns/` or `solvers/`.
- **New rubric.** Add a YAML under `registry/modelgraded/` and reference it with `modelgraded_spec`.
- **CI smoke test.** The repository's own workflow runs `oaieval dummy <new-eval> --max_samples 10` for every YAML file a pull request adds, which checks wiring and data loading without calling a model ([test_eval.yaml](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/.github/workflows/test_eval.yaml#L38-L59)). The same trick works for your private registry.
- **Private registries.** Pass `--registry_path` (repeatable) or drop files in `~/.evals` to keep proprietary evals out of the repo.

## Running it

- `pip install -e .` (or `pip install evals`), set `OPENAI_API_KEY`, and run `git lfs pull` to fetch sample data, since the JSONL files are LFS pointers.
- `oaieval <completion_fn> <eval> [--max_samples N] [--record_path ...]`. Use `EVALS_THREADS` for concurrency and `EVALS_SEQUENTIAL=1` for debugging.
- `oaievalset <model> <eval_set>` runs a named set with resume.
- `--http-run --http-run-url ...` posts event batches to your own endpoint, with a local fallback file. `--no-local-run` without `--http-run` selects the Snowflake recorder, which needs Snowflake credentials.
- The dependency list is broad (Playwright, Snowflake connector, spaCy encoder, LangChain, Gymnasium and more), so expect a heavy install even for a simple match eval.

## Strengths and caveats

- **Strength: transparent and reproducible.** Fixed shuffle seed, per-sample seeded RNG, and a full JSONL event log of every prompt and output make runs easy to audit and diff.
- **Strength: a large, readable corpus.** Hundreds of eval specs and rubrics show how real capability and safety evals are built, and they are useful as templates even if you never run the CLI.
- **Strength: the solver layer** supports multi-turn and tool-using evals with per-sample isolation, up to Docker-hosted WebArena-style web environments in `multistep_web_tasks`.
- **Caveat: OpenAI-era model routing.** `is_chat_model` only recognises `gpt-3.5-turbo*` and `gpt-4-*` names, and the context table says "last updated 2023-10-24" ([registry.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry.py#L37-L96)). A bare newer name such as `gpt-4o` falls through to the legacy completions class. Use a registered solver instead.
- **Caveat: `--no-cache` does nothing.** The CLI builds `api_extra_options["cache_level"] = 0` for `--no-cache` but never passes it anywhere ([oaieval.py](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L210-L212)).
- **Caveat: no analysis layer.** There is no run comparison, regression check, dashboard or tracing. You get a JSONL file and a logged dict.
- **Caveat: brittle judge parsing.** Model-graded verdicts are found by line-wise string matching, and unparseable answers silently score as the worst choice.
- **Caveat: low activity.** The upstream team now steers users to the hosted product, so new model support has to come from you.

*Sources: code at 8eac7a7, deepwiki-open wiki (11 pages), verified Q&A.*

## How openai/evals answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

**Built-in metrics** are implemented as standalone functions in `evals/metrics.py`. `get_accuracy(events)` computes `sum(correct)/total` from `match` events (`evals/metrics.py:12-18`). `get_bootstrap_accuracy_std` resamples with replacement for uncertainty (`evals/metrics.py:21-23`). `get_confusion_matrix` builds an N×(N+1) array from `expected` vs `picked` labels (`evals/metrics.py:26-40`), then `compute_matthew_corr`, `compute_precision`, `compute_recall`, `compute_f_score`, and `compute_averaged_f_score` derive classification metrics from it (`evals/metrics.py:43-73`).

**Custom per-sample scoring** uses `evals.record.record_metrics(**kwargs)` — a free-form dict attached to a `"metrics"` event on the `RecorderBase` (`evals/record.py:248-249`). Standard evals like `FuzzyMatch` record per-sample `accuracy` (float) and `f1_score` (`evals/elsuite/basic/fuzzy_match.py:48-50`). Others record custom keys: `Includes` records correctness per sample (`evals/elsuite/basic/includes.py:45-47`), `JsonValidator` records `accuracy` alone, and `ModelBasedClassify` records `choice`, `score`, and optionally `metascore` (`evals/elsuite/modelgraded/classify.py:93-98`).

**Per-task aggregation** happens in each eval's `run()` method, which returns a `dict[str, float]`. `Match.run()` averages `accuracy` and `bootstrap_std` from all `match` events via `recorder.get_events("match")` (`evals/elsuite/basic/match.py:58-65`). `MultipleChoice.run()` calls `evals.metrics.get_accuracy()` on match events (`evals/elsuite/multiple_choice.py:95-100`). `ModelBasedClassify.run()` additionally computes choice-count distributions, per-choice score averages, and metascore (`evals/elsuite/modelgraded/classify.py:104-127`). The `BaseEvalSpec` declares a `metrics` list (e.g. `[accuracy]`) and a `higher_is_better` flag (`evals/base.py:36-44`).

> **Editor's note.** Correction: `get_bootstrap_accuracy_std` is not a with-replacement bootstrap. It returns the standard deviation of accuracy over 1,000 random half-size subsamples drawn without replacement (evals/metrics.py L21-L23).

Citations: [evals/elsuite/basic/fuzzy_match.py:40-60](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/basic/fuzzy_match.py#L40-L60) · [evals/elsuite/modelgraded/classify.py:93-127](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L93-L127) · [evals/record.py:44-70](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L44-L70) · [evals/base.py:30-48](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/base.py#L30-L48)

### How is LLM-as-a-judge implemented? (answered)

**Judge prompts and rubrics** are specified via `ModelGradedSpec`, a pydantic dataclass with fields `prompt`, `choice_strings`, `input_outputs`, `eval_type`, `choice_scores`, and `output_template` (`evals/elsuite/modelgraded/base.py:11-25`). Rubrics live as YAML files in `evals/registry/modelgraded/` — e.g. `fact.yaml` lets the judge compare a submission to an expert answer on correctness using a 5-option rubric (A–E) (`evals/registry/modelgraded/fact.yaml:1-21`), and `battle.yaml` compares two model outputs head-to-head with a Yes/No preference (`evals/registry/modelgraded/battle.yaml:1-23`).

**Structured output** is extracted by `classify()` in `evals/elsuite/modelgraded/classify_utils.py`. The judge model receives the rubric prompt with format kwargs filled in, plus an appended answer prompt depending on `eval_type` — `classify`, `classify_cot`, `cot_classify`, or `cot_classify_jp` — each of which instructs the LLM to output exactly one choice string from the allowed set (`classify_utils.py:13-28`). `get_choice()` then parses the raw text by trying each line against `choice_strings` using a `match_fn` (one of `include`, `exact`, `endswith`, `starts_or_endswith`), returning `"__invalid__"` on failure (`classify_utils.py:110-128`). `get_choice_score` maps parsed choices to numeric scores via `choice_scores` (`classify_utils.py:90-102`).

**Multi-sample and consensus** is supported via `multicomp_n`. When `multicomp_n > 1`, `sample_and_concat_n_completions` runs the subject model N times (either with N separate model instances or the same model N times), then concatenates the outputs into a single text using `output_template` before passing to the judge (`classify_utils.py:152-187`). When `multicomp_n == "from_models"`, N is derived from the number of completion functions (`evals/elsuite/modelgraded/classify.py:42-48`).

**Judge model choice** is determined by the last `completion_fn` in the list — the eval splits off `self.eval_completion_fn` as the judge, while the earlier ones serve as the subject model(s) (`evals/elsuite/modelgraded/classify.py:29-32`). There are no built-in calibration or bias controls.


Citations: [evals/registry/modelgraded/fact.yaml:1-22](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/modelgraded/fact.yaml#L1-L22) · [evals/registry/modelgraded/battle.yaml:1-23](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/modelgraded/battle.yaml#L1-L23) · [evals/elsuite/modelgraded/classify.py:29-52](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L29-L52)

### How are test datasets and cases defined, generated and versioned? (answered)

**File formats and DSL.** Test data uses the JSONL format with standard fields `"input"` (a chat-message list or plain text) and `"ideal"` (a string or list of accepted answers). For example, `test_fuzzy_match/samples.jsonl` entries contain OpenAI-format chat arrays with example few-shot messages and an `"ideal"` answer list (`evals/registry/data/test_fuzzy_match/samples.jsonl:1-3`). Eval configurations are YAML files in `evals/registry/evals/` — each file names a base spec (e.g. `ab:` with `metrics: [accuracy]`) and one or more versioned split entries (e.g. `ab.dev.v0:`) linking to a class, `samples_jsonl` path, `eval_type`, and `modelgraded_spec` (`evals/registry/evals/ab.yaml:1-11`). The `Registry` class loads these YAML files from `evals/registry/evals/`, `completion_fns/`, `solvers/`, `modelgraded/`, and `eval_sets/` directories (`evals/registry.py:103-331`).

**Synthetic data generation.** Custom generators live in `evals/registry/data/*/` — e.g. `simple_physics_engine/samples_generator.py`, `poker_analysis/poker_analysis_sample_generator.py`, `mazes/nxn_maze_eval_generator.py`, and `solve-for-variable/tools/main.py`. These emit JSONL files consumed by the evals.

**External dataset integration.** The `MultipleChoice` eval class loads from HuggingFace datasets (HellaSwag, Hendrycks MMLU) via `datasets.load_dataset()` using `hf://` URLs (`evals/elsuite/multiple_choice.py:20-48`). The `Lambada` eval similarly loads `EleutherAI/lambada_openai` from HuggingFace (`evals/elsuite/lambada.py:42-44`).

**Versioning.** Versioning follows a `{base_eval}.{split}.v{N}` convention — e.g. `ab.dev.v0`, `prompt-injection.dev.v0`, `human-safety.test.v0`. Splits (dev, test, etc.) are freeform strings. The `registry_path` parameter supports loading multiple registry directories, and `~/.evals` is a secondary path.

**Benchmark registries.** There are 463 eval YAML files in `evals/registry/evals/` and 472 data directories in `evals/registry/data/`. Eval sets like `test-all` and `test-basic` list multiple eval names to run together (`evals/registry/eval_sets/test-all.yaml:1-21`).


Citations: [evals/elsuite/multiple_choice.py:20-48](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/multiple_choice.py#L20-L48) · [evals/data.py:47-108](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/data.py#L47-L108) · [evals/registry/data/test_fuzzy_match/samples.jsonl:1-3](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/data/test_fuzzy_match/samples.jsonl#L1-L3)

### How are evals executed and reported? (answered)

**CLI entry points.** `oaieval` (`evals/cli/oaieval.py:297-311`) is the primary runner. It takes a `completion_fn` (model ID or registry key), an `eval` name, and optional flags (`--max_samples`, `--cache`, `--seed`, `--extra_eval_params`, etc.). `oaievalset` (`evals/cli/oaievalset.py:134-141`) runs an ordered set of evals as subprocess calls to `oaieval`, with checkpoint/resume via a progress file.

**Runners and parallelism.** The `Eval` base class (`evals/eval.py:46-147`) defines `eval_sample()` (abstract, per-sample logic) and `run()` (abstract, orchestrates loading + scoring). `eval_all_samples()` shuffles samples with a fixed seed, then processes them via `ThreadPool` (default 10 threads, controlled by `EVALS_THREADS` env var) or sequentially (`EVALS_SEQUENTIAL=1`). `async_eval_all_samples()` provides an asyncio-based path with configurable concurrency and a semaphore (`evals/eval.py:112-147`). The `SolverEval` subclass copies the solver per sample for state isolation (`evals/eval.py:168-255`). Gentle interrupt (`EVALS_GENTLE_INTERRUPT`) allows early stopping with partial results.

**Caching.** `--cache` controls an API-level cache_level; disabled with `--no-cache` (`evals/cli/oaieval.py:211-212`).

**Result storage.** Three recorders implement `RecorderBase`: `LocalRecorder` writes JSONL locally (default, at `/tmp/evallogs/...`) (`evals/record.py:316-371`); `Recorder` writes to Snowflake plus local fallback (`evals/record.py:468-581`); `HttpRecorder` POSTs batches to a URL with failover to local storage (`evals/record.py:374-465`). `DummyRecorder` logs to console for dry-runs (`evals/record.py:274-314`). All events (`match`, `sampling`, `metrics`, `error`, `extra`, etc.) share the same flush-and-batch infrastructure with configurable thresholds (`MIN_FLUSH_EVENTS=100`, `MIN_FLUSH_SECONDS=10`).

**Comparison and dashboards.** There is no built-in comparison, regression detection, or dashboard. Results are raw JSONL files or Snowflake tables — analysis is left to external tooling. Final reports are logged to console and stored, showing per-metric values (`evals/cli/oaieval.py:236-238`).

> **Editor's note.** Correction: `--no-cache` has no effect at this commit. `oaieval.py` puts `cache_level = 0` into a local `api_extra_options` dict (L210-L212) that is never passed to the eval or the completion functions.

Citations: [evals/cli/oaieval.py:117-240](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L117-L240) · [evals/eval.py:46-167](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L46-L167) · [evals/record.py:316-466](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L316-L466) · [evals/eval.py:112-147](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L112-L147) · [evals/cli/oaievalset.py:80-141](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaievalset.py#L80-L141)

### How are traces or production data captured and linked to evaluations? (answered)

**SDK instrumentation and OpenTelemetry.** This repository does not implement OpenTelemetry or any SDK-level instrumentation. There is no tracing, no span export, and no OpenTelemetry dependency. Evals run as isolated batch processes, not as instrumented services.

**Online vs offline evals.** All evals are offline batch workloads. The `oaieval` CLI runs a single eval process that loads samples, queries a model, and records results. There is no online evaluation mode, no streaming pipeline, and no live traffic evaluation.

**What is captured.** During execution, `RecorderBase` stores typed events in memory and flushes them periodically: `match` (correct/incorrect comparisons), `sampling` (prompts and model outputs), `embedding`, `cond_logp`, `pick_option`, `function_call`, `metrics`, `error`, and `extra` (`evals/record.py:44-70`). Each event carries `run_id`, `sample_id`, `type`, `data`, `created_by`, and `created_at`. Token usage from sampling events is extracted after the run and folded into the final result dict (`evals/cli/oaieval.py:269-294`).

**Feedback and annotation.** There is no human feedback or annotation system. The `record_extra` method (`evals/record.py:259-260`) could be used as a generic escape hatch to log arbitrary data during an eval, but no tooling exists to collect, view, or manage annotations.

**Linked production data.** There is no mechanism to import or link production traces to evaluation runs. The Snowflake `Recorder` stores runs and events in relational tables with a `run_id` foreign key (`evals/record.py:492-511`), but this is only for the evaluation run results, not production traffic.


Citations: [evals/record.py:44-90](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L44-L90) · [evals/record.py:260-270](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L260-L270) · [evals/record.py:468-512](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L468-L512) · [evals/cli/oaieval.py:269-295](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L269-L295)

### Does it support red-teaming or safety testing, and how? (answered)

**The repo includes a few safety-oriented evals but no structured red-teaming framework.** The key examples: `prompt-injection` tests a model's ability to distinguish instructions from data using log-conversion tasks where injected content tries to override the system instruction — implemented as a `Match` eval with adversarial JSONL samples (`evals/registry/evals/prompt-injection.yaml:1-9`, `evals/registry/data/prompt-injection/samples.jsonl:1-3`). `human-safety` evaluates logical reasoning about real-life situations involving physical danger (self-driving car dilemmas, medical robot decisions) using `FuzzyMatch` (`evals/registry/evals/human-safety.yaml:1-8`, `evals/registry/data/human-safety/human-safety.jsonl:1-3`).

**Beyond basic safety evals**, `Sandbagging` (`evals/elsuite/sandbagging/sandbagging_eval.py:16-44`) tests whether a model can deliberately underperform to a target accuracy on MMLU — a capability concern rather than an adversarial probe. `Steganography` (`evals/elsuite/steganography/eval.py:12-35`) evaluates whether a model can hide payload text within otherwise-normal task outputs, with a monitor model attempting to detect the hidden content. `Already_said_that` tests for unwanted repetition.

**What is absent.** There are no adversarial probe generators, no attack plugins (gradient-based, token-manipulation, or suffix-injection), no automated jailbreak discovery, no prompt-injection benchmark suites (like MITRE ATLAS), and no structured vulnerability reporting workflow beyond the general `SECURITY.md` linking to OpenAI's CVD policy (`SECURITY.md:1-4`). The existing `prompt-injection` eval covers a narrow class of prompt-override attacks in a controlled setting, not a systematic red-teaming suite.


Citations: [evals/registry/evals/prompt-injection.yaml:1-9](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/evals/prompt-injection.yaml#L1-L9) · [evals/registry/evals/human-safety.yaml:1-8](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/evals/human-safety.yaml#L1-L8) · [evals/elsuite/sandbagging/sandbagging_eval.py:16-44](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/sandbagging/sandbagging_eval.py#L16-L44) · [evals/elsuite/steganography/eval.py:12-35](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/steganography/eval.py#L12-L35)
