openai/evals
Python CLI and YAML registry of ~460 benchmark evals that run models over JSONL samples and log every match and sample as events.
Overview
OpenAI Evals is a batch benchmark runner plus a large registry of eval definitions. You name a model (or a registered “completion function” or “solver”) and an eval, and the oaieval CLI loads the eval’s YAML spec, reads its JSONL samples, runs the model over every sample in a thread pool, records each comparison as an event, and prints a final metrics dict. The registry at the pinned commit holds 463 eval YAML files under evals/registry/evals/, plus model-graded rubrics, solver configs and eval sets.
The design is small and explicit. An eval is a Python class with two methods: eval_sample (score one sample, record events) and run (load samples, fan out, aggregate). Most of the registry reuses a handful of generic classes (Match, Includes, FuzzyMatch, ModelBasedClassify). The larger custom suites under evals/elsuite/ (sandbagging, steganography, make-me-pay, multistep web tasks and others) are full Python evals built on the newer Solver interface.
Contributions are restricted: the README says the project is not accepting evals with custom code, only YAML-plus-data evals on the existing templates (README.md). This is a research harness, not an eval platform. There is no tracing, no dashboard, no regression gate and no human-annotation tool. Results go to a JSONL file by default. The README now points users to the hosted Evals product in the OpenAI dashboard, and the most recent commit only pins pre-commit hooks, so treat the repository as a stable reference implementation rather than an actively growing tool.
Architecture
flowchart LR
CLI["oaieval CLI"] --> REG["Registry (YAML dirs)"]
REG --> SPEC["EvalSpec + JSONL samples"]
REG --> CFN["CompletionFn / Solver"]
CLI --> EV["Eval class (elsuite)"]
EV --> POOL["ThreadPool over samples"]
POOL --> CFN
CFN --> API["Model API"]
EV --> MG["Model-graded classify"]
MG --> CFN
POOL --> REC["Recorder (events)"]
REC --> OUT["JSONL / HTTP / Snowflake"]
EV --> REP["Final report dict"]
| Component | Path | Role |
|---|---|---|
| CLI | evals/cli/oaieval.py, evals/cli/oaievalset.py |
Parse args, build completion functions and recorder, run one eval or a resumable set |
| Registry | evals/registry.py |
Load YAML from registry/{evals,completion_fns,solvers,modelgraded,eval_sets}, resolve aliases, map model names to completion classes |
| Eval base | evals/eval.py |
Eval and SolverEval: seeded shuffling, thread-pool fan-out, per-sample recorder context |
| Generic evals | evals/elsuite/basic/, evals/elsuite/modelgraded/ |
Match, Includes, FuzzyMatch, JsonValidator, ModelBasedClassify |
| Custom suites | evals/elsuite/*/ |
Bespoke evals (sandbagging, steganography, bluff, theory_of_mind, …) |
| Completion functions | evals/api.py, evals/completion_fns/ |
CompletionFn protocol; OpenAI, LangChain, CoT and retrieval wrappers |
| Solvers | evals/solvers/ |
Stateful Solver over a TaskState; providers for OpenAI, Anthropic, Gemini, Together |
| Recorder | evals/record.py |
Typed events, batched flushes, Local/HTTP/Snowflake/Dummy backends |
| Metrics | evals/metrics.py |
Accuracy, bootstrap std, confusion matrix, precision/recall/F1 |
| Data | evals/registry/data/ |
JSONL samples, stored in Git LFS |
How a request flows
Take oaieval gpt-3.5-turbo ab:
- Parse and resolve.
run()asks theRegistryfor the eval spec, merges--extra_eval_params, and builds one completion function per comma-separated name (oaieval.py). The registry loads every YAML file under the default paths, the package’s ownregistry/and~/.evals(registry.py, L287-L310).abis an alias whoseidpoints atab.dev.v0, and_dereferencefollows such aliases until it reaches a concrete spec (registry.py). - Pick a model class.
make_completion_fnreturns a dummy, anOpenAIChatCompletionFnfor names thatis_chat_modelrecognises, anOpenAICompletionFnfor other API model ids, or instantiates a registry entry by itsclasspath (registry.py). - Build the run. A
RunSpecgets a time-plus-randomrun_id(base.py).build_recorderpicksDummyRecorder(dry run),LocalRecorder(default,/tmp/evallogs/<run_id>_....jsonl),HttpRecorderor the SnowflakeRecorder(oaieval.py). - Fan out. The eval class is instantiated and
eval.run(recorder)is called. Forab.dev.v0that isModelBasedClassify.runloads the JSONL and callseval_all_samples, which shuffles indices with a fixed seed, truncates to--max_samples, and mapseval_sampleover aThreadPoolofEVALS_THREADS(default 10) workers. Each worker runs insiderecorder.as_default_recorder(sample_id)with an RNG seeded from the sample id (eval.py, L112-L147). - Sample and judge.
eval_samplegets the subject model’s completion, then callsclassify()with the judge model, which is the last completion function on the command line (classify.py). Every OpenAI call records asamplingevent with prompt, output and usage (openai.py). - Record.
record_metrics(choice=..., score=...)appends an event.record_eventflushes in batches, at most every 100 events or 10 seconds (record.py). - Aggregate and report.
runreads themetricsevents back, counts each choice and averages scores (classify.py). The CLI sums token usage fromsamplingevents into the result, writes afinal_reportline and logs every key (oaieval.py).
Key components
Registry and naming
Everything is a YAML entry keyed by name. Eval names must be <base>.<split>.<version> (the constructor rejects names without a split), and a bare base name is an alias with metadata such as metrics: [accuracy] and higher_is_better (base.py). Duplicate keys across registry paths fail with an assertion, and key, group and cls are reserved. oaievalset expands an eval set into one oaieval subprocess per eval and keeps a progress file in /tmp/oaievalset/ so a set can resume (oaievalset.py).
Match-style evals
Match sends the sample’s input (optionally with few-shot examples spliced in before the last message) at temperature 0. record_and_check_match then checks whether the output starts with any ideal string (match.py, api.py). The run reports accuracy plus a bootstrap spread. That spread is the standard deviation of means over 1,000 random half-size subsamples drawn without replacement, not a classic with-replacement bootstrap (metrics.py).
Model-graded evals
A ModelGradedSpec holds a prompt template, the allowed choice_strings, a map from sample fields to completions, and optional choice_scores (base.py). classify() appends an answer instruction for one of four modes (classify, classify_cot, cot_classify, plus a Japanese CoT variant). get_choice strips punctuation and scans lines (in reverse for CoT modes) for the first matching choice, returning __invalid__ otherwise. Invalid answers score as the lowest choice (classify_utils.py). This is plain text parsing, with no structured output or schema validation. multicomp_n concatenates several subject completions into one judge input for best-of or diversity rubrics.
Solvers
Solver is the newer abstraction for agent-style evals. It receives a TaskState (task description plus message history), returns a SolverResult, and can chain postprocessors that each emit a postprocessor event (solver.py). SolverEval deep-copies the solver for every sample so that stateful solvers do not leak memory between samples, and supports a “gentle interrupt” that reports partial results on Ctrl-C (eval.py). Plain completion functions and solvers are wrapped into each other as needed (utils.py).
Recorder
Events are (run_id, event_id, sample_id, type, data, created_by, created_at) records (record.py). The sample id comes from a ContextVar, which is why helpers like evals.record.record_metrics work from anywhere inside eval_sample. LocalRecorder writes through blobfile, so the log path can be local, GCS or Azure, and it can drop hidden_data_fields before writing (record.py).
Extending it
- New eval from data only. Add JSONL under
registry/data/<name>/and a YAML entry pointing a generic class (Match,Includes,FuzzyMatch,ModelBasedClassify) at it (ab.yaml). - New eval class. Subclass
Eval(orSolverEval), implementeval_sampleandrun, and reference it asmodule:Classin YAML (eval.py). - New model or agent. Implement the
CompletionFnprotocol, a callable that returns an object withget_completions()(api.py), or aSolver, then register it undercompletion_fns/orsolvers/. - New rubric. Add a YAML under
registry/modelgraded/and reference it withmodelgraded_spec. - CI smoke test. The repository’s own workflow runs
oaieval dummy <new-eval> --max_samples 10for every YAML file a pull request adds, which checks wiring and data loading without calling a model (test_eval.yaml). The same trick works for your private registry. - Private registries. Pass
--registry_path(repeatable) or drop files in~/.evalsto keep proprietary evals out of the repo.
Running it
pip install -e .(orpip install evals), setOPENAI_API_KEY, and rungit lfs pullto fetch sample data, since the JSONL files are LFS pointers.oaieval <completion_fn> <eval> [--max_samples N] [--record_path ...]. UseEVALS_THREADSfor concurrency andEVALS_SEQUENTIAL=1for debugging.oaievalset <model> <eval_set>runs a named set with resume.--http-run --http-run-url ...posts event batches to your own endpoint, with a local fallback file.--no-local-runwithout--http-runselects the Snowflake recorder, which needs Snowflake credentials.- The dependency list is broad (Playwright, Snowflake connector, spaCy encoder, LangChain, Gymnasium and more), so expect a heavy install even for a simple match eval.
Strengths and caveats
- Strength: transparent and reproducible. Fixed shuffle seed, per-sample seeded RNG, and a full JSONL event log of every prompt and output make runs easy to audit and diff.
- Strength: a large, readable corpus. Hundreds of eval specs and rubrics show how real capability and safety evals are built, and they are useful as templates even if you never run the CLI.
- Strength: the solver layer supports multi-turn and tool-using evals with per-sample isolation, up to Docker-hosted WebArena-style web environments in
multistep_web_tasks. - Caveat: OpenAI-era model routing.
is_chat_modelonly recognisesgpt-3.5-turbo*andgpt-4-*names, and the context table says “last updated 2023-10-24” (registry.py). A bare newer name such asgpt-4ofalls through to the legacy completions class. Use a registered solver instead. - Caveat:
--no-cachedoes nothing. The CLI buildsapi_extra_options["cache_level"] = 0for--no-cachebut never passes it anywhere (oaieval.py). - Caveat: no analysis layer. There is no run comparison, regression check, dashboard or tracing. You get a JSONL file and a logged dict.
- Caveat: brittle judge parsing. Model-graded verdicts are found by line-wise string matching, and unparseable answers silently score as the worst choice.
- Caveat: low activity. The upstream team now steers users to the hosted product, so new model support has to come from you.
Sources: code at 8eac7a7, deepwiki-open wiki (11 pages), verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit 8eac7a7. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredBuilt-in metrics are implemented as standalone functions in evals/metrics.py. get_accuracy(events) computes sum(correct)/total from match events (evals/metrics.py:12-18). get_bootstrap_accuracy_std resamples with replacement for uncertainty (evals/metrics.py:21-23). get_confusion_matrix builds an N×(N+1) array from expected vs picked labels (evals/metrics.py:26-40), then compute_matthew_corr, compute_precision, compute_recall, compute_f_score, and compute_averaged_f_score derive classification metrics from it (evals/metrics.py:43-73).
Custom per-sample scoring uses evals.record.record_metrics(**kwargs) — a free-form dict attached to a "metrics" event on the RecorderBase (evals/record.py:248-249). Standard evals like FuzzyMatch record per-sample accuracy (float) and f1_score (evals/elsuite/basic/fuzzy_match.py:48-50). Others record custom keys: Includes records correctness per sample (evals/elsuite/basic/includes.py:45-47), JsonValidator records accuracy alone, and ModelBasedClassify records choice, score, and optionally metascore (evals/elsuite/modelgraded/classify.py:93-98).
Per-task aggregation happens in each eval's run() method, which returns a dict[str, float]. Match.run() averages accuracy and bootstrap_std from all match events via recorder.get_events("match") (evals/elsuite/basic/match.py:58-65). MultipleChoice.run() calls evals.metrics.get_accuracy() on match events (evals/elsuite/multiple_choice.py:95-100). ModelBasedClassify.run() additionally computes choice-count distributions, per-choice score averages, and metascore (evals/elsuite/modelgraded/classify.py:104-127). The BaseEvalSpec declares a metrics list (e.g. [accuracy]) and a higher_is_better flag (evals/base.py:36-44).
get_bootstrap_accuracy_std is not a with-replacement bootstrap. It returns the standard deviation of accuracy over 1,000 random half-size subsamples drawn without replacement (evals/metrics.py L21-L23).How is LLM-as-a-judge implemented?
answeredJudge prompts and rubrics are specified via ModelGradedSpec, a pydantic dataclass with fields prompt, choice_strings, input_outputs, eval_type, choice_scores, and output_template (evals/elsuite/modelgraded/base.py:11-25). Rubrics live as YAML files in evals/registry/modelgraded/ — e.g. fact.yaml lets the judge compare a submission to an expert answer on correctness using a 5-option rubric (A–E) (evals/registry/modelgraded/fact.yaml:1-21), and battle.yaml compares two model outputs head-to-head with a Yes/No preference (evals/registry/modelgraded/battle.yaml:1-23).
Structured output is extracted by classify() in evals/elsuite/modelgraded/classify_utils.py. The judge model receives the rubric prompt with format kwargs filled in, plus an appended answer prompt depending on eval_type — classify, classify_cot, cot_classify, or cot_classify_jp — each of which instructs the LLM to output exactly one choice string from the allowed set (classify_utils.py:13-28). get_choice() then parses the raw text by trying each line against choice_strings using a match_fn (one of include, exact, endswith, starts_or_endswith), returning "__invalid__" on failure (classify_utils.py:110-128). get_choice_score maps parsed choices to numeric scores via choice_scores (classify_utils.py:90-102).
Multi-sample and consensus is supported via multicomp_n. When multicomp_n > 1, sample_and_concat_n_completions runs the subject model N times (either with N separate model instances or the same model N times), then concatenates the outputs into a single text using output_template before passing to the judge (classify_utils.py:152-187). When multicomp_n == "from_models", N is derived from the number of completion functions (evals/elsuite/modelgraded/classify.py:42-48).
Judge model choice is determined by the last completion_fn in the list — the eval splits off self.eval_completion_fn as the judge, while the earlier ones serve as the subject model(s) (evals/elsuite/modelgraded/classify.py:29-32). There are no built-in calibration or bias controls.
How are test datasets and cases defined, generated and versioned?
answeredFile formats and DSL. Test data uses the JSONL format with standard fields "input" (a chat-message list or plain text) and "ideal" (a string or list of accepted answers). For example, test_fuzzy_match/samples.jsonl entries contain OpenAI-format chat arrays with example few-shot messages and an "ideal" answer list (evals/registry/data/test_fuzzy_match/samples.jsonl:1-3). Eval configurations are YAML files in evals/registry/evals/ — each file names a base spec (e.g. ab: with metrics: [accuracy]) and one or more versioned split entries (e.g. ab.dev.v0:) linking to a class, samples_jsonl path, eval_type, and modelgraded_spec (evals/registry/evals/ab.yaml:1-11). The Registry class loads these YAML files from evals/registry/evals/, completion_fns/, solvers/, modelgraded/, and eval_sets/ directories (evals/registry.py:103-331).
Synthetic data generation. Custom generators live in evals/registry/data/*/ — e.g. simple_physics_engine/samples_generator.py, poker_analysis/poker_analysis_sample_generator.py, mazes/nxn_maze_eval_generator.py, and solve-for-variable/tools/main.py. These emit JSONL files consumed by the evals.
External dataset integration. The MultipleChoice eval class loads from HuggingFace datasets (HellaSwag, Hendrycks MMLU) via datasets.load_dataset() using hf:// URLs (evals/elsuite/multiple_choice.py:20-48). The Lambada eval similarly loads EleutherAI/lambada_openai from HuggingFace (evals/elsuite/lambada.py:42-44).
Versioning. Versioning follows a {base_eval}.{split}.v{N} convention — e.g. ab.dev.v0, prompt-injection.dev.v0, human-safety.test.v0. Splits (dev, test, etc.) are freeform strings. The registry_path parameter supports loading multiple registry directories, and ~/.evals is a secondary path.
Benchmark registries. There are 463 eval YAML files in evals/registry/evals/ and 472 data directories in evals/registry/data/. Eval sets like test-all and test-basic list multiple eval names to run together (evals/registry/eval_sets/test-all.yaml:1-21).
How are evals executed and reported?
answeredCLI entry points. oaieval (evals/cli/oaieval.py:297-311) is the primary runner. It takes a completion_fn (model ID or registry key), an eval name, and optional flags (--max_samples, --cache, --seed, --extra_eval_params, etc.). oaievalset (evals/cli/oaievalset.py:134-141) runs an ordered set of evals as subprocess calls to oaieval, with checkpoint/resume via a progress file.
Runners and parallelism. The Eval base class (evals/eval.py:46-147) defines eval_sample() (abstract, per-sample logic) and run() (abstract, orchestrates loading + scoring). eval_all_samples() shuffles samples with a fixed seed, then processes them via ThreadPool (default 10 threads, controlled by EVALS_THREADS env var) or sequentially (EVALS_SEQUENTIAL=1). async_eval_all_samples() provides an asyncio-based path with configurable concurrency and a semaphore (evals/eval.py:112-147). The SolverEval subclass copies the solver per sample for state isolation (evals/eval.py:168-255). Gentle interrupt (EVALS_GENTLE_INTERRUPT) allows early stopping with partial results.
Caching. --cache controls an API-level cache_level; disabled with --no-cache (evals/cli/oaieval.py:211-212).
Result storage. Three recorders implement RecorderBase: LocalRecorder writes JSONL locally (default, at /tmp/evallogs/...) (evals/record.py:316-371); Recorder writes to Snowflake plus local fallback (evals/record.py:468-581); HttpRecorder POSTs batches to a URL with failover to local storage (evals/record.py:374-465). DummyRecorder logs to console for dry-runs (evals/record.py:274-314). All events (match, sampling, metrics, error, extra, etc.) share the same flush-and-batch infrastructure with configurable thresholds (MIN_FLUSH_EVENTS=100, MIN_FLUSH_SECONDS=10).
Comparison and dashboards. There is no built-in comparison, regression detection, or dashboard. Results are raw JSONL files or Snowflake tables — analysis is left to external tooling. Final reports are logged to console and stored, showing per-metric values (evals/cli/oaieval.py:236-238).
--no-cache has no effect at this commit. oaieval.py puts cache_level = 0 into a local api_extra_options dict (L210-L212) that is never passed to the eval or the completion functions.How are traces or production data captured and linked to evaluations?
answeredSDK instrumentation and OpenTelemetry. This repository does not implement OpenTelemetry or any SDK-level instrumentation. There is no tracing, no span export, and no OpenTelemetry dependency. Evals run as isolated batch processes, not as instrumented services.
Online vs offline evals. All evals are offline batch workloads. The oaieval CLI runs a single eval process that loads samples, queries a model, and records results. There is no online evaluation mode, no streaming pipeline, and no live traffic evaluation.
What is captured. During execution, RecorderBase stores typed events in memory and flushes them periodically: match (correct/incorrect comparisons), sampling (prompts and model outputs), embedding, cond_logp, pick_option, function_call, metrics, error, and extra (evals/record.py:44-70). Each event carries run_id, sample_id, type, data, created_by, and created_at. Token usage from sampling events is extracted after the run and folded into the final result dict (evals/cli/oaieval.py:269-294).
Feedback and annotation. There is no human feedback or annotation system. The record_extra method (evals/record.py:259-260) could be used as a generic escape hatch to log arbitrary data during an eval, but no tooling exists to collect, view, or manage annotations.
Linked production data. There is no mechanism to import or link production traces to evaluation runs. The Snowflake Recorder stores runs and events in relational tables with a run_id foreign key (evals/record.py:492-511), but this is only for the evaluation run results, not production traffic.
Does it support red-teaming or safety testing, and how?
answeredThe repo includes a few safety-oriented evals but no structured red-teaming framework. The key examples: prompt-injection tests a model's ability to distinguish instructions from data using log-conversion tasks where injected content tries to override the system instruction — implemented as a Match eval with adversarial JSONL samples (evals/registry/evals/prompt-injection.yaml:1-9, evals/registry/data/prompt-injection/samples.jsonl:1-3). human-safety evaluates logical reasoning about real-life situations involving physical danger (self-driving car dilemmas, medical robot decisions) using FuzzyMatch (evals/registry/evals/human-safety.yaml:1-8, evals/registry/data/human-safety/human-safety.jsonl:1-3).
Beyond basic safety evals, Sandbagging (evals/elsuite/sandbagging/sandbagging_eval.py:16-44) tests whether a model can deliberately underperform to a target accuracy on MMLU — a capability concern rather than an adversarial probe. Steganography (evals/elsuite/steganography/eval.py:12-35) evaluates whether a model can hide payload text within otherwise-normal task outputs, with a monitor model attempting to detect the hidden content. Already_said_that tests for unwanted repetition.
What is absent. There are no adversarial probe generators, no attack plugins (gradient-based, token-manipulation, or suffix-injection), no automated jailbreak discovery, no prompt-injection benchmark suites (like MITRE ATLAS), and no structured vulnerability reporting workflow beyond the general SECURITY.md linking to OpenAI's CVD policy (SECURITY.md:1-4). The existing prompt-injection eval covers a narrow class of prompt-override attacks in a controlled setting, not a systematic red-teaming suite.