# EleutherAI/lm-evaluation-harness

> Benchmark harness that scores a model on YAML-defined academic tasks via loglikelihood or generation requests, with bootstrap stderr.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/EleutherAI/lm-evaluation-harness (reviewed at commit `d6de81643928d653435c431bae19945d41d32520`, 2026-09-14)
- Stars: 14143 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/lm-evaluation-harness/

## Overview

EleutherAI's lm-evaluation-harness (`lm_eval`) is the reference tool for scoring a language model on public benchmarks: MMLU, HellaSwag, ARC, GSM8K, IFEval, the Open LLM Leaderboard suite and a long tail of multilingual and domain sets. At this commit `lm_eval/tasks/` holds about 220 benchmark folders and close to 14,000 YAML files. You name a model backend and a list of tasks. The harness builds prompts from a Hugging Face dataset, sends the model a batch of low-level requests, and reduces the responses to metrics such as `acc`, `acc_norm`, `exact_match` or perplexity, each with a bootstrap standard error.

It evaluates the **model**, not an application. The unit of work is a document from a fixed dataset turned into a prompt by a template. The model is only ever asked three things: the loglikelihood of a continuation, a rolling loglikelihood over a text, or a greedy generation until a stop sequence. There are no traces, no tool calls, no retrieval step and no notion of a user session. For an LLM app team it answers one question well: "which base or fine-tuned model should sit under my app, and did my fine-tune or quantisation regress general capability?" You can point it at an OpenAI-compatible endpoint, so it can also check a served model. It will not score your prompt chain, your RAG pipeline or your agent.

Scoring is reference-based by design. Metrics compare against a gold label or string, and there is no LLM-as-judge in the core. A few individual tasks call a judge from their own Python hooks (see Key components).

## Architecture

```mermaid
flowchart LR
  CLI["lm-eval run"] --> CFG["EvaluatorConfig.from_cli"]
  CFG --> SE["simple_evaluate()"]
  SE --> REG["registry.get_model()"]
  REG --> LM["LM backend (hf, vllm, api...)"]
  SE --> TM["TaskManager.load()"]
  TM --> YAML["Task YAML + utils.py"]
  SE --> EV["evaluate()"]
  EV --> REQ["build_all_requests -> Instances"]
  REQ --> LM
  LM --> FIL["apply_filters()"]
  FIL --> PR["process_results() per doc"]
  PR --> AGG["aggregation + bootstrap stderr"]
  AGG --> OUT["EvaluationTracker: JSON, HF Hub, WandB"]
```

| Component | Path | Role |
|---|---|---|
| CLI | `lm_eval/_cli/` | `lm-eval run / ls / validate`; legacy calls without a subcommand get `run` inserted |
| Orchestrator | `lm_eval/evaluator.py` | `simple_evaluate` (setup, seeds, cache, config overrides) and `evaluate` (request dispatch, scoring) |
| Model API | `lm_eval/api/model.py` | `LM` base class with `loglikelihood`, `loglikelihood_rolling`, `generate_until`; `CachingLM` SQLite wrapper |
| Backends | `lm_eval/models/` | HF, vLLM, SGLang, GGUF, OpenAI/Anthropic/LiteLLM APIs, local OpenAI-compatible servers, and more; lazy alias map |
| Tasks | `lm_eval/api/task.py`, `lm_eval/config/task.py` | `ConfigurableTask` turns a YAML config into requests and per-document metrics |
| Task discovery | `lm_eval/tasks/manager.py` | `TaskManager` indexes YAML files, groups and tags |
| Metrics | `lm_eval/api/metrics.py` | Registered metrics and aggregations, bootstrap stderr |
| Filters | `lm_eval/filters/` | Post-process generations (regex extract, take-first, majority vote) before scoring |
| Registry | `lm_eval/api/registry.py` | Lazy registries for models, metrics, filters; entry-point plugins |
| Output | `lm_eval/loggers/` | `EvaluationTracker` (JSON, HF Hub), `WandbLogger`, `TrackioLogger` |

## How a request flows

Take `lm-eval run --model hf --model_args pretrained=gpt2 --tasks arc_easy`:

1. **Parse.** `HarnessCLI` dispatches to the `run` subcommand ([harness.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/_cli/harness.py#L10-L60)). `Run._execute` builds an `EvaluatorConfig`, imports any `--plugins` modules so their decorators register, and creates the `EvaluationTracker` and `TaskManager` ([run.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/_cli/run.py#L361-L405)).
2. **Load the model.** `simple_evaluate` seeds Python, NumPy and torch, then resolves the backend with `registry.get_model(model).create_from_arg_string(...)` ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L232-L267)). With `--use_cache`, the model is wrapped in `CachingLM`, one SQLite file per rank ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L270-L291)).
3. **Load tasks.** `TaskManager.load` expands names, groups and tags into a flat `{tasks, groups, group_map}` dict ([manager.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/manager.py#L179-L240)). CLI `gen_kwargs`, `--num_fewshot` and `--predict_only` are then applied to each task ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L300-L324)).
4. **Build requests.** `evaluate` refuses tasks marked `UNSAFE_CODE` unless confirmed, then calls `task.build_all_requests` per task, sharded by rank ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L515-L560)). For `arc_easy` (`output_type: multiple_choice`) `construct_requests` emits one `loglikelihood` `Instance` per answer choice ([task.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/task.py#L1362-L1453)).
5. **Run the model.** Requests are grouped by type, duplicated by `repeats`, padded for distributed runs, and sent with `getattr(lm, reqtype)(cloned_reqs)` ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L586-L607)).
6. **Filter and score.** `task.apply_filters()` produces `filtered_resps`. For each document and filter, `task.process_results` returns per-document metric values; with `log_samples` the prompt, response and hashes are kept ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L611-L667)). For multiple choice that means argmax over loglikelihoods, raw, length-normalised and byte-normalised ([task.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/task.py#L1455-L1560)).
7. **Aggregate.** `_compute_task_aggregations` looks up each metric's aggregation (falling back to `mean`) and computes a bootstrap stderr, capped at 100 iterations for BLEU/chrF/TER ([evaluator_utils.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator_utils.py#L193-L224)).
8. **Record.** `simple_evaluate` adds model config, seeds, git hash, environment and tokenizer info ([evaluator.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L384-L425)). The CLI then saves aggregated JSON and per-sample JSONL, optionally pushes to the Hugging Face Hub and W&B/Trackio, and prints the table ([run.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/_cli/run.py#L455-L518)).

## Key components

### Task configs

A task is mostly YAML. [arc_easy.yaml](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/arc/arc_easy.yaml#L1-L23) shows the shape: `dataset_path`, splits, Jinja templates for `doc_to_text`, `doc_to_target` and `doc_to_choice`, an `output_type`, a `metric_list` and a `metadata.version`. Anything a template can't express goes into a sibling `utils.py` referenced with `!function`, for example a custom `process_docs` or `process_results`. Groups aggregate subtasks, which is how MMLU or the leaderboard report one number.

### Request types and backends

Every backend implements three methods on `LM` ([model.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/model.py#L25-L40)). That small surface is why so many backends exist. It also limits which tasks a backend can run. API backends inherit from `TemplateAPI` (concurrency, retries, tokenizer choice; [api_models.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/models/api_models.py#L105-L140)). Chat-completion endpoints raise `NotImplementedError` for loglikelihood ([openai_completions.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/models/openai_completions.py#L324-L327)), so against a chat API only `generate_until` tasks work. Most classic multiple-choice benchmarks are out of reach unless you serve the model yourself (vLLM, SGLang, or a local completions server that returns prompt logprobs).

### Metrics and statistics

Metrics are registered with `@register_metric`, which records the output types they apply to, an aggregation and direction ([metrics.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L176-L184)). `bootstrap_stderr` resamples in a multiprocessing pool in chunks of 1,000; the comment admits the estimate is slightly biased ([metrics.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L550-L586)). `simple_evaluate` defaults to 100,000 bootstrap iterations, which is noticeable on big runs. Passing `bootstrap_iters=0` from Python skips it; the `run` CLI does not expose the setting.

### Judges inside tasks

The core has no judge abstraction. Tasks can still call one from `process_results`. The `pisa_*_llm_judged` tasks send each generation to an OpenAI chat model and map a one-token reply to `acc` ([pisa/utils.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/pisa/utils.py#L85-L111), [L312-L329](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/pisa/utils.py#L312-L329)). That pattern works, but you get no judge configuration, caching or bias control from the framework.

### Caching

There are two caches. `CachingLM` stores model responses in `sqlitedict`, keyed by request hash ([model.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/model.py#L275-L300)), so a re-run with the same prompts skips inference. `--cache_requests` pickles the built `Instance` lists with `dill` so large tasks don't rebuild prompts.

## Extending it

- **New task:** drop a YAML (plus optional `utils.py`) into `lm_eval/tasks/` or pass `--include_path` to a folder of your own. `lm-eval validate --tasks ...` checks it. This is the practical way to put a private, reference-labelled dataset in front of candidate models.
- **New backend:** subclass `LM` or `TemplateAPI`, decorate with `@register_model`, and add it to `MODEL_MAPPING` ([models/__init__.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/models/__init__.py#L1-L60)).
- **Out-of-tree plugins:** ship an entry point in the `lm_eval.models` (or metrics/filters) group; `load_plugins` registers it lazily and never overrides built-ins. `--plugins my_pkg.mod` imports a module directly ([registry.py](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/registry.py#L189-L268)).
- **Library use:** call `simple_evaluate(model=HFLM(pretrained=my_model), tasks=[...])` with an in-memory model object.

## Running it

- `pip install lm-eval` installs the core. Backends and some task families are extras: `hf`, `vllm`, `api`, `litellm`, `ifeval`, `math`, `wandb` and others ([pyproject.toml](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/pyproject.toml#L52-L96)).
- Local GPUs: `--model hf` (with `accelerate launch` for data parallel) or `--model vllm`. Remote: `--model local-completions` / `local-chat-completions` with `base_url=`, or a hosted API with its key in the environment.
- Useful flags: `--limit` for smoke tests, `--log_samples --output_path` to keep per-sample records, `--apply_chat_template` for instruct models (the harness warns if it looks like you forgot), `--use_cache` for resumable runs.
- No server or database is needed. Outputs are files, optionally mirrored to the HF Hub.

## Strengths and caveats

- **Strength: comparability.** Versioned task YAML, fixed seeds, and recorded git hash, task hashes and tokenizer info make numbers reproducible and comparable with published leaderboards.
- **Strength: breadth.** Hundreds of benchmarks and dozens of backends behind one three-method interface.
- **Strength: honest statistics.** Every metric ships with a bootstrap stderr, and the aggregation code is short and readable.
- **Caveat: model-level only.** No traces, multi-turn sessions, tools, retrieval or production data. To evaluate the app built on the model you need another tool.
- **Caveat: chat APIs are second-class.** Loglikelihood tasks need prompt logprobs, which chat-completion APIs don't return.
- **Caveat: reference-based scoring.** Open-ended quality needs a custom `process_results`; the few judge-based tasks hard-code their judge.
- **Caveat: no regression tooling.** Comparing runs means diffing JSON files or using an external tracker. The harness has no baselines or thresholds.

*Sources: code at d6de816, verified Q&A.*

## How EleutherAI/lm-evaluation-harness answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Built-in metrics are defined in `lm_eval/api/metrics.py` via registration decorators. **Per-sample metrics** (passthrough functions) include `acc`, `acc_norm`, `acc_mutual_info`, `acc_bytes`, `acc_all`, `exact_match`, `perplexity`, `likelihood`, `word_perplexity`, `byte_perplexity`, `bits_per_byte`, `brier_score`, `mcc`, `f1`, `bleu`, `chrf`, `chrf++`, `ter`, and `bypass`. Each is registered with `@register_metric(metric=..., higher_is_better=..., output_type=[...], aggregation=...)` — the decorator links the metric name to its aggregation, specifies which output types it applies to, and records directionality. Example: `acc` is registered at `lm_eval/api/metrics.py:176-183` with `aggregation="mean"` and `higher_is_better=True`; `exact_match` at line 272-279 uses `aggregation="mean"` on `generate_until` output. **Aggregation functions** (registered with `@register_aggregation`) include `mean`, `median`, `nanmean`, `perplexity`, `weighted_perplexity`, `bits_per_byte`, `f1`, `matthews_corrcoef`, `bleu`, `chrf`, `chrf++`, `ter`, `brier_score`, and `bypass`. These functions take per-document metric values and reduce them to a single task-level score. Translation metrics (bleu, chrf, ter) use `sacrebleu` for corpus-level computation. **Custom per-task metrics** — any YAML task can declare a `metric_list` with arbitrary metric names, aggregation functions, and `higher_is_better` flags; arbitrary metrics without a defined aggregation default to `mean`. **Per-task vs per-trace scoring**: Per-sample metrics are computed in `task.process_results(doc, results)` — called per document in `evaluate()` at `lm_eval/evaluator.py:639-641` — which returns a dict of metric values for that document. These raw per-document values are collected into `raw_metrics[(metric, filter_key)]` lists, then aggregated at `lm_eval/evaluator_utils.py:193-204` by looking up `task.aggregation()[metric]` and calling the registered aggregation function. **Bootstrap stderr** is computed via `bootstrap_stderr()` at `lm_eval/api/metrics.py:550-586`, using multiprocessing (Pool.imap) to generate bootstrap resamples of aggregated metrics. **Filters** operate on outputs before metrics: `RegexFilter`, `TakeFirstFilter`, `MajorityVoteFilter`, `LowercaseFilter`, `MapFilter`, etc. in `lm_eval/filters/`. Each filter pipeline produces a `filtered_resps` key that metrics are computed over, so the same task can report multiple filter variants.


Citations: [lm_eval/api/metrics.py:176-184](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L176-L184) · [lm_eval/api/metrics.py:22-72](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L22-L72) · [lm_eval/evaluator.py:636-670](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L636-L670) · [lm_eval/evaluator_utils.py:193-224](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator_utils.py#L193-L224) · [lm_eval/api/metrics.py:550-587](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L550-L587)

### How is LLM-as-a-judge implemented? (insufficient evidence)

The repository does **not** implement LLM-as-a-judge evaluation. After searching the entire codebase for judge prompts, rubric-based scoring, structured output parsing for evaluation, consensus-based multi-sample judgment, and calibration/bias controls, none were found. The harness's evaluation model is reference-based: each task provides gold-standard answers (target strings or multiple-choice labels), and the harness computes deterministic metrics (accuracy, perplexity, BLEU, etc.) by comparing model outputs to these references. Tasks like `toxigen` at `lm_eval/tasks/toxigen/toxigen.yaml` and `truthfulqa` use multiple-choice or generation tasks with regex extraction and exact-match accuracy — not LLM judges. The `exact_match` metric at `lm_eval/api/metrics.py:234-266` implements string-level comparison with normalization options (case, punctuation, number, regex ignore). The `EvalResults` schema at `lm_eval/result_schema.py` defines results as typed dicts with numeric scores per task — no provision for judge verdicts, rubrics, or model-graded outputs. The project instead relies on the well-known Open LLM Leaderboard tasks (in `lm_eval/tasks/leaderboard/`) and hundreds of standard academic benchmarks, all of which use reference-based evaluation. For generation tasks, output parsing is handled by regex-based filters (`RegexFilter`, `MultiChoiceRegexFilter` in `lm_eval/filters/extraction.py`) to extract answers from free-text completions for comparison against targets, not for judgment.

> **Editor's note.** Correction: the core harness has no judge abstraction, but some tasks do use one. The pisa_*_llm_judged tasks call an OpenAI chat model from their process_results hook (lm_eval/tasks/pisa/utils.py:85-111, 312-329) and map its one-token reply to acc.

Citations: [lm_eval/api/metrics.py:234-266](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L234-L266) · [lm_eval/filters/extraction.py:15-63](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/filters/extraction.py#L15-L63) · [lm_eval/result_schema.py:1-108](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/result_schema.py#L1-L108)

### How are test datasets and cases defined, generated and versioned? (answered)

Tasks are defined as **YAML config files** in `lm_eval/tasks/`, organized into subdirectories by benchmark (arc, mmlu, hellaswag, etc.). Each YAML declares `task` name, `dataset_path` (HuggingFace dataset identifier), `dataset_name` (subset), split mappings (`training_split`, `validation_split`, `test_split`), `output_type` (`loglikelihood`, `multiple_choice`, `loglikelihood_rolling`, `generate_until`), prompt templates (`doc_to_text`, `doc_to_target`, `doc_to_choice` with Jinja2-style `{{}}` interpolation), and `metric_list`. Example: `lm_eval/tasks/arc/arc_easy.yaml` with `dataset_path: allenai/ai2_arc`. **Custom dataset functions** can be specified via `process_docs` as an inline Python callable (serialized in output). **Task discovery** uses `TaskManager` (`lm_eval/tasks/manager.py`) and `TaskIndex` which scans all YAML files in the tasks directory (and optional `include_path`), building a registry of tasks, groups, and tags. **Versioning** is per-task via `metadata.version` in each YAML (e.g. `version: 1.0`), stored in `task.VERSION`. The `TaskConfig` dataclass at `lm_eval/config/task.py:82-168` defines all configuration fields. **Groups** are defined in YAML too (e.g. `lm_eval/tasks/leaderboard/leaderboard.yaml`) with `group` name, `task` list, and `aggregate_metric_list` for hierarchical aggregation. Tags allow task selection by category. **Synthetic data generation** is not a built-in feature; tasks load from HuggingFace datasets hub or local scripts. The `bear` and `toxigen` tasks use published datasets. The `model_written_evals/` directory contains tasks from the "advanced AI risk" benchmark (EleutherAI/advanced_ai_risk dataset), auto-generated by `_generate_configs.py` scripts. There is no formal golden set mechanism beyond the per-task split definitions.


Citations: [lm_eval/tasks/arc/arc_easy.yaml:1-20](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/arc/arc_easy.yaml#L1-L20) · [lm_eval/tasks/manager.py:37-100](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/manager.py#L37-L100) · [lm_eval/config/task.py:82-170](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/config/task.py#L82-L170) · [lm_eval/tasks/leaderboard/leaderboard.yaml:1-22](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/leaderboard/leaderboard.yaml#L1-L22)

### How are evals executed and reported? (answered)

Evals are executed via the `lm-eval` CLI (`lm_eval/_cli/harness.py`), which provides `run`, `ls` (list tasks), and `validate` subcommands. The core entry point is `simple_evaluate()` in `lm_eval/evaluator.py:55-425`, which initializes the model, loads tasks via `TaskManager`, then calls `evaluate()` (line 429-714). **Runners**: For each task, `task.build_all_requests()` generates `Instance` objects across output types (`loglikelihood`, `generate_until`, etc.). Requests are grouped by type and dispatched via `getattr(lm, reqtype)(cloned_reqs)` — the abstract `LM` base class at `lm_eval/api/model.py:25` defines `loglikelihood()`, `loglikelihood_rolling()`, and `generate_until()`. Concrete backend models (HF, vLLM, API) implement these with batching. **Parallelism**: For data-parallel (DP) models, `lm.rank`/`lm.world_size` partition documents; for tensor-parallel (TP), padding equalizes batch sizes across ranks (`evaluator.py:568-584`). Results are gathered via `lm.gather_object()` across ranks. **Caching**: `lm_eval/caching/cache.py` provides pickle-based disk caching of built requests (keyed by task+num_fewshot+rank) using dill. An SQLite-based request-response cache (`CachingLM`) is available via `--use_cache` (`evaluator.py:276-289`). **Result storage**: `EvaluationTracker` at `lm_eval/loggers/evaluation_tracker.py:123-230` handles saving aggregated results as JSON (with timestamped filenames) and optionally pushing to HuggingFace hub datasets repos. The WandbLogger (`lm_eval/loggers/wandb_logger.py`) logs metrics as wandb runs. The TrackioLogger (`lm_eval/loggers/trackio_logger.py`) logs per-sample traces. **CI integration**: GitHub Actions workflows in `.github/workflows/` run tests. **Comparison/regression/dashboards**: Not built-in — results are JSON files pushed to HF hub, lacking formal regression testing or dashboards within the harness itself.


Citations: [lm_eval/evaluator.py:55-88](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L55-L88) · [lm_eval/evaluator.py:428-605](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L428-L605) · [lm_eval/api/model.py:25-100](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/model.py#L25-L100) · [lm_eval/caching/cache.py:1-80](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/caching/cache.py#L1-L80) · [lm_eval/loggers/evaluation_tracker.py:123-230](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/evaluation_tracker.py#L123-L230)

### How are traces or production data captured and linked to evaluations? (answered)

The harness does **not** have built-in OpenTelemetry instrumentation, SDK tracing, or automated production trace capture. Evaluation observability is limited to: **Logging** — the `eval_logger` (standard Python logging) throughout the codebase, plus environment info via `add_env_info()` and tokenizer info via `add_tokenizer_info()` in `lm_eval/evaluator.py:421-422`. **WandbLogger** at `lm_eval/loggers/wandb_logger.py:24-60` integrates with Weights & Biases for experiment tracking, logging aggregated metrics and config. **TrackioLogger** at `lm_eval/loggers/trackio_logger.py` provides a lightweight local-first alternative with per-sample `Trace` objects that capture prompt/response pairs as conversational traces with gold targets and metric values attached as metadata. The `_sample_to_trace()` helper at `lm_eval/loggers/trackio_logger.py:15-54` converts eval samples into `trackio.Trace` objects with the natural prompt/response for each output type (loglikelihood, multiple_choice, generate_until). **Result files** are JSON with full config, git hash, environment, tokenizer info, task versions, and per-task hashes for reproducibility (`result_schema.py` defines the schema). **Online vs offline evals**: The harness is fundamentally offline — tasks are loaded from datasets, inference is run, metrics are computed. No production data ingestion or online eval triggers exist. **Feedback/annotation**: Not present — there is no annotation UI, no feedback API, and no mechanism to incorporate human feedback into eval results within the framework itself.


Citations: [lm_eval/loggers/trackio_logger.py:15-80](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/trackio_logger.py#L15-L80) · [lm_eval/loggers/wandb_logger.py:24-70](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/wandb_logger.py#L24-L70) · [lm_eval/loggers/evaluation_tracker.py:230-300](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/evaluation_tracker.py#L230-L300) · [lm_eval/evaluator.py:420-425](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L420-L425)

### Does it support red-teaming or safety testing, and how? (answered)

The harness does **not** have built-in adversarial probes, jailbreak/prompt-injection test generators, or automated red-teaming tooling. Instead, it includes several published safety-relevant benchmarks as standard tasks that can be run like any other eval. **ToxiGen** (`lm_eval/tasks/toxigen/toxigen.yaml`) tests hate speech detection as a multiple-choice task: given a statement, classify it as hateful or not. **BEAR** (`lm_eval/tasks/bear/bear.yaml`) tests the tendency to repeat misinformation about protected groups. **Advanced AI Risk** (`lm_eval/tasks/model_written_evals/advanced_ai_risk/`) includes 18 subtasks generated from the EleutherAI/advanced_ai_risk dataset, evaluating models on corrigibility, coordination with other AIs, myopic reward-seeking, power-seeking inclination, survival instinct, and self-awareness. Each subtask (e.g., `fewshot-corrigible-less-HHH.yaml`) is a multiple-choice task where one answer matches a behavior of concern and the other does not. The `_template_yaml` defines the shared prompt format: `"Human: {{question}}

Assistant:"` with `answer_matching_behavior` vs `answer_not_matching_behavior` as choices. **Model-written persona evals** (`lm_eval/tasks/model_written_evals/persona/`) assess desire to remove safety precautions. **Sycophancy** and **winogenerated** tasks are also in `model_written_evals/`. There is **no** jailbreak/prompt-injection library, no adversarial attack plugin system, and no automated prompt mutation or red-team reporting pipeline. A task-level `UNSAFE_CODE` flag (`lm_eval/evaluator.py:522-524`) gates tasks that execute generated code, requiring explicit `--confirm_run_unsafe_code`. Vulnerability reporting is not a feature; tasks are standard benchmarks rather than adversarial discovery tools.


Citations: [lm_eval/tasks/bear/bear.yaml:1-16](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/bear/bear.yaml#L1-L16) · [lm_eval/tasks/model_written_evals/advanced_ai_risk/fewshot-corrigible-less-HHH.yaml:1-4](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/model_written_evals/advanced_ai_risk/fewshot-corrigible-less-HHH.yaml#L1-L4) · [lm_eval/evaluator.py:520-525](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L520-L525)
