LLMs Technical Reviews
Home / LLM evals and testing / lm-evaluation-harness

EleutherAI/lm-evaluation-harness

Benchmark harness that scores a model on YAML-defined academic tasks via loglikelihood or generation requests, with bootstrap stderr.

GitHub ↗★ 14kPythonMITcommit d6de816 · 2026-09-14homepage ↗

Overview

EleutherAI’s lm-evaluation-harness (lm_eval) is the reference tool for scoring a language model on public benchmarks: MMLU, HellaSwag, ARC, GSM8K, IFEval, the Open LLM Leaderboard suite and a long tail of multilingual and domain sets. At this commit lm_eval/tasks/ holds about 220 benchmark folders and close to 14,000 YAML files. You name a model backend and a list of tasks. The harness builds prompts from a Hugging Face dataset, sends the model a batch of low-level requests, and reduces the responses to metrics such as acc, acc_norm, exact_match or perplexity, each with a bootstrap standard error.

It evaluates the model, not an application. The unit of work is a document from a fixed dataset turned into a prompt by a template. The model is only ever asked three things: the loglikelihood of a continuation, a rolling loglikelihood over a text, or a greedy generation until a stop sequence. There are no traces, no tool calls, no retrieval step and no notion of a user session. For an LLM app team it answers one question well: “which base or fine-tuned model should sit under my app, and did my fine-tune or quantisation regress general capability?” You can point it at an OpenAI-compatible endpoint, so it can also check a served model. It will not score your prompt chain, your RAG pipeline or your agent.

Scoring is reference-based by design. Metrics compare against a gold label or string, and there is no LLM-as-judge in the core. A few individual tasks call a judge from their own Python hooks (see Key components).

Architecture

flowchart LR
  CLI["lm-eval run"] --> CFG["EvaluatorConfig.from_cli"]
  CFG --> SE["simple_evaluate()"]
  SE --> REG["registry.get_model()"]
  REG --> LM["LM backend (hf, vllm, api...)"]
  SE --> TM["TaskManager.load()"]
  TM --> YAML["Task YAML + utils.py"]
  SE --> EV["evaluate()"]
  EV --> REQ["build_all_requests -> Instances"]
  REQ --> LM
  LM --> FIL["apply_filters()"]
  FIL --> PR["process_results() per doc"]
  PR --> AGG["aggregation + bootstrap stderr"]
  AGG --> OUT["EvaluationTracker: JSON, HF Hub, WandB"]
Component Path Role
CLI lm_eval/_cli/ lm-eval run / ls / validate; legacy calls without a subcommand get run inserted
Orchestrator lm_eval/evaluator.py simple_evaluate (setup, seeds, cache, config overrides) and evaluate (request dispatch, scoring)
Model API lm_eval/api/model.py LM base class with loglikelihood, loglikelihood_rolling, generate_until; CachingLM SQLite wrapper
Backends lm_eval/models/ HF, vLLM, SGLang, GGUF, OpenAI/Anthropic/LiteLLM APIs, local OpenAI-compatible servers, and more; lazy alias map
Tasks lm_eval/api/task.py, lm_eval/config/task.py ConfigurableTask turns a YAML config into requests and per-document metrics
Task discovery lm_eval/tasks/manager.py TaskManager indexes YAML files, groups and tags
Metrics lm_eval/api/metrics.py Registered metrics and aggregations, bootstrap stderr
Filters lm_eval/filters/ Post-process generations (regex extract, take-first, majority vote) before scoring
Registry lm_eval/api/registry.py Lazy registries for models, metrics, filters; entry-point plugins
Output lm_eval/loggers/ EvaluationTracker (JSON, HF Hub), WandbLogger, TrackioLogger

How a request flows

Take lm-eval run --model hf --model_args pretrained=gpt2 --tasks arc_easy:

  1. Parse. HarnessCLI dispatches to the run subcommand (harness.py). Run._execute builds an EvaluatorConfig, imports any --plugins modules so their decorators register, and creates the EvaluationTracker and TaskManager (run.py).
  2. Load the model. simple_evaluate seeds Python, NumPy and torch, then resolves the backend with registry.get_model(model).create_from_arg_string(...) (evaluator.py). With --use_cache, the model is wrapped in CachingLM, one SQLite file per rank (evaluator.py).
  3. Load tasks. TaskManager.load expands names, groups and tags into a flat {tasks, groups, group_map} dict (manager.py). CLI gen_kwargs, --num_fewshot and --predict_only are then applied to each task (evaluator.py).
  4. Build requests. evaluate refuses tasks marked UNSAFE_CODE unless confirmed, then calls task.build_all_requests per task, sharded by rank (evaluator.py). For arc_easy (output_type: multiple_choice) construct_requests emits one loglikelihood Instance per answer choice (task.py).
  5. Run the model. Requests are grouped by type, duplicated by repeats, padded for distributed runs, and sent with getattr(lm, reqtype)(cloned_reqs) (evaluator.py).
  6. Filter and score. task.apply_filters() produces filtered_resps. For each document and filter, task.process_results returns per-document metric values; with log_samples the prompt, response and hashes are kept (evaluator.py). For multiple choice that means argmax over loglikelihoods, raw, length-normalised and byte-normalised (task.py).
  7. Aggregate. _compute_task_aggregations looks up each metric’s aggregation (falling back to mean) and computes a bootstrap stderr, capped at 100 iterations for BLEU/chrF/TER (evaluator_utils.py).
  8. Record. simple_evaluate adds model config, seeds, git hash, environment and tokenizer info (evaluator.py). The CLI then saves aggregated JSON and per-sample JSONL, optionally pushes to the Hugging Face Hub and W&B/Trackio, and prints the table (run.py).

Key components

Task configs

A task is mostly YAML. arc_easy.yaml shows the shape: dataset_path, splits, Jinja templates for doc_to_text, doc_to_target and doc_to_choice, an output_type, a metric_list and a metadata.version. Anything a template can’t express goes into a sibling utils.py referenced with !function, for example a custom process_docs or process_results. Groups aggregate subtasks, which is how MMLU or the leaderboard report one number.

Request types and backends

Every backend implements three methods on LM (model.py). That small surface is why so many backends exist. It also limits which tasks a backend can run. API backends inherit from TemplateAPI (concurrency, retries, tokenizer choice; api_models.py). Chat-completion endpoints raise NotImplementedError for loglikelihood (openai_completions.py), so against a chat API only generate_until tasks work. Most classic multiple-choice benchmarks are out of reach unless you serve the model yourself (vLLM, SGLang, or a local completions server that returns prompt logprobs).

Metrics and statistics

Metrics are registered with @register_metric, which records the output types they apply to, an aggregation and direction (metrics.py). bootstrap_stderr resamples in a multiprocessing pool in chunks of 1,000; the comment admits the estimate is slightly biased (metrics.py). simple_evaluate defaults to 100,000 bootstrap iterations, which is noticeable on big runs. Passing bootstrap_iters=0 from Python skips it; the run CLI does not expose the setting.

Judges inside tasks

The core has no judge abstraction. Tasks can still call one from process_results. The pisa_*_llm_judged tasks send each generation to an OpenAI chat model and map a one-token reply to acc (pisa/utils.py, L312-L329). That pattern works, but you get no judge configuration, caching or bias control from the framework.

Caching

There are two caches. CachingLM stores model responses in sqlitedict, keyed by request hash (model.py), so a re-run with the same prompts skips inference. --cache_requests pickles the built Instance lists with dill so large tasks don’t rebuild prompts.

Extending it

  • New task: drop a YAML (plus optional utils.py) into lm_eval/tasks/ or pass --include_path to a folder of your own. lm-eval validate --tasks ... checks it. This is the practical way to put a private, reference-labelled dataset in front of candidate models.
  • New backend: subclass LM or TemplateAPI, decorate with @register_model, and add it to MODEL_MAPPING (models/init.py).
  • Out-of-tree plugins: ship an entry point in the lm_eval.models (or metrics/filters) group; load_plugins registers it lazily and never overrides built-ins. --plugins my_pkg.mod imports a module directly (registry.py).
  • Library use: call simple_evaluate(model=HFLM(pretrained=my_model), tasks=[...]) with an in-memory model object.

Running it

  • pip install lm-eval installs the core. Backends and some task families are extras: hf, vllm, api, litellm, ifeval, math, wandb and others (pyproject.toml).
  • Local GPUs: --model hf (with accelerate launch for data parallel) or --model vllm. Remote: --model local-completions / local-chat-completions with base_url=, or a hosted API with its key in the environment.
  • Useful flags: --limit for smoke tests, --log_samples --output_path to keep per-sample records, --apply_chat_template for instruct models (the harness warns if it looks like you forgot), --use_cache for resumable runs.
  • No server or database is needed. Outputs are files, optionally mirrored to the HF Hub.

Strengths and caveats

  • Strength: comparability. Versioned task YAML, fixed seeds, and recorded git hash, task hashes and tokenizer info make numbers reproducible and comparable with published leaderboards.
  • Strength: breadth. Hundreds of benchmarks and dozens of backends behind one three-method interface.
  • Strength: honest statistics. Every metric ships with a bootstrap stderr, and the aggregation code is short and readable.
  • Caveat: model-level only. No traces, multi-turn sessions, tools, retrieval or production data. To evaluate the app built on the model you need another tool.
  • Caveat: chat APIs are second-class. Loglikelihood tasks need prompt logprobs, which chat-completion APIs don’t return.
  • Caveat: reference-based scoring. Open-ended quality needs a custom process_results; the few judge-based tasks hard-code their judge.
  • Caveat: no regression tooling. Comparing runs means diffing JSON files or using an external tracker. The harness has no baselines or thresholds.

Sources: code at d6de816, verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit d6de816. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Built-in metrics are defined in lm_eval/api/metrics.py via registration decorators. Per-sample metrics (passthrough functions) include acc, acc_norm, acc_mutual_info, acc_bytes, acc_all, exact_match, perplexity, likelihood, word_perplexity, byte_perplexity, bits_per_byte, brier_score, mcc, f1, bleu, chrf, chrf++, ter, and bypass. Each is registered with @register_metric(metric=..., higher_is_better=..., output_type=[...], aggregation=...) — the decorator links the metric name to its aggregation, specifies which output types it applies to, and records directionality. Example: acc is registered at lm_eval/api/metrics.py:176-183 with aggregation="mean" and higher_is_better=True; exact_match at line 272-279 uses aggregation="mean" on generate_until output. Aggregation functions (registered with @register_aggregation) include mean, median, nanmean, perplexity, weighted_perplexity, bits_per_byte, f1, matthews_corrcoef, bleu, chrf, chrf++, ter, brier_score, and bypass. These functions take per-document metric values and reduce them to a single task-level score. Translation metrics (bleu, chrf, ter) use sacrebleu for corpus-level computation. Custom per-task metrics — any YAML task can declare a metric_list with arbitrary metric names, aggregation functions, and higher_is_better flags; arbitrary metrics without a defined aggregation default to mean. Per-task vs per-trace scoring: Per-sample metrics are computed in task.process_results(doc, results) — called per document in evaluate() at lm_eval/evaluator.py:639-641 — which returns a dict of metric values for that document. These raw per-document values are collected into raw_metrics[(metric, filter_key)] lists, then aggregated at lm_eval/evaluator_utils.py:193-204 by looking up task.aggregation()[metric] and calling the registered aggregation function. Bootstrap stderr is computed via bootstrap_stderr() at lm_eval/api/metrics.py:550-586, using multiprocessing (Pool.imap) to generate bootstrap resamples of aggregated metrics. Filters operate on outputs before metrics: RegexFilter, TakeFirstFilter, MajorityVoteFilter, LowercaseFilter, MapFilter, etc. in lm_eval/filters/. Each filter pipeline produces a filtered_resps key that metrics are computed over, so the same task can report multiple filter variants.

How is LLM-as-a-judge implemented?

insufficient evidence

The repository does not implement LLM-as-a-judge evaluation. After searching the entire codebase for judge prompts, rubric-based scoring, structured output parsing for evaluation, consensus-based multi-sample judgment, and calibration/bias controls, none were found. The harness's evaluation model is reference-based: each task provides gold-standard answers (target strings or multiple-choice labels), and the harness computes deterministic metrics (accuracy, perplexity, BLEU, etc.) by comparing model outputs to these references. Tasks like toxigen at lm_eval/tasks/toxigen/toxigen.yaml and truthfulqa use multiple-choice or generation tasks with regex extraction and exact-match accuracy — not LLM judges. The exact_match metric at lm_eval/api/metrics.py:234-266 implements string-level comparison with normalization options (case, punctuation, number, regex ignore). The EvalResults schema at lm_eval/result_schema.py defines results as typed dicts with numeric scores per task — no provision for judge verdicts, rubrics, or model-graded outputs. The project instead relies on the well-known Open LLM Leaderboard tasks (in lm_eval/tasks/leaderboard/) and hundreds of standard academic benchmarks, all of which use reference-based evaluation. For generation tasks, output parsing is handled by regex-based filters (RegexFilter, MultiChoiceRegexFilter in lm_eval/filters/extraction.py) to extract answers from free-text completions for comparison against targets, not for judgment.

Editor's note. Correction: the core harness has no judge abstraction, but some tasks do use one. The pisa_*_llm_judged tasks call an OpenAI chat model from their process_results hook (lm_eval/tasks/pisa/utils.py:85-111, 312-329) and map its one-token reply to acc.

How are test datasets and cases defined, generated and versioned?

answered

Tasks are defined as YAML config files in lm_eval/tasks/, organized into subdirectories by benchmark (arc, mmlu, hellaswag, etc.). Each YAML declares task name, dataset_path (HuggingFace dataset identifier), dataset_name (subset), split mappings (training_split, validation_split, test_split), output_type (loglikelihood, multiple_choice, loglikelihood_rolling, generate_until), prompt templates (doc_to_text, doc_to_target, doc_to_choice with Jinja2-style {{}} interpolation), and metric_list. Example: lm_eval/tasks/arc/arc_easy.yaml with dataset_path: allenai/ai2_arc. Custom dataset functions can be specified via process_docs as an inline Python callable (serialized in output). Task discovery uses TaskManager (lm_eval/tasks/manager.py) and TaskIndex which scans all YAML files in the tasks directory (and optional include_path), building a registry of tasks, groups, and tags. Versioning is per-task via metadata.version in each YAML (e.g. version: 1.0), stored in task.VERSION. The TaskConfig dataclass at lm_eval/config/task.py:82-168 defines all configuration fields. Groups are defined in YAML too (e.g. lm_eval/tasks/leaderboard/leaderboard.yaml) with group name, task list, and aggregate_metric_list for hierarchical aggregation. Tags allow task selection by category. Synthetic data generation is not a built-in feature; tasks load from HuggingFace datasets hub or local scripts. The bear and toxigen tasks use published datasets. The model_written_evals/ directory contains tasks from the "advanced AI risk" benchmark (EleutherAI/advanced_ai_risk dataset), auto-generated by _generate_configs.py scripts. There is no formal golden set mechanism beyond the per-task split definitions.

How are evals executed and reported?

answered

Evals are executed via the lm-eval CLI (lm_eval/_cli/harness.py), which provides run, ls (list tasks), and validate subcommands. The core entry point is simple_evaluate() in lm_eval/evaluator.py:55-425, which initializes the model, loads tasks via TaskManager, then calls evaluate() (line 429-714). Runners: For each task, task.build_all_requests() generates Instance objects across output types (loglikelihood, generate_until, etc.). Requests are grouped by type and dispatched via getattr(lm, reqtype)(cloned_reqs) — the abstract LM base class at lm_eval/api/model.py:25 defines loglikelihood(), loglikelihood_rolling(), and generate_until(). Concrete backend models (HF, vLLM, API) implement these with batching. Parallelism: For data-parallel (DP) models, lm.rank/lm.world_size partition documents; for tensor-parallel (TP), padding equalizes batch sizes across ranks (evaluator.py:568-584). Results are gathered via lm.gather_object() across ranks. Caching: lm_eval/caching/cache.py provides pickle-based disk caching of built requests (keyed by task+num_fewshot+rank) using dill. An SQLite-based request-response cache (CachingLM) is available via --use_cache (evaluator.py:276-289). Result storage: EvaluationTracker at lm_eval/loggers/evaluation_tracker.py:123-230 handles saving aggregated results as JSON (with timestamped filenames) and optionally pushing to HuggingFace hub datasets repos. The WandbLogger (lm_eval/loggers/wandb_logger.py) logs metrics as wandb runs. The TrackioLogger (lm_eval/loggers/trackio_logger.py) logs per-sample traces. CI integration: GitHub Actions workflows in .github/workflows/ run tests. Comparison/regression/dashboards: Not built-in — results are JSON files pushed to HF hub, lacking formal regression testing or dashboards within the harness itself.

How are traces or production data captured and linked to evaluations?

answered

The harness does not have built-in OpenTelemetry instrumentation, SDK tracing, or automated production trace capture. Evaluation observability is limited to: Logging — the eval_logger (standard Python logging) throughout the codebase, plus environment info via add_env_info() and tokenizer info via add_tokenizer_info() in lm_eval/evaluator.py:421-422. WandbLogger at lm_eval/loggers/wandb_logger.py:24-60 integrates with Weights & Biases for experiment tracking, logging aggregated metrics and config. TrackioLogger at lm_eval/loggers/trackio_logger.py provides a lightweight local-first alternative with per-sample Trace objects that capture prompt/response pairs as conversational traces with gold targets and metric values attached as metadata. The _sample_to_trace() helper at lm_eval/loggers/trackio_logger.py:15-54 converts eval samples into trackio.Trace objects with the natural prompt/response for each output type (loglikelihood, multiple_choice, generate_until). Result files are JSON with full config, git hash, environment, tokenizer info, task versions, and per-task hashes for reproducibility (result_schema.py defines the schema). Online vs offline evals: The harness is fundamentally offline — tasks are loaded from datasets, inference is run, metrics are computed. No production data ingestion or online eval triggers exist. Feedback/annotation: Not present — there is no annotation UI, no feedback API, and no mechanism to incorporate human feedback into eval results within the framework itself.

Does it support red-teaming or safety testing, and how?

answered

The harness does not have built-in adversarial probes, jailbreak/prompt-injection test generators, or automated red-teaming tooling. Instead, it includes several published safety-relevant benchmarks as standard tasks that can be run like any other eval. ToxiGen (lm_eval/tasks/toxigen/toxigen.yaml) tests hate speech detection as a multiple-choice task: given a statement, classify it as hateful or not. BEAR (lm_eval/tasks/bear/bear.yaml) tests the tendency to repeat misinformation about protected groups. Advanced AI Risk (lm_eval/tasks/model_written_evals/advanced_ai_risk/) includes 18 subtasks generated from the EleutherAI/advanced_ai_risk dataset, evaluating models on corrigibility, coordination with other AIs, myopic reward-seeking, power-seeking inclination, survival instinct, and self-awareness. Each subtask (e.g., fewshot-corrigible-less-HHH.yaml) is a multiple-choice task where one answer matches a behavior of concern and the other does not. The _template_yaml defines the shared prompt format: `"Human: {{question}}

Assistant:"withanswer_matching_behaviorvsanswer_not_matching_behavior as choices. **Model-written persona evals** (lm_eval/tasks/model_written_evals/persona/) assess desire to remove safety precautions. **Sycophancy** and **winogenerated** tasks are also in model_written_evals/. There is **no** jailbreak/prompt-injection library, no adversarial attack plugin system, and no automated prompt mutation or red-team reporting pipeline. A task-level UNSAFE_CODE flag (lm_eval/evaluator.py:522-524) gates tasks that execute generated code, requiring explicit --confirm_run_unsafe_code`. Vulnerability reporting is not a feature; tasks are standard benchmarks rather than adversarial discovery tools.