# Which evaluation metrics and scorers are provided, and how are they implemented?

> LLM evals and testing — a good answer covers: Built-in metrics (heuristic, statistical, model-based); custom scorers; per-task vs per-trace scoring.

Canonical page: https://llms-technical-reviews.com/evals/q/metrics/

## Verdict

[DeepEval](/p/deepeval/) and [Opik](/p/opik/) ship the broadest metric libraries for LLM apps, and [Ragas](/p/ragas/) remains the reference for RAG metrics. For benchmark-style scoring with honest error bars, [Inspect](/p/inspect_ai/) and [lm-evaluation-harness](/p/lm-evaluation-harness/) are the strongest.

**Metric libraries for applications.** DeepEval has about 50 metric packages. Most are judge pipelines: extract claims, ask for a verdict per claim, return a ratio, and pass when `score >= threshold`. Ragas works the same way for faithfulness (statements, then NLI verdicts). `FaithfulnesswithHHEM` swaps the judge for Vectara's HHEM classifier, and non-LLM context precision and recall variants use embeddings. Opik pairs LLM judges with the widest set of classic metrics here: BLEU, ROUGE, METEOR, BERTScore, Levenshtein and divergence metrics. A `RagasMetricWrapper` lets you run Ragas metrics inside Opik. [Phoenix](/p/phoenix/)'s `phoenix-evals` turns judge labels into numbers through a `label_score_map` and adds code metrics such as `PrecisionRecallFScore`.

**Assertion scorecards.** [promptfoo](/p/promptfoo/) has about 70 assertion types. They range from `equals` and `is-json` through `rouge-n` and embedding similarity to `llm-rubric` and checks over recorded trace spans. Any type can be negated with `not-`, and weights and named metrics combine them into one score.

**Platforms without a metric library.** [Langfuse](/p/langfuse/) scores with LLM judges, user code or decision models, and ships no BLEU, ROUGE or embedding metrics. Code evaluators run only on the observation and experiment path, and only when a dispatcher is configured.

**Task and benchmark scoring.** lm-evaluation-harness registers each metric with an aggregation and a direction, and reports a bootstrap stderr (100,000 iterations by default). Inspect separates per-sample scorers from metrics (clustered stderr, bootstrap and Wilson CIs) and from epoch reducers such as `pass_at`. [OpenAI Evals](/p/openai-evals/) reports accuracy with a "bootstrap" spread. The editor found that this is the standard deviation over half-size subsamples drawn without replacement, not a classic bootstrap. [garak](/p/garak/) is different in kind. Its detectors score each output from 0 to 1, a 0.5 threshold decides pass or fail, and the result is an attack success rate, not a quality score.

Pick: DeepEval or Opik for a large ready-made set of app metrics.
Pick: Ragas when RAG faithfulness and context metrics are the main need.
Pick: Inspect or lm-evaluation-harness when you need confidence intervals on task accuracy.

## Per-project answers

### langfuse/langfuse (answered)

Langfuse provides **three evaluator types** that produce scores, plus deterministic sampling and blocking. **LLM-as-a-judge** (`evalService.ts` lines 941–1233) — the primary mechanism — constructs a prompt from a template, calls a configurable LLM (OpenAI, Anthropic, etc.) with Zod-validated structured output, and normalizes the response into **NUMERIC**, **BOOLEAN**, or **CATEGORICAL** scores (`outputDefinition.ts` lines 1–428). The output definition is a persisted schema (`PersistedEvalOutputDefinitionSchema`) supporting numeric ranges (`minValue`/`maxValue`), boolean verdicts, and categorical choices (single or multi-match). The `toNormalizedScores` helper (`evalService.ts` lines 1235–1270) converts the LLM response — e.g. a numeric 0–1, a boolean true/false, or a string array — into `CodeEvalScoreWithName[]` records with a shared comment (the model's reasoning). Each score carries a `dataType`, `value`, and optionally `configId` and `metadata`. **Code-based evaluators** (`codeEvalExecution.ts`) dispatch user-supplied code (JS/Python) to either an AWS Lambda dispatcher or a local process, returning scores in a `{scores: [...]}` JSON contract — NO built-in heuristic or statistical metrics exist outside these three types. **Decision-model evaluators** (`decisionModelEvaluatorExecution.ts`) call a separate "TypeSafe" API for choice/score/boolean questions and map responses into scores. **Deterministic sampling** (`deterministicSampling.ts` lines 1–31) controls what fraction of traces are scored, using a SHA‑256 over target ID and sampling rate. **Per-trace vs per-observation scoring** is handled via `targetObject`: `TRACE` scores the whole trace at the trace level, `DATASET` ties scores to dataset items, and `EVENT` (observation-level) schedules per-observation evaluations (`observationEval/`). All scores are persisted with `source: "EVAL"` and share an `EvalScoreWritePayload` structure (`evalScoreEvent.ts`).

> **Editor's note.** Correction: code and decision-model evaluators only run in the observation-level and experiment path (`processObservationEval`); trace- and dataset-level rules load only LLM-as-judge evaluators. The `insecure-local` code dispatcher runs TypeScript only, in a `node:vm` context inside the worker, and production has no code-eval dispatcher unless `LANGFUSE_CODE_EVAL_DISPATCHER` is set.

Citations: [worker/src/features/evaluation/evalService.ts:941-1270](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L941-L1270) · [packages/shared/src/server/evals/codeEvalExecution.ts:210-320](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/codeEvalExecution.ts#L210-L320) · [packages/shared/src/server/evals/decisionModelEvaluatorExecution.ts:252-307](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/decisionModelEvaluatorExecution.ts#L252-L307) · [worker/src/features/evaluation/evalScoreEvent.ts:1-68](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalScoreEvent.ts#L1-L68)

### promptfoo/promptfoo (answered)

Built-in metrics registered via ASSERTION_HANDLERS map. Heuristic: equals, contains/contains-all/contains-any, regex, starts-with, is-json, is-sql, is-html, is-xml, is-refusal, finish-reason, cost. Statistical/distance: levenshtein, bleu, gleu, rouge-n, meteor, perplexity-score, similar (cosine/dot/euclidean), latency, word-count, tool-call-f1. Model-based: llm-rubric, model-graded-closedqa, factuality, g-eval, agent-rubric, search-rubric (src/matchers/llmGrading.ts). Trace-aware: trace-error-spans, trace-span-count, trace-span-duration, five trajectory types load spans from SQLite via loadTraceData. Custom: javascript/python/ruby file:// references return GradingResult. weight=0 forces metric-only pass. Custom per-test assertScoringFunction overrides default aggregation.


Citations: [src/assertions/index.ts:229-320](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L229-L320) · [src/assertions/index.ts:125-137](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L125-L137) · [src/assertions/index.ts:670-678](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L670-L678) · [src/evaluator.ts:2632-2642](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L2632-L2642)

### comet-ml/opik (answered)

Opik provides a three-tier metrics system built on an abstract `BaseMetric` class (in `sdks/python/src/opik/evaluation/metrics/base_metric.py:34-104`). All metrics return a `ScoreResult` (value 0.0–1.0, plus name/reason/metadata). 

**Heuristic/statistical metrics** — no LLM calls needed. Includes: `Equals` (exact/case-insensitive string match), `Contains` (substring detection), `RegexMatch`, `IsJson`, `LevenshteinRatio`, `SentenceBLEU`/`CorpusBLEU` (wrapping NLTK), `ROUGE` (rouge1/2/L/Lsum, wrapping `rouge_score`), `ChrF`, `GLEU`, `METEOR`, `BERTScore` (wrapping `bert-score`), `JSDivergence`/`KLDivergence`, `SpearmanRanking`, `Readability`, `Tone`, `Sentiment`, `VADERSentiment`, and `LanguageAdherenceMetric`. All live in `sdks/python/src/opik/evaluation/metrics/heuristics/`. 

**LLM-judge metrics** — use a judge LLM (configurable model name, defaults via `OPIK_DEFAULT_LLM`) with structured JSON output via Pydantic response-format schemas. Includes: `Hallucination` (1.0 if hallucination detected), `AnswerRelevance`, `ContextPrecision`, `ContextRecall`, `Factuality` (per-claim scoring), `Moderation` (0.0–1.0 content-appropriateness), `Usefulness`, `TrajectoryAccuracy` (ReAct agent trajectory quality), `SycEval` (sycophancy detection with rebuttal generation), `StructuredOutputCompliance`, and `LLMJuriesJudge`. Each lives under `sdks/python/src/opik/evaluation/metrics/llm_judges/<name>/`. The `GEval` metric (`llm_judges/g_eval/metric.py:32-288`) is a generalised LLM-as-a-judge that builds a reusable chain-of-thought prompt from `task_introduction` and `evaluation_criteria`, caches it (LRU, max 128), and scores each output. `GEvalPreset` wraps pre-built rubrics including: `QARelevanceJudge`, `DemographicBiasJudge`, `GenderBiasJudge`, `PoliticalBiasJudge`, `RegionalBiasJudge`, `ReligiousBiasJudge`, `ComplianceRiskJudge`, `DialogueHelpfulnessJudge`, `PromptUncertaintyJudge`, `AgentTaskCompletionJudge`, `AgentToolCorrectnessJudge`, `SummarizationCoherenceJudge`, `SummarizationConsistencyJudge` (all from `sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval_presets.py`).

**RAGAS integration**: `RagasMetricWrapper` (`metrics/ragas_metric.py:13-77`) wraps any `ragas.SingleTurnMetric` as an Opik `BaseMetric`, mapping 'input'/'output' keys to RAGAS's 'user_input'/'response'.

**Custom scoring**: `ScorerFunction` (`scorers/scorer_function.py:8-27`) is a protocol taking `(dataset_item, task_outputs, task_span?)` returning ScoreResults — wrapped into metrics via `scorer_wrapper_metric.wrap_scorer_functions()`.

**Conversation metrics** under `metrics/conversation/` assess multi-turn threads: `ConversationThreadMetric` wraps a per-turn metric into a thread-level score; `GEvalConversationMetric` variant judges entire conversations. Specific LLM-judge conversation metrics include `ConversationalCoherenceMetric`, `SessionCompletenessQuality`, `UserFrustrationMetric`, and heuristic ones like `KnowledgeRetentionMetric` and `ConversationDegenerationMetric`.

**AggregatedMetric** (`metrics/aggregated_metric.py`) wraps metrics and computes statistics across all items.


Citations: [sdks/python/src/opik/evaluation/metrics/__init__.py:1-152](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/__init__.py#L1-L152) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:32-288](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L32-L288) · [sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py:139-213](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py#L139-L213) · [sdks/python/src/opik/evaluation/metrics/ragas_metric.py:13-77](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/ragas_metric.py#L13-L77)

### openai/evals (answered)

**Built-in metrics** are implemented as standalone functions in `evals/metrics.py`. `get_accuracy(events)` computes `sum(correct)/total` from `match` events (`evals/metrics.py:12-18`). `get_bootstrap_accuracy_std` resamples with replacement for uncertainty (`evals/metrics.py:21-23`). `get_confusion_matrix` builds an N×(N+1) array from `expected` vs `picked` labels (`evals/metrics.py:26-40`), then `compute_matthew_corr`, `compute_precision`, `compute_recall`, `compute_f_score`, and `compute_averaged_f_score` derive classification metrics from it (`evals/metrics.py:43-73`).

**Custom per-sample scoring** uses `evals.record.record_metrics(**kwargs)` — a free-form dict attached to a `"metrics"` event on the `RecorderBase` (`evals/record.py:248-249`). Standard evals like `FuzzyMatch` record per-sample `accuracy` (float) and `f1_score` (`evals/elsuite/basic/fuzzy_match.py:48-50`). Others record custom keys: `Includes` records correctness per sample (`evals/elsuite/basic/includes.py:45-47`), `JsonValidator` records `accuracy` alone, and `ModelBasedClassify` records `choice`, `score`, and optionally `metascore` (`evals/elsuite/modelgraded/classify.py:93-98`).

**Per-task aggregation** happens in each eval's `run()` method, which returns a `dict[str, float]`. `Match.run()` averages `accuracy` and `bootstrap_std` from all `match` events via `recorder.get_events("match")` (`evals/elsuite/basic/match.py:58-65`). `MultipleChoice.run()` calls `evals.metrics.get_accuracy()` on match events (`evals/elsuite/multiple_choice.py:95-100`). `ModelBasedClassify.run()` additionally computes choice-count distributions, per-choice score averages, and metascore (`evals/elsuite/modelgraded/classify.py:104-127`). The `BaseEvalSpec` declares a `metrics` list (e.g. `[accuracy]`) and a `higher_is_better` flag (`evals/base.py:36-44`).

> **Editor's note.** Correction: `get_bootstrap_accuracy_std` is not a with-replacement bootstrap. It returns the standard deviation of accuracy over 1,000 random half-size subsamples drawn without replacement (evals/metrics.py L21-L23).

Citations: [evals/elsuite/basic/fuzzy_match.py:40-60](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/basic/fuzzy_match.py#L40-L60) · [evals/elsuite/modelgraded/classify.py:93-127](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L93-L127) · [evals/record.py:44-70](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L44-L70) · [evals/base.py:30-48](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/base.py#L30-L48)

### confident-ai/deepeval (answered)

DeepEval provides ~50+ built-in metrics across four categories. **Heuristic/statistical** metrics need no LLM: `ExactMatchMetric` does strict string comparison (`deepeval/metrics/exact_match/exact_match.py:12-88`), `PatternMatchMetric` applies regex, and the `Scorer` class offers ROUGE, BLEU, BERTScore, and exact/quasi-exact match utilities (`deepeval/scorer/scorer.py:11-160`). **LLM-based metrics** use a judge LLM: `FaithfulnessMetric` extracts claims from `actual_output` and checks each against `retrieval_context` via a yes/no verdict loop (`deepeval/metrics/faithfulness/faithfulness.py:66-100`), `HallucinationMetric` and `AnswerRelevancyMetric` follow the same QAG (question-answer generation) pattern — calling `generate_qag_verdicts` to classify statements, then `score_qag_verdicts` to aggregate. `ToxicityMetric` and `BiasMetric` extract opinions/statements and judge each for toxicity/bias. **Configurable LLM judges**: `GEval` accepts `criteria`, `evaluation_steps`, and `rubric` (score-range → outcome mappings), renders a structured prompt, and returns a Pydantic-validated `ReasonScore` (`deepeval/metrics/g_eval/g_eval.py:53-120`, `deepeval/metrics/g_eval/schema.py:5-13`). `ArenaGEval` compares two contestants side-by-side. `DAGMetric` composes multiple judgement and task nodes into an acyclic graph for multi-step evaluation (`deepeval/metrics/dag/dag.py:25-70`). **Agentic/conversational metrics**: `ToolCorrectnessMetric`, `GoalAccuracyMetric`, `AgentLoopDetectionMetric`, turn-level variants (`TurnFaithfulnessMetric`), and `ConversationCompletenessMetric`. Each metric is a subclass of `BaseMetric` (single-turn `LLMTestCase`), `BaseConversationalMetric` (multi-turn `ConversationalTestCase`), or `BaseArenaMetric` (pairwise `ArenaTestCase`) (`deepeval/metrics/base_metric.py:111-350`). Metrics define `_required_params` (e.g. `[INPUT, ACTUAL_OUTPUT, RETRIEVAL_CONTEXT]`), implement `measure()` and `a_measure()`, and set `self.score`, `self.reason`, `self.success` (score ≥ threshold). Per-test-case scoring: `execute_test_cases` iterates test cases and runs each metric against each case individually (`deepeval/evaluate/execute/e2e.py:144-270`). Custom metrics are created by subclassing a base class and providing the `measure` method.


Citations: [deepeval/metrics/__init__.py:1-166](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/__init__.py#L1-L166) · [deepeval/scorer/scorer.py:11-160](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/scorer/scorer.py#L11-L160) · [deepeval/metrics/faithfulness/faithfulness.py:66-100](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/faithfulness/faithfulness.py#L66-L100) · [deepeval/metrics/g_eval/g_eval.py:53-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L53-L80)

### vibrantlabsai/ragas (answered)

Ragas provides ~30+ built-in metrics across three implementation categories.

**Heuristic/statistical metrics** (no LLM call): `ExactMatch` (string equality), `StringPresence` (substring match), `NonLLMStringSimilarity` (Levenshtein/Hamming/Jaro/Jaro-Winkler via rapidfuzz), `BleuScore`, `RougeScore`, `ChrfScore`, and `DataCompyScore`. These are pure Python computations implementing SingleTurnMetric. Example: `ExactMatch` returns `float(sample.reference == sample.response)` (metrics/_string.py:29-35).

**LLM-based metrics**: `Faithfulness` (statement decomposition → NLI verdicts), `AnswerRelevancy`, `AnswerCorrectness`, `FactualCorrectness`, `ContextPrecision` (multiple variants: LLM, NonLLM, ID-based), `ContextRecall`, `ContextEntityRecall`, `NoiseSensitivity`, `SummarizationScore`, `TopicAdherenceScore`, `ToolCallAccuracy`, `ToolCallF1`, `MultiModalFaithfulness`, `MultiModalRelevance`, `SQLSemanticEquivalence`, `GoalAccuracy` (agent). These extend `MetricWithLLM` and implement `_single_turn_ascore()` which calls a `PydanticPrompt.generate()` against an LLM. Example: Faithfulness breaks the response into atomic statements using `StatementGeneratorPrompt`, then judges each against retrieved contexts via `NLIStatementPrompt` (metrics/_faithfulness.py:152-214).

**Hybrid/embedding-based**: `FaithfulnesswithHHEM` replaces the LLM NLI stage with a HuggingFace transformer (`vectara/hallucination_evaluation_model`) (metrics/_faithfulness.py:217-273). `SemanticSimilarity` uses embeddings. `NonLLM variants` of ContextPrecision/ContextRecall use embedding cosine similarity instead of LLM calls.

**Custom metrics**: Three decorator-based primitives: `@discrete_metric` (categorical output, e.g. 'pass'/'fail'), `@numeric_metric` (continuous float), `@ranking_metric` (ranked output). The `SimpleLLMMetric` class accepts a custom prompt string and Pydantic response model for ad-hoc LLM-as-a-judge metrics with save/load/reproducibility (metrics/base.py:846-935). `AspectCritic` and `SimpleCriteriaScore` take user-defined rubric strings for binary/discrete scoring with automatic strictness-based majority voting.

**Per-trace vs per-task**: All metrics implement `SingleTurnMetric` (single QA pair) or `MultiTurnMetric` (conversation). The `Metric` base class declares `required_columns` per type, e.g. Faithfulness needs `{user_input, response, retrieved_contexts}` for single-turn. Collection metrics (e.g. `AnswerCorrectness` composite) internally orchestrate multiple sub-metrics per trace.


Citations: [src/ragas/metrics/__init__.py:1-99](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/__init__.py#L1-L99) · [src/ragas/metrics/base.py:74-94](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L74-L94) · [src/ragas/metrics/_string.py:20-99](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_string.py#L20-L99) · [src/ragas/metrics/_faithfulness.py:134-215](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_faithfulness.py#L134-L215) · [src/ragas/metrics/base.py:688-935](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L688-L935) · [src/ragas/metrics/discrete.py:1-50](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/discrete.py#L1-L50)

### EleutherAI/lm-evaluation-harness (answered)

Built-in metrics are defined in `lm_eval/api/metrics.py` via registration decorators. **Per-sample metrics** (passthrough functions) include `acc`, `acc_norm`, `acc_mutual_info`, `acc_bytes`, `acc_all`, `exact_match`, `perplexity`, `likelihood`, `word_perplexity`, `byte_perplexity`, `bits_per_byte`, `brier_score`, `mcc`, `f1`, `bleu`, `chrf`, `chrf++`, `ter`, and `bypass`. Each is registered with `@register_metric(metric=..., higher_is_better=..., output_type=[...], aggregation=...)` — the decorator links the metric name to its aggregation, specifies which output types it applies to, and records directionality. Example: `acc` is registered at `lm_eval/api/metrics.py:176-183` with `aggregation="mean"` and `higher_is_better=True`; `exact_match` at line 272-279 uses `aggregation="mean"` on `generate_until` output. **Aggregation functions** (registered with `@register_aggregation`) include `mean`, `median`, `nanmean`, `perplexity`, `weighted_perplexity`, `bits_per_byte`, `f1`, `matthews_corrcoef`, `bleu`, `chrf`, `chrf++`, `ter`, `brier_score`, and `bypass`. These functions take per-document metric values and reduce them to a single task-level score. Translation metrics (bleu, chrf, ter) use `sacrebleu` for corpus-level computation. **Custom per-task metrics** — any YAML task can declare a `metric_list` with arbitrary metric names, aggregation functions, and `higher_is_better` flags; arbitrary metrics without a defined aggregation default to `mean`. **Per-task vs per-trace scoring**: Per-sample metrics are computed in `task.process_results(doc, results)` — called per document in `evaluate()` at `lm_eval/evaluator.py:639-641` — which returns a dict of metric values for that document. These raw per-document values are collected into `raw_metrics[(metric, filter_key)]` lists, then aggregated at `lm_eval/evaluator_utils.py:193-204` by looking up `task.aggregation()[metric]` and calling the registered aggregation function. **Bootstrap stderr** is computed via `bootstrap_stderr()` at `lm_eval/api/metrics.py:550-586`, using multiprocessing (Pool.imap) to generate bootstrap resamples of aggregated metrics. **Filters** operate on outputs before metrics: `RegexFilter`, `TakeFirstFilter`, `MajorityVoteFilter`, `LowercaseFilter`, `MapFilter`, etc. in `lm_eval/filters/`. Each filter pipeline produces a `filtered_resps` key that metrics are computed over, so the same task can report multiple filter variants.


Citations: [lm_eval/api/metrics.py:176-184](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L176-L184) · [lm_eval/api/metrics.py:22-72](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L22-L72) · [lm_eval/evaluator.py:636-670](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L636-L670) · [lm_eval/evaluator_utils.py:193-224](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator_utils.py#L193-L224) · [lm_eval/api/metrics.py:550-587](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L550-L587)

### Arize-ai/phoenix (answered)

Phoenix provides three categories of evaluators: code/heuristic, LLM-based, and statistical. **Code evaluators** include `exact_match` (string equality check) at `packages/phoenix-evals/src/phoenix/evals/metrics/exact_match.py:37`, `MatchesRegex` (regex matching) and `PrecisionRecallFScore` (macro/micro/weighted precision, recall, F-beta with binary one-vs-rest) at `packages/phoenix-evals/src/phoenix/evals/metrics/precision_recall.py:66-424`. **LLM-based evaluators** subclass `ClassificationEvaluator` and include `CorrectnessEvaluator`, `FaithfulnessEvaluator`, `HallucinationEvaluator`, `CompletenessEvaluator`, `ConcisenessEvaluator`, `DocumentRelevanceEvaluator`, `RetrievalRelevanceEvaluator`, `ToxicityEvaluator`, `RefusalEvaluator`, `PiiDetectionEvaluator`, `UserFrictionEvaluator`, and three tool-oriented evaluators (`ToolInvocationEvaluator`, `ToolSelectionEvaluator`, `ToolResponseHandlingEvaluator`) — all in `packages/phoenix-evals/src/phoenix/evals/metrics/`. Each LLM evaluator loads a generated prompt config (e.g. `_correctness_classification_evaluator_config.py` defines a rubric string, model `choices` like `{"correct": 1.0, "incorrect": 0.0}`, and `optimization_direction`). The base `Evaluator` class (`packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506`) handles input remapping via `input_mapping` (field-name or JSONPath-based), Pydantic schema validation, and returns `Score` objects. The `Score` dataclass (`evaluators.py:158-271`) carries `name`, `score` (float/int), `label` (string), `explanation`, `metadata`, `kind` ("human"|"llm"|"code"), and `direction` ("maximize"|"minimize"|"neutral"). **Per-trace vs per-task**: Evaluators operate on a single `EvalInput` dict (one row = one trace or span); `evaluate_dataframe()` and `async_evaluate_dataframe()` apply an evaluator list over every row and add score columns back to the DataFrame. **Custom evaluators** are created via the `@create_evaluator(name=..., kind=...)` decorator which wraps any sync/async function into an `Evaluator` instance with auto-generated Pydantic input schema from function signature, and `_convert_to_score` handles returns of type `Score`, `bool`, `float`, `str`, `dict`, or `tuple`.


Citations: [packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L303-L506) · [packages/phoenix-evals/src/phoenix/evals/metrics/__init__.py:1-37](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/metrics/__init__.py#L1-L37) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:828-1128](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L828-L1128)

### NVIDIA/garak (answered)

Scoring lives in `garak/evaluators/base.py`. The `Evaluator` base class evaluates detector scores per-probe, per-detector. Each `Attempt` carries `detector_results` — a dict keyed by detector name, mapping to a list of floats (one per generation). `Evaluator.evaluate()` iterates attempts, groups results by detector, and calls `_evaluate_one_detector()` which counts passes, fails, and nulls using a pluggable `test()` method. Two concrete evaluators ship: `ThresholdEvaluator` (pass if score < threshold, default 0.5) and `ZeroToleranceEvaluator` (pass only if score == 0.0; `evaluators/base.py:481-488`). The default threshold of 0.5 is set in `garak.core.yaml` as `eval_threshold`.

Confidence intervals use non-parametric bootstrap: `calculate_bootstrap_ci()` in `analyze/bootstrap_ci.py:90-105` resamples results with replacement (default 10,000 iterations, 95% confidence) and corrects for detector sensitivity/specificity loaded from `data/detectors_eval/detector_metrics_summary.json` via `analyze/detector_metrics.py`. Calibration Z-scores: `Calibration` class (`analyze/calibration.py:17-99`) loads prior-run distributions from `data/calibration/calibration.json` and computes `(score - mu) / sigma`, gated by `MINIMUM_STD_DEV = 1/30`. These feed "defcon" ratings (1-5, `analyze/__init__.py:48-58`).

Detectors range from heuristic (`StringDetector` — substring/word/prefix matching with optional Unicode normalization, `detectors/base.py:197-272`) to model-based (`HFDetector` wraps HuggingFace text-classification pipelines, normalizing logits to 0-1, `detectors/base.py:82-194`), and trigger-list matching (`TriggerListDetector`, `detectors/base.py:275-304`). Scoring is per-trace (each generation output independently), then aggregated to per-probe/per-detector summaries in the eval record.


Citations: [garak/evaluators/base.py:52-62](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/evaluators/base.py#L52-L62) · [garak/evaluators/base.py:224-311](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/evaluators/base.py#L224-L311) · [garak/evaluators/base.py:481-501](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/evaluators/base.py#L481-L501) · [garak/analyze/calibration.py:17-99](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/analyze/calibration.py#L17-L99) · [garak/detectors/base.py:197-272](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/detectors/base.py#L197-L272)

### UKGovernmentBEIS/inspect_ai (answered)

Inspect provides three tiers of scoring: per-sample scorers, per-task metrics, and epoch reducers.

**Built-in scorers** (`src/inspect_ai/scorer/`) each return a `Score` object with a `value` (str/int/float/bool/list/dict). They include:
- **match** — string matching at begin/end/any/exact locations, with case-insensitive and numeric options
- **includes** — substring containment check
- **exact** — normalized exact-match for QA
- **f1** — SQuAD-style F1 token overlap
- **choice** — multiple-choice letter grading with unshuffle support
- **math** — symbolic math via SymPy/LaTeX parsing with sandboxed expression validation, timeout, and complexity limits
- **answer** — extracts ANSWER:-prefixed answers by letter/word/line pattern
- **model_graded_qa / model_graded_fact** — LLM-as-a-judge using configurable templates and grader models (covered under llm-judge)
- **perplexity** — scores via prompt logprobs NLL
- **cascade** — chains scorers cheapest-first, short-circuiting when threshold met
- **multi_scorer** — runs multiple scorers in parallel and reduces via majority/mode/mean
- **precomputed_scores** — loads externally computed scores from JSON/JSONL by sample ID

**Metrics** (`src/inspect_ai/scorer/_metrics/`) aggregate per-sample scores into eval-level values:
- **accuracy** — proportion correct, with CORRECT/INCORRECT/PARTIAL/NOANSWER sentinels and a pluggable ValueToFloat converter
- **mean / std / var** — arithmetic mean, sample standard deviation, variance
- **stderr / bootstrap_stderr / ci / ci_wilson** — standard error (plain or clustered by sample metadata), confidence intervals via t-distribution or percentile bootstrap, Wilson score for binary proportions
- **frequency / categorical** — categorical score distribution (counts or proportions), with StrEnum integration and zero-fill for unobserved categories
- **aggregate** — extracts one key from dict-valued scores and runs another metric on it
- **grouped** — partitions scores by metadata key and applies a metric per group
- **krippendorff_alpha** — inter-rater agreement across multiple judges

**Custom scorers** are created with the `@scorer` decorator. The `Scorer` protocol requires an async callable `(state: TaskState, target: Target) -> Score | None`. Custom metrics use `@metric` and accept `list[SampleScore] -> Value`. Metrics can declare a `scores` mode: "auto" (reduced), "reduced", or "unreduced" (per-epoch).

**Score reducers** (`src/inspect_ai/scorer/_reducer/`) aggregate multi-epoch samples: `mean_score`, `median_score`, `mode_score`, `max_score`, `majority_score`, `pass_at`, `pass_k`, `at_least`, `collect_score`.


Citations: [src/inspect_ai/scorer/__init__.py:1-123](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/__init__.py#L1-L123) · [src/inspect_ai/scorer/_scorer.py:34-210](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_scorer.py#L34-L210) · [src/inspect_ai/scorer/_metric.py:42-552](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metric.py#L42-L552) · [src/inspect_ai/scorer/_metrics/accuracy.py:14-39](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metrics/accuracy.py#L14-L39) · [src/inspect_ai/scorer/_metrics/std.py:17-604](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_metrics/std.py#L17-L604) · [src/inspect_ai/scorer/_cascade.py:13-80](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_cascade.py#L13-L80)
