Which evaluation metrics and scorers are provided, and how are they implemented?
Built-in metrics (heuristic, statistical, model-based); custom scorers; per-task vs per-trace scoring.
Verdict
DeepEval and Opik ship the broadest metric libraries for LLM apps, and Ragas remains the reference for RAG metrics. For benchmark-style scoring with honest error bars, Inspect and lm-evaluation-harness are the strongest.
Metric libraries for applications. DeepEval has about 50 metric packages. Most are judge pipelines: extract claims, ask for a verdict per claim, return a ratio, and pass when score >= threshold. Ragas works the same way for faithfulness (statements, then NLI verdicts). FaithfulnesswithHHEM swaps the judge for Vectara’s HHEM classifier, and non-LLM context precision and recall variants use embeddings. Opik pairs LLM judges with the widest set of classic metrics here: BLEU, ROUGE, METEOR, BERTScore, Levenshtein and divergence metrics. A RagasMetricWrapper lets you run Ragas metrics inside Opik. Phoenix’s phoenix-evals turns judge labels into numbers through a label_score_map and adds code metrics such as PrecisionRecallFScore.
Assertion scorecards. promptfoo has about 70 assertion types. They range from equals and is-json through rouge-n and embedding similarity to llm-rubric and checks over recorded trace spans. Any type can be negated with not-, and weights and named metrics combine them into one score.
Platforms without a metric library. Langfuse scores with LLM judges, user code or decision models, and ships no BLEU, ROUGE or embedding metrics. Code evaluators run only on the observation and experiment path, and only when a dispatcher is configured.
Task and benchmark scoring. lm-evaluation-harness registers each metric with an aggregation and a direction, and reports a bootstrap stderr (100,000 iterations by default). Inspect separates per-sample scorers from metrics (clustered stderr, bootstrap and Wilson CIs) and from epoch reducers such as pass_at. OpenAI Evals reports accuracy with a “bootstrap” spread. The editor found that this is the standard deviation over half-size subsamples drawn without replacement, not a classic bootstrap. garak is different in kind. Its detectors score each output from 0 to 1, a 0.5 threshold decides pass or fail, and the result is an attack success rate, not a quality score.
Pick: DeepEval or Opik for a large ready-made set of app metrics. Pick: Ragas when RAG faithfulness and context metrics are the main need. Pick: Inspect or lm-evaluation-harness when you need confidence intervals on task accuracy.
Per-project answers
langfuse/langfuse
answeredLangfuse provides three evaluator types that produce scores, plus deterministic sampling and blocking. LLM-as-a-judge (evalService.ts lines 941–1233) — the primary mechanism — constructs a prompt from a template, calls a configurable LLM (OpenAI, Anthropic, etc.) with Zod-validated structured output, and normalizes the response into NUMERIC, BOOLEAN, or CATEGORICAL scores (outputDefinition.ts lines 1–428). The output definition is a persisted schema (PersistedEvalOutputDefinitionSchema) supporting numeric ranges (minValue/maxValue), boolean verdicts, and categorical choices (single or multi-match). The toNormalizedScores helper (evalService.ts lines 1235–1270) converts the LLM response — e.g. a numeric 0–1, a boolean true/false, or a string array — into CodeEvalScoreWithName[] records with a shared comment (the model's reasoning). Each score carries a dataType, value, and optionally configId and metadata. Code-based evaluators (codeEvalExecution.ts) dispatch user-supplied code (JS/Python) to either an AWS Lambda dispatcher or a local process, returning scores in a {scores: [...]} JSON contract — NO built-in heuristic or statistical metrics exist outside these three types. Decision-model evaluators (decisionModelEvaluatorExecution.ts) call a separate "TypeSafe" API for choice/score/boolean questions and map responses into scores. Deterministic sampling (deterministicSampling.ts lines 1–31) controls what fraction of traces are scored, using a SHA‑256 over target ID and sampling rate. Per-trace vs per-observation scoring is handled via targetObject: TRACE scores the whole trace at the trace level, DATASET ties scores to dataset items, and EVENT (observation-level) schedules per-observation evaluations (observationEval/). All scores are persisted with source: "EVAL" and share an EvalScoreWritePayload structure (evalScoreEvent.ts).
processObservationEval); trace- and dataset-level rules load only LLM-as-judge evaluators. The insecure-local code dispatcher runs TypeScript only, in a node:vm context inside the worker, and production has no code-eval dispatcher unless LANGFUSE_CODE_EVAL_DISPATCHER is set.promptfoo/promptfoo
answeredBuilt-in metrics registered via ASSERTION_HANDLERS map. Heuristic: equals, contains/contains-all/contains-any, regex, starts-with, is-json, is-sql, is-html, is-xml, is-refusal, finish-reason, cost. Statistical/distance: levenshtein, bleu, gleu, rouge-n, meteor, perplexity-score, similar (cosine/dot/euclidean), latency, word-count, tool-call-f1. Model-based: llm-rubric, model-graded-closedqa, factuality, g-eval, agent-rubric, search-rubric (src/matchers/llmGrading.ts). Trace-aware: trace-error-spans, trace-span-count, trace-span-duration, five trajectory types load spans from SQLite via loadTraceData. Custom: javascript/python/ruby file:// references return GradingResult. weight=0 forces metric-only pass. Custom per-test assertScoringFunction overrides default aggregation.
comet-ml/opik
answeredOpik provides a three-tier metrics system built on an abstract BaseMetric class (in sdks/python/src/opik/evaluation/metrics/base_metric.py:34-104). All metrics return a ScoreResult (value 0.0–1.0, plus name/reason/metadata).
Heuristic/statistical metrics — no LLM calls needed. Includes: Equals (exact/case-insensitive string match), Contains (substring detection), RegexMatch, IsJson, LevenshteinRatio, SentenceBLEU/CorpusBLEU (wrapping NLTK), ROUGE (rouge1/2/L/Lsum, wrapping rouge_score), ChrF, GLEU, METEOR, BERTScore (wrapping bert-score), JSDivergence/KLDivergence, SpearmanRanking, Readability, Tone, Sentiment, VADERSentiment, and LanguageAdherenceMetric. All live in sdks/python/src/opik/evaluation/metrics/heuristics/.
LLM-judge metrics — use a judge LLM (configurable model name, defaults via OPIK_DEFAULT_LLM) with structured JSON output via Pydantic response-format schemas. Includes: Hallucination (1.0 if hallucination detected), AnswerRelevance, ContextPrecision, ContextRecall, Factuality (per-claim scoring), Moderation (0.0–1.0 content-appropriateness), Usefulness, TrajectoryAccuracy (ReAct agent trajectory quality), SycEval (sycophancy detection with rebuttal generation), StructuredOutputCompliance, and LLMJuriesJudge. Each lives under sdks/python/src/opik/evaluation/metrics/llm_judges/<name>/. The GEval metric (llm_judges/g_eval/metric.py:32-288) is a generalised LLM-as-a-judge that builds a reusable chain-of-thought prompt from task_introduction and evaluation_criteria, caches it (LRU, max 128), and scores each output. GEvalPreset wraps pre-built rubrics including: QARelevanceJudge, DemographicBiasJudge, GenderBiasJudge, PoliticalBiasJudge, RegionalBiasJudge, ReligiousBiasJudge, ComplianceRiskJudge, DialogueHelpfulnessJudge, PromptUncertaintyJudge, AgentTaskCompletionJudge, AgentToolCorrectnessJudge, SummarizationCoherenceJudge, SummarizationConsistencyJudge (all from sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval_presets.py).
RAGAS integration: RagasMetricWrapper (metrics/ragas_metric.py:13-77) wraps any ragas.SingleTurnMetric as an Opik BaseMetric, mapping 'input'/'output' keys to RAGAS's 'user_input'/'response'.
Custom scoring: ScorerFunction (scorers/scorer_function.py:8-27) is a protocol taking (dataset_item, task_outputs, task_span?) returning ScoreResults — wrapped into metrics via scorer_wrapper_metric.wrap_scorer_functions().
Conversation metrics under metrics/conversation/ assess multi-turn threads: ConversationThreadMetric wraps a per-turn metric into a thread-level score; GEvalConversationMetric variant judges entire conversations. Specific LLM-judge conversation metrics include ConversationalCoherenceMetric, SessionCompletenessQuality, UserFrustrationMetric, and heuristic ones like KnowledgeRetentionMetric and ConversationDegenerationMetric.
AggregatedMetric (metrics/aggregated_metric.py) wraps metrics and computes statistics across all items.
openai/evals
answeredBuilt-in metrics are implemented as standalone functions in evals/metrics.py. get_accuracy(events) computes sum(correct)/total from match events (evals/metrics.py:12-18). get_bootstrap_accuracy_std resamples with replacement for uncertainty (evals/metrics.py:21-23). get_confusion_matrix builds an N×(N+1) array from expected vs picked labels (evals/metrics.py:26-40), then compute_matthew_corr, compute_precision, compute_recall, compute_f_score, and compute_averaged_f_score derive classification metrics from it (evals/metrics.py:43-73).
Custom per-sample scoring uses evals.record.record_metrics(**kwargs) — a free-form dict attached to a "metrics" event on the RecorderBase (evals/record.py:248-249). Standard evals like FuzzyMatch record per-sample accuracy (float) and f1_score (evals/elsuite/basic/fuzzy_match.py:48-50). Others record custom keys: Includes records correctness per sample (evals/elsuite/basic/includes.py:45-47), JsonValidator records accuracy alone, and ModelBasedClassify records choice, score, and optionally metascore (evals/elsuite/modelgraded/classify.py:93-98).
Per-task aggregation happens in each eval's run() method, which returns a dict[str, float]. Match.run() averages accuracy and bootstrap_std from all match events via recorder.get_events("match") (evals/elsuite/basic/match.py:58-65). MultipleChoice.run() calls evals.metrics.get_accuracy() on match events (evals/elsuite/multiple_choice.py:95-100). ModelBasedClassify.run() additionally computes choice-count distributions, per-choice score averages, and metascore (evals/elsuite/modelgraded/classify.py:104-127). The BaseEvalSpec declares a metrics list (e.g. [accuracy]) and a higher_is_better flag (evals/base.py:36-44).
get_bootstrap_accuracy_std is not a with-replacement bootstrap. It returns the standard deviation of accuracy over 1,000 random half-size subsamples drawn without replacement (evals/metrics.py L21-L23).confident-ai/deepeval
answeredDeepEval provides ~50+ built-in metrics across four categories. Heuristic/statistical metrics need no LLM: ExactMatchMetric does strict string comparison (deepeval/metrics/exact_match/exact_match.py:12-88), PatternMatchMetric applies regex, and the Scorer class offers ROUGE, BLEU, BERTScore, and exact/quasi-exact match utilities (deepeval/scorer/scorer.py:11-160). LLM-based metrics use a judge LLM: FaithfulnessMetric extracts claims from actual_output and checks each against retrieval_context via a yes/no verdict loop (deepeval/metrics/faithfulness/faithfulness.py:66-100), HallucinationMetric and AnswerRelevancyMetric follow the same QAG (question-answer generation) pattern — calling generate_qag_verdicts to classify statements, then score_qag_verdicts to aggregate. ToxicityMetric and BiasMetric extract opinions/statements and judge each for toxicity/bias. Configurable LLM judges: GEval accepts criteria, evaluation_steps, and rubric (score-range → outcome mappings), renders a structured prompt, and returns a Pydantic-validated ReasonScore (deepeval/metrics/g_eval/g_eval.py:53-120, deepeval/metrics/g_eval/schema.py:5-13). ArenaGEval compares two contestants side-by-side. DAGMetric composes multiple judgement and task nodes into an acyclic graph for multi-step evaluation (deepeval/metrics/dag/dag.py:25-70). Agentic/conversational metrics: ToolCorrectnessMetric, GoalAccuracyMetric, AgentLoopDetectionMetric, turn-level variants (TurnFaithfulnessMetric), and ConversationCompletenessMetric. Each metric is a subclass of BaseMetric (single-turn LLMTestCase), BaseConversationalMetric (multi-turn ConversationalTestCase), or BaseArenaMetric (pairwise ArenaTestCase) (deepeval/metrics/base_metric.py:111-350). Metrics define _required_params (e.g. [INPUT, ACTUAL_OUTPUT, RETRIEVAL_CONTEXT]), implement measure() and a_measure(), and set self.score, self.reason, self.success (score ≥ threshold). Per-test-case scoring: execute_test_cases iterates test cases and runs each metric against each case individually (deepeval/evaluate/execute/e2e.py:144-270). Custom metrics are created by subclassing a base class and providing the measure method.
vibrantlabsai/ragas
answeredRagas provides ~30+ built-in metrics across three implementation categories.
Heuristic/statistical metrics (no LLM call): ExactMatch (string equality), StringPresence (substring match), NonLLMStringSimilarity (Levenshtein/Hamming/Jaro/Jaro-Winkler via rapidfuzz), BleuScore, RougeScore, ChrfScore, and DataCompyScore. These are pure Python computations implementing SingleTurnMetric. Example: ExactMatch returns float(sample.reference == sample.response) (metrics/_string.py:29-35).
LLM-based metrics: Faithfulness (statement decomposition → NLI verdicts), AnswerRelevancy, AnswerCorrectness, FactualCorrectness, ContextPrecision (multiple variants: LLM, NonLLM, ID-based), ContextRecall, ContextEntityRecall, NoiseSensitivity, SummarizationScore, TopicAdherenceScore, ToolCallAccuracy, ToolCallF1, MultiModalFaithfulness, MultiModalRelevance, SQLSemanticEquivalence, GoalAccuracy (agent). These extend MetricWithLLM and implement _single_turn_ascore() which calls a PydanticPrompt.generate() against an LLM. Example: Faithfulness breaks the response into atomic statements using StatementGeneratorPrompt, then judges each against retrieved contexts via NLIStatementPrompt (metrics/_faithfulness.py:152-214).
Hybrid/embedding-based: FaithfulnesswithHHEM replaces the LLM NLI stage with a HuggingFace transformer (vectara/hallucination_evaluation_model) (metrics/_faithfulness.py:217-273). SemanticSimilarity uses embeddings. NonLLM variants of ContextPrecision/ContextRecall use embedding cosine similarity instead of LLM calls.
Custom metrics: Three decorator-based primitives: @discrete_metric (categorical output, e.g. 'pass'/'fail'), @numeric_metric (continuous float), @ranking_metric (ranked output). The SimpleLLMMetric class accepts a custom prompt string and Pydantic response model for ad-hoc LLM-as-a-judge metrics with save/load/reproducibility (metrics/base.py:846-935). AspectCritic and SimpleCriteriaScore take user-defined rubric strings for binary/discrete scoring with automatic strictness-based majority voting.
Per-trace vs per-task: All metrics implement SingleTurnMetric (single QA pair) or MultiTurnMetric (conversation). The Metric base class declares required_columns per type, e.g. Faithfulness needs {user_input, response, retrieved_contexts} for single-turn. Collection metrics (e.g. AnswerCorrectness composite) internally orchestrate multiple sub-metrics per trace.
EleutherAI/lm-evaluation-harness
answeredBuilt-in metrics are defined in lm_eval/api/metrics.py via registration decorators. Per-sample metrics (passthrough functions) include acc, acc_norm, acc_mutual_info, acc_bytes, acc_all, exact_match, perplexity, likelihood, word_perplexity, byte_perplexity, bits_per_byte, brier_score, mcc, f1, bleu, chrf, chrf++, ter, and bypass. Each is registered with @register_metric(metric=..., higher_is_better=..., output_type=[...], aggregation=...) — the decorator links the metric name to its aggregation, specifies which output types it applies to, and records directionality. Example: acc is registered at lm_eval/api/metrics.py:176-183 with aggregation="mean" and higher_is_better=True; exact_match at line 272-279 uses aggregation="mean" on generate_until output. Aggregation functions (registered with @register_aggregation) include mean, median, nanmean, perplexity, weighted_perplexity, bits_per_byte, f1, matthews_corrcoef, bleu, chrf, chrf++, ter, brier_score, and bypass. These functions take per-document metric values and reduce them to a single task-level score. Translation metrics (bleu, chrf, ter) use sacrebleu for corpus-level computation. Custom per-task metrics — any YAML task can declare a metric_list with arbitrary metric names, aggregation functions, and higher_is_better flags; arbitrary metrics without a defined aggregation default to mean. Per-task vs per-trace scoring: Per-sample metrics are computed in task.process_results(doc, results) — called per document in evaluate() at lm_eval/evaluator.py:639-641 — which returns a dict of metric values for that document. These raw per-document values are collected into raw_metrics[(metric, filter_key)] lists, then aggregated at lm_eval/evaluator_utils.py:193-204 by looking up task.aggregation()[metric] and calling the registered aggregation function. Bootstrap stderr is computed via bootstrap_stderr() at lm_eval/api/metrics.py:550-586, using multiprocessing (Pool.imap) to generate bootstrap resamples of aggregated metrics. Filters operate on outputs before metrics: RegexFilter, TakeFirstFilter, MajorityVoteFilter, LowercaseFilter, MapFilter, etc. in lm_eval/filters/. Each filter pipeline produces a filtered_resps key that metrics are computed over, so the same task can report multiple filter variants.
Arize-ai/phoenix
answeredPhoenix provides three categories of evaluators: code/heuristic, LLM-based, and statistical. Code evaluators include exact_match (string equality check) at packages/phoenix-evals/src/phoenix/evals/metrics/exact_match.py:37, MatchesRegex (regex matching) and PrecisionRecallFScore (macro/micro/weighted precision, recall, F-beta with binary one-vs-rest) at packages/phoenix-evals/src/phoenix/evals/metrics/precision_recall.py:66-424. LLM-based evaluators subclass ClassificationEvaluator and include CorrectnessEvaluator, FaithfulnessEvaluator, HallucinationEvaluator, CompletenessEvaluator, ConcisenessEvaluator, DocumentRelevanceEvaluator, RetrievalRelevanceEvaluator, ToxicityEvaluator, RefusalEvaluator, PiiDetectionEvaluator, UserFrictionEvaluator, and three tool-oriented evaluators (ToolInvocationEvaluator, ToolSelectionEvaluator, ToolResponseHandlingEvaluator) — all in packages/phoenix-evals/src/phoenix/evals/metrics/. Each LLM evaluator loads a generated prompt config (e.g. _correctness_classification_evaluator_config.py defines a rubric string, model choices like {"correct": 1.0, "incorrect": 0.0}, and optimization_direction). The base Evaluator class (packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506) handles input remapping via input_mapping (field-name or JSONPath-based), Pydantic schema validation, and returns Score objects. The Score dataclass (evaluators.py:158-271) carries name, score (float/int), label (string), explanation, metadata, kind ("human"|"llm"|"code"), and direction ("maximize"|"minimize"|"neutral"). Per-trace vs per-task: Evaluators operate on a single EvalInput dict (one row = one trace or span); evaluate_dataframe() and async_evaluate_dataframe() apply an evaluator list over every row and add score columns back to the DataFrame. Custom evaluators are created via the @create_evaluator(name=..., kind=...) decorator which wraps any sync/async function into an Evaluator instance with auto-generated Pydantic input schema from function signature, and _convert_to_score handles returns of type Score, bool, float, str, dict, or tuple.
NVIDIA/garak
answeredScoring lives in garak/evaluators/base.py. The Evaluator base class evaluates detector scores per-probe, per-detector. Each Attempt carries detector_results — a dict keyed by detector name, mapping to a list of floats (one per generation). Evaluator.evaluate() iterates attempts, groups results by detector, and calls _evaluate_one_detector() which counts passes, fails, and nulls using a pluggable test() method. Two concrete evaluators ship: ThresholdEvaluator (pass if score < threshold, default 0.5) and ZeroToleranceEvaluator (pass only if score == 0.0; evaluators/base.py:481-488). The default threshold of 0.5 is set in garak.core.yaml as eval_threshold.
Confidence intervals use non-parametric bootstrap: calculate_bootstrap_ci() in analyze/bootstrap_ci.py:90-105 resamples results with replacement (default 10,000 iterations, 95% confidence) and corrects for detector sensitivity/specificity loaded from data/detectors_eval/detector_metrics_summary.json via analyze/detector_metrics.py. Calibration Z-scores: Calibration class (analyze/calibration.py:17-99) loads prior-run distributions from data/calibration/calibration.json and computes (score - mu) / sigma, gated by MINIMUM_STD_DEV = 1/30. These feed "defcon" ratings (1-5, analyze/__init__.py:48-58).
Detectors range from heuristic (StringDetector — substring/word/prefix matching with optional Unicode normalization, detectors/base.py:197-272) to model-based (HFDetector wraps HuggingFace text-classification pipelines, normalizing logits to 0-1, detectors/base.py:82-194), and trigger-list matching (TriggerListDetector, detectors/base.py:275-304). Scoring is per-trace (each generation output independently), then aggregated to per-probe/per-detector summaries in the eval record.
UKGovernmentBEIS/inspect_ai
answeredInspect provides three tiers of scoring: per-sample scorers, per-task metrics, and epoch reducers.
Built-in scorers (src/inspect_ai/scorer/) each return a Score object with a value (str/int/float/bool/list/dict). They include:
- match — string matching at begin/end/any/exact locations, with case-insensitive and numeric options
- includes — substring containment check
- exact — normalized exact-match for QA
- f1 — SQuAD-style F1 token overlap
- choice — multiple-choice letter grading with unshuffle support
- math — symbolic math via SymPy/LaTeX parsing with sandboxed expression validation, timeout, and complexity limits
- answer — extracts ANSWER:-prefixed answers by letter/word/line pattern
- model_graded_qa / model_graded_fact — LLM-as-a-judge using configurable templates and grader models (covered under llm-judge)
- perplexity — scores via prompt logprobs NLL
- cascade — chains scorers cheapest-first, short-circuiting when threshold met
- multi_scorer — runs multiple scorers in parallel and reduces via majority/mode/mean
- precomputed_scores — loads externally computed scores from JSON/JSONL by sample ID
Metrics (src/inspect_ai/scorer/_metrics/) aggregate per-sample scores into eval-level values:
- accuracy — proportion correct, with CORRECT/INCORRECT/PARTIAL/NOANSWER sentinels and a pluggable ValueToFloat converter
- mean / std / var — arithmetic mean, sample standard deviation, variance
- stderr / bootstrap_stderr / ci / ci_wilson — standard error (plain or clustered by sample metadata), confidence intervals via t-distribution or percentile bootstrap, Wilson score for binary proportions
- frequency / categorical — categorical score distribution (counts or proportions), with StrEnum integration and zero-fill for unobserved categories
- aggregate — extracts one key from dict-valued scores and runs another metric on it
- grouped — partitions scores by metadata key and applies a metric per group
- krippendorff_alpha — inter-rater agreement across multiple judges
Custom scorers are created with the @scorer decorator. The Scorer protocol requires an async callable (state: TaskState, target: Target) -> Score | None. Custom metrics use @metric and accept list[SampleScore] -> Value. Metrics can declare a scores mode: "auto" (reduced), "reduced", or "unreduced" (per-epoch).
Score reducers (src/inspect_ai/scorer/_reducer/) aggregate multi-epoch samples: mean_score, median_score, mode_score, max_score, majority_score, pass_at, pass_k, at_least, collect_score.