# LLM evals and testing: comparison

> Frameworks and platforms that test LLM apps and models: metrics and LLM judges, datasets, eval runners, tracing and red-teaming.

Canonical page: https://llms-technical-reviews.com/compare/evals/

## Which evaluation metrics and scorers are provided, and how are they implemented?

[DeepEval](/p/deepeval/) and [Opik](/p/opik/) ship the broadest metric libraries for LLM apps, and [Ragas](/p/ragas/) remains the reference for RAG metrics. For benchmark-style scoring with honest error bars, [Inspect](/p/inspect_ai/) and [lm-evaluation-harness](/p/lm-evaluation-harness/) are the strongest.

**Metric libraries for applications.** DeepEval has about 50 metric packages. Most are judge pipelines: extract claims, ask for a verdict per claim, return a ratio, and pass when `score >= threshold`. Ragas works the same way for faithfulness (statements, then NLI verdicts). `FaithfulnesswithHHEM` swaps the judge for Vectara's HHEM classifier, and non-LLM context precision and recall variants use embeddings. Opik pairs LLM judges with the widest set of classic metrics here: BLEU, ROUGE, METEOR, BERTScore, Levenshtein and divergence metrics. A `RagasMetricWrapper` lets you run Ragas metrics inside Opik. [Phoenix](/p/phoenix/)'s `phoenix-evals` turns judge labels into numbers through a `label_score_map` and adds code metrics such as `PrecisionRecallFScore`.

**Assertion scorecards.** [promptfoo](/p/promptfoo/) has about 70 assertion types. They range from `equals` and `is-json` through `rouge-n` and embedding similarity to `llm-rubric` and checks over recorded trace spans. Any type can be negated with `not-`, and weights and named metrics combine them into one score.

**Platforms without a metric library.** [Langfuse](/p/langfuse/) scores with LLM judges, user code or decision models, and ships no BLEU, ROUGE or embedding metrics. Code evaluators run only on the observation and experiment path, and only when a dispatcher is configured.

**Task and benchmark scoring.** lm-evaluation-harness registers each metric with an aggregation and a direction, and reports a bootstrap stderr (100,000 iterations by default). Inspect separates per-sample scorers from metrics (clustered stderr, bootstrap and Wilson CIs) and from epoch reducers such as `pass_at`. [OpenAI Evals](/p/openai-evals/) reports accuracy with a "bootstrap" spread. The editor found that this is the standard deviation over half-size subsamples drawn without replacement, not a classic bootstrap. [garak](/p/garak/) is different in kind. Its detectors score each output from 0 to 1, a 0.5 threshold decides pass or fail, and the result is an attack success rate, not a quality score.

Pick: DeepEval or Opik for a large ready-made set of app metrics.
Pick: Ragas when RAG faithfulness and context metrics are the main need.
Pick: Inspect or lm-evaluation-harness when you need confidence intervals on task accuracy.

Per-project answers: https://llms-technical-reviews.com/evals/q/metrics/index.md

## How is LLM-as-a-judge implemented?

[Inspect](/p/inspect_ai/) has the most careful judge: grader panels with majority vote, verdict parsing that resists injection, and parse failures kept separate from wrong answers. [DeepEval](/p/deepeval/), [Ragas](/p/ragas/) and [Opik](/p/opik/) have the most structured judge input and output. [Langfuse](/p/langfuse/) runs judges as production infrastructure.

**Schema-enforced judges.** DeepEval's `GEval` takes criteria, steps or a rubric of score ranges. When the model returns logprobs, it replaces the sampled score with a probability-weighted average. Its optional `hybrid` mode lets a hosted "System One" model answer the decision points and falls back to the LLM only when that call fails. Ragas validates every judge call into a Pydantic model through Instructor, and retries with a fix-format prompt. The editor found its `strictness` voting inert at the reviewed commit, so `AspectCritic` is a single-sample judge. Opik passes a Pydantic `response_format`, retries parse failures three times, and merges identical `LLMJudge` instances into one call. Its default judge is `gpt-5-nano`. [Phoenix](/p/phoenix/) builds a JSON schema with an enum of allowed labels. On OpenAI it tries native structured output and falls back to tool calling. Langfuse validates numeric, boolean or categorical output with Zod and records every judge call as its own trace. It pauses an evaluator when the judge fails with an auth or billing error.

**Text-parsed judges.** Inspect takes the *last* `GRADE: X` in the reply and escapes `[BEGIN DATA]` markers in dataset text. [promptfoo](/p/promptfoo/) parses a JSON `{pass, score, reason}`. Red-team configs without an explicit grader prefer promptfoo's remote grading endpoint. [OpenAI Evals](/p/openai-evals/) matches choice strings line by line, scores an unparseable answer as the lowest choice, and uses the last completion function on the command line as the judge. [garak](/p/garak/)'s `ModelAsJudge` asks for a `[[rating]]` from 1 to 10 and counts 7 or more as a hit. Its default judge is Llama 3 70B through NVIDIA NIM.

**No judge in the core.** [lm-evaluation-harness](/p/lm-evaluation-harness/) is reference-based. Only the `pisa_*_llm_judged` tasks call an OpenAI model from their own hooks.

Langfuse, Phoenix, Ragas, OpenAI Evals, Inspect and garak have no calibration or position-bias controls.

Pick: Inspect when judge robustness matters, for example graded agent tasks.
Pick: DeepEval or Ragas for rubric and claim-level judging in Python.
Pick: Langfuse or Opik to run judges continuously on live traces.

Per-project answers: https://llms-technical-reviews.com/evals/q/llm-judge/index.md

## How are test datasets and cases defined, generated and versioned?

[Phoenix](/p/phoenix/), [Opik](/p/opik/) and [Langfuse](/p/langfuse/) are the only projects that version test data themselves. [DeepEval](/p/deepeval/) and [Ragas](/p/ragas/) are best at generating test data. [lm-evaluation-harness](/p/lm-evaluation-harness/) and [OpenAI Evals](/p/openai-evals/) are benchmark registries.

**Versioned datasets on a server.** Phoenix stores `DatasetVersion`, per-example revisions and splits, and each experiment is pinned to a dataset version. Langfuse versions items in time with `validFrom`/`validTo`. A dataset-level eval reads each item at its `validFrom` version, and items can point back to the trace they came from. Opik has dataset versions and test suites whose items carry their own evaluator configs and pass policies (`runs_per_item`, `pass_threshold`). None of the three generates synthetic data.

**Datasets in code, with generators.** DeepEval's `Synthesizer` rewrites seed inputs with evolutions (reasoning, multi-context, hypothetical and others). It also wraps about 15 standard benchmarks, but dataset versions live on the vendor's hosted platform. Ragas's `TestsetGenerator` builds a knowledge graph from your documents, invents personas and writes single-hop and multi-hop questions. Its `version_experiment` commits your working tree to a git branch. [promptfoo](/p/promptfoo/) reads tests from YAML, CSV, XLSX, Google Sheets or Hugging Face, and expands array variables into every combination. Versioning is git plus a config snapshot per run. [Inspect](/p/inspect_ai/) `Sample`s can carry sandbox files and setup. Versioning is left to the source, such as a Hugging Face `revision`.

**Benchmark and attack registries.** lm-evaluation-harness has about 220 benchmark folders and close to 14,000 task YAML files, each with a `metadata.version`. OpenAI Evals has 463 eval YAML files named `<base>.<split>.v<N>`, and its JSONL samples are stored in Git LFS. [garak](/p/garak/) ships attack payloads as JSON under `garak/data/`. A copy in the user data directory overrides them. It has no versioning and no generation.

Pick: Phoenix or Opik when experiments must be pinned to a dataset version.
Pick: Ragas or DeepEval to generate questions from your own documents.
Pick: lm-evaluation-harness for public benchmarks with versioned task configs.

Per-project answers: https://llms-technical-reviews.com/evals/q/datasets/index.md

## How are evals executed and reported?

[promptfoo](/p/promptfoo/) and [DeepEval](/p/deepeval/) fit most easily into CI. [Inspect](/p/inspect_ai/) has the most robust runner for long or agentic evals. [Langfuse](/p/langfuse/) and [Opik](/p/opik/) run evaluations as continuous server jobs.

**Test runners for CI.** promptfoo expands prompt × provider × test into a matrix. It runs 4 steps at a time by default, or 1 when a test carries conversation state. Responses are cached in memory and on disk with a 14-day TTL, and results go to a local SQLite database with a web viewer. DeepEval's `deepeval test run` wraps pytest, supports xdist and returns pytest's exit code. Cases marked `flaky` only warn. [Ragas](/p/ragas/) runs jobs with 16 workers by default and turns failed rows into `NaN`. Its `ragas evals` CLI depends on a Project class the code marks as not implemented, so CI gates should use the Python API.

**Research runners with logs.** Inspect stores a full event transcript per sample in a `.eval` zip archive. `eval_set` retries and resumes from logs, and `inspect view` browses and compares runs. [lm-evaluation-harness](/p/lm-evaluation-harness/) batches requests by type, shards them across data-parallel ranks, and can cache responses in SQLite. It has no baselines or thresholds. [OpenAI Evals](/p/openai-evals/) runs a 10-thread pool and writes JSONL to `/tmp/evallogs`. The editor found that its `--no-cache` flag has no effect. [garak](/p/garak/) streams a JSONL report and a hitlog of successful attacks, then builds an HTML digest.

**Platform experiments.** Langfuse schedules eval jobs on BullMQ queues with deterministic job ids, so a trace is not scored twice. Its CI integration is whatever you build on its API. Opik's `evaluate()` uses 16 threads, can resume an interrupted run, and compares experiments in its UI. [Phoenix](/p/phoenix/)'s `run_experiment` pins a dataset version and lines runs up per example in the UI. The editor found that the CI gate in its `evals/pxi` folder is the team's internal suite and is not shipped with the package.

Pick: promptfoo or DeepEval for a pass/fail check in CI.
Pick: Inspect for long agent runs that must resume and stay auditable.
Pick: Langfuse or Opik to keep scoring after deployment.

Per-project answers: https://llms-technical-reviews.com/evals/q/execution/index.md

## How are traces or production data captured and linked to evaluations?

[Langfuse](/p/langfuse/) and [Opik](/p/opik/) are the only projects that score production traces as they arrive. [Phoenix](/p/phoenix/) has the cleanest OpenTelemetry ingest, but you script online scoring yourself. The other seven tools trace only their own eval runs, or nothing.

**Trace platforms.** Langfuse accepts OTLP and its own SDK events. Evaluation rules pick traces with a deterministic SHA-256 sample, so the same trace is always in or out. Each judge call is itself a trace, linked from the score by `executionTraceId`. Human annotations share the score table with `source: ANNOTATION`. Note that the SDKs are in separate repositories. Opik captures spans with `@opik.track`, an OTel span processor or its OTLP endpoint. A sampler using `SecureRandom` sends traces through Redis streams to LLM-judge or Python-metric scorers. Experiment items are ordinary traces, so offline and online scores sit side by side. Phoenix receives OTLP over gRPC or HTTP with OpenInference attributes. Scores, labels and code checks all become span annotations, and `log_span_annotations` attaches offline results to production spans. Its online runner in `evals/pxi` is internal and not shipped.

**Libraries that trace evaluations.** [DeepEval](/p/deepeval/) has `@observe` and OpenAI and Anthropic patchers, plus an OTel processor that forwards spans to its vendor's hosted platform. Online evaluation of production traffic is a feature of that platform. [promptfoo](/p/promptfoo/) starts a local OTLP receiver during a run so that assertions can check the spans your app emitted. It can also read traces from Langfuse, Braintrust or Tempo, but it never captures production traffic. [Ragas](/p/ragas/) records an evaluation → row → metric → prompt callback tree and has Langfuse and MLflow helpers. It has no OTel support.

**Offline logs only.** [Inspect](/p/inspect_ai/) records typed events per sample and exposes lifecycle hooks, but has no OTel. [lm-evaluation-harness](/p/lm-evaluation-harness/) logs to W&B or Trackio. [OpenAI Evals](/p/openai-evals/) records prompt and output events to JSONL. For [garak](/p/garak/) the question is not applicable. It is a scanner with a report file.

Pick: Langfuse for rule-based online scoring with auditable judges.
Pick: Opik when you also want a Python metric library on the same traces.
Pick: Phoenix for standard OTLP ingest and annotation-driven review.

Per-project answers: https://llms-technical-reviews.com/evals/q/observability/index.md

## Does it support red-teaming or safety testing, and how?

Only [garak](/p/garak/) and [promptfoo](/p/promptfoo/) generate attacks. garak is a self-contained scanner for a model or endpoint. promptfoo puts red-teaming into the same config and report as its quality tests, but leans on a hosted API.

**Attack generators.** garak has about 44 probe modules: DAN variants, prompt injection, latent injection hidden in documents, encoding and smuggling tricks, and attacker-model loops (`goat`, `tap`, `atkgen`). `TreeSearchProbe` branches on detector scores. Its REST and function generators let you scan a deployed app rather than only the bare model, and results export to AVID. The editor found its buffs limited to Base64, CharCode, lowercase, low-resource-language translation and paraphrase, with no leetspeak or typo buffs. The tool-attacking `agent_breaker` probe is off by default. promptfoo has dozens of plugins, with over 30 harm categories in the `harmful` set alone, and about 30 strategies (crescendo, GOAT, Hydra, GCG, Base64, leetspeak and others). It writes the generated cases to `redteam.yaml`, then evaluates them like any other test. Plugins in `REMOTE_ONLY_PLUGIN_IDS` exist only on promptfoo's API. Disabling remote generation removes them, and a 100,000-probe monthly limit applies unless you log in to its cloud.

**Safety metrics, not attacks.** These tools judge outputs you already have. [Opik](/p/opik/) has a regex `PromptInjection` heuristic, `Moderation`, a `SycEval` sycophancy test, bias presets and guardrails. [DeepEval](/p/deepeval/) moved red-teaming to the separate DeepTeam project, so the question is not applicable to it. Toxicity, bias and PII-leakage judges plus a prompt-injection classifier remain. [Langfuse](/p/langfuse/) ships two managed judge templates, "Detect Prompt Injection" and "Check Rule Adherence". [Ragas](/p/ragas/) has `harmfulness` and `maliciousness` critics. [Phoenix](/p/phoenix/) has toxicity and PII evaluators. It has no red-teaming tooling.

**Safety benchmarks and building blocks.** [lm-evaluation-harness](/p/lm-evaluation-harness/) runs ToxiGen, BEAR and the advanced-AI-risk persona tasks as multiple-choice benchmarks. [OpenAI Evals](/p/openai-evals/) includes a narrow `prompt-injection` eval plus sandbagging and steganography suites. [Inspect](/p/inspect_ai/) provides Docker sandboxes, approval policies and agents for writing your own attack tasks, but ships no attack library.

Pick: garak to scan a model or endpoint with no hosted dependency.
Pick: promptfoo for app-level red-teaming next to your regression tests.
Pick: Inspect to build custom agentic safety evaluations.

Per-project answers: https://llms-technical-reviews.com/evals/q/red-teaming/index.md

## Projects

- [langfuse/langfuse](https://llms-technical-reviews.com/p/langfuse/index.md) — Self-hostable LLM tracing and eval platform: Next.js API, BullMQ worker, ClickHouse traces, and queued LLM-judge and code evaluators.
- [promptfoo/promptfoo](https://llms-technical-reviews.com/p/promptfoo/index.md) — Local-first CLI and Node library that runs prompt x provider x test matrices, grades them with ~70 assertion types, and red-teams targets.
- [comet-ml/opik](https://llms-technical-reviews.com/p/opik/index.md) — LLM tracing and eval platform: Python/TS SDKs with ~30 metrics, a Java backend on ClickHouse and MySQL, and Redis-streamed online scoring.
- [openai/evals](https://llms-technical-reviews.com/p/openai-evals/index.md) — Python CLI and YAML registry of ~460 benchmark evals that run models over JSONL samples and log every match and sample as events.
- [confident-ai/deepeval](https://llms-technical-reviews.com/p/deepeval/index.md) — Pytest-style LLM evaluation library with ~50 judge-based metrics, G-Eval rubrics, tracing, dataset synthesis and Confident AI upload.
- [vibrantlabsai/ragas](https://llms-technical-reviews.com/p/ragas/index.md) — Python library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.
- [EleutherAI/lm-evaluation-harness](https://llms-technical-reviews.com/p/lm-evaluation-harness/index.md) — Benchmark harness that scores a model on YAML-defined academic tasks via loglikelihood or generation requests, with bootstrap stderr.
- [Arize-ai/phoenix](https://llms-technical-reviews.com/p/phoenix/index.md) — Self-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.
- [NVIDIA/garak](https://llms-technical-reviews.com/p/garak/index.md) — Vulnerability scanner that fires adversarial probe prompts at a model or app endpoint and scores responses with detectors.
- [UKGovernmentBEIS/inspect_ai](https://llms-technical-reviews.com/p/inspect_ai/index.md) — Python eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.