LLM evals and testing: langfuse vs promptfoo vs opik vs evals vs deepeval vs ragas vs lm-evaluation-harness vs phoenix vs garak vs inspect_ai
Frameworks and platforms that test LLM apps and models: metrics and LLM judges, datasets, eval runners, tracing and red-teaming. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.
At a glance
● answered from code · — not applicable (the project does not do this) · ? insufficient evidence
Which evaluation metrics and scorers are provided, and how are they implemented?
DeepEval and Opik ship the broadest metric libraries for LLM apps, and Ragas remains the reference for RAG metrics. For benchmark-style scoring with honest error bars, Inspect and lm-evaluation-harness are the strongest.
Metric libraries for applications. DeepEval has about 50 metric packages. Most are judge pipelines: extract claims, ask for a verdict per claim, return a ratio, and pass when score >= threshold. Ragas works the same way for faithfulness (statements, then NLI verdicts). FaithfulnesswithHHEM swaps the judge for Vectara's HHEM classifier, and non-LLM context precision and recall variants use embeddings. Opik pairs LLM judges with the widest set of classic metrics here: BLEU, ROUGE, METEOR, BERTScore, Levenshtein and divergence metrics. A RagasMetricWrapper lets you run Ragas metrics inside Opik. Phoenix's phoenix-evals turns judge labels into numbers through a label_score_map and adds code metrics such as PrecisionRecallFScore.
Assertion scorecards. promptfoo has about 70 assertion types. They range from equals and is-json through rouge-n and embedding similarity to llm-rubric and checks over recorded trace spans. Any type can be negated with not-, and weights and named metrics combine them into one score.
Platforms without a metric library. Langfuse scores with LLM judges, user code or decision models, and ships no BLEU, ROUGE or embedding metrics. Code evaluators run only on the observation and experiment path, and only when a dispatcher is configured.
Task and benchmark scoring. lm-evaluation-harness registers each metric with an aggregation and a direction, and reports a bootstrap stderr (100,000 iterations by default). Inspect separates per-sample scorers from metrics (clustered stderr, bootstrap and Wilson CIs) and from epoch reducers such as pass_at. OpenAI Evals reports accuracy with a "bootstrap" spread. The editor found that this is the standard deviation over half-size subsamples drawn without replacement, not a classic bootstrap. garak is different in kind. Its detectors score each output from 0 to 1, a 0.5 threshold decides pass or fail, and the result is an attack success rate, not a quality score.
Pick: DeepEval or Opik for a large ready-made set of app metrics. Pick: Ragas when RAG faithfulness and context metrics are the main need. Pick: Inspect or lm-evaluation-harness when you need confidence intervals on task accuracy.
How is LLM-as-a-judge implemented?
Inspect has the most careful judge: grader panels with majority vote, verdict parsing that resists injection, and parse failures kept separate from wrong answers. DeepEval, Ragas and Opik have the most structured judge input and output. Langfuse runs judges as production infrastructure.
Schema-enforced judges. DeepEval's GEval takes criteria, steps or a rubric of score ranges. When the model returns logprobs, it replaces the sampled score with a probability-weighted average. Its optional hybrid mode lets a hosted "System One" model answer the decision points and falls back to the LLM only when that call fails. Ragas validates every judge call into a Pydantic model through Instructor, and retries with a fix-format prompt. The editor found its strictness voting inert at the reviewed commit, so AspectCritic is a single-sample judge. Opik passes a Pydantic response_format, retries parse failures three times, and merges identical LLMJudge instances into one call. Its default judge is gpt-5-nano. Phoenix builds a JSON schema with an enum of allowed labels. On OpenAI it tries native structured output and falls back to tool calling. Langfuse validates numeric, boolean or categorical output with Zod and records every judge call as its own trace. It pauses an evaluator when the judge fails with an auth or billing error.
Text-parsed judges. Inspect takes the last GRADE: X in the reply and escapes [BEGIN DATA] markers in dataset text. promptfoo parses a JSON {pass, score, reason}. Red-team configs without an explicit grader prefer promptfoo's remote grading endpoint. OpenAI Evals matches choice strings line by line, scores an unparseable answer as the lowest choice, and uses the last completion function on the command line as the judge. garak's ModelAsJudge asks for a [[rating]] from 1 to 10 and counts 7 or more as a hit. Its default judge is Llama 3 70B through NVIDIA NIM.
No judge in the core. lm-evaluation-harness is reference-based. Only the pisa_*_llm_judged tasks call an OpenAI model from their own hooks.
Langfuse, Phoenix, Ragas, OpenAI Evals, Inspect and garak have no calibration or position-bias controls.
Pick: Inspect when judge robustness matters, for example graded agent tasks. Pick: DeepEval or Ragas for rubric and claim-level judging in Python. Pick: Langfuse or Opik to run judges continuously on live traces.
How are test datasets and cases defined, generated and versioned?
Phoenix, Opik and Langfuse are the only projects that version test data themselves. DeepEval and Ragas are best at generating test data. lm-evaluation-harness and OpenAI Evals are benchmark registries.
Versioned datasets on a server. Phoenix stores DatasetVersion, per-example revisions and splits, and each experiment is pinned to a dataset version. Langfuse versions items in time with validFrom/validTo. A dataset-level eval reads each item at its validFrom version, and items can point back to the trace they came from. Opik has dataset versions and test suites whose items carry their own evaluator configs and pass policies (runs_per_item, pass_threshold). None of the three generates synthetic data.
Datasets in code, with generators. DeepEval's Synthesizer rewrites seed inputs with evolutions (reasoning, multi-context, hypothetical and others). It also wraps about 15 standard benchmarks, but dataset versions live on the vendor's hosted platform. Ragas's TestsetGenerator builds a knowledge graph from your documents, invents personas and writes single-hop and multi-hop questions. Its version_experiment commits your working tree to a git branch. promptfoo reads tests from YAML, CSV, XLSX, Google Sheets or Hugging Face, and expands array variables into every combination. Versioning is git plus a config snapshot per run. Inspect Samples can carry sandbox files and setup. Versioning is left to the source, such as a Hugging Face revision.
Benchmark and attack registries. lm-evaluation-harness has about 220 benchmark folders and close to 14,000 task YAML files, each with a metadata.version. OpenAI Evals has 463 eval YAML files named <base>.<split>.v<N>, and its JSONL samples are stored in Git LFS. garak ships attack payloads as JSON under garak/data/. A copy in the user data directory overrides them. It has no versioning and no generation.
Pick: Phoenix or Opik when experiments must be pinned to a dataset version. Pick: Ragas or DeepEval to generate questions from your own documents. Pick: lm-evaluation-harness for public benchmarks with versioned task configs.
How are evals executed and reported?
promptfoo and DeepEval fit most easily into CI. Inspect has the most robust runner for long or agentic evals. Langfuse and Opik run evaluations as continuous server jobs.
Test runners for CI. promptfoo expands prompt × provider × test into a matrix. It runs 4 steps at a time by default, or 1 when a test carries conversation state. Responses are cached in memory and on disk with a 14-day TTL, and results go to a local SQLite database with a web viewer. DeepEval's deepeval test run wraps pytest, supports xdist and returns pytest's exit code. Cases marked flaky only warn. Ragas runs jobs with 16 workers by default and turns failed rows into NaN. Its ragas evals CLI depends on a Project class the code marks as not implemented, so CI gates should use the Python API.
Research runners with logs. Inspect stores a full event transcript per sample in a .eval zip archive. eval_set retries and resumes from logs, and inspect view browses and compares runs. lm-evaluation-harness batches requests by type, shards them across data-parallel ranks, and can cache responses in SQLite. It has no baselines or thresholds. OpenAI Evals runs a 10-thread pool and writes JSONL to /tmp/evallogs. The editor found that its --no-cache flag has no effect. garak streams a JSONL report and a hitlog of successful attacks, then builds an HTML digest.
Platform experiments. Langfuse schedules eval jobs on BullMQ queues with deterministic job ids, so a trace is not scored twice. Its CI integration is whatever you build on its API. Opik's evaluate() uses 16 threads, can resume an interrupted run, and compares experiments in its UI. Phoenix's run_experiment pins a dataset version and lines runs up per example in the UI. The editor found that the CI gate in its evals/pxi folder is the team's internal suite and is not shipped with the package.
Pick: promptfoo or DeepEval for a pass/fail check in CI. Pick: Inspect for long agent runs that must resume and stay auditable. Pick: Langfuse or Opik to keep scoring after deployment.
How are traces or production data captured and linked to evaluations?
Langfuse and Opik are the only projects that score production traces as they arrive. Phoenix has the cleanest OpenTelemetry ingest, but you script online scoring yourself. The other seven tools trace only their own eval runs, or nothing.
Trace platforms. Langfuse accepts OTLP and its own SDK events. Evaluation rules pick traces with a deterministic SHA-256 sample, so the same trace is always in or out. Each judge call is itself a trace, linked from the score by executionTraceId. Human annotations share the score table with source: ANNOTATION. Note that the SDKs are in separate repositories. Opik captures spans with @opik.track, an OTel span processor or its OTLP endpoint. A sampler using SecureRandom sends traces through Redis streams to LLM-judge or Python-metric scorers. Experiment items are ordinary traces, so offline and online scores sit side by side. Phoenix receives OTLP over gRPC or HTTP with OpenInference attributes. Scores, labels and code checks all become span annotations, and log_span_annotations attaches offline results to production spans. Its online runner in evals/pxi is internal and not shipped.
Libraries that trace evaluations. DeepEval has @observe and OpenAI and Anthropic patchers, plus an OTel processor that forwards spans to its vendor's hosted platform. Online evaluation of production traffic is a feature of that platform. promptfoo starts a local OTLP receiver during a run so that assertions can check the spans your app emitted. It can also read traces from Langfuse, Braintrust or Tempo, but it never captures production traffic. Ragas records an evaluation → row → metric → prompt callback tree and has Langfuse and MLflow helpers. It has no OTel support.
Offline logs only. Inspect records typed events per sample and exposes lifecycle hooks, but has no OTel. lm-evaluation-harness logs to W&B or Trackio. OpenAI Evals records prompt and output events to JSONL. For garak the question is not applicable. It is a scanner with a report file.
Pick: Langfuse for rule-based online scoring with auditable judges. Pick: Opik when you also want a Python metric library on the same traces. Pick: Phoenix for standard OTLP ingest and annotation-driven review.
Does it support red-teaming or safety testing, and how?
Only garak and promptfoo generate attacks. garak is a self-contained scanner for a model or endpoint. promptfoo puts red-teaming into the same config and report as its quality tests, but leans on a hosted API.
Attack generators. garak has about 44 probe modules: DAN variants, prompt injection, latent injection hidden in documents, encoding and smuggling tricks, and attacker-model loops (goat, tap, atkgen). TreeSearchProbe branches on detector scores. Its REST and function generators let you scan a deployed app rather than only the bare model, and results export to AVID. The editor found its buffs limited to Base64, CharCode, lowercase, low-resource-language translation and paraphrase, with no leetspeak or typo buffs. The tool-attacking agent_breaker probe is off by default. promptfoo has dozens of plugins, with over 30 harm categories in the harmful set alone, and about 30 strategies (crescendo, GOAT, Hydra, GCG, Base64, leetspeak and others). It writes the generated cases to redteam.yaml, then evaluates them like any other test. Plugins in REMOTE_ONLY_PLUGIN_IDS exist only on promptfoo's API. Disabling remote generation removes them, and a 100,000-probe monthly limit applies unless you log in to its cloud.
Safety metrics, not attacks. These tools judge outputs you already have. Opik has a regex PromptInjection heuristic, Moderation, a SycEval sycophancy test, bias presets and guardrails. DeepEval moved red-teaming to the separate DeepTeam project, so the question is not applicable to it. Toxicity, bias and PII-leakage judges plus a prompt-injection classifier remain. Langfuse ships two managed judge templates, "Detect Prompt Injection" and "Check Rule Adherence". Ragas has harmfulness and maliciousness critics. Phoenix has toxicity and PII evaluators. It has no red-teaming tooling.
Safety benchmarks and building blocks. lm-evaluation-harness runs ToxiGen, BEAR and the advanced-AI-risk persona tasks as multiple-choice benchmarks. OpenAI Evals includes a narrow prompt-injection eval plus sandbagging and steganography suites. Inspect provides Docker sandboxes, approval policies and agents for writing your own attack tasks, but ships no attack library.
Pick: garak to scan a model or endpoint with no hosted dependency. Pick: promptfoo for app-level red-teaming next to your regression tests. Pick: Inspect to build custom agentic safety evaluations.
The projects
- langfuse/langfuse: Self-hostable LLM tracing and eval platform: Next.js API, BullMQ worker, ClickHouse traces, and queued LLM-judge and code evaluators.
- promptfoo/promptfoo: Local-first CLI and Node library that runs prompt x provider x test matrices, grades them with ~70 assertion types, and red-teams targets.
- comet-ml/opik: LLM tracing and eval platform: Python/TS SDKs with ~30 metrics, a Java backend on ClickHouse and MySQL, and Redis-streamed online scoring.
- openai/evals: Python CLI and YAML registry of ~460 benchmark evals that run models over JSONL samples and log every match and sample as events.
- confident-ai/deepeval: Pytest-style LLM evaluation library with ~50 judge-based metrics, G-Eval rubrics, tracing, dataset synthesis and Confident AI upload.
- vibrantlabsai/ragas: Python library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.
- EleutherAI/lm-evaluation-harness: Benchmark harness that scores a model on YAML-defined academic tasks via loglikelihood or generation requests, with bootstrap stderr.
- Arize-ai/phoenix: Self-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.
- NVIDIA/garak: Vulnerability scanner that fires adversarial probe prompts at a model or app endpoint and scores responses with detectors.
- UKGovernmentBEIS/inspect_ai: Python eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.