LLMs Technical Reviews
Home / LLM evals and testing

LLM evals and testing

Frameworks and platforms that test LLM apps and models: metrics and LLM judges, datasets, eval runners, tracing and red-teaming.

LLM eval tools measure whether a model or an LLM application behaves as intended. The ten projects here fall into three kinds. Langfuse, Opik, Phoenix and promptfoo are platforms that store runs and traces and show them in a UI. DeepEval, Ragas, OpenAI Evals and Inspect are libraries and runners that you call from code or CI. lm-evaluation-harness benchmarks base models, and garak scans models and endpoints for security weaknesses. Four axes separate them. The first is the target: a model on public benchmarks, an application on your own data, or a deployed endpoint under attack. The second is scoring: reference-based metrics, LLM judges, or detectors. The third is timing: offline runs against datasets, or online scoring of production traces. The fourth is how much depends on a hosted service. When choosing, check how judge output is parsed and validated, and whether datasets are versioned. Check that results can fail a CI job, and whether you must run a server.

Projects (10)

ProjectStarsLanguageLicense
langfuse/langfuseSelf-hostable LLM tracing and eval platform: Next.js API, BullMQ worker, ClickHouse traces, and queued LLM-judge and code evaluators.★ 35kTypeScriptMIT (ee/ folders under a commercial licence)
promptfoo/promptfooLocal-first CLI and Node library that runs prompt x provider x test matrices, grades them with ~70 assertion types, and red-teams targets.★ 26kTypeScriptMIT
comet-ml/opikLLM tracing and eval platform: Python/TS SDKs with ~30 metrics, a Java backend on ClickHouse and MySQL, and Redis-streamed online scoring.★ 22kPythonApache-2.0
openai/evalsPython CLI and YAML registry of ~460 benchmark evals that run models over JSONL samples and log every match and sample as events.★ 20kPython—
confident-ai/deepevalPytest-style LLM evaluation library with ~50 judge-based metrics, G-Eval rubrics, tracing, dataset synthesis and Confident AI upload.★ 19kPythonApache-2.0
vibrantlabsai/ragasPython library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.★ 16kPythonApache-2.0
EleutherAI/lm-evaluation-harnessBenchmark harness that scores a model on YAML-defined academic tasks via loglikelihood or generation requests, with bootstrap stderr.★ 14kPythonMIT
Arize-ai/phoenixSelf-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.★ 12kPythonElastic-2.0
NVIDIA/garakVulnerability scanner that fires adversarial probe prompts at a model or app endpoint and scores responses with detectors.★ 9.5kPythonApache-2.0
UKGovernmentBEIS/inspect_aiPython eval framework where a Task wires a dataset to solvers or agents and scorers, run async with sandboxes and logged per sample.★ 2.9kPythonMIT

Comparison questions

Each question is answered separately for every project in this category, from that project's source code.

All verdicts on one page →

  1. Which evaluation metrics and scorers are provided, and how are they implemented?Built-in metrics (heuristic, statistical, model-based); custom scorers; per-task vs per-trace scoring.
  2. How is LLM-as-a-judge implemented?Judge prompts and rubrics; structured output; multi-sample or consensus; judge model choice; calibration or bias controls.
  3. How are test datasets and cases defined, generated and versioned?File formats/DSL; synthetic data generation; versioning; golden sets; benchmark task registries.
  4. How are evals executed and reported?Runners; parallelism and caching; CI integration; result storage; comparison, regression and dashboards.
  5. How are traces or production data captured and linked to evaluations?SDK instrumentation; OpenTelemetry; online vs offline evals; feedback and annotation; if absent, say so.
  6. Does it support red-teaming or safety testing, and how?Adversarial probes and attack plugins; jailbreak/prompt-injection tests; vulnerability reporting; if absent, say so.