How are evals executed and reported?
Runners; parallelism and caching; CI integration; result storage; comparison, regression and dashboards.
Verdict
promptfoo and DeepEval fit most easily into CI. Inspect has the most robust runner for long or agentic evals. Langfuse and Opik run evaluations as continuous server jobs.
Test runners for CI. promptfoo expands prompt × provider × test into a matrix. It runs 4 steps at a time by default, or 1 when a test carries conversation state. Responses are cached in memory and on disk with a 14-day TTL, and results go to a local SQLite database with a web viewer. DeepEval’s deepeval test run wraps pytest, supports xdist and returns pytest’s exit code. Cases marked flaky only warn. Ragas runs jobs with 16 workers by default and turns failed rows into NaN. Its ragas evals CLI depends on a Project class the code marks as not implemented, so CI gates should use the Python API.
Research runners with logs. Inspect stores a full event transcript per sample in a .eval zip archive. eval_set retries and resumes from logs, and inspect view browses and compares runs. lm-evaluation-harness batches requests by type, shards them across data-parallel ranks, and can cache responses in SQLite. It has no baselines or thresholds. OpenAI Evals runs a 10-thread pool and writes JSONL to /tmp/evallogs. The editor found that its --no-cache flag has no effect. garak streams a JSONL report and a hitlog of successful attacks, then builds an HTML digest.
Platform experiments. Langfuse schedules eval jobs on BullMQ queues with deterministic job ids, so a trace is not scored twice. Its CI integration is whatever you build on its API. Opik’s evaluate() uses 16 threads, can resume an interrupted run, and compares experiments in its UI. Phoenix’s run_experiment pins a dataset version and lines runs up per example in the UI. The editor found that the CI gate in its evals/pxi folder is the team’s internal suite and is not shipped with the package.
Pick: promptfoo or DeepEval for a pass/fail check in CI. Pick: Inspect for long agent runs that must resume and stay auditable. Pick: Langfuse or Opik to keep scoring after deployment.
Per-project answers
langfuse/langfuse
answeredEval execution is a job-queue pipeline built on BullMQ (Redis). Scheduling (evalService.ts createEvalJobs lines 302–913): when a trace is upserted or a dataset run item created, the system fetches active evaluation rules for the project, filters them against the trace (filter state and sampling), deduplicates via deterministic job IDs, and enqueues EvalExecutionQueue jobs. Three parallel queues exist for trace-level and observation-level evals, routed by EvalExecutionQueue.getInstance({shardingKey}) (line 841). Parallelism is handled by BullMQ worker concurrency; no built-in batching per model provider. Caching (deterministicSampling.ts) uses SHA‑256 hashing of the target ID to produce a deterministic [0,1) value for sampling decisions. A trace-cache optimization (evalService.ts lines 389–442) fetches trace data once for multiple configs. Execution calls into runLLMAsJudgeEvaluation (or code/decision-model executors), which calls the LLM, validates the structured output, and persists scores. Results flow through completeEvalExecution (evalCompletion.ts lines 23–98): scores are uploaded to S3 as JSON event files, queued for ClickHouse ingestion via IngestionQueue, and the JobExecution row is updated to COMPLETED with the primary score ID and execution trace ID. Error handling (evalExecutionMetrics.ts) classifies outcomes into success, platform_error, upstream_error, customer_error, or cancelled. LLM errors are categorized and can auto-block the evaluator. CI integration is user-defined (no built-in harness); users call the public API or run evaluations ad-hoc via the UI's "Run evaluation" batch dialog. Results appear as scores on traces/observations in the Langfuse UI, filterable and searchable via the Eval Logs table (eval-log.tsx). Dashboards and regression are user-built on top of the score data — the platform stores all scores with timestamps, trace IDs, and metadata, enabling external charting or the built-in analytics views.
deterministicSampling.ts is sampling, not caching: it hashes a domain prefix plus the target id with SHA-256 and compares the result to the rule's rate. Trace-level evaluate only executes LLM-as-judge templates; code and decision-model evaluators run through the observation-eval queues.promptfoo/promptfoo
answeredevaluate() (evaluator.ts L5576-5627) creates Evaluator instance that builds RunEvalOptions from provider-prompt-test-vars cross product. runEval (L1618) renders prompt via Nunjucks, calls provider, applies transforms, runs assertions. Parallelism via async.forEachOfLimit with configurable max-concurrency (default 4, max 20). Caching (src/cache.ts) uses two-level memory+disk store (KeyvFile) with 14-day TTL, namespace isolation via withCacheNamespace for repeat runs. CI: CIProgressReporter and cli-progress SingleBar. SQLite persistence via Drizzle: evalsTable (id, author, config snapshot, results JSON, prompts, isRedteam) and evalResultsTable. Also writes JSON/JSONL output files. Comparison assertions: select-best and max-score compare outputs across providers. Web dashboard (src/app/) displays history and side-by-side results.
_conversation, storeOutputAs or a persistent browser session is used.comet-ml/opik
answeredEvaluation execution is orchestrated by EvaluationEngine (evaluation/engine/engine.py:67-676).
Runners and parallelism: The primary entry point is evaluate() in evaluator.py:138-382, which creates an experiment, resolves dataset items, then delegates to EvaluationEngine.run_and_score(). Parallelism is managed by StreamingExecutor (evaluation_tasks_executor.py), which accepts a workers parameter (default 16). Each dataset item gets its own task submission; multiple trial_count runs per item are executed with the same executor, grouped by item ID for progress tracking. task_threads=1 runs sequentially. A separate score_test_cases() method re-scores existing experiments without task execution.
Execution policies: Each item can have a per-item ExecutionPolicy (runs_per_item, pass_threshold) merged with a suite-level default. The get_item_execution_policy() function merges item-level overrides (engine.py:33-64).
Caching: GEval has a thread-safe LRU chain-of-thought cache (128 entries, keyed by task_introduction + criteria + model_name + model fingerprint) — shared across all instances via class-level OrderedDict (metrics/llm_judges/g_eval/metric.py:66-173). No general response caching is built in.
CI integration: No built-in CI adapter. The SDK can be used in any Python CI pipeline via evaluate() or run_tests(). JSON report generation is available via report_output_path.
Result storage: Feedback scores are logged to the Opik backend during evaluation via rest_operations.log_test_result_feedback_scores(). Each score is written to a trace-level feedback score on the experiment trace. EvaluationResult contains all test results, experiment metadata, and an optional experiment URL.
Comparison, regression and dashboards: run_tests() runs test suites with pass/fail aggregation based on execution policy thresholds. evaluate() creates experiments linked to dataset versions. The backend provides experiment comparison views (via the Java backend + React frontend). Dashboard objects (opik.api_objects.dashboard.Dashboard) are SDK-level wrappers for the UI dashboards. evaluate_resume() (evaluator.py:1510-1650) restores interrupted evaluations by reading experiment state and local checkpoints — skipping already-completed items and re-executing only pending ones, then merging results.
Error tolerance: ErrorTolerance enum controls how many scoring failures abort the run. METRIC_ERRORS (default, 10) tolerates exceptions raised inside score(); ALL_SCORING_ERRORS (20) additionally tolerates metrics that cannot even be built. Tolerated failures produce ScoreResult(scoring_failed=True) with error details, excluded from aggregates.
openai/evals
answeredCLI entry points. oaieval (evals/cli/oaieval.py:297-311) is the primary runner. It takes a completion_fn (model ID or registry key), an eval name, and optional flags (--max_samples, --cache, --seed, --extra_eval_params, etc.). oaievalset (evals/cli/oaievalset.py:134-141) runs an ordered set of evals as subprocess calls to oaieval, with checkpoint/resume via a progress file.
Runners and parallelism. The Eval base class (evals/eval.py:46-147) defines eval_sample() (abstract, per-sample logic) and run() (abstract, orchestrates loading + scoring). eval_all_samples() shuffles samples with a fixed seed, then processes them via ThreadPool (default 10 threads, controlled by EVALS_THREADS env var) or sequentially (EVALS_SEQUENTIAL=1). async_eval_all_samples() provides an asyncio-based path with configurable concurrency and a semaphore (evals/eval.py:112-147). The SolverEval subclass copies the solver per sample for state isolation (evals/eval.py:168-255). Gentle interrupt (EVALS_GENTLE_INTERRUPT) allows early stopping with partial results.
Caching. --cache controls an API-level cache_level; disabled with --no-cache (evals/cli/oaieval.py:211-212).
Result storage. Three recorders implement RecorderBase: LocalRecorder writes JSONL locally (default, at /tmp/evallogs/...) (evals/record.py:316-371); Recorder writes to Snowflake plus local fallback (evals/record.py:468-581); HttpRecorder POSTs batches to a URL with failover to local storage (evals/record.py:374-465). DummyRecorder logs to console for dry-runs (evals/record.py:274-314). All events (match, sampling, metrics, error, extra, etc.) share the same flush-and-batch infrastructure with configurable thresholds (MIN_FLUSH_EVENTS=100, MIN_FLUSH_SECONDS=10).
Comparison and dashboards. There is no built-in comparison, regression detection, or dashboard. Results are raw JSONL files or Snowflake tables — analysis is left to external tooling. Final reports are logged to console and stored, showing per-metric values (evals/cli/oaieval.py:236-238).
--no-cache has no effect at this commit. oaieval.py puts cache_level = 0 into a local api_extra_options dict (L210-L212) that is never passed to the eval or the completion functions.confident-ai/deepeval
answeredThe primary entry point is evaluate() in deepeval/evaluate/evaluate.py:1-80, which accepts an EvaluationDataset (or list of test cases) and a list of metrics/classifiers. For individual test assertions, assert_test() wraps a single test case and runs metrics synchronously or asynchronously (deepeval/evaluate/evaluate.py:84-158). Under the hood, execute_test_cases() / a_execute_test_cases() iterate test cases, evaluate each metric, and collect TestResult objects (deepeval/evaluate/execute/e2e.py:144-270). Parallelism: async_mode uses asyncio with configurable max_concurrent throttling via AsyncConfig. Caching: TestRunCacheManager caches per-test-case metric results based on hyperparameters; use_cache / write_cache flags in CacheConfig control disk persistence (deepeval/evaluate/execute/e2e.py:260-270). Result storage: Runs are exported as local JSON files at {HIDDEN_DIR}/.temp_test_run_data.json and .latest_test_run.json (deepeval/test_run/test_run.py:64-68). When integrated with Confident AI, results are uploaded via API (APIEvaluate model) for cloud persistence. Console reporting: EvaluationConsoleReport uses Rich to display per-metric scores, pass/fail, and aggregate summaries (deepeval/evaluate/console_report.py). Comparison: compare() (deepeval/evaluate/compare.py:48-80) runs ArenaGEval pairwise across ArenaTestCase contestants and produces a win/loss dictionary. CI integration: assert_test() raises AssertionError on failure (metric.success=False), making it pytest-compatible. Agentic trace execution: execute_agentic_test_cases_from_loop() evaluates traces collected during agent runs, matching traces to golden expectations. Hyperparameters are captured from model config and prompt hash/alias, linked to test run for full reproducibility (deepeval/test_run/hyperparameters.py:10-58).
vibrantlabsai/ragas
answeredEvaluation execution is centered on the Executor class and the newer @experiment decorator pattern.
Runners: The Executor (executor.py:18-216) manages async job queues. Jobs are submitted via .submit(callable, *args, name=...) and executed concurrently via as_completed() with a configurable max_workers limit from RunConfig. It wraps each callable with error handling (returns np.nan on failure unless raise_exceptions=True). Both synchronous .results() and async .aresults() paths are supported. In the evaluate()/aevaluate() flow (now deprecated in favor of @experiment), one Executor is created per evaluation run. Each metric's single_turn_ascore() or multi_turn_ascore() is submitted as a separate job per dataset row (evaluation.py:244-276).
Parallelism and caching: Concurrency is controlled by RunConfig.max_workers (default unlimited). Batching is optional via batch_size — batches process as nested progress bars. LLM responses are cacheable via CacheInterface: DiskCacheBackend (diskcache) stores results keyed by SHA-256 of the function + arguments. The cacher decorator wraps generate_text/agenerate_text (llms/base.py:60-66). This provides ~60x speedup for repeated evaluations with identical inputs.
CI integration: No native CI integration — tests use standard pytest via make test. The CLI (ragas command via typer) can run experiments from the command line, making it scriptable in CI pipelines.
Result storage: Results are returned as EvaluationResult (dataset_schema.py:411-552), which exposes per-metric scores as dict-of-lists, mean-aggregated scores via _repr_dict, cost tracking via CostCallbackHandler (total tokens, total cost), and full run traces via RagasTracer (callback tree of evaluation → rows → metrics → prompts). Conversion to pandas is available via .to_pandas(). The Experiment/@experiment decorator pattern (experiment.py:103-232) saves results to a backend automatically.
Comparison, regression, dashboards: The CLI (cli.py:54-99) implements baseline comparison: it renders Rich tables showing current, baseline, delta values and pass/fail gates with small-regression tolerance (±0.01). No web dashboard exists. The CLI supports ragas evaluate for running experiments and ragas compare for baseline comparisons with categorical frequency tables.
RunConfig.max_workers defaults to 16, not unlimited. The CLI has no evaluate or compare commands, only evals (baseline gate tables via --baseline), quickstart and hello_world, and evals depends on a Project class the code marks as not implemented.EleutherAI/lm-evaluation-harness
answeredEvals are executed via the lm-eval CLI (lm_eval/_cli/harness.py), which provides run, ls (list tasks), and validate subcommands. The core entry point is simple_evaluate() in lm_eval/evaluator.py:55-425, which initializes the model, loads tasks via TaskManager, then calls evaluate() (line 429-714). Runners: For each task, task.build_all_requests() generates Instance objects across output types (loglikelihood, generate_until, etc.). Requests are grouped by type and dispatched via getattr(lm, reqtype)(cloned_reqs) — the abstract LM base class at lm_eval/api/model.py:25 defines loglikelihood(), loglikelihood_rolling(), and generate_until(). Concrete backend models (HF, vLLM, API) implement these with batching. Parallelism: For data-parallel (DP) models, lm.rank/lm.world_size partition documents; for tensor-parallel (TP), padding equalizes batch sizes across ranks (evaluator.py:568-584). Results are gathered via lm.gather_object() across ranks. Caching: lm_eval/caching/cache.py provides pickle-based disk caching of built requests (keyed by task+num_fewshot+rank) using dill. An SQLite-based request-response cache (CachingLM) is available via --use_cache (evaluator.py:276-289). Result storage: EvaluationTracker at lm_eval/loggers/evaluation_tracker.py:123-230 handles saving aggregated results as JSON (with timestamped filenames) and optionally pushing to HuggingFace hub datasets repos. The WandbLogger (lm_eval/loggers/wandb_logger.py) logs metrics as wandb runs. The TrackioLogger (lm_eval/loggers/trackio_logger.py) logs per-sample traces. CI integration: GitHub Actions workflows in .github/workflows/ run tests. Comparison/regression/dashboards: Not built-in — results are JSON files pushed to HF hub, lacking formal regression testing or dashboards within the harness itself.
Arize-ai/phoenix
answeredEvaluation execution uses a producer-consumer pattern with two executor classes in packages/phoenix-evals/src/phoenix/evals/executors.py. SyncExecutor (executors.py:464-573) is a synchronous loop with configurable max_retries (default 10), exit_on_error (default True), tqdm progress bar, and SIGINT handling. AsyncExecutor (executors.py:170-462) is an async producer-consumer with configurable concurrency (default 3), retries with requeue into a priority queue, per-task timeout (default 60s), and a dynamic concurrency controller using an AIMD (Additive Increase/Multiplicative Decrease) algorithm (executors.py:73-167) that adapts target concurrency based on error rates — automatically scaling down on failures and back up during success windows, with a collapse mode that drops concurrency to 1 on multiple errors within a window. The high-level evaluate_dataframe() (evaluators.py:1408-1555) and async_evaluate_dataframe() (evaluators.py:1558-1735) accept a DataFrame and list of evaluators, create the task list as a Cartesian product of (row × evaluator), execute, and add {score_name}_score and {evaluator_name}_execution_details columns. The get_executor_on_sync_context() helper (executors.py:576-645) automatically selects between sync and async execution based on the current thread and event loop. Online (production) evals: The PXI online-eval runner (evals/pxi/online_evals/run.py:445-483) is a CLI script that discovers candidate spans via client.spans.get_spans(), deduplicates against existing annotations, applies deterministic per-trace sampling (_sampled() using SHA256 of artifact ID), fetches full trace spans, evaluates with up to 8 concurrent evaluations, and writes SpanAnnotationData back to the Phoenix server in batches of 100. The runner produces a RunSummary with discovered/evaluated/sampled-out/already-annotated/error counts and supports GitHub Step Summary output. CI integration: no built-in CI integration — users run evals as scripts or notebooks and can compare results manually. Result storage: scores are either columns in a DataFrame (offline) or annotations persisted to the Phoenix server via client.spans.log_span_annotations(). The to_annotation_dataframe() utility (utils.py:410-491) reformats DataFrame scores into annotation format for logging.
NVIDIA/garak
answeredTwo harnesses coordinate execution. ProbewiseHarness (harnesses/probewise.py:18-111) runs probes one-by-one, each with its own recommended detectors (primary + extended). PxD (harnesses/pxd.py:22-61) runs all probe × detector combinations. Both call Harness.run() (harnesses/base.py:159-290), which: loads buffs, iterates probes, calls probe.probe(model) to get attempts, runs each detector on each attempt via _run_detector(), writes attempts to report JSONL, and calls evaluator.evaluate().
Parallelism: Generator.generate() (generators/base.py:138-244) uses multiprocessing.Pool for parallel requests when parallel_requests > 1 and the generator doesn't natively support multi-generation. Probe._execute_all() (probes/base.py:321-383) uses Pool.imap_unordered() for parallel attempts when parallel_attempts > 1.
Result storage: everything streams to garak.<uuid>.report.jsonl (JSONL, one JSON object per line). Entry types include start_run setup, attempt, eval, probe_summary, plugin_cache, completion. Hit logs go to a parallel .hitlog.jsonl when detectors fail. Report (report.py:15-174) loads a report and exports to AVID format. report_digest.py:581-689 builds HTML digests with results stored in an in-memory SQLite3 database, enabling grouping by MISP taxonomy tags and per-group ASR aggregation (default: lower quartile, garak.core.yaml:39). Calibration Z-scores and bootstrap CIs are computed during digest building. No CI/CD integration or regression dashboards are built into the code.
UKGovernmentBEIS/inspect_ai
answeredRunners. Evals are launched via eval() / eval_async() (src/inspect_ai/_eval/eval.py). The top-level orchestrator eval_set() (src/inspect_ai/_eval/evalset.py) manages multiple tasks with shared configuration, sample dispatch, retry logic, and logging. Inside a task, task_run() sets up sandbox lifecycle and prepares task options. Samples are fanned out by SampleScheduler, a live fanout loop that accepts new samples mid-run.
Parallelism. parallel= controls concurrent samples, running in an anyio TaskGroup. Model-level parallelism uses adaptive connection pools (DEFAULT_MAX_CONNECTIONS / DEFAULT_MAX_CONNECTIONS_BATCH). max_tasks controls parallel task execution within eval sets.
Caching. Model outputs are cached with configurable TTL. The inspect cache CLI command (src/inspect_ai/_cli/cache.py) provides clear, prune, list, and size operations.
CI integration. Standard Python library approach — pip install inspect-ai && python your_eval.py. The CLI supports --detach for long-running evals, ACP for cross-machine dispatch, and JSON output modes.
Result storage. Logs use a compact binary .eval format (or JSON) via the Recorder hierarchy (src/inspect_ai/log/_recorders/). FileRecorder writes to local/remote filesystems via fsspec. BufferSampleStore provides crash-recovery through a SQLite buffer.
Comparison / regression. Eval results are structured as EvalResults with EvalScore objects per scorer and EvalMetric dicts per metric (src/inspect_ai/log/_log.py:799-904). The inspect view command starts a web dashboard for browsing logs and comparing runs. There is no built-in regression test framework — users compare metrics programmatically from returned EvalLog objects.
← How are test datasets and cases defined, generated and versioned? · How are traces or production data captured and linked to evaluations? →