# How are evals executed and reported?

> LLM evals and testing — a good answer covers: Runners; parallelism and caching; CI integration; result storage; comparison, regression and dashboards.

Canonical page: https://llms-technical-reviews.com/evals/q/execution/

## Verdict

[promptfoo](/p/promptfoo/) and [DeepEval](/p/deepeval/) fit most easily into CI. [Inspect](/p/inspect_ai/) has the most robust runner for long or agentic evals. [Langfuse](/p/langfuse/) and [Opik](/p/opik/) run evaluations as continuous server jobs.

**Test runners for CI.** promptfoo expands prompt × provider × test into a matrix. It runs 4 steps at a time by default, or 1 when a test carries conversation state. Responses are cached in memory and on disk with a 14-day TTL, and results go to a local SQLite database with a web viewer. DeepEval's `deepeval test run` wraps pytest, supports xdist and returns pytest's exit code. Cases marked `flaky` only warn. [Ragas](/p/ragas/) runs jobs with 16 workers by default and turns failed rows into `NaN`. Its `ragas evals` CLI depends on a Project class the code marks as not implemented, so CI gates should use the Python API.

**Research runners with logs.** Inspect stores a full event transcript per sample in a `.eval` zip archive. `eval_set` retries and resumes from logs, and `inspect view` browses and compares runs. [lm-evaluation-harness](/p/lm-evaluation-harness/) batches requests by type, shards them across data-parallel ranks, and can cache responses in SQLite. It has no baselines or thresholds. [OpenAI Evals](/p/openai-evals/) runs a 10-thread pool and writes JSONL to `/tmp/evallogs`. The editor found that its `--no-cache` flag has no effect. [garak](/p/garak/) streams a JSONL report and a hitlog of successful attacks, then builds an HTML digest.

**Platform experiments.** Langfuse schedules eval jobs on BullMQ queues with deterministic job ids, so a trace is not scored twice. Its CI integration is whatever you build on its API. Opik's `evaluate()` uses 16 threads, can resume an interrupted run, and compares experiments in its UI. [Phoenix](/p/phoenix/)'s `run_experiment` pins a dataset version and lines runs up per example in the UI. The editor found that the CI gate in its `evals/pxi` folder is the team's internal suite and is not shipped with the package.

Pick: promptfoo or DeepEval for a pass/fail check in CI.
Pick: Inspect for long agent runs that must resume and stay auditable.
Pick: Langfuse or Opik to keep scoring after deployment.

## Per-project answers

### langfuse/langfuse (answered)

Eval execution is a **job-queue pipeline** built on BullMQ (Redis). **Scheduling** (`evalService.ts` `createEvalJobs` lines 302–913): when a trace is upserted or a dataset run item created, the system fetches active evaluation rules for the project, filters them against the trace (filter state and sampling), deduplicates via deterministic job IDs, and enqueues `EvalExecutionQueue` jobs. **Three parallel queues** exist for trace-level and observation-level evals, routed by `EvalExecutionQueue.getInstance({shardingKey})` (line 841). **Parallelism** is handled by BullMQ worker concurrency; no built-in batching per model provider. **Caching** (`deterministicSampling.ts`) uses SHA‑256 hashing of the target ID to produce a deterministic `[0,1)` value for sampling decisions. A trace-cache optimization (`evalService.ts` lines 389–442) fetches trace data once for multiple configs. **Execution** calls into `runLLMAsJudgeEvaluation` (or code/decision-model executors), which calls the LLM, validates the structured output, and persists scores. Results flow through `completeEvalExecution` (`evalCompletion.ts` lines 23–98): scores are uploaded to S3 as JSON event files, queued for ClickHouse ingestion via `IngestionQueue`, and the `JobExecution` row is updated to `COMPLETED` with the primary score ID and execution trace ID. **Error handling** (`evalExecutionMetrics.ts`) classifies outcomes into `success`, `platform_error`, `upstream_error`, `customer_error`, or `cancelled`. LLM errors are categorized and can auto-block the evaluator. **CI integration** is user-defined (no built-in harness); users call the public API or run evaluations ad-hoc via the UI's "Run evaluation" batch dialog. **Results** appear as scores on traces/observations in the Langfuse UI, filterable and searchable via the Eval Logs table (`eval-log.tsx`). **Dashboards and regression** are user-built on top of the score data — the platform stores all scores with timestamps, trace IDs, and metadata, enabling external charting or the built-in analytics views.

> **Editor's note.** Correction: `deterministicSampling.ts` is sampling, not caching: it hashes a domain prefix plus the target id with SHA-256 and compares the result to the rule's rate. Trace-level `evaluate` only executes LLM-as-judge templates; code and decision-model evaluators run through the observation-eval queues.

Citations: [worker/src/features/evaluation/evalService.ts:302-913](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L302-L913) · [worker/src/features/evaluation/evalCompletion.ts:23-98](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalCompletion.ts#L23-L98) · [worker/src/features/evaluation/evalExecutionMetrics.ts:1-141](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionMetrics.ts#L1-L141) · [web/src/features/evals/components/eval-log.tsx:1-60](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/components/eval-log.tsx#L1-L60)

### promptfoo/promptfoo (answered)

evaluate() (evaluator.ts L5576-5627) creates Evaluator instance that builds RunEvalOptions from provider-prompt-test-vars cross product. runEval (L1618) renders prompt via Nunjucks, calls provider, applies transforms, runs assertions. Parallelism via async.forEachOfLimit with configurable max-concurrency (default 4, max 20). Caching (src/cache.ts) uses two-level memory+disk store (KeyvFile) with 14-day TTL, namespace isolation via withCacheNamespace for repeat runs. CI: CIProgressReporter and cli-progress SingleBar. SQLite persistence via Drizzle: evalsTable (id, author, config snapshot, results JSON, prompts, isRedteam) and evalResultsTable. Also writes JSON/JSONL output files. Comparison assertions: select-best and max-score compare outputs across providers. Web dashboard (src/app/) displays history and side-by-side results.

> **Editor's note.** Correction: the default max concurrency is 4, but there is no maximum of 20; concurrency is forced to 1 when `_conversation`, `storeOutputAs` or a persistent browser session is used.

Citations: [src/evaluator.ts:5576-5627](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L5576-L5627) · [src/evaluator.ts:1618-1684](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1618-L1684) · [src/cache.ts:1-80](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/cache.ts#L1-L80) · [src/database/tables.ts:58-98](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/database/tables.ts#L58-L98) · [src/evaluator.ts:880-892](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L880-L892)

### comet-ml/opik (answered)

Evaluation execution is orchestrated by `EvaluationEngine` (`evaluation/engine/engine.py:67-676`).

**Runners and parallelism**: The primary entry point is `evaluate()` in `evaluator.py:138-382`, which creates an experiment, resolves dataset items, then delegates to `EvaluationEngine.run_and_score()`. Parallelism is managed by `StreamingExecutor` (`evaluation_tasks_executor.py`), which accepts a `workers` parameter (default 16). Each dataset item gets its own task submission; multiple `trial_count` runs per item are executed with the same executor, grouped by item ID for progress tracking. `task_threads=1` runs sequentially. A separate `score_test_cases()` method re-scores existing experiments without task execution.

**Execution policies**: Each item can have a per-item `ExecutionPolicy` (runs_per_item, pass_threshold) merged with a suite-level default. The `get_item_execution_policy()` function merges item-level overrides (`engine.py:33-64`).

**Caching**: `GEval` has a thread-safe LRU chain-of-thought cache (128 entries, keyed by task_introduction + criteria + model_name + model fingerprint) — shared across all instances via class-level OrderedDict (`metrics/llm_judges/g_eval/metric.py:66-173`). No general response caching is built in.

**CI integration**: No built-in CI adapter. The SDK can be used in any Python CI pipeline via `evaluate()` or `run_tests()`. JSON report generation is available via `report_output_path`.

**Result storage**: Feedback scores are logged to the Opik backend during evaluation via `rest_operations.log_test_result_feedback_scores()`. Each score is written to a trace-level feedback score on the experiment trace. `EvaluationResult` contains all test results, experiment metadata, and an optional experiment URL.

**Comparison, regression and dashboards**: `run_tests()` runs test suites with pass/fail aggregation based on execution policy thresholds. `evaluate()` creates experiments linked to dataset versions. The backend provides experiment comparison views (via the Java backend + React frontend). Dashboard objects (`opik.api_objects.dashboard.Dashboard`) are SDK-level wrappers for the UI dashboards. `evaluate_resume()` (`evaluator.py:1510-1650`) restores interrupted evaluations by reading experiment state and local checkpoints — skipping already-completed items and re-executing only pending ones, then merging results.

**Error tolerance**: `ErrorTolerance` enum controls how many scoring failures abort the run. `METRIC_ERRORS` (default, 10) tolerates exceptions raised inside `score()`; `ALL_SCORING_ERRORS` (20) additionally tolerates metrics that cannot even be built. Tolerated failures produce `ScoreResult(scoring_failed=True)` with error details, excluded from aggregates.


Citations: [sdks/python/src/opik/evaluation/evaluator.py:138-700](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/evaluator.py#L138-L700) · [sdks/python/src/opik/evaluation/evaluator.py:1510-1650](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/evaluator.py#L1510-L1650) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:66-173](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L66-L173)

### openai/evals (answered)

**CLI entry points.** `oaieval` (`evals/cli/oaieval.py:297-311`) is the primary runner. It takes a `completion_fn` (model ID or registry key), an `eval` name, and optional flags (`--max_samples`, `--cache`, `--seed`, `--extra_eval_params`, etc.). `oaievalset` (`evals/cli/oaievalset.py:134-141`) runs an ordered set of evals as subprocess calls to `oaieval`, with checkpoint/resume via a progress file.

**Runners and parallelism.** The `Eval` base class (`evals/eval.py:46-147`) defines `eval_sample()` (abstract, per-sample logic) and `run()` (abstract, orchestrates loading + scoring). `eval_all_samples()` shuffles samples with a fixed seed, then processes them via `ThreadPool` (default 10 threads, controlled by `EVALS_THREADS` env var) or sequentially (`EVALS_SEQUENTIAL=1`). `async_eval_all_samples()` provides an asyncio-based path with configurable concurrency and a semaphore (`evals/eval.py:112-147`). The `SolverEval` subclass copies the solver per sample for state isolation (`evals/eval.py:168-255`). Gentle interrupt (`EVALS_GENTLE_INTERRUPT`) allows early stopping with partial results.

**Caching.** `--cache` controls an API-level cache_level; disabled with `--no-cache` (`evals/cli/oaieval.py:211-212`).

**Result storage.** Three recorders implement `RecorderBase`: `LocalRecorder` writes JSONL locally (default, at `/tmp/evallogs/...`) (`evals/record.py:316-371`); `Recorder` writes to Snowflake plus local fallback (`evals/record.py:468-581`); `HttpRecorder` POSTs batches to a URL with failover to local storage (`evals/record.py:374-465`). `DummyRecorder` logs to console for dry-runs (`evals/record.py:274-314`). All events (`match`, `sampling`, `metrics`, `error`, `extra`, etc.) share the same flush-and-batch infrastructure with configurable thresholds (`MIN_FLUSH_EVENTS=100`, `MIN_FLUSH_SECONDS=10`).

**Comparison and dashboards.** There is no built-in comparison, regression detection, or dashboard. Results are raw JSONL files or Snowflake tables — analysis is left to external tooling. Final reports are logged to console and stored, showing per-metric values (`evals/cli/oaieval.py:236-238`).

> **Editor's note.** Correction: `--no-cache` has no effect at this commit. `oaieval.py` puts `cache_level = 0` into a local `api_extra_options` dict (L210-L212) that is never passed to the eval or the completion functions.

Citations: [evals/cli/oaieval.py:117-240](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L117-L240) · [evals/eval.py:46-167](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L46-L167) · [evals/record.py:316-466](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L316-L466) · [evals/eval.py:112-147](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/eval.py#L112-L147) · [evals/cli/oaievalset.py:80-141](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaievalset.py#L80-L141)

### confident-ai/deepeval (answered)

The primary entry point is `evaluate()` in `deepeval/evaluate/evaluate.py:1-80`, which accepts an `EvaluationDataset` (or list of test cases) and a list of metrics/classifiers. For individual test assertions, `assert_test()` wraps a single test case and runs metrics synchronously or asynchronously (`deepeval/evaluate/evaluate.py:84-158`). Under the hood, `execute_test_cases()` / `a_execute_test_cases()` iterate test cases, evaluate each metric, and collect `TestResult` objects (`deepeval/evaluate/execute/e2e.py:144-270`). **Parallelism**: `async_mode` uses `asyncio` with configurable `max_concurrent` throttling via `AsyncConfig`. **Caching**: `TestRunCacheManager` caches per-test-case metric results based on hyperparameters; `use_cache` / `write_cache` flags in `CacheConfig` control disk persistence (`deepeval/evaluate/execute/e2e.py:260-270`). **Result storage**: Runs are exported as local JSON files at `{HIDDEN_DIR}/.temp_test_run_data.json` and `.latest_test_run.json` (`deepeval/test_run/test_run.py:64-68`). When integrated with Confident AI, results are uploaded via API (`APIEvaluate` model) for cloud persistence. **Console reporting**: `EvaluationConsoleReport` uses Rich to display per-metric scores, pass/fail, and aggregate summaries (`deepeval/evaluate/console_report.py`). **Comparison**: `compare()` (`deepeval/evaluate/compare.py:48-80`) runs `ArenaGEval` pairwise across `ArenaTestCase` contestants and produces a win/loss dictionary. **CI integration**: `assert_test()` raises `AssertionError` on failure (metric.success=False), making it pytest-compatible. **Agentic trace execution**: `execute_agentic_test_cases_from_loop()` evaluates traces collected during agent runs, matching traces to golden expectations. **Hyperparameters** are captured from model config and prompt hash/alias, linked to test run for full reproducibility (`deepeval/test_run/hyperparameters.py:10-58`).


Citations: [deepeval/evaluate/evaluate.py:1-158](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/evaluate.py#L1-L158) · [deepeval/evaluate/execute/e2e.py:144-270](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/execute/e2e.py#L144-L270) · [deepeval/evaluate/compare.py:48-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/compare.py#L48-L80) · [deepeval/test_run/test_run.py:64-75](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/test_run/test_run.py#L64-L75) · [deepeval/test_run/hyperparameters.py:10-58](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/test_run/hyperparameters.py#L10-L58)

### vibrantlabsai/ragas (answered)

Evaluation execution is centered on the `Executor` class and the newer `@experiment` decorator pattern.

**Runners**: The `Executor` (executor.py:18-216) manages async job queues. Jobs are submitted via `.submit(callable, *args, name=...)` and executed concurrently via `as_completed()` with a configurable `max_workers` limit from `RunConfig`. It wraps each callable with error handling (returns `np.nan` on failure unless `raise_exceptions=True`). Both synchronous `.results()` and async `.aresults()` paths are supported. In the `evaluate()`/`aevaluate()` flow (now deprecated in favor of `@experiment`), one `Executor` is created per evaluation run. Each metric's `single_turn_ascore()` or `multi_turn_ascore()` is submitted as a separate job per dataset row (evaluation.py:244-276).

**Parallelism and caching**: Concurrency is controlled by `RunConfig.max_workers` (default unlimited). Batching is optional via `batch_size` — batches process as nested progress bars. LLM responses are cacheable via `CacheInterface`: `DiskCacheBackend` (diskcache) stores results keyed by SHA-256 of the function + arguments. The `cacher` decorator wraps `generate_text`/`agenerate_text` (llms/base.py:60-66). This provides ~60x speedup for repeated evaluations with identical inputs.

**CI integration**: No native CI integration — tests use standard `pytest` via `make test`. The CLI (`ragas` command via typer) can run experiments from the command line, making it scriptable in CI pipelines.

**Result storage**: Results are returned as `EvaluationResult` (dataset_schema.py:411-552), which exposes per-metric scores as dict-of-lists, mean-aggregated scores via `_repr_dict`, cost tracking via `CostCallbackHandler` (total tokens, total cost), and full run traces via `RagasTracer` (callback tree of evaluation → rows → metrics → prompts). Conversion to pandas is available via `.to_pandas()`. The `Experiment`/`@experiment` decorator pattern (experiment.py:103-232) saves results to a backend automatically.

**Comparison, regression, dashboards**: The CLI (cli.py:54-99) implements baseline comparison: it renders Rich tables showing current, baseline, delta values and pass/fail gates with small-regression tolerance (±0.01). No web dashboard exists. The CLI supports `ragas evaluate` for running experiments and `ragas compare` for baseline comparisons with categorical frequency tables.

> **Editor's note.** Correction: `RunConfig.max_workers` defaults to 16, not unlimited. The CLI has no `evaluate` or `compare` commands, only `evals` (baseline gate tables via `--baseline`), `quickstart` and `hello_world`, and `evals` depends on a Project class the code marks as not implemented.

Citations: [src/ragas/executor.py:18-216](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/executor.py#L18-L216) · [src/ragas/evaluation.py:59-345](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L59-L345) · [src/ragas/experiment.py:103-232](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L103-L232) · [src/ragas/cli.py:54-100](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/cli.py#L54-L100) · [src/ragas/dataset_schema.py:411-552](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L411-L552)

### EleutherAI/lm-evaluation-harness (answered)

Evals are executed via the `lm-eval` CLI (`lm_eval/_cli/harness.py`), which provides `run`, `ls` (list tasks), and `validate` subcommands. The core entry point is `simple_evaluate()` in `lm_eval/evaluator.py:55-425`, which initializes the model, loads tasks via `TaskManager`, then calls `evaluate()` (line 429-714). **Runners**: For each task, `task.build_all_requests()` generates `Instance` objects across output types (`loglikelihood`, `generate_until`, etc.). Requests are grouped by type and dispatched via `getattr(lm, reqtype)(cloned_reqs)` — the abstract `LM` base class at `lm_eval/api/model.py:25` defines `loglikelihood()`, `loglikelihood_rolling()`, and `generate_until()`. Concrete backend models (HF, vLLM, API) implement these with batching. **Parallelism**: For data-parallel (DP) models, `lm.rank`/`lm.world_size` partition documents; for tensor-parallel (TP), padding equalizes batch sizes across ranks (`evaluator.py:568-584`). Results are gathered via `lm.gather_object()` across ranks. **Caching**: `lm_eval/caching/cache.py` provides pickle-based disk caching of built requests (keyed by task+num_fewshot+rank) using dill. An SQLite-based request-response cache (`CachingLM`) is available via `--use_cache` (`evaluator.py:276-289`). **Result storage**: `EvaluationTracker` at `lm_eval/loggers/evaluation_tracker.py:123-230` handles saving aggregated results as JSON (with timestamped filenames) and optionally pushing to HuggingFace hub datasets repos. The WandbLogger (`lm_eval/loggers/wandb_logger.py`) logs metrics as wandb runs. The TrackioLogger (`lm_eval/loggers/trackio_logger.py`) logs per-sample traces. **CI integration**: GitHub Actions workflows in `.github/workflows/` run tests. **Comparison/regression/dashboards**: Not built-in — results are JSON files pushed to HF hub, lacking formal regression testing or dashboards within the harness itself.


Citations: [lm_eval/evaluator.py:55-88](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L55-L88) · [lm_eval/evaluator.py:428-605](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L428-L605) · [lm_eval/api/model.py:25-100](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/model.py#L25-L100) · [lm_eval/caching/cache.py:1-80](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/caching/cache.py#L1-L80) · [lm_eval/loggers/evaluation_tracker.py:123-230](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/evaluation_tracker.py#L123-L230)

### Arize-ai/phoenix (answered)

Evaluation execution uses a **producer-consumer pattern** with two executor classes in `packages/phoenix-evals/src/phoenix/evals/executors.py`. `SyncExecutor` (`executors.py:464-573`) is a synchronous loop with configurable `max_retries` (default 10), `exit_on_error` (default True), tqdm progress bar, and SIGINT handling. `AsyncExecutor` (`executors.py:170-462`) is an async producer-consumer with configurable `concurrency` (default 3), retries with requeue into a priority queue, per-task `timeout` (default 60s), and a **dynamic concurrency controller** using an AIMD (Additive Increase/Multiplicative Decrease) algorithm (`executors.py:73-167`) that adapts target concurrency based on error rates — automatically scaling down on failures and back up during success windows, with a collapse mode that drops concurrency to 1 on multiple errors within a window. The high-level `evaluate_dataframe()` (`evaluators.py:1408-1555`) and `async_evaluate_dataframe()` (`evaluators.py:1558-1735`) accept a DataFrame and list of evaluators, create the task list as a Cartesian product of (row × evaluator), execute, and add `{score_name}_score` and `{evaluator_name}_execution_details` columns. The `get_executor_on_sync_context()` helper (`executors.py:576-645`) automatically selects between sync and async execution based on the current thread and event loop. **Online (production) evals**: The PXI online-eval runner (`evals/pxi/online_evals/run.py:445-483`) is a CLI script that discovers candidate spans via `client.spans.get_spans()`, deduplicates against existing annotations, applies deterministic per-trace sampling (`_sampled()` using SHA256 of artifact ID), fetches full trace spans, evaluates with up to 8 concurrent evaluations, and writes `SpanAnnotationData` back to the Phoenix server in batches of 100. The runner produces a `RunSummary` with discovered/evaluated/sampled-out/already-annotated/error counts and supports GitHub Step Summary output. **CI integration**: no built-in CI integration — users run evals as scripts or notebooks and can compare results manually. **Result storage**: scores are either columns in a DataFrame (offline) or annotations persisted to the Phoenix server via `client.spans.log_span_annotations()`. The `to_annotation_dataframe()` utility (`utils.py:410-491`) reformats DataFrame scores into annotation format for logging.

> **Editor's note.** Correction: evals/pxi (online runner, CI gate) is Phoenix's internal eval suite for its own assistant and is not shipped in the package. User-facing execution is client.experiments.run_experiment or the server-side ExperimentRunner daemon, with results compared per example in the UI.

Citations: [packages/phoenix-evals/src/phoenix/evals/executors.py:170-462](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/executors.py#L170-L462) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:1408-1735](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L1408-L1735) · [evals/pxi/online_evals/run.py:39-120](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L39-L120) · [evals/pxi/online_evals/run.py:210-342](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L210-L342) · [packages/phoenix-evals/src/phoenix/evals/executors.py:73-167](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/executors.py#L73-L167) · [packages/phoenix-evals/src/phoenix/evals/utils.py:410-491](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L410-L491)

### NVIDIA/garak (answered)

Two harnesses coordinate execution. `ProbewiseHarness` (`harnesses/probewise.py:18-111`) runs probes one-by-one, each with its own recommended detectors (primary + extended). `PxD` (`harnesses/pxd.py:22-61`) runs all probe × detector combinations. Both call `Harness.run()` (`harnesses/base.py:159-290`), which: loads buffs, iterates probes, calls `probe.probe(model)` to get attempts, runs each detector on each attempt via `_run_detector()`, writes attempts to report JSONL, and calls `evaluator.evaluate()`.

Parallelism: `Generator.generate()` (`generators/base.py:138-244`) uses `multiprocessing.Pool` for parallel requests when `parallel_requests > 1` and the generator doesn't natively support multi-generation. `Probe._execute_all()` (`probes/base.py:321-383`) uses `Pool.imap_unordered()` for parallel attempts when `parallel_attempts > 1`.

Result storage: everything streams to `garak.<uuid>.report.jsonl` (JSONL, one JSON object per line). Entry types include `start_run setup`, `attempt`, `eval`, `probe_summary`, `plugin_cache`, `completion`. Hit logs go to a parallel `.hitlog.jsonl` when detectors fail. `Report` (`report.py:15-174`) loads a report and exports to AVID format. `report_digest.py:581-689` builds HTML digests with results stored in an in-memory SQLite3 database, enabling grouping by MISP taxonomy tags and per-group ASR aggregation (default: lower quartile, `garak.core.yaml:39`). Calibration Z-scores and bootstrap CIs are computed during digest building. No CI/CD integration or regression dashboards are built into the code.


Citations: [garak/harnesses/base.py:159-290](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/harnesses/base.py#L159-L290) · [garak/harnesses/probewise.py:18-111](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/harnesses/probewise.py#L18-L111) · [garak/harnesses/pxd.py:22-61](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/harnesses/pxd.py#L22-L61) · [garak/generators/base.py:138-244](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/generators/base.py#L138-L244) · [garak/probes/base.py:321-383](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/probes/base.py#L321-L383) · [garak/analyze/report_digest.py:58-101](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/analyze/report_digest.py#L58-L101)

### UKGovernmentBEIS/inspect_ai (answered)

**Runners.** Evals are launched via `eval()` / `eval_async()` (`src/inspect_ai/_eval/eval.py`). The top-level orchestrator `eval_set()` (`src/inspect_ai/_eval/evalset.py`) manages multiple tasks with shared configuration, sample dispatch, retry logic, and logging. Inside a task, `task_run()` sets up sandbox lifecycle and prepares task options. Samples are fanned out by `SampleScheduler`, a live fanout loop that accepts new samples mid-run.

**Parallelism.** `parallel=` controls concurrent samples, running in an anyio `TaskGroup`. Model-level parallelism uses adaptive connection pools (`DEFAULT_MAX_CONNECTIONS` / `DEFAULT_MAX_CONNECTIONS_BATCH`). `max_tasks` controls parallel task execution within eval sets.

**Caching.** Model outputs are cached with configurable TTL. The `inspect cache` CLI command (`src/inspect_ai/_cli/cache.py`) provides `clear`, `prune`, `list`, and `size` operations.

**CI integration.** Standard Python library approach — `pip install inspect-ai && python your_eval.py`. The CLI supports `--detach` for long-running evals, ACP for cross-machine dispatch, and JSON output modes.

**Result storage.** Logs use a compact binary `.eval` format (or JSON) via the `Recorder` hierarchy (`src/inspect_ai/log/_recorders/`). `FileRecorder` writes to local/remote filesystems via fsspec. `BufferSampleStore` provides crash-recovery through a SQLite buffer.

**Comparison / regression.** Eval results are structured as `EvalResults` with `EvalScore` objects per scorer and `EvalMetric` dicts per metric (`src/inspect_ai/log/_log.py:799-904`). The `inspect view` command starts a web dashboard for browsing logs and comparing runs. There is no built-in regression test framework — users compare metrics programmatically from returned `EvalLog` objects.

> **Editor's note.** Correction: --acp-server is not cross-machine dispatch. It exposes a running eval over the Agent Client Protocol so clients (inspect acp, editors) can attach for human-in-the-loop approvals. The .eval log format is a zip archive of JSON entries, not a custom binary format.

Citations: [src/inspect_ai/_eval/eval.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/eval.py#L1-L50) · [src/inspect_ai/_eval/task/run.py:139-300](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/run.py#L139-L300) · [src/inspect_ai/_eval/task/scheduler.py:1-100](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_eval/task/scheduler.py#L1-L100) · [src/inspect_ai/_cli/cache.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/cache.py#L1-L50) · [src/inspect_ai/log/_log.py:799-904](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/log/_log.py#L799-L904) · [src/inspect_ai/_cli/view.py:1-60](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/view.py#L1-L60)
