# confident-ai/deepeval

> Pytest-style LLM evaluation library with ~50 judge-based metrics, G-Eval rubrics, tracing, dataset synthesis and Confident AI upload.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/confident-ai/deepeval (reviewed at commit `ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159`, 2026-10-05)
- Stars: 18669 · Language: Python · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/deepeval/

## Overview

DeepEval is a Python library for unit-testing LLM applications. You build `LLMTestCase` objects (input, actual output, expected output, retrieval context, tool calls and so on), attach metrics with thresholds, and either call `assert_test()` inside a pytest test or `evaluate()` over a list. Each metric is usually an LLM-as-a-judge pipeline: extract claims or statements, ask the judge for a verdict per item, compute a ratio, and ask for a reason. A metric passes when `score >= threshold`.

The package is much bigger than its core loop. At the pinned commit (v4.2.8) it ships about 50 metric packages (RAG, agentic, conversational, multimodal, voice, safety), a G-Eval implementation with logprob-weighted scores, DAG metrics, a set of label classifiers, a `Synthesizer` for generating goldens, a conversation simulator, a prompt optimizer, standard benchmarks (MMLU, GSM8K, HumanEval and others), and a tracing SDK with `@observe` and OpenTelemetry bridges.

Most of the persistence and visualisation is designed around Confident AI, the vendor's hosted platform. The open-source library runs fully locally and writes JSON under a hidden `.deepeval` folder. Dataset push and pull, metric collections, trace dashboards and annotation live on the platform. Red-teaming was split out into the separate DeepTeam project.

## Architecture

```mermaid
flowchart LR
  T["pytest test / script"] --> AT["assert_test / evaluate"]
  CLI["deepeval test run"] --> PY["pytest + deepeval plugin"]
  PY --> T
  AT --> EX["a_execute_test_cases"]
  EX --> CACHE["Test-run cache"]
  EX --> M["Metrics (BaseMetric)"]
  M --> S1["System One (Jev)"]
  M --> J["Judge LLM (DeepEvalBaseLLM)"]
  EX --> TRM["TestRunManager"]
  TRM --> LOCAL["Local JSON + console report"]
  TRM --> CAI["Confident AI upload"]
  OBS["@observe / OTel"] --> TM["TraceManager"]
  TM --> CAI
```

| Component | Path | Role |
|---|---|---|
| Entry points | `deepeval/evaluate/evaluate.py` | `assert_test` (pytest-friendly) and `evaluate` (batch) |
| Executors | `deepeval/evaluate/execute/` | Sync/async loops over test cases, agentic trace loops, per-task timeouts |
| Configs | `deepeval/evaluate/configs.py` | `AsyncConfig`, `DisplayConfig`, `CacheConfig`, `ErrorConfig` |
| Metrics | `deepeval/metrics/` | ~50 metric packages on `BaseMetric` / `BaseConversationalMetric` / `BaseArenaMetric` |
| Templates | `deepeval/templates/` | Judge prompt bundles resolved by `PromptMixin` |
| Models | `deepeval/models/` | `DeepEvalBaseLLM` plus OpenAI, Anthropic, Azure, Gemini, Bedrock, Ollama, LiteLLM and other wrappers; embeddings, TTS/STT |
| System One | `deepeval/models/system_one/`, `deepeval/metrics/utils/system_one.py` | Optional hosted "Jev" decision model and eval modes |
| Test cases | `deepeval/test_case/` | `LLMTestCase`, `ConversationalTestCase`, `ArenaTestCase`, MCP types |
| Test runs | `deepeval/test_run/` | Aggregation, cache, local files, upload |
| Datasets | `deepeval/dataset/`, `deepeval/synthesizer/` | Goldens, CSV/JSON loaders, `evals_iterator`, synthetic generation |
| Tracing | `deepeval/tracing/` | `TraceManager`, `@observe`, OpenAI/Anthropic patchers, OTel exporter/processor |
| CLI and plugin | `deepeval/cli/`, `deepeval/plugins/plugin.py` | `deepeval test run` wrapper around pytest |

## How a request flows

Take `deepeval test run test_rag.py`, where a test calls `assert_test(test_case, [FaithfulnessMetric(threshold=0.7)])`:

1. **CLI.** `run` sets process flags (running-deepeval, cache, ignore-errors, verbose), resets the global test run, and calls `pytest.main` with `-p deepeval` plus `-n` for xdist or `--count` for repeats. When pytest finishes it calls `wrap_up_test_run` and exits with pytest's code, so CI sees failures ([command.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/cli/test/command.py#L125-L199)).
2. **Plugin.** `pytest_sessionstart` creates the test run, and `pytest_runtest_call` wraps each test in an `Observer` span so that `@observe` spans attach to the run ([plugin.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/plugins/plugin.py#L51-L107)).
3. **Assert.** `assert_test` builds configs (async, up to 100 concurrent) and runs `a_execute_test_cases` on the single case ([evaluate.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/evaluate.py#L84-L158)).
4. **Execute.** The executor splits metrics into single-turn and conversational, guards each task with an `asyncio.Semaphore` and an outer deadline, and copies metrics per test case ([e2e.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/execute/e2e.py#L483-L560)). For each case it looks up a cached result keyed on the test case and hyperparameters, runs any System One batch, then measures the remaining metrics and classifiers ([e2e.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/execute/e2e.py#L741-L793)).
5. **Measure.** `FaithfulnessMetric.a_measure` first offers the case to System One. Otherwise it extracts truths from `retrieval_context` and claims from `actual_output` in parallel, generates a verdict per claim, computes the score, writes a reason and sets `success` ([faithfulness.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/faithfulness/faithfulness.py#L161-L199)). Each judge call goes through `generate_with_schema_and_extract`. Native models return a Pydantic instance plus cost, and custom models may return raw text that `trimAndLoadJson` repairs, for example by stripping trailing commas ([generation.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/utils/generation.py#L13-L97)).
6. **Record and cache.** In the `finally` block each metric becomes `MetricData`, is added to the run, and is written to the cache with cost zeroed ([e2e.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/execute/e2e.py#L819-L872)).
7. **Assert or warn.** If the result failed, `assert_test` raises an `AssertionError` that lists every failing metric's score, threshold and reason. Test cases marked `flaky` only emit a warning ([evaluate.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/evaluate.py#L160-L209)).
8. **Wrap up.** The run is saved to `.deepeval/.latest_test_run.json` and related files ([test_run.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/test_run/test_run.py#L64-L67)), printed as a Rich table, and uploaded if a Confident AI key is set.

## Key components

### Metric base and verdicts

`BaseMetric` carries the threshold, score, reason, cost, token counts and the eval-mode fields. `is_successful` is simply `score >= threshold`, and an error forces failure ([base_metric.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/base_metric.py#L111-L182)). Yes/no metrics share a `Verdict` enum with an optional `borderline` bucket, and they still accept the legacy `idk` answer ([base_metric.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/base_metric.py#L44-L69)). Prompts are not hard-coded strings. `PromptMixin._get_prompt` renders a named template from a bundle, and a user-supplied `evaluation_template` class overrides it ([base_metric.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/base_metric.py#L72-L108)).

### G-Eval

`GEval` takes a `name`, the test-case fields to show, and either `criteria` (from which it generates evaluation steps), explicit `evaluation_steps`, or a `rubric` of score ranges ([g_eval.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L53-L98)). The judge returns a `ReasonScore`, and the raw score is normalised into the rubric range ([g_eval.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L166-L211)). When the model exposes logprobs, `calculate_weighted_summed_score` replaces the sampled integer with a probability-weighted average over the top score tokens and drops tokens under 1%. This is the G-Eval paper's trick for smoother scores ([utils.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/utils.py#L337-L375)).

### Eval modes and System One

Every built-in judge metric can run in one of three modes, set by `eval_mode` or `DEEPEVAL_EVAL_MODE`. `llm` (the default) lets the judge LLM do everything. `hybrid` lets the LLM extract and explain while a hosted "System One" model called Jev answers the decision points, falling back to the LLM only when a Jev call fails at runtime. `system_one` runs the whole metric as one Jev request with a deterministic reason and no fallback ([eval_mode.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/config/eval_mode.py#L1-L21), [system_one.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/utils/system_one.py#L1-L26)). Jev is a third-party API with a 64k-token request budget ([constants.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/models/system_one/constants.py#L1-L27)). It is opt-in, and nothing turns it on implicitly.

### Judge models

`initialize_model` accepts a model name string or any `DeepEvalBaseLLM`. With no model it picks a native provider from settings (OpenAI first, then Gemini, LiteLLM, Portkey and others), and under `system_one` mode it builds no LLM at all ([models.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/utils/models.py#L170-L210)). Native wrappers report cost and tokens. A custom subclass works too, but cost tracking stops.

### Tracing

`@observe(type="llm"|"retriever"|"tool"|"agent", metrics=...)` turns functions into spans, and `metrics` can be attached per span for component-level evals ([tracing.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/tracing.py#L1408-L1428)). For OpenTelemetry users, `ContextAwareSpanProcessor` sends OTel spans either into the REST trace manager (inside a deepeval trace or evaluation) or straight to Confident AI's OTLP endpoint. `ConfidentSpanExporter` translates OTel spans into deepeval traces ([context_aware_processor.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/context_aware_processor.py#L82-L136), [exporter.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/exporter.py#L100-L134)). Integrations exist for LangChain, LlamaIndex, CrewAI, PydanticAI, Google ADK, Strands, AgentCore and OpenAI Agents.

## Extending it

- **Custom metric.** Subclass `BaseMetric`, set `_required_params`, and implement `measure`/`a_measure` to set `score`, `reason` and `success`. The base `is_successful` compares the score with the threshold.
- **Custom rubric without code.** Use `GEval` with `criteria` or `evaluation_steps`, `DAGMetric` for decision trees, or `ArenaGEval` with `compare()` for pairwise A/B.
- **Custom judge.** Subclass `DeepEvalBaseLLM` (`load_model`, `generate`, `a_generate`, `get_model_name`) to use any provider or a local model.
- **Prompt overrides.** Pass an `evaluation_template` class to a metric to replace individual judge prompts.
- **Agent evals.** Decorate the app with `@observe` and loop over `dataset.evals_iterator(metrics=...)`, which runs agentic test cases built from the traces it collects ([dataset.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/dataset.py#L1668-L1700)).

## Running it

- `pip install deepeval`, set a judge provider key (or configure one with the CLI), and write pytest files that call `assert_test`.
- `deepeval test run <path> [-n 4] [--repeat N] [--use-cache]` runs them. `--use-cache` (`-c`) reuses cached metric results, and caching is disabled automatically with `--repeat`.
- In notebooks or scripts, call `evaluate(test_cases, metrics, async_config=AsyncConfig(max_concurrent=20))` ([configs.py](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/configs.py#L7-L47)). `DisplayConfig.results_folder` exports each full run as JSON.
- `CONFIDENT_API_KEY` (or `deepeval login`) enables uploads, dataset push/pull and hosted metric collections. Anonymous usage telemetry is on by default and is disabled with `DEEPEVAL_TELEMETRY_OPT_OUT=1`.

## Strengths and caveats

- **Strength: CI-native.** The pytest plugin, `assert_test`, flaky-case handling, xdist and real exit codes make it the most natural fit of the evals tools for a test suite.
- **Strength: breadth.** RAG, agent (tool correctness, plan quality, step efficiency), conversation, MCP, multimodal and voice metrics, plus synthesis, simulation and benchmarks, all come from one package.
- **Strength: careful judge plumbing.** Schema-validated verdicts, logprob-weighted G-Eval, per-task timeouts, and cost and token accounting on native models.
- **Caveat: scores depend on the judge.** Nearly every metric that is not exact match is a chain of LLM calls. Results vary with the judge model and prompt version, and a metric costs several calls per test case.
- **Caveat: platform gravity.** Run history, dataset versioning, dashboards, annotation and online evaluation of production traces are Confident AI features. Locally you get JSON files and a console table.
- **Caveat: surface area.** Two eval modes beyond `llm`, a third-party System One model, classifiers, voice simulation and a large settings surface make the package heavy, and the repository root even carries ad-hoc scratch scripts (`a.py`, `b.py`, `c.py`).
- **Caveat: no attack generation.** Safety coverage is limited to judge metrics (toxicity, bias, PII leakage, misuse) and classifiers such as prompt-injection detection. Adversarial probing lives in DeepTeam ([README.md](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/red_teaming/README.md#L1-L3)).

*Sources: code at ec9b983, verified Q&A.*

## How confident-ai/deepeval answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

DeepEval provides ~50+ built-in metrics across four categories. **Heuristic/statistical** metrics need no LLM: `ExactMatchMetric` does strict string comparison (`deepeval/metrics/exact_match/exact_match.py:12-88`), `PatternMatchMetric` applies regex, and the `Scorer` class offers ROUGE, BLEU, BERTScore, and exact/quasi-exact match utilities (`deepeval/scorer/scorer.py:11-160`). **LLM-based metrics** use a judge LLM: `FaithfulnessMetric` extracts claims from `actual_output` and checks each against `retrieval_context` via a yes/no verdict loop (`deepeval/metrics/faithfulness/faithfulness.py:66-100`), `HallucinationMetric` and `AnswerRelevancyMetric` follow the same QAG (question-answer generation) pattern — calling `generate_qag_verdicts` to classify statements, then `score_qag_verdicts` to aggregate. `ToxicityMetric` and `BiasMetric` extract opinions/statements and judge each for toxicity/bias. **Configurable LLM judges**: `GEval` accepts `criteria`, `evaluation_steps`, and `rubric` (score-range → outcome mappings), renders a structured prompt, and returns a Pydantic-validated `ReasonScore` (`deepeval/metrics/g_eval/g_eval.py:53-120`, `deepeval/metrics/g_eval/schema.py:5-13`). `ArenaGEval` compares two contestants side-by-side. `DAGMetric` composes multiple judgement and task nodes into an acyclic graph for multi-step evaluation (`deepeval/metrics/dag/dag.py:25-70`). **Agentic/conversational metrics**: `ToolCorrectnessMetric`, `GoalAccuracyMetric`, `AgentLoopDetectionMetric`, turn-level variants (`TurnFaithfulnessMetric`), and `ConversationCompletenessMetric`. Each metric is a subclass of `BaseMetric` (single-turn `LLMTestCase`), `BaseConversationalMetric` (multi-turn `ConversationalTestCase`), or `BaseArenaMetric` (pairwise `ArenaTestCase`) (`deepeval/metrics/base_metric.py:111-350`). Metrics define `_required_params` (e.g. `[INPUT, ACTUAL_OUTPUT, RETRIEVAL_CONTEXT]`), implement `measure()` and `a_measure()`, and set `self.score`, `self.reason`, `self.success` (score ≥ threshold). Per-test-case scoring: `execute_test_cases` iterates test cases and runs each metric against each case individually (`deepeval/evaluate/execute/e2e.py:144-270`). Custom metrics are created by subclassing a base class and providing the `measure` method.


Citations: [deepeval/metrics/__init__.py:1-166](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/__init__.py#L1-L166) · [deepeval/scorer/scorer.py:11-160](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/scorer/scorer.py#L11-L160) · [deepeval/metrics/faithfulness/faithfulness.py:66-100](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/faithfulness/faithfulness.py#L66-L100) · [deepeval/metrics/g_eval/g_eval.py:53-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L53-L80)

### How is LLM-as-a-judge implemented? (answered)

The LLM-as-a-judge system centers on `GEval` (`deepeval/metrics/g_eval/g_eval.py:53-120`). The judge is prompted with a `name`, optional `criteria` and `evaluation_steps` (a numbered list of checks), and an optional `rubric` — a list of `Rubric(score_range, expected_outcome)` that maps score intervals to natural-language descriptions (`deepeval/metrics/g_eval/utils.py:35-50`). The rendered prompt inserts the test case fields (input, actual output, context, etc.) as template variables and asks the LLM to output a structured `ReasonScore` (reason + score) via Pydantic schema enforcement (`deepeval/metrics/g_eval/schema.py:5-8`). Multi-turn conversation evaluation uses `ConversationalGEval`, which reads the full turn list. **Structured output** is enforced through `generate_with_schema_and_extract`, which passes a Pydantic model class as the response format to the underlying LLM API call (`deepeval/metrics/utils/__init__.py:19-26`). For **multi-sample / consensus** evaluation, `top_logprobs` can be set to request log-probability data from the judge model. **Judge model choice** is fully configurable via the `model` parameter (accepting a string model name or a `DeepEvalBaseLLM` instance wrapping any provider — OpenAI, Anthropic, Azure, local). The `eval_mode` system introduces **System One (Jev)** — a small-model pre-screener that runs first; only ambiguous or low-confidence cases are escalated to the full LLM judge, controlled via `system_one_model` and `eval_mode` (`hybrid`, `llm`, `system_one`). **Calibration controls**: `threshold` sets pass/fail, `strict_mode` overrides threshold to 1.0, and the rubric's score range can be customised. `JevEval` runs the judge prompt against a small / deterministic model for speed, and `ArenaGEval` compares two LLM outputs pairwise (`deepeval/metrics/__init__.py:21-23`). All judge metrics inherit template rendering from `PromptMixin._get_prompt`, which loads Jinja templates from the `templates/metrics/` bundle (`deepeval/metrics/base_metric.py:72-108`).

> **Editor's note.** Correction: System One is not a pre-screener with confidence-based escalation. In `hybrid` mode the LLM extracts and explains while Jev answers decision points, falling back to the LLM only on a runtime Jev error; `system_one` mode has no fallback (deepeval/config/eval_mode.py L1-L21). `top_logprobs` drives logprob-weighted G-Eval scoring, not multi-sample consensus.

Citations: [deepeval/metrics/g_eval/g_eval.py:53-120](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L53-L120) · [deepeval/metrics/g_eval/utils.py:35-50](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/utils.py#L35-L50) · [deepeval/metrics/g_eval/schema.py:1-17](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/schema.py#L1-L17) · [deepeval/metrics/base_metric.py:72-108](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/base_metric.py#L72-L108) · [deepeval/metrics/utils/__init__.py:1-26](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/utils/__init__.py#L1-L26) · [deepeval/metrics/__init__.py:21-23](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/__init__.py#L21-L23)

### How are test datasets and cases defined, generated and versioned? (answered)

**`EvaluationDataset`** (`deepeval/dataset/dataset.py:90-250`) is the core data container, holding either `Golden` objects (single-turn) or `ConversationalGolden` objects (multi-turn). Each `Golden` wraps `input`, `actual_output`, `expected_output`, `context`, `retrieval_context`, `tools_called`, `expected_tools`, `additional_metadata`, and optional `persona`/`scenario`/`expected_outcome` fields (`deepeval/dataset/golden.py:1-180`). **File formats**: `add_test_cases_from_csv_file()` parses CSV with configurable column names and list delimiters (`deepeval/dataset/dataset.py:266-340`). JSON/JSONL loading is supported through `add_test_cases_from_json_file()`. Column mappings support `input`, `actual_output`, `expected_output`, `context`, `retrieval_context` and tool call fields. **Synthetic generation**: The `Synthesizer` class (`deepeval/synthesizer/synthesizer.py:1-80`) produces `SyntheticData` / `ConversationalScenario` from source documents via evolution techniques — `Reasoning`, `Multi-context`, `Concretizing`, `Constrained`, `Comparative`, `Hypothetical`, `In-Breadth` (`deepeval/synthesizer/schema.py:14-20`). These use LLM-driven evolution templates to re-write seed inputs, and `FiltrationConfig` / `EvolutionConfig` / `StylingConfig` control quality and diversity. Conversational goldens support `Persona` (demographics, speaking style, interruption behavior) and `BackgroundNoiseSettings` for voice simulations (`deepeval/dataset/golden.py:78-150`). **Versioning**: Datasets have `_alias`, `_id`, and `_version` fields, and the `DatasetVersion` / `CreateDatasetVersionHttpResponse` API types enable versioned snapshots on Confident AI (`deepeval/dataset/api.py:49-61`). **Benchmark task registries**: The `benchmarks/` package contains ~15 standard benchmarks (MMLU, HellaSwag, BIG-Bench-Hard, GSM8K, DROP, SQuAD, HumanEval, etc.) each as a `DeepEvalBaseBenchmark` subclass that loads HuggingFace datasets and evaluates via `load_benchmark_dataset()` → `evaluate()` (`deepeval/benchmarks/base_benchmark.py:16-32`).


Citations: [deepeval/dataset/dataset.py:90-250](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/dataset.py#L90-L250) · [deepeval/dataset/golden.py:1-180](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/golden.py#L1-L180) · [deepeval/synthesizer/synthesizer.py:1-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/synthesizer/synthesizer.py#L1-L80) · [deepeval/synthesizer/schema.py:14-20](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/synthesizer/schema.py#L14-L20) · [deepeval/dataset/api.py:49-61](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/api.py#L49-L61) · [deepeval/benchmarks/base_benchmark.py:16-32](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/benchmarks/base_benchmark.py#L16-L32)

### How are evals executed and reported? (answered)

The primary entry point is `evaluate()` in `deepeval/evaluate/evaluate.py:1-80`, which accepts an `EvaluationDataset` (or list of test cases) and a list of metrics/classifiers. For individual test assertions, `assert_test()` wraps a single test case and runs metrics synchronously or asynchronously (`deepeval/evaluate/evaluate.py:84-158`). Under the hood, `execute_test_cases()` / `a_execute_test_cases()` iterate test cases, evaluate each metric, and collect `TestResult` objects (`deepeval/evaluate/execute/e2e.py:144-270`). **Parallelism**: `async_mode` uses `asyncio` with configurable `max_concurrent` throttling via `AsyncConfig`. **Caching**: `TestRunCacheManager` caches per-test-case metric results based on hyperparameters; `use_cache` / `write_cache` flags in `CacheConfig` control disk persistence (`deepeval/evaluate/execute/e2e.py:260-270`). **Result storage**: Runs are exported as local JSON files at `{HIDDEN_DIR}/.temp_test_run_data.json` and `.latest_test_run.json` (`deepeval/test_run/test_run.py:64-68`). When integrated with Confident AI, results are uploaded via API (`APIEvaluate` model) for cloud persistence. **Console reporting**: `EvaluationConsoleReport` uses Rich to display per-metric scores, pass/fail, and aggregate summaries (`deepeval/evaluate/console_report.py`). **Comparison**: `compare()` (`deepeval/evaluate/compare.py:48-80`) runs `ArenaGEval` pairwise across `ArenaTestCase` contestants and produces a win/loss dictionary. **CI integration**: `assert_test()` raises `AssertionError` on failure (metric.success=False), making it pytest-compatible. **Agentic trace execution**: `execute_agentic_test_cases_from_loop()` evaluates traces collected during agent runs, matching traces to golden expectations. **Hyperparameters** are captured from model config and prompt hash/alias, linked to test run for full reproducibility (`deepeval/test_run/hyperparameters.py:10-58`).


Citations: [deepeval/evaluate/evaluate.py:1-158](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/evaluate.py#L1-L158) · [deepeval/evaluate/execute/e2e.py:144-270](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/execute/e2e.py#L144-L270) · [deepeval/evaluate/compare.py:48-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/evaluate/compare.py#L48-L80) · [deepeval/test_run/test_run.py:64-75](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/test_run/test_run.py#L64-L75) · [deepeval/test_run/hyperparameters.py:10-58](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/test_run/hyperparameters.py#L10-L58)

### How are traces or production data captured and linked to evaluations? (answered)

DeepEval has a comprehensive tracing subsystem that captures LLM calls, agent steps, retrievals, and tool invocations as spans within traces. The `TraceManager` (`deepeval/tracing/tracing.py:1-80`) manages a tree of spans of types `BaseSpan`, `LlmSpan`, `AgentSpan`, `RetrieverSpan`, and `ToolSpan` (`deepeval/tracing/types.py`). **SDK instrumentation** is automatic: `patch_openai_client()` and `patch_anthropic_client()` monkey-patch the SDK client methods (`chat.completions.create`, etc.) to wrap each call in a span (`deepeval/tracing/patchers.py:16-60`). The `@observe` decorator or `with trace(...)` context manager mark arbitrary Python functions as traceable spans. **OpenTelemetry integration** is two-way: `ConfidentSpanExporter` implements `SpanExporter` to push deepeval spans into the OTel ecosystem, and `ContextAwareSpanProcessor` routes incoming OTel spans either through the REST-based trace manager path (when inside a deepeval trace context) or directly to Confident AI's OTLP endpoint (`deepeval/tracing/otel/exporter.py:1-60`, `deepeval/tracing/otel/context_aware_processor.py:1-60`). Supported frameworks (LangChain, CrewAI, LlamaIndex, OpenAI Agents, PydanticAI, Google ADK, Mastra) get automatic OTel integration (`deepeval/tracing/integrations.py:6-22`). **Online evals**: traces collected during a `trace()` session are automatically linked to evaluation when `assert_test()` is called from within the context. **Offline evals**: `evaluate_trace()`, `evaluate_span()`, and `evaluate_thread()` in `deepeval/tracing/offline_evals/` can retrospectively evaluate a stored trace against a metric collection (`deepeval/tracing/offline_evals/trace.py:6-36`). Trace data (input, output, timestamps, metadata) is serialized as `BaseApiSpan` Pydantic models and uploaded to Confident AI for dashboards. Feedback and labeling happen through the Confident AI platform rather than the open-source library itself.


Citations: [deepeval/tracing/tracing.py:1-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/tracing.py#L1-L80) · [deepeval/tracing/patchers.py:16-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/patchers.py#L16-L60) · [deepeval/tracing/otel/exporter.py:1-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/exporter.py#L1-L60) · [deepeval/tracing/otel/context_aware_processor.py:1-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/context_aware_processor.py#L1-L60) · [deepeval/tracing/integrations.py:1-22](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/integrations.py#L1-L22) · [deepeval/tracing/offline_evals/trace.py:6-36](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/offline_evals/trace.py#L6-L36)

### Does it support red-teaming or safety testing, and how? (not applicable)

Red-teaming was previously part of this repository but has been extracted into a separate project. A README at `deepeval/red_teaming/README.md` states: "The Red Teaming module is now in DeepTeam for deepeval-v3.0 onwards" and directs users to https://github.com/confident-ai/deepteam. The current codebase has no adversarial probes, jailbreak tests, prompt-injection detectors, or attack plugins. Instead, safety-related evaluation is handled through metrics like `ToxicityMetric`, `BiasMetric`, `MisuseMetric`, `RoleViolationMetric`, `NonAdviceMetric`, and `PIILeakageMetric`, which apply LLM judges to classify outputs as toxic, biased, non-compliant, etc., but these are evaluation metrics, not adversarial generation or vulnerability disclosure tooling.


