LLMs Technical Reviews

confident-ai/deepeval

Pytest-style LLM evaluation library with ~50 judge-based metrics, G-Eval rubrics, tracing, dataset synthesis and Confident AI upload.

GitHub ↗★ 19kPythonApache-2.0commit ec9b983 · 2026-10-05homepage ↗

Overview

DeepEval is a Python library for unit-testing LLM applications. You build LLMTestCase objects (input, actual output, expected output, retrieval context, tool calls and so on), attach metrics with thresholds, and either call assert_test() inside a pytest test or evaluate() over a list. Each metric is usually an LLM-as-a-judge pipeline: extract claims or statements, ask the judge for a verdict per item, compute a ratio, and ask for a reason. A metric passes when score >= threshold.

The package is much bigger than its core loop. At the pinned commit (v4.2.8) it ships about 50 metric packages (RAG, agentic, conversational, multimodal, voice, safety), a G-Eval implementation with logprob-weighted scores, DAG metrics, a set of label classifiers, a Synthesizer for generating goldens, a conversation simulator, a prompt optimizer, standard benchmarks (MMLU, GSM8K, HumanEval and others), and a tracing SDK with @observe and OpenTelemetry bridges.

Most of the persistence and visualisation is designed around Confident AI, the vendor’s hosted platform. The open-source library runs fully locally and writes JSON under a hidden .deepeval folder. Dataset push and pull, metric collections, trace dashboards and annotation live on the platform. Red-teaming was split out into the separate DeepTeam project.

Architecture

flowchart LR
  T["pytest test / script"] --> AT["assert_test / evaluate"]
  CLI["deepeval test run"] --> PY["pytest + deepeval plugin"]
  PY --> T
  AT --> EX["a_execute_test_cases"]
  EX --> CACHE["Test-run cache"]
  EX --> M["Metrics (BaseMetric)"]
  M --> S1["System One (Jev)"]
  M --> J["Judge LLM (DeepEvalBaseLLM)"]
  EX --> TRM["TestRunManager"]
  TRM --> LOCAL["Local JSON + console report"]
  TRM --> CAI["Confident AI upload"]
  OBS["@observe / OTel"] --> TM["TraceManager"]
  TM --> CAI
Component Path Role
Entry points deepeval/evaluate/evaluate.py assert_test (pytest-friendly) and evaluate (batch)
Executors deepeval/evaluate/execute/ Sync/async loops over test cases, agentic trace loops, per-task timeouts
Configs deepeval/evaluate/configs.py AsyncConfig, DisplayConfig, CacheConfig, ErrorConfig
Metrics deepeval/metrics/ ~50 metric packages on BaseMetric / BaseConversationalMetric / BaseArenaMetric
Templates deepeval/templates/ Judge prompt bundles resolved by PromptMixin
Models deepeval/models/ DeepEvalBaseLLM plus OpenAI, Anthropic, Azure, Gemini, Bedrock, Ollama, LiteLLM and other wrappers; embeddings, TTS/STT
System One deepeval/models/system_one/, deepeval/metrics/utils/system_one.py Optional hosted “Jev” decision model and eval modes
Test cases deepeval/test_case/ LLMTestCase, ConversationalTestCase, ArenaTestCase, MCP types
Test runs deepeval/test_run/ Aggregation, cache, local files, upload
Datasets deepeval/dataset/, deepeval/synthesizer/ Goldens, CSV/JSON loaders, evals_iterator, synthetic generation
Tracing deepeval/tracing/ TraceManager, @observe, OpenAI/Anthropic patchers, OTel exporter/processor
CLI and plugin deepeval/cli/, deepeval/plugins/plugin.py deepeval test run wrapper around pytest

How a request flows

Take deepeval test run test_rag.py, where a test calls assert_test(test_case, [FaithfulnessMetric(threshold=0.7)]):

  1. CLI. run sets process flags (running-deepeval, cache, ignore-errors, verbose), resets the global test run, and calls pytest.main with -p deepeval plus -n for xdist or --count for repeats. When pytest finishes it calls wrap_up_test_run and exits with pytest’s code, so CI sees failures (command.py).
  2. Plugin. pytest_sessionstart creates the test run, and pytest_runtest_call wraps each test in an Observer span so that @observe spans attach to the run (plugin.py).
  3. Assert. assert_test builds configs (async, up to 100 concurrent) and runs a_execute_test_cases on the single case (evaluate.py).
  4. Execute. The executor splits metrics into single-turn and conversational, guards each task with an asyncio.Semaphore and an outer deadline, and copies metrics per test case (e2e.py). For each case it looks up a cached result keyed on the test case and hyperparameters, runs any System One batch, then measures the remaining metrics and classifiers (e2e.py).
  5. Measure. FaithfulnessMetric.a_measure first offers the case to System One. Otherwise it extracts truths from retrieval_context and claims from actual_output in parallel, generates a verdict per claim, computes the score, writes a reason and sets success (faithfulness.py). Each judge call goes through generate_with_schema_and_extract. Native models return a Pydantic instance plus cost, and custom models may return raw text that trimAndLoadJson repairs, for example by stripping trailing commas (generation.py).
  6. Record and cache. In the finally block each metric becomes MetricData, is added to the run, and is written to the cache with cost zeroed (e2e.py).
  7. Assert or warn. If the result failed, assert_test raises an AssertionError that lists every failing metric’s score, threshold and reason. Test cases marked flaky only emit a warning (evaluate.py).
  8. Wrap up. The run is saved to .deepeval/.latest_test_run.json and related files (test_run.py), printed as a Rich table, and uploaded if a Confident AI key is set.

Key components

Metric base and verdicts

BaseMetric carries the threshold, score, reason, cost, token counts and the eval-mode fields. is_successful is simply score >= threshold, and an error forces failure (base_metric.py). Yes/no metrics share a Verdict enum with an optional borderline bucket, and they still accept the legacy idk answer (base_metric.py). Prompts are not hard-coded strings. PromptMixin._get_prompt renders a named template from a bundle, and a user-supplied evaluation_template class overrides it (base_metric.py).

G-Eval

GEval takes a name, the test-case fields to show, and either criteria (from which it generates evaluation steps), explicit evaluation_steps, or a rubric of score ranges (g_eval.py). The judge returns a ReasonScore, and the raw score is normalised into the rubric range (g_eval.py). When the model exposes logprobs, calculate_weighted_summed_score replaces the sampled integer with a probability-weighted average over the top score tokens and drops tokens under 1%. This is the G-Eval paper’s trick for smoother scores (utils.py).

Eval modes and System One

Every built-in judge metric can run in one of three modes, set by eval_mode or DEEPEVAL_EVAL_MODE. llm (the default) lets the judge LLM do everything. hybrid lets the LLM extract and explain while a hosted “System One” model called Jev answers the decision points, falling back to the LLM only when a Jev call fails at runtime. system_one runs the whole metric as one Jev request with a deterministic reason and no fallback (eval_mode.py, system_one.py). Jev is a third-party API with a 64k-token request budget (constants.py). It is opt-in, and nothing turns it on implicitly.

Judge models

initialize_model accepts a model name string or any DeepEvalBaseLLM. With no model it picks a native provider from settings (OpenAI first, then Gemini, LiteLLM, Portkey and others), and under system_one mode it builds no LLM at all (models.py). Native wrappers report cost and tokens. A custom subclass works too, but cost tracking stops.

Tracing

@observe(type="llm"|"retriever"|"tool"|"agent", metrics=...) turns functions into spans, and metrics can be attached per span for component-level evals (tracing.py). For OpenTelemetry users, ContextAwareSpanProcessor sends OTel spans either into the REST trace manager (inside a deepeval trace or evaluation) or straight to Confident AI’s OTLP endpoint. ConfidentSpanExporter translates OTel spans into deepeval traces (context_aware_processor.py, exporter.py). Integrations exist for LangChain, LlamaIndex, CrewAI, PydanticAI, Google ADK, Strands, AgentCore and OpenAI Agents.

Extending it

  • Custom metric. Subclass BaseMetric, set _required_params, and implement measure/a_measure to set score, reason and success. The base is_successful compares the score with the threshold.
  • Custom rubric without code. Use GEval with criteria or evaluation_steps, DAGMetric for decision trees, or ArenaGEval with compare() for pairwise A/B.
  • Custom judge. Subclass DeepEvalBaseLLM (load_model, generate, a_generate, get_model_name) to use any provider or a local model.
  • Prompt overrides. Pass an evaluation_template class to a metric to replace individual judge prompts.
  • Agent evals. Decorate the app with @observe and loop over dataset.evals_iterator(metrics=...), which runs agentic test cases built from the traces it collects (dataset.py).

Running it

  • pip install deepeval, set a judge provider key (or configure one with the CLI), and write pytest files that call assert_test.
  • deepeval test run <path> [-n 4] [--repeat N] [--use-cache] runs them. --use-cache (-c) reuses cached metric results, and caching is disabled automatically with --repeat.
  • In notebooks or scripts, call evaluate(test_cases, metrics, async_config=AsyncConfig(max_concurrent=20)) (configs.py). DisplayConfig.results_folder exports each full run as JSON.
  • CONFIDENT_API_KEY (or deepeval login) enables uploads, dataset push/pull and hosted metric collections. Anonymous usage telemetry is on by default and is disabled with DEEPEVAL_TELEMETRY_OPT_OUT=1.

Strengths and caveats

  • Strength: CI-native. The pytest plugin, assert_test, flaky-case handling, xdist and real exit codes make it the most natural fit of the evals tools for a test suite.
  • Strength: breadth. RAG, agent (tool correctness, plan quality, step efficiency), conversation, MCP, multimodal and voice metrics, plus synthesis, simulation and benchmarks, all come from one package.
  • Strength: careful judge plumbing. Schema-validated verdicts, logprob-weighted G-Eval, per-task timeouts, and cost and token accounting on native models.
  • Caveat: scores depend on the judge. Nearly every metric that is not exact match is a chain of LLM calls. Results vary with the judge model and prompt version, and a metric costs several calls per test case.
  • Caveat: platform gravity. Run history, dataset versioning, dashboards, annotation and online evaluation of production traces are Confident AI features. Locally you get JSON files and a console table.
  • Caveat: surface area. Two eval modes beyond llm, a third-party System One model, classifiers, voice simulation and a large settings surface make the package heavy, and the repository root even carries ad-hoc scratch scripts (a.py, b.py, c.py).
  • Caveat: no attack generation. Safety coverage is limited to judge metrics (toxicity, bias, PII leakage, misuse) and classifiers such as prompt-injection detection. Adversarial probing lives in DeepTeam (README.md).

Sources: code at ec9b983, verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit ec9b983. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

DeepEval provides ~50+ built-in metrics across four categories. Heuristic/statistical metrics need no LLM: ExactMatchMetric does strict string comparison (deepeval/metrics/exact_match/exact_match.py:12-88), PatternMatchMetric applies regex, and the Scorer class offers ROUGE, BLEU, BERTScore, and exact/quasi-exact match utilities (deepeval/scorer/scorer.py:11-160). LLM-based metrics use a judge LLM: FaithfulnessMetric extracts claims from actual_output and checks each against retrieval_context via a yes/no verdict loop (deepeval/metrics/faithfulness/faithfulness.py:66-100), HallucinationMetric and AnswerRelevancyMetric follow the same QAG (question-answer generation) pattern — calling generate_qag_verdicts to classify statements, then score_qag_verdicts to aggregate. ToxicityMetric and BiasMetric extract opinions/statements and judge each for toxicity/bias. Configurable LLM judges: GEval accepts criteria, evaluation_steps, and rubric (score-range → outcome mappings), renders a structured prompt, and returns a Pydantic-validated ReasonScore (deepeval/metrics/g_eval/g_eval.py:53-120, deepeval/metrics/g_eval/schema.py:5-13). ArenaGEval compares two contestants side-by-side. DAGMetric composes multiple judgement and task nodes into an acyclic graph for multi-step evaluation (deepeval/metrics/dag/dag.py:25-70). Agentic/conversational metrics: ToolCorrectnessMetric, GoalAccuracyMetric, AgentLoopDetectionMetric, turn-level variants (TurnFaithfulnessMetric), and ConversationCompletenessMetric. Each metric is a subclass of BaseMetric (single-turn LLMTestCase), BaseConversationalMetric (multi-turn ConversationalTestCase), or BaseArenaMetric (pairwise ArenaTestCase) (deepeval/metrics/base_metric.py:111-350). Metrics define _required_params (e.g. [INPUT, ACTUAL_OUTPUT, RETRIEVAL_CONTEXT]), implement measure() and a_measure(), and set self.score, self.reason, self.success (score ≥ threshold). Per-test-case scoring: execute_test_cases iterates test cases and runs each metric against each case individually (deepeval/evaluate/execute/e2e.py:144-270). Custom metrics are created by subclassing a base class and providing the measure method.

How is LLM-as-a-judge implemented?

answered

The LLM-as-a-judge system centers on GEval (deepeval/metrics/g_eval/g_eval.py:53-120). The judge is prompted with a name, optional criteria and evaluation_steps (a numbered list of checks), and an optional rubric — a list of Rubric(score_range, expected_outcome) that maps score intervals to natural-language descriptions (deepeval/metrics/g_eval/utils.py:35-50). The rendered prompt inserts the test case fields (input, actual output, context, etc.) as template variables and asks the LLM to output a structured ReasonScore (reason + score) via Pydantic schema enforcement (deepeval/metrics/g_eval/schema.py:5-8). Multi-turn conversation evaluation uses ConversationalGEval, which reads the full turn list. Structured output is enforced through generate_with_schema_and_extract, which passes a Pydantic model class as the response format to the underlying LLM API call (deepeval/metrics/utils/__init__.py:19-26). For multi-sample / consensus evaluation, top_logprobs can be set to request log-probability data from the judge model. Judge model choice is fully configurable via the model parameter (accepting a string model name or a DeepEvalBaseLLM instance wrapping any provider — OpenAI, Anthropic, Azure, local). The eval_mode system introduces System One (Jev) — a small-model pre-screener that runs first; only ambiguous or low-confidence cases are escalated to the full LLM judge, controlled via system_one_model and eval_mode (hybrid, llm, system_one). Calibration controls: threshold sets pass/fail, strict_mode overrides threshold to 1.0, and the rubric's score range can be customised. JevEval runs the judge prompt against a small / deterministic model for speed, and ArenaGEval compares two LLM outputs pairwise (deepeval/metrics/__init__.py:21-23). All judge metrics inherit template rendering from PromptMixin._get_prompt, which loads Jinja templates from the templates/metrics/ bundle (deepeval/metrics/base_metric.py:72-108).

Editor's note. Correction: System One is not a pre-screener with confidence-based escalation. In hybrid mode the LLM extracts and explains while Jev answers decision points, falling back to the LLM only on a runtime Jev error; system_one mode has no fallback (deepeval/config/eval_mode.py L1-L21). top_logprobs drives logprob-weighted G-Eval scoring, not multi-sample consensus.

How are test datasets and cases defined, generated and versioned?

answered

EvaluationDataset (deepeval/dataset/dataset.py:90-250) is the core data container, holding either Golden objects (single-turn) or ConversationalGolden objects (multi-turn). Each Golden wraps input, actual_output, expected_output, context, retrieval_context, tools_called, expected_tools, additional_metadata, and optional persona/scenario/expected_outcome fields (deepeval/dataset/golden.py:1-180). File formats: add_test_cases_from_csv_file() parses CSV with configurable column names and list delimiters (deepeval/dataset/dataset.py:266-340). JSON/JSONL loading is supported through add_test_cases_from_json_file(). Column mappings support input, actual_output, expected_output, context, retrieval_context and tool call fields. Synthetic generation: The Synthesizer class (deepeval/synthesizer/synthesizer.py:1-80) produces SyntheticData / ConversationalScenario from source documents via evolution techniques — Reasoning, Multi-context, Concretizing, Constrained, Comparative, Hypothetical, In-Breadth (deepeval/synthesizer/schema.py:14-20). These use LLM-driven evolution templates to re-write seed inputs, and FiltrationConfig / EvolutionConfig / StylingConfig control quality and diversity. Conversational goldens support Persona (demographics, speaking style, interruption behavior) and BackgroundNoiseSettings for voice simulations (deepeval/dataset/golden.py:78-150). Versioning: Datasets have _alias, _id, and _version fields, and the DatasetVersion / CreateDatasetVersionHttpResponse API types enable versioned snapshots on Confident AI (deepeval/dataset/api.py:49-61). Benchmark task registries: The benchmarks/ package contains ~15 standard benchmarks (MMLU, HellaSwag, BIG-Bench-Hard, GSM8K, DROP, SQuAD, HumanEval, etc.) each as a DeepEvalBaseBenchmark subclass that loads HuggingFace datasets and evaluates via load_benchmark_dataset() → evaluate() (deepeval/benchmarks/base_benchmark.py:16-32).

How are evals executed and reported?

answered

The primary entry point is evaluate() in deepeval/evaluate/evaluate.py:1-80, which accepts an EvaluationDataset (or list of test cases) and a list of metrics/classifiers. For individual test assertions, assert_test() wraps a single test case and runs metrics synchronously or asynchronously (deepeval/evaluate/evaluate.py:84-158). Under the hood, execute_test_cases() / a_execute_test_cases() iterate test cases, evaluate each metric, and collect TestResult objects (deepeval/evaluate/execute/e2e.py:144-270). Parallelism: async_mode uses asyncio with configurable max_concurrent throttling via AsyncConfig. Caching: TestRunCacheManager caches per-test-case metric results based on hyperparameters; use_cache / write_cache flags in CacheConfig control disk persistence (deepeval/evaluate/execute/e2e.py:260-270). Result storage: Runs are exported as local JSON files at {HIDDEN_DIR}/.temp_test_run_data.json and .latest_test_run.json (deepeval/test_run/test_run.py:64-68). When integrated with Confident AI, results are uploaded via API (APIEvaluate model) for cloud persistence. Console reporting: EvaluationConsoleReport uses Rich to display per-metric scores, pass/fail, and aggregate summaries (deepeval/evaluate/console_report.py). Comparison: compare() (deepeval/evaluate/compare.py:48-80) runs ArenaGEval pairwise across ArenaTestCase contestants and produces a win/loss dictionary. CI integration: assert_test() raises AssertionError on failure (metric.success=False), making it pytest-compatible. Agentic trace execution: execute_agentic_test_cases_from_loop() evaluates traces collected during agent runs, matching traces to golden expectations. Hyperparameters are captured from model config and prompt hash/alias, linked to test run for full reproducibility (deepeval/test_run/hyperparameters.py:10-58).

How are traces or production data captured and linked to evaluations?

answered

DeepEval has a comprehensive tracing subsystem that captures LLM calls, agent steps, retrievals, and tool invocations as spans within traces. The TraceManager (deepeval/tracing/tracing.py:1-80) manages a tree of spans of types BaseSpan, LlmSpan, AgentSpan, RetrieverSpan, and ToolSpan (deepeval/tracing/types.py). SDK instrumentation is automatic: patch_openai_client() and patch_anthropic_client() monkey-patch the SDK client methods (chat.completions.create, etc.) to wrap each call in a span (deepeval/tracing/patchers.py:16-60). The @observe decorator or with trace(...) context manager mark arbitrary Python functions as traceable spans. OpenTelemetry integration is two-way: ConfidentSpanExporter implements SpanExporter to push deepeval spans into the OTel ecosystem, and ContextAwareSpanProcessor routes incoming OTel spans either through the REST-based trace manager path (when inside a deepeval trace context) or directly to Confident AI's OTLP endpoint (deepeval/tracing/otel/exporter.py:1-60, deepeval/tracing/otel/context_aware_processor.py:1-60). Supported frameworks (LangChain, CrewAI, LlamaIndex, OpenAI Agents, PydanticAI, Google ADK, Mastra) get automatic OTel integration (deepeval/tracing/integrations.py:6-22). Online evals: traces collected during a trace() session are automatically linked to evaluation when assert_test() is called from within the context. Offline evals: evaluate_trace(), evaluate_span(), and evaluate_thread() in deepeval/tracing/offline_evals/ can retrospectively evaluate a stored trace against a metric collection (deepeval/tracing/offline_evals/trace.py:6-36). Trace data (input, output, timestamps, metadata) is serialized as BaseApiSpan Pydantic models and uploaded to Confident AI for dashboards. Feedback and labeling happen through the Confident AI platform rather than the open-source library itself.

Does it support red-teaming or safety testing, and how?

not applicable

Red-teaming was previously part of this repository but has been extracted into a separate project. A README at deepeval/red_teaming/README.md states: "The Red Teaming module is now in DeepTeam for deepeval-v3.0 onwards" and directs users to https://github.com/confident-ai/deepteam. The current codebase has no adversarial probes, jailbreak tests, prompt-injection detectors, or attack plugins. Instead, safety-related evaluation is handled through metrics like ToxicityMetric, BiasMetric, MisuseMetric, RoleViolationMetric, NonAdviceMetric, and PIILeakageMetric, which apply LLM judges to classify outputs as toxic, biased, non-compliant, etc., but these are evaluation metrics, not adversarial generation or vulnerability disclosure tooling.