# comet-ml/opik

> LLM tracing and eval platform: Python/TS SDKs with ~30 metrics, a Java backend on ClickHouse and MySQL, and Redis-streamed online scoring.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/comet-ml/opik (reviewed at commit `f217a863ea3cbfcef7f9ab8217c543049bdff095`, 2026-10-06)
- Stars: 22412 · Language: Python · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/opik/

## Overview

Opik is Comet's open-source platform for tracing and evaluating LLM applications. It has two halves that are useful on their own. The Python SDK (`sdks/python`) is a full offline evaluation library: datasets, `evaluate()`, around 30 heuristic and LLM-judge metrics, test suites and conversation simulation. The server (`apps/`) stores traces, spans, feedback scores, datasets and experiments, shows them in a React UI, and runs "online" evaluation rules against incoming traces.

The two halves meet at the trace. `@opik.track` turns your functions into spans and traces. `evaluate()` runs your task under that same decorator, so every experiment item is a normal trace with feedback scores attached. Online rules on the server score production traces with LLM judges or user-written Python metrics.

The whole repository is Apache-2.0: the backend, the frontend, the Python and TypeScript SDKs and the `opik_optimizer` package. There is no `ee/` folder or licence-gated code path. The open-core boundary is authentication instead. With `authentication.enabled` false (the default), every request maps to a single `default` workspace and user ([AuthService.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/infrastructure/auth/AuthService.java#L30-L55), [config.yml](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/config.yml#L431-L434)). Turning it on requires `authentication.reactService.url`, an external account service that is not in this repository ([AuthModule.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/infrastructure/auth/AuthModule.java#L38-L66)). In practice, multi-user workspaces, API keys and permissions come from Comet's hosted product.

## Architecture

```mermaid
flowchart LR
  APP["Your code + @opik.track"] --> Q["SDK message queue + batcher"]
  EVAL["opik.evaluate()"] --> APP
  EVAL --> MET["SDK metrics (heuristic, LLM judge)"]
  Q --> REST["Backend REST /traces/batch"]
  OTEL["OTel exporters"] --> OTLP["Backend /otel/v1"]
  REST --> SVC["TraceService"]
  OTLP --> SVC
  SVC --> CH["ClickHouse: traces, spans, scores"]
  SVC --> BUS["EventBus: TracesCreated"]
  BUS --> SAMP["OnlineScoringSampler"]
  SAMP --> RS["Redis streams"]
  RS --> JUDGE["LLM-judge scorer"]
  RS --> PYS["Python metric scorer"]
  PYS --> PYB["opik-python-backend"]
  JUDGE --> CH
  PYS --> CH
  UI["React frontend"] --> REST
```

| Component | Path | Role |
|---|---|---|
| Python SDK tracing | `sdks/python/src/opik/decorator/`, `message_processing/` | `@track`, context managers, background queue, batching, REST upload |
| Python SDK evaluation | `sdks/python/src/opik/evaluation/` | `evaluate`, `EvaluationEngine`, metrics, suite evaluators, resume |
| Integrations | `sdks/python/src/opik/integrations/` | OpenAI, Anthropic, LangChain, LlamaIndex, DSPy, ADK, Bedrock, an OTel span processor and more |
| TypeScript SDK | `sdks/typescript/` | Tracing and evaluation client for Node |
| Backend | `apps/opik-backend/` | Dropwizard (Java) REST API; ClickHouse for traces/spans/scores, MySQL for state |
| Online scoring | `apps/opik-backend/.../api/resources/v1/events/` | Samplers, Redis-stream publishers and scorers for LLM-judge and Python rules |
| Python evaluator service | `apps/opik-python-backend/` | Flask service that executes user Python metrics |
| Guardrails service | `apps/opik-guardrails-backend/` | Optional PII and topic guard models (GPU or CPU profile) |
| Optimizer | `sdks/opik_optimizer/` | Prompt optimizers (evolutionary, few-shot Bayesian, GEPA, meta-prompt and others) |

## How a request flows

A traced call in production, scored by an online LLM-judge rule:

1. **Capture.** `@track` wraps the function. When it returns, `_after_call` fills end time, output or error on the span and trace data. It then hands them to the global client's internal `span`/`trace` methods ([base_track_decorator.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/decorator/base_track_decorator.py#L331-L382), [L563-L588](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/decorator/base_track_decorator.py#L563-L588)).
2. **Queue and batch.** The client turns those into messages. `Streamer.put` extracts embedded attachments, lets the batcher absorb the message, and otherwise puts it on a bounded in-memory queue. If the queue is full, the oldest message is dropped with a warning ([streamer.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/message_processing/streamer.py#L52-L80)).
3. **Upload.** A consumer thread truncates oversized payloads and posts batches with `spans.create_spans` / `traces.create_traces` ([online_message_processor.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/message_processing/processors/online_message_processor.py#L430-L453)).
4. **Store.** `POST /v1/private/traces/batch` calls `TraceService.create` ([TracesResource.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/priv/TracesResource.java#L308-L325)). The service dedups, resolves or creates projects, strips large attachments to blob storage, batch-inserts into ClickHouse, and posts `TracesCreated` on the in-process event bus ([TraceService.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/domain/TraceService.java#L215-L243)).
5. **Sample.** `OnlineScoringSampler.onTracesCreated` keeps only traces with an `end_time` ([OnlineScoringSampler.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/OnlineScoringSampler.java#L151-L165)). It loads the project's rules and, per rule, checks enabled state, filters and a `SecureRandom` roll against the sampling rate ([L361-L384](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/OnlineScoringSampler.java#L361-L384)). Sampled traces become messages per evaluator type ([L243-L290](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/OnlineScoringSampler.java#L243-L290)), which `OnlineScorePublisher` adds to a Redis stream ([OnlineScorePublisher.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/domain/evaluators/OnlineScorePublisher.java#L126-L160)).
6. **Score.** The LLM-judge scorer reads the stream. It fetches spans only when the prompt needs them, then runs either a chat judge (with optional tool calls to inspect the trace) or a single decision-model call. It stores the result as feedback scores ([OnlineScoringLlmAsJudgeScorer.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/OnlineScoringLlmAsJudgeScorer.java#L183-L222)). When a toggle is on, the judge's own calls are recorded as a hidden evaluator trace.

## Key components

### SDK evaluation engine

`evaluate(dataset, task, scoring_metrics, ...)` creates an experiment and runs the task on each item with 16 threads by default. It supports `trial_count`, dataset filters and samplers, and an `error_tolerance` level ([evaluator.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/evaluator.py#L138-L160)). `EvaluationEngine.run_and_score` splits metrics into regular metrics and task-span metrics, which score the spans the task produced ([engine.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/engine/engine.py#L595-L636)). For each item, the engine wraps an untracked task in `opik.track`, opens a trace whose input is the item, runs the task and attaches the output ([L269-L330](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/engine/engine.py#L269-L330)). Scores are logged as feedback scores on that trace and linked to the experiment item. `evaluate_resume` can pick up an interrupted run.

### Metrics

Heuristic metrics need no LLM: `Equals`, `Contains`, `RegexMatch`, `IsJson`, `LevenshteinRatio`, BLEU, ROUGE, chrF, GLEU, METEOR, BERTScore, divergence and ranking metrics, sentiment, readability and a regex `PromptInjection` detector. LLM-judge metrics (`Hallucination`, `AnswerRelevance`, `ContextPrecision`/`Recall`, `Factuality`, `Moderation`, `Usefulness`, `TrajectoryAccuracy`, `SycEval`, `GEval` and its presets, `LLMJuriesJudge`) build a prompt, call the model with a Pydantic `response_format`, and parse a `ScoreResult` ([hallucination/metric.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/hallucination/metric.py#L80-L110)). Judge models resolve through LiteLLM, with a per-process cache of model instances ([models_factory.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/models/models_factory.py#L31-L58)). The default is `openai/gpt-5-nano` ([config.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/config.py#L124-L125)). `GEval` caches its generated chain-of-thought in a 128-entry class-level LRU ([g_eval/metric.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L66-L70)).

### Python metrics on the server

Online "user-defined metric" rules send the rule's Python code and mapped trace fields to the scorer ([OnlineScoringUserDefinedMetricPythonScorer.java](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-backend/src/main/java/com/comet/opik/api/resources/v1/events/OnlineScoringUserDefinedMetricPythonScorer.java#L90-L112)). The scorer calls the Python backend's `POST /python` ([evaluator.py](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-python-backend/src/opik_backend/evaluator.py#L54-L105)). That service runs the code with either a `process` executor (a local process pool) or a `docker` executor, chosen by `PYTHON_CODE_EXECUTOR_STRATEGY` ([L23-L40](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/apps/opik-python-backend/src/opik_backend/evaluator.py#L23-L40)). The bundled compose file defaults to `process` ([docker-compose.yaml](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/deployment/docker-compose/docker-compose.yaml#L298-L298)).

### Test suites and simulation

Datasets can carry per-item evaluator configs and execution policies (`runs_per_item`, `pass_threshold`). `run_tests()` turns that into pass/fail test suites judged by an `LLMJudge` that checks lists of assertions, merging identical judges into one call. `SimulatedUser` drives multi-turn conversations for thread-level metrics.

## Extending it

- **Metrics.** Subclass `BaseMetric` and implement `score`/`ascore`, or pass plain scorer functions that receive `dataset_item`, `task_outputs` and optionally `task_span`. `RagasMetricWrapper` adapts Ragas metrics.
- **Judge models.** Pass any LiteLLM model string or an `OpikBaseModel` subclass.
- **Tracing.** Use framework integrations, `opik.track` on anything, or `OpikSpanProcessor` for OTel-instrumented code. The backend also exposes an OTLP endpoint.
- **Online rules.** LLM-judge prompts with variable mappings, or Python metric code, scoped to traces, spans or threads with filters and a sampling rate.

## Running it

`./opik.sh` (or `opik.ps1`) starts the Docker Compose stack: MySQL, Redis, ClickHouse with ZooKeeper, MinIO, the Java backend, the Python backend, the frontend, and optional guardrails and observability profiles. A Helm chart is in `deployment/helm_chart`. The SDK is `pip install opik`. Run `opik configure` and point it at a local instance or Comet's cloud. LLM-judge metrics need an API key for the chosen provider. Heuristic metrics need no model key.

## Strengths and caveats

- **Strength: a real eval library.** The SDK ships the broadest built-in metric set of the trace-first platforms, including classic NLG metrics. Metrics are plain Python classes you can call without the server. `evaluate()` itself logs experiments to a backend.
- **Strength: one trace model for experiments and production.** Experiment items, online scores, annotations and human feedback all land as feedback scores on traces, so comparisons in the UI cover both.
- **Strength: fully Apache-2.0.** No feature folders are held back under another licence in this repository.
- **Caveat: auth is external.** Self-hosted Opik is single-tenant with no login unless you supply the external account service. Do not expose the default install publicly.
- **Caveat: random, lossy edges.** Online sampling uses `SecureRandom`, so re-ingesting a trace can change whether it is scored. The SDK queue drops the oldest messages under pressure.
- **Caveat: user Python runs in a process pool by default.** Isolation depends on choosing the Docker executor.
- **Caveat: no red-teaming.** Safety coverage is the `PromptInjection` heuristic, `Moderation`, bias presets and guardrails, not attack generation.
- **Caveat: large polyglot stack.** Java, Python and TypeScript services plus four datastores make the self-hosted footprint heavier than the SDK suggests.

*Sources: code at f217a86, verified Q&A.*

## How comet-ml/opik answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Opik provides a three-tier metrics system built on an abstract `BaseMetric` class (in `sdks/python/src/opik/evaluation/metrics/base_metric.py:34-104`). All metrics return a `ScoreResult` (value 0.0–1.0, plus name/reason/metadata). 

**Heuristic/statistical metrics** — no LLM calls needed. Includes: `Equals` (exact/case-insensitive string match), `Contains` (substring detection), `RegexMatch`, `IsJson`, `LevenshteinRatio`, `SentenceBLEU`/`CorpusBLEU` (wrapping NLTK), `ROUGE` (rouge1/2/L/Lsum, wrapping `rouge_score`), `ChrF`, `GLEU`, `METEOR`, `BERTScore` (wrapping `bert-score`), `JSDivergence`/`KLDivergence`, `SpearmanRanking`, `Readability`, `Tone`, `Sentiment`, `VADERSentiment`, and `LanguageAdherenceMetric`. All live in `sdks/python/src/opik/evaluation/metrics/heuristics/`. 

**LLM-judge metrics** — use a judge LLM (configurable model name, defaults via `OPIK_DEFAULT_LLM`) with structured JSON output via Pydantic response-format schemas. Includes: `Hallucination` (1.0 if hallucination detected), `AnswerRelevance`, `ContextPrecision`, `ContextRecall`, `Factuality` (per-claim scoring), `Moderation` (0.0–1.0 content-appropriateness), `Usefulness`, `TrajectoryAccuracy` (ReAct agent trajectory quality), `SycEval` (sycophancy detection with rebuttal generation), `StructuredOutputCompliance`, and `LLMJuriesJudge`. Each lives under `sdks/python/src/opik/evaluation/metrics/llm_judges/<name>/`. The `GEval` metric (`llm_judges/g_eval/metric.py:32-288`) is a generalised LLM-as-a-judge that builds a reusable chain-of-thought prompt from `task_introduction` and `evaluation_criteria`, caches it (LRU, max 128), and scores each output. `GEvalPreset` wraps pre-built rubrics including: `QARelevanceJudge`, `DemographicBiasJudge`, `GenderBiasJudge`, `PoliticalBiasJudge`, `RegionalBiasJudge`, `ReligiousBiasJudge`, `ComplianceRiskJudge`, `DialogueHelpfulnessJudge`, `PromptUncertaintyJudge`, `AgentTaskCompletionJudge`, `AgentToolCorrectnessJudge`, `SummarizationCoherenceJudge`, `SummarizationConsistencyJudge` (all from `sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval_presets.py`).

**RAGAS integration**: `RagasMetricWrapper` (`metrics/ragas_metric.py:13-77`) wraps any `ragas.SingleTurnMetric` as an Opik `BaseMetric`, mapping 'input'/'output' keys to RAGAS's 'user_input'/'response'.

**Custom scoring**: `ScorerFunction` (`scorers/scorer_function.py:8-27`) is a protocol taking `(dataset_item, task_outputs, task_span?)` returning ScoreResults — wrapped into metrics via `scorer_wrapper_metric.wrap_scorer_functions()`.

**Conversation metrics** under `metrics/conversation/` assess multi-turn threads: `ConversationThreadMetric` wraps a per-turn metric into a thread-level score; `GEvalConversationMetric` variant judges entire conversations. Specific LLM-judge conversation metrics include `ConversationalCoherenceMetric`, `SessionCompletenessQuality`, `UserFrustrationMetric`, and heuristic ones like `KnowledgeRetentionMetric` and `ConversationDegenerationMetric`.

**AggregatedMetric** (`metrics/aggregated_metric.py`) wraps metrics and computes statistics across all items.


Citations: [sdks/python/src/opik/evaluation/metrics/__init__.py:1-152](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/__init__.py#L1-L152) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:32-288](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L32-L288) · [sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py:139-213](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py#L139-L213) · [sdks/python/src/opik/evaluation/metrics/ragas_metric.py:13-77](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/ragas_metric.py#L13-L77)

### How is LLM-as-a-judge implemented? (answered)

Opik implements LLM-as-a-judge through two main paths: the `BaseMetric`-based LLM-judge metrics and the `LLMJudge` suite evaluator, both using structured output via Pydantic response-format models.

**Judge prompts and rubrics**: The test-suite `LLMJudge` (`suite_evaluators/llm_judge/metric.py:35-533`) has a fixed system prompt ("You are an expert judge tasked with evaluating if an AI agent's output satisfies a set of assertions") and a user template that formats input, output, and assertion criteria. Each assertion produces a boolean pass/fail with confidence. `GEval` (`metrics/llm_judges/g_eval/metric.py:32-288`) uses a two-stage process: first a chain-of-thought generation prompt (cached per judge+criteria+model combination in a 128-entry LRU), then a scoring query that includes the CoT in the prompt. Judge prompt templates are per-metric — `Hallucination`, `AnswerRelevance`, `Factuality`, `Moderation`, etc. each have their own template modules with few-shot examples.

**Structured output**: All LLM judges use Pydantic models as response-format schemas passed to the model via `response_format=<PydanticModel>` (e.g. `HallucinationResponseFormat`, `GEvalScoreFormat`, `AnswerRelevanceResponseFormat`). For LiteLLM models, structured output is enforced via `provider_kwargs['response_format']`; for others, the response string is parsed with dedicated parser modules.

**Judge model choice**: Configurable per metric instance via a `model` parameter (string name or `OpikBaseModel` instance). The factory in `models_factory.py` resolves names — defaults to `gpt-5-nano` (overridable via `OPIK_DEFAULT_LLM`). Supports `LiteLLMChatModel` (multi-provider), `OpenAI`/`Anthropic` direct models, and custom subclasses. Per-metric parameters include `temperature`, `seed`, `reasoning_effort`.

**Agentic vs one-shot scoring**: The suite-evaluator `LLMJudge` (`metric.py:291-348`) supports two scoring modes controlled by a `scoring_tool_strategy` selector ('auto', 'always', 'never', or custom). When a trace context is available and the strategy permits, it runs an `AgenticLLMJudge` (imported from `suite_evaluators/agentic/judge.py`) that can look up spans and tool calls from the full trace tree via an in-process emulator.

**Retry and resilience**: `_generate_and_parse` is wrapped with Tenacity retry (3 attempts, retries `LLMJudgeParseError` and `EmptyLLMResponseError`, reraise on exhaustion). On final parse failure, returns partial results with `LLMJudgeParseError.results`.

**Merge optimization**: Multiple `LLMJudge` instances with identical settings (model, temperature, seed, track) are merged into one via `LLMJudge.merged()` — assertions are deduplicated and sent in a single LLM call.


Citations: [sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py:35-533](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py#L35-L533) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:32-288](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L32-L288) · [sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py:350-365](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py#L350-L365)

### How are test datasets and cases defined, generated and versioned? (answered)

Opik treats datasets as first-class objects managed through a backend API with versioning built in.

**File formats/DSL**: Datasets are collections of items (dicts with arbitrary JSON-serializable content). Created via the SDK (`client.create_dataset(name)` → `dataset.insert(items)`) or the UI. Items can contain arbitrary data fields, tags, metadata, and trace/span references. The `DatasetItem` model (`api_objects/dataset/dataset_item.py:40-65`) has an `EvaluatorItem` list — each item can carry per-item evaluator configs (type + config dict, currently only 'llm_judge' supported) and per-item execution policies (runs_per_item, pass_threshold). Datasets export to pandas via `to_pandas()`, to JSON via `to_json()`, and stream items as chunks or individually through `stream_items()` / `get_items()`.

**Versioning**: `Dataset` objects have a `get_version_info()` returning `DatasetVersionPublic` (id, version_name, etc.). `DatasetVersion` provides a read-only snapshot at a specific version. Test suites (`api_objects/dataset/test_suite/test_suite.py`) are a wrapper around datasets: a `TestSuite` uses a `Dataset` internally and `TestSuiteVersion` wraps a `DatasetVersion`. The `TestSuite` API exposes `get_version_view()` for pinning to a specific version. 

**Synthetic data generation**: No built-in synthetic data generators exist in the SDK, but `SimulatedUser` (`simulation/simulated_user.py:10-100`) generates synthetic user messages via an LLM based on a persona prompt — used in multi-turn simulation but not for direct dataset creation.

**Benchmark task registry**: No static benchmark registry is included. Tasks are callables (`LLMTask = Callable[[Dict[str, Any]], Any]`) passed at evaluation time. Opik integrates with Ragas via `RagasMetricWrapper` but does not ship standard benchmark datasets.

**Dataset filtering**: Items can be filtered at retrieval time via OQL filter strings supporting fields like `data.field_name`, `tags`, `created_at` etc. Filters also support `dataset_filter_string` in `evaluate()` calls.


Citations: [sdks/python/src/opik/api_objects/dataset/dataset_item.py:9-65](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/dataset_item.py#L9-L65) · [sdks/python/src/opik/api_objects/dataset/dataset.py:100-300](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/dataset.py#L100-L300) · [sdks/python/src/opik/api_objects/dataset/test_suite/test_suite.py:84-120](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/test_suite/test_suite.py#L84-L120) · [sdks/python/src/opik/simulation/simulated_user.py:10-100](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/simulation/simulated_user.py#L10-L100)

### How are evals executed and reported? (answered)

Evaluation execution is orchestrated by `EvaluationEngine` (`evaluation/engine/engine.py:67-676`).

**Runners and parallelism**: The primary entry point is `evaluate()` in `evaluator.py:138-382`, which creates an experiment, resolves dataset items, then delegates to `EvaluationEngine.run_and_score()`. Parallelism is managed by `StreamingExecutor` (`evaluation_tasks_executor.py`), which accepts a `workers` parameter (default 16). Each dataset item gets its own task submission; multiple `trial_count` runs per item are executed with the same executor, grouped by item ID for progress tracking. `task_threads=1` runs sequentially. A separate `score_test_cases()` method re-scores existing experiments without task execution.

**Execution policies**: Each item can have a per-item `ExecutionPolicy` (runs_per_item, pass_threshold) merged with a suite-level default. The `get_item_execution_policy()` function merges item-level overrides (`engine.py:33-64`).

**Caching**: `GEval` has a thread-safe LRU chain-of-thought cache (128 entries, keyed by task_introduction + criteria + model_name + model fingerprint) — shared across all instances via class-level OrderedDict (`metrics/llm_judges/g_eval/metric.py:66-173`). No general response caching is built in.

**CI integration**: No built-in CI adapter. The SDK can be used in any Python CI pipeline via `evaluate()` or `run_tests()`. JSON report generation is available via `report_output_path`.

**Result storage**: Feedback scores are logged to the Opik backend during evaluation via `rest_operations.log_test_result_feedback_scores()`. Each score is written to a trace-level feedback score on the experiment trace. `EvaluationResult` contains all test results, experiment metadata, and an optional experiment URL.

**Comparison, regression and dashboards**: `run_tests()` runs test suites with pass/fail aggregation based on execution policy thresholds. `evaluate()` creates experiments linked to dataset versions. The backend provides experiment comparison views (via the Java backend + React frontend). Dashboard objects (`opik.api_objects.dashboard.Dashboard`) are SDK-level wrappers for the UI dashboards. `evaluate_resume()` (`evaluator.py:1510-1650`) restores interrupted evaluations by reading experiment state and local checkpoints — skipping already-completed items and re-executing only pending ones, then merging results.

**Error tolerance**: `ErrorTolerance` enum controls how many scoring failures abort the run. `METRIC_ERRORS` (default, 10) tolerates exceptions raised inside `score()`; `ALL_SCORING_ERRORS` (20) additionally tolerates metrics that cannot even be built. Tolerated failures produce `ScoreResult(scoring_failed=True)` with error details, excluded from aggregates.


Citations: [sdks/python/src/opik/evaluation/evaluator.py:138-700](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/evaluator.py#L138-L700) · [sdks/python/src/opik/evaluation/evaluator.py:1510-1650](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/evaluator.py#L1510-L1650) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:66-173](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L66-L173)

### How are traces or production data captured and linked to evaluations? (answered)

Opik has deep observability built around its `@opik.track` decorator (`decorator/tracker.py:11-99` and `decorator/base_track_decorator.py:41-80`).

**SDK instrumentation**: The `@track` decorator wraps any function (LLM calls, internal steps, entire tasks) to automatically create spans and traces. It captures inputs/outputs, start/end times, tags, metadata, project name, and error info. The `OpikTrackDecorator` has `_start_span_inputs_preprocessor` and `_end_span_inputs_preprocessor` hooks. Spans and traces have IDs, can be nested, and support `update_current_span()`/`update_current_trace()` for mid-execution metadata. `@track` can be applied transparently to any callable; the evaluation system auto-wraps task functions that are not already tracked (`engine.py:279-281`).

**Context managers**: `start_as_current_span()` and `start_as_current_trace()` provide explicit context-manager access for fine-grained manual tracing.

**LLM provider integrations**: Opik has dedicated integration modules under `integrations/` covering: OpenAI, Anthropic, LangChain, LlamaIndex, Haystack, CrewAI, DSPy, Bedrock, Mistral, Ollama, Groq, Cerebras, AISuite, ADK, AgentSpec, Harbor, SageMaker, and Guardrails. Each logs the provider's request/response as Opik spans.

**OpenTelemetry**: The `integrations/otel/` module provides `OpikSpanProcessor` — an OpenTelemetry span processor that converts OTel spans into Opik spans, enabling integration with any OTel-instrumented code.

**Online vs offline evals**: The default path connects to the Opik backend to log experiments and feedback scores. For offline use, `record_traces_locally()` (`api_objects/local_recording.py`) captures traces to an in-memory handle that can be inspected post-hoc without a backend.

**Feedback and annotation**: `TracesAnnotationQueue` and `ThreadsAnnotationQueue` (`api_objects/annotation_queue.py`, linked from `__init__.py:2-4`) support human annotation workflows. Feedback scores are the primary mechanism for recording evaluation results on traces.

**Guardrails tracing**: The `guardrails/` module has its own `GuardrailsTrackDecorator` that logs guardrail validation results as special spans.


Citations: [sdks/python/src/opik/decorator/base_track_decorator.py:41-80](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/decorator/base_track_decorator.py#L41-L80) · [sdks/python/src/opik/integrations/otel/__init__.py:1-8](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/integrations/otel/__init__.py#L1-L8) · [sdks/python/src/opik/__init__.py:2-6](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/__init__.py#L2-L6)

### Does it support red-teaming or safety testing, and how? (answered)

Opik has multiple red-teaming-adjacent features, though it does not have a dedicated red-teaming module.

**`PromptInjection` metric** (`metrics/heuristics/prompt_injection.py:139-213`): A heuristic regex-based metric that scans LLM outputs for prompt injection and system-prompt leakage patterns. It uses 30+ compiled regex patterns (ignore/disregard/override/pretend/expose patterns) and 30+ suspicious keyword substrings. Returns 1.0 for regex matches (strong injection signal), 0.5 for keyword-only hits, 0.0 otherwise. Patterns cover: ignore/disregard/override instructions, pretend role-playing ("pretend to be the assistant/system/DAN"), expose/leak system prompt phrases, developer mode/DAN mode/Jailbreak patterns, and "no longer bound/restricted" escape clauses.

**`Moderation` LLM judge** (`metrics/llm_judges/moderation/metric.py:14-123`): An LLM-based metric scoring output content-appropriateness from 0.0 to 1.0. Configurable with few-shot examples.

**`SycEval` metric** (`metrics/llm_judges/syc_eval/metric.py:18-261`): Implements the SycEval protocol from arxiv 2502.08177 to detect sycophantic behavior. Generates rebuttals of varying rhetorical strength (simple/ethos/justification/citation) via a separate rebuttal model (prevents contamination), then classifies whether the model changes its answers under pressure.

**Guardrails API** (`guardrails/guardrail.py:28-79`): A runtime validation layer with built-in guards including `Topic` (restricted topic detection), `PII` (entity blocking with thresholds), `LLMJudge`, and `PromptInjection`. Guardrail results can be logged as Opik trace spans.

**Bias judge presets**: `GEvalPreset` includes built-in bias evaluation judges: `DemographicBiasJudge`, `GenderBiasJudge`, `PoliticalBiasJudge`, `RegionalBiasJudge`, `ReligiousBiasJudge` (from `metrics/llm_judges/g_eval_presets.py`). These are pre-configured GEval instances with bias-specific evaluation criteria.

**Limitations**: There is no adversarial attack generator, no jailbreak test harness, no automated vulnerability reporting pipeline, and no integration with dedicated red-teaming frameworks. The simulation module (`SimulatedUser`) can generate adversarial persona-based user messages but is not specifically designed for red-teaming.


Citations: [sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py:1-213](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/heuristics/prompt_injection.py#L1-L213) · [sdks/python/src/opik/guardrails/guardrail.py:28-79](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/guardrails/guardrail.py#L28-L79) · [sdks/python/src/opik/evaluation/metrics/__init__.py:53-68](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/__init__.py#L53-L68)
