comet-ml/opik
LLM tracing and eval platform: Python/TS SDKs with ~30 metrics, a Java backend on ClickHouse and MySQL, and Redis-streamed online scoring.
Overview
Opik is Comet’s open-source platform for tracing and evaluating LLM applications. It has two halves that are useful on their own. The Python SDK (sdks/python) is a full offline evaluation library: datasets, evaluate(), around 30 heuristic and LLM-judge metrics, test suites and conversation simulation. The server (apps/) stores traces, spans, feedback scores, datasets and experiments, shows them in a React UI, and runs “online” evaluation rules against incoming traces.
The two halves meet at the trace. @opik.track turns your functions into spans and traces. evaluate() runs your task under that same decorator, so every experiment item is a normal trace with feedback scores attached. Online rules on the server score production traces with LLM judges or user-written Python metrics.
The whole repository is Apache-2.0: the backend, the frontend, the Python and TypeScript SDKs and the opik_optimizer package. There is no ee/ folder or licence-gated code path. The open-core boundary is authentication instead. With authentication.enabled false (the default), every request maps to a single default workspace and user (AuthService.java, config.yml). Turning it on requires authentication.reactService.url, an external account service that is not in this repository (AuthModule.java). In practice, multi-user workspaces, API keys and permissions come from Comet’s hosted product.
Architecture
flowchart LR
APP["Your code + @opik.track"] --> Q["SDK message queue + batcher"]
EVAL["opik.evaluate()"] --> APP
EVAL --> MET["SDK metrics (heuristic, LLM judge)"]
Q --> REST["Backend REST /traces/batch"]
OTEL["OTel exporters"] --> OTLP["Backend /otel/v1"]
REST --> SVC["TraceService"]
OTLP --> SVC
SVC --> CH["ClickHouse: traces, spans, scores"]
SVC --> BUS["EventBus: TracesCreated"]
BUS --> SAMP["OnlineScoringSampler"]
SAMP --> RS["Redis streams"]
RS --> JUDGE["LLM-judge scorer"]
RS --> PYS["Python metric scorer"]
PYS --> PYB["opik-python-backend"]
JUDGE --> CH
PYS --> CH
UI["React frontend"] --> REST
| Component | Path | Role |
|---|---|---|
| Python SDK tracing | sdks/python/src/opik/decorator/, message_processing/ |
@track, context managers, background queue, batching, REST upload |
| Python SDK evaluation | sdks/python/src/opik/evaluation/ |
evaluate, EvaluationEngine, metrics, suite evaluators, resume |
| Integrations | sdks/python/src/opik/integrations/ |
OpenAI, Anthropic, LangChain, LlamaIndex, DSPy, ADK, Bedrock, an OTel span processor and more |
| TypeScript SDK | sdks/typescript/ |
Tracing and evaluation client for Node |
| Backend | apps/opik-backend/ |
Dropwizard (Java) REST API; ClickHouse for traces/spans/scores, MySQL for state |
| Online scoring | apps/opik-backend/.../api/resources/v1/events/ |
Samplers, Redis-stream publishers and scorers for LLM-judge and Python rules |
| Python evaluator service | apps/opik-python-backend/ |
Flask service that executes user Python metrics |
| Guardrails service | apps/opik-guardrails-backend/ |
Optional PII and topic guard models (GPU or CPU profile) |
| Optimizer | sdks/opik_optimizer/ |
Prompt optimizers (evolutionary, few-shot Bayesian, GEPA, meta-prompt and others) |
How a request flows
A traced call in production, scored by an online LLM-judge rule:
- Capture.
@trackwraps the function. When it returns,_after_callfills end time, output or error on the span and trace data. It then hands them to the global client’s internalspan/tracemethods (base_track_decorator.py, L563-L588). - Queue and batch. The client turns those into messages.
Streamer.putextracts embedded attachments, lets the batcher absorb the message, and otherwise puts it on a bounded in-memory queue. If the queue is full, the oldest message is dropped with a warning (streamer.py). - Upload. A consumer thread truncates oversized payloads and posts batches with
spans.create_spans/traces.create_traces(online_message_processor.py). - Store.
POST /v1/private/traces/batchcallsTraceService.create(TracesResource.java). The service dedups, resolves or creates projects, strips large attachments to blob storage, batch-inserts into ClickHouse, and postsTracesCreatedon the in-process event bus (TraceService.java). - Sample.
OnlineScoringSampler.onTracesCreatedkeeps only traces with anend_time(OnlineScoringSampler.java). It loads the project’s rules and, per rule, checks enabled state, filters and aSecureRandomroll against the sampling rate (L361-L384). Sampled traces become messages per evaluator type (L243-L290), whichOnlineScorePublisheradds to a Redis stream (OnlineScorePublisher.java). - Score. The LLM-judge scorer reads the stream. It fetches spans only when the prompt needs them, then runs either a chat judge (with optional tool calls to inspect the trace) or a single decision-model call. It stores the result as feedback scores (OnlineScoringLlmAsJudgeScorer.java). When a toggle is on, the judge’s own calls are recorded as a hidden evaluator trace.
Key components
SDK evaluation engine
evaluate(dataset, task, scoring_metrics, ...) creates an experiment and runs the task on each item with 16 threads by default. It supports trial_count, dataset filters and samplers, and an error_tolerance level (evaluator.py). EvaluationEngine.run_and_score splits metrics into regular metrics and task-span metrics, which score the spans the task produced (engine.py). For each item, the engine wraps an untracked task in opik.track, opens a trace whose input is the item, runs the task and attaches the output (L269-L330). Scores are logged as feedback scores on that trace and linked to the experiment item. evaluate_resume can pick up an interrupted run.
Metrics
Heuristic metrics need no LLM: Equals, Contains, RegexMatch, IsJson, LevenshteinRatio, BLEU, ROUGE, chrF, GLEU, METEOR, BERTScore, divergence and ranking metrics, sentiment, readability and a regex PromptInjection detector. LLM-judge metrics (Hallucination, AnswerRelevance, ContextPrecision/Recall, Factuality, Moderation, Usefulness, TrajectoryAccuracy, SycEval, GEval and its presets, LLMJuriesJudge) build a prompt, call the model with a Pydantic response_format, and parse a ScoreResult (hallucination/metric.py). Judge models resolve through LiteLLM, with a per-process cache of model instances (models_factory.py). The default is openai/gpt-5-nano (config.py). GEval caches its generated chain-of-thought in a 128-entry class-level LRU (g_eval/metric.py).
Python metrics on the server
Online “user-defined metric” rules send the rule’s Python code and mapped trace fields to the scorer (OnlineScoringUserDefinedMetricPythonScorer.java). The scorer calls the Python backend’s POST /python (evaluator.py). That service runs the code with either a process executor (a local process pool) or a docker executor, chosen by PYTHON_CODE_EXECUTOR_STRATEGY (L23-L40). The bundled compose file defaults to process (docker-compose.yaml).
Test suites and simulation
Datasets can carry per-item evaluator configs and execution policies (runs_per_item, pass_threshold). run_tests() turns that into pass/fail test suites judged by an LLMJudge that checks lists of assertions, merging identical judges into one call. SimulatedUser drives multi-turn conversations for thread-level metrics.
Extending it
- Metrics. Subclass
BaseMetricand implementscore/ascore, or pass plain scorer functions that receivedataset_item,task_outputsand optionallytask_span.RagasMetricWrapperadapts Ragas metrics. - Judge models. Pass any LiteLLM model string or an
OpikBaseModelsubclass. - Tracing. Use framework integrations,
opik.trackon anything, orOpikSpanProcessorfor OTel-instrumented code. The backend also exposes an OTLP endpoint. - Online rules. LLM-judge prompts with variable mappings, or Python metric code, scoped to traces, spans or threads with filters and a sampling rate.
Running it
./opik.sh (or opik.ps1) starts the Docker Compose stack: MySQL, Redis, ClickHouse with ZooKeeper, MinIO, the Java backend, the Python backend, the frontend, and optional guardrails and observability profiles. A Helm chart is in deployment/helm_chart. The SDK is pip install opik. Run opik configure and point it at a local instance or Comet’s cloud. LLM-judge metrics need an API key for the chosen provider. Heuristic metrics need no model key.
Strengths and caveats
- Strength: a real eval library. The SDK ships the broadest built-in metric set of the trace-first platforms, including classic NLG metrics. Metrics are plain Python classes you can call without the server.
evaluate()itself logs experiments to a backend. - Strength: one trace model for experiments and production. Experiment items, online scores, annotations and human feedback all land as feedback scores on traces, so comparisons in the UI cover both.
- Strength: fully Apache-2.0. No feature folders are held back under another licence in this repository.
- Caveat: auth is external. Self-hosted Opik is single-tenant with no login unless you supply the external account service. Do not expose the default install publicly.
- Caveat: random, lossy edges. Online sampling uses
SecureRandom, so re-ingesting a trace can change whether it is scored. The SDK queue drops the oldest messages under pressure. - Caveat: user Python runs in a process pool by default. Isolation depends on choosing the Docker executor.
- Caveat: no red-teaming. Safety coverage is the
PromptInjectionheuristic,Moderation, bias presets and guardrails, not attack generation. - Caveat: large polyglot stack. Java, Python and TypeScript services plus four datastores make the self-hosted footprint heavier than the SDK suggests.
Sources: code at f217a86, verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit f217a86. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredOpik provides a three-tier metrics system built on an abstract BaseMetric class (in sdks/python/src/opik/evaluation/metrics/base_metric.py:34-104). All metrics return a ScoreResult (value 0.0–1.0, plus name/reason/metadata).
Heuristic/statistical metrics — no LLM calls needed. Includes: Equals (exact/case-insensitive string match), Contains (substring detection), RegexMatch, IsJson, LevenshteinRatio, SentenceBLEU/CorpusBLEU (wrapping NLTK), ROUGE (rouge1/2/L/Lsum, wrapping rouge_score), ChrF, GLEU, METEOR, BERTScore (wrapping bert-score), JSDivergence/KLDivergence, SpearmanRanking, Readability, Tone, Sentiment, VADERSentiment, and LanguageAdherenceMetric. All live in sdks/python/src/opik/evaluation/metrics/heuristics/.
LLM-judge metrics — use a judge LLM (configurable model name, defaults via OPIK_DEFAULT_LLM) with structured JSON output via Pydantic response-format schemas. Includes: Hallucination (1.0 if hallucination detected), AnswerRelevance, ContextPrecision, ContextRecall, Factuality (per-claim scoring), Moderation (0.0–1.0 content-appropriateness), Usefulness, TrajectoryAccuracy (ReAct agent trajectory quality), SycEval (sycophancy detection with rebuttal generation), StructuredOutputCompliance, and LLMJuriesJudge. Each lives under sdks/python/src/opik/evaluation/metrics/llm_judges/<name>/. The GEval metric (llm_judges/g_eval/metric.py:32-288) is a generalised LLM-as-a-judge that builds a reusable chain-of-thought prompt from task_introduction and evaluation_criteria, caches it (LRU, max 128), and scores each output. GEvalPreset wraps pre-built rubrics including: QARelevanceJudge, DemographicBiasJudge, GenderBiasJudge, PoliticalBiasJudge, RegionalBiasJudge, ReligiousBiasJudge, ComplianceRiskJudge, DialogueHelpfulnessJudge, PromptUncertaintyJudge, AgentTaskCompletionJudge, AgentToolCorrectnessJudge, SummarizationCoherenceJudge, SummarizationConsistencyJudge (all from sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval_presets.py).
RAGAS integration: RagasMetricWrapper (metrics/ragas_metric.py:13-77) wraps any ragas.SingleTurnMetric as an Opik BaseMetric, mapping 'input'/'output' keys to RAGAS's 'user_input'/'response'.
Custom scoring: ScorerFunction (scorers/scorer_function.py:8-27) is a protocol taking (dataset_item, task_outputs, task_span?) returning ScoreResults — wrapped into metrics via scorer_wrapper_metric.wrap_scorer_functions().
Conversation metrics under metrics/conversation/ assess multi-turn threads: ConversationThreadMetric wraps a per-turn metric into a thread-level score; GEvalConversationMetric variant judges entire conversations. Specific LLM-judge conversation metrics include ConversationalCoherenceMetric, SessionCompletenessQuality, UserFrustrationMetric, and heuristic ones like KnowledgeRetentionMetric and ConversationDegenerationMetric.
AggregatedMetric (metrics/aggregated_metric.py) wraps metrics and computes statistics across all items.
How is LLM-as-a-judge implemented?
answeredOpik implements LLM-as-a-judge through two main paths: the BaseMetric-based LLM-judge metrics and the LLMJudge suite evaluator, both using structured output via Pydantic response-format models.
Judge prompts and rubrics: The test-suite LLMJudge (suite_evaluators/llm_judge/metric.py:35-533) has a fixed system prompt ("You are an expert judge tasked with evaluating if an AI agent's output satisfies a set of assertions") and a user template that formats input, output, and assertion criteria. Each assertion produces a boolean pass/fail with confidence. GEval (metrics/llm_judges/g_eval/metric.py:32-288) uses a two-stage process: first a chain-of-thought generation prompt (cached per judge+criteria+model combination in a 128-entry LRU), then a scoring query that includes the CoT in the prompt. Judge prompt templates are per-metric — Hallucination, AnswerRelevance, Factuality, Moderation, etc. each have their own template modules with few-shot examples.
Structured output: All LLM judges use Pydantic models as response-format schemas passed to the model via response_format=<PydanticModel> (e.g. HallucinationResponseFormat, GEvalScoreFormat, AnswerRelevanceResponseFormat). For LiteLLM models, structured output is enforced via provider_kwargs['response_format']; for others, the response string is parsed with dedicated parser modules.
Judge model choice: Configurable per metric instance via a model parameter (string name or OpikBaseModel instance). The factory in models_factory.py resolves names — defaults to gpt-5-nano (overridable via OPIK_DEFAULT_LLM). Supports LiteLLMChatModel (multi-provider), OpenAI/Anthropic direct models, and custom subclasses. Per-metric parameters include temperature, seed, reasoning_effort.
Agentic vs one-shot scoring: The suite-evaluator LLMJudge (metric.py:291-348) supports two scoring modes controlled by a scoring_tool_strategy selector ('auto', 'always', 'never', or custom). When a trace context is available and the strategy permits, it runs an AgenticLLMJudge (imported from suite_evaluators/agentic/judge.py) that can look up spans and tool calls from the full trace tree via an in-process emulator.
Retry and resilience: _generate_and_parse is wrapped with Tenacity retry (3 attempts, retries LLMJudgeParseError and EmptyLLMResponseError, reraise on exhaustion). On final parse failure, returns partial results with LLMJudgeParseError.results.
Merge optimization: Multiple LLMJudge instances with identical settings (model, temperature, seed, track) are merged into one via LLMJudge.merged() — assertions are deduplicated and sent in a single LLM call.
How are test datasets and cases defined, generated and versioned?
answeredOpik treats datasets as first-class objects managed through a backend API with versioning built in.
File formats/DSL: Datasets are collections of items (dicts with arbitrary JSON-serializable content). Created via the SDK (client.create_dataset(name) → dataset.insert(items)) or the UI. Items can contain arbitrary data fields, tags, metadata, and trace/span references. The DatasetItem model (api_objects/dataset/dataset_item.py:40-65) has an EvaluatorItem list — each item can carry per-item evaluator configs (type + config dict, currently only 'llm_judge' supported) and per-item execution policies (runs_per_item, pass_threshold). Datasets export to pandas via to_pandas(), to JSON via to_json(), and stream items as chunks or individually through stream_items() / get_items().
Versioning: Dataset objects have a get_version_info() returning DatasetVersionPublic (id, version_name, etc.). DatasetVersion provides a read-only snapshot at a specific version. Test suites (api_objects/dataset/test_suite/test_suite.py) are a wrapper around datasets: a TestSuite uses a Dataset internally and TestSuiteVersion wraps a DatasetVersion. The TestSuite API exposes get_version_view() for pinning to a specific version.
Synthetic data generation: No built-in synthetic data generators exist in the SDK, but SimulatedUser (simulation/simulated_user.py:10-100) generates synthetic user messages via an LLM based on a persona prompt — used in multi-turn simulation but not for direct dataset creation.
Benchmark task registry: No static benchmark registry is included. Tasks are callables (LLMTask = Callable[[Dict[str, Any]], Any]) passed at evaluation time. Opik integrates with Ragas via RagasMetricWrapper but does not ship standard benchmark datasets.
Dataset filtering: Items can be filtered at retrieval time via OQL filter strings supporting fields like data.field_name, tags, created_at etc. Filters also support dataset_filter_string in evaluate() calls.
How are evals executed and reported?
answeredEvaluation execution is orchestrated by EvaluationEngine (evaluation/engine/engine.py:67-676).
Runners and parallelism: The primary entry point is evaluate() in evaluator.py:138-382, which creates an experiment, resolves dataset items, then delegates to EvaluationEngine.run_and_score(). Parallelism is managed by StreamingExecutor (evaluation_tasks_executor.py), which accepts a workers parameter (default 16). Each dataset item gets its own task submission; multiple trial_count runs per item are executed with the same executor, grouped by item ID for progress tracking. task_threads=1 runs sequentially. A separate score_test_cases() method re-scores existing experiments without task execution.
Execution policies: Each item can have a per-item ExecutionPolicy (runs_per_item, pass_threshold) merged with a suite-level default. The get_item_execution_policy() function merges item-level overrides (engine.py:33-64).
Caching: GEval has a thread-safe LRU chain-of-thought cache (128 entries, keyed by task_introduction + criteria + model_name + model fingerprint) — shared across all instances via class-level OrderedDict (metrics/llm_judges/g_eval/metric.py:66-173). No general response caching is built in.
CI integration: No built-in CI adapter. The SDK can be used in any Python CI pipeline via evaluate() or run_tests(). JSON report generation is available via report_output_path.
Result storage: Feedback scores are logged to the Opik backend during evaluation via rest_operations.log_test_result_feedback_scores(). Each score is written to a trace-level feedback score on the experiment trace. EvaluationResult contains all test results, experiment metadata, and an optional experiment URL.
Comparison, regression and dashboards: run_tests() runs test suites with pass/fail aggregation based on execution policy thresholds. evaluate() creates experiments linked to dataset versions. The backend provides experiment comparison views (via the Java backend + React frontend). Dashboard objects (opik.api_objects.dashboard.Dashboard) are SDK-level wrappers for the UI dashboards. evaluate_resume() (evaluator.py:1510-1650) restores interrupted evaluations by reading experiment state and local checkpoints — skipping already-completed items and re-executing only pending ones, then merging results.
Error tolerance: ErrorTolerance enum controls how many scoring failures abort the run. METRIC_ERRORS (default, 10) tolerates exceptions raised inside score(); ALL_SCORING_ERRORS (20) additionally tolerates metrics that cannot even be built. Tolerated failures produce ScoreResult(scoring_failed=True) with error details, excluded from aggregates.
How are traces or production data captured and linked to evaluations?
answeredOpik has deep observability built around its @opik.track decorator (decorator/tracker.py:11-99 and decorator/base_track_decorator.py:41-80).
SDK instrumentation: The @track decorator wraps any function (LLM calls, internal steps, entire tasks) to automatically create spans and traces. It captures inputs/outputs, start/end times, tags, metadata, project name, and error info. The OpikTrackDecorator has _start_span_inputs_preprocessor and _end_span_inputs_preprocessor hooks. Spans and traces have IDs, can be nested, and support update_current_span()/update_current_trace() for mid-execution metadata. @track can be applied transparently to any callable; the evaluation system auto-wraps task functions that are not already tracked (engine.py:279-281).
Context managers: start_as_current_span() and start_as_current_trace() provide explicit context-manager access for fine-grained manual tracing.
LLM provider integrations: Opik has dedicated integration modules under integrations/ covering: OpenAI, Anthropic, LangChain, LlamaIndex, Haystack, CrewAI, DSPy, Bedrock, Mistral, Ollama, Groq, Cerebras, AISuite, ADK, AgentSpec, Harbor, SageMaker, and Guardrails. Each logs the provider's request/response as Opik spans.
OpenTelemetry: The integrations/otel/ module provides OpikSpanProcessor — an OpenTelemetry span processor that converts OTel spans into Opik spans, enabling integration with any OTel-instrumented code.
Online vs offline evals: The default path connects to the Opik backend to log experiments and feedback scores. For offline use, record_traces_locally() (api_objects/local_recording.py) captures traces to an in-memory handle that can be inspected post-hoc without a backend.
Feedback and annotation: TracesAnnotationQueue and ThreadsAnnotationQueue (api_objects/annotation_queue.py, linked from __init__.py:2-4) support human annotation workflows. Feedback scores are the primary mechanism for recording evaluation results on traces.
Guardrails tracing: The guardrails/ module has its own GuardrailsTrackDecorator that logs guardrail validation results as special spans.
Does it support red-teaming or safety testing, and how?
answeredOpik has multiple red-teaming-adjacent features, though it does not have a dedicated red-teaming module.
PromptInjection metric (metrics/heuristics/prompt_injection.py:139-213): A heuristic regex-based metric that scans LLM outputs for prompt injection and system-prompt leakage patterns. It uses 30+ compiled regex patterns (ignore/disregard/override/pretend/expose patterns) and 30+ suspicious keyword substrings. Returns 1.0 for regex matches (strong injection signal), 0.5 for keyword-only hits, 0.0 otherwise. Patterns cover: ignore/disregard/override instructions, pretend role-playing ("pretend to be the assistant/system/DAN"), expose/leak system prompt phrases, developer mode/DAN mode/Jailbreak patterns, and "no longer bound/restricted" escape clauses.
Moderation LLM judge (metrics/llm_judges/moderation/metric.py:14-123): An LLM-based metric scoring output content-appropriateness from 0.0 to 1.0. Configurable with few-shot examples.
SycEval metric (metrics/llm_judges/syc_eval/metric.py:18-261): Implements the SycEval protocol from arxiv 2502.08177 to detect sycophantic behavior. Generates rebuttals of varying rhetorical strength (simple/ethos/justification/citation) via a separate rebuttal model (prevents contamination), then classifies whether the model changes its answers under pressure.
Guardrails API (guardrails/guardrail.py:28-79): A runtime validation layer with built-in guards including Topic (restricted topic detection), PII (entity blocking with thresholds), LLMJudge, and PromptInjection. Guardrail results can be logged as Opik trace spans.
Bias judge presets: GEvalPreset includes built-in bias evaluation judges: DemographicBiasJudge, GenderBiasJudge, PoliticalBiasJudge, RegionalBiasJudge, ReligiousBiasJudge (from metrics/llm_judges/g_eval_presets.py). These are pre-configured GEval instances with bias-specific evaluation criteria.
Limitations: There is no adversarial attack generator, no jailbreak test harness, no automated vulnerability reporting pipeline, and no integration with dedicated red-teaming frameworks. The simulation module (SimulatedUser) can generate adversarial persona-based user messages but is not specifically designed for red-teaming.