# How are traces or production data captured and linked to evaluations?

> LLM evals and testing — a good answer covers: SDK instrumentation; OpenTelemetry; online vs offline evals; feedback and annotation; if absent, say so.

Canonical page: https://llms-technical-reviews.com/evals/q/observability/

## Verdict

[Langfuse](/p/langfuse/) and [Opik](/p/opik/) are the only projects that score production traces as they arrive. [Phoenix](/p/phoenix/) has the cleanest OpenTelemetry ingest, but you script online scoring yourself. The other seven tools trace only their own eval runs, or nothing.

**Trace platforms.** Langfuse accepts OTLP and its own SDK events. Evaluation rules pick traces with a deterministic SHA-256 sample, so the same trace is always in or out. Each judge call is itself a trace, linked from the score by `executionTraceId`. Human annotations share the score table with `source: ANNOTATION`. Note that the SDKs are in separate repositories. Opik captures spans with `@opik.track`, an OTel span processor or its OTLP endpoint. A sampler using `SecureRandom` sends traces through Redis streams to LLM-judge or Python-metric scorers. Experiment items are ordinary traces, so offline and online scores sit side by side. Phoenix receives OTLP over gRPC or HTTP with OpenInference attributes. Scores, labels and code checks all become span annotations, and `log_span_annotations` attaches offline results to production spans. Its online runner in `evals/pxi` is internal and not shipped.

**Libraries that trace evaluations.** [DeepEval](/p/deepeval/) has `@observe` and OpenAI and Anthropic patchers, plus an OTel processor that forwards spans to its vendor's hosted platform. Online evaluation of production traffic is a feature of that platform. [promptfoo](/p/promptfoo/) starts a local OTLP receiver during a run so that assertions can check the spans your app emitted. It can also read traces from Langfuse, Braintrust or Tempo, but it never captures production traffic. [Ragas](/p/ragas/) records an evaluation → row → metric → prompt callback tree and has Langfuse and MLflow helpers. It has no OTel support.

**Offline logs only.** [Inspect](/p/inspect_ai/) records typed events per sample and exposes lifecycle hooks, but has no OTel. [lm-evaluation-harness](/p/lm-evaluation-harness/) logs to W&B or Trackio. [OpenAI Evals](/p/openai-evals/) records prompt and output events to JSONL. For [garak](/p/garak/) the question is not applicable. It is a scanner with a report file.

Pick: Langfuse for rule-based online scoring with auditable judges.
Pick: Opik when you also want a Python metric library on the same traces.
Pick: Phoenix for standard OTLP ingest and annotation-driven review.

## Per-project answers

### langfuse/langfuse (answered)

Langfuse **is** an observability platform, so traces are captured via a first-party Python/JS/TS SDK (`langfuse` pip package, JS SDK in `packages/shared/src/server/llm/`), and it also ingests OpenTelemetry data through the OTel ingestion pipeline (`OtelIngestionProcessor.ts`). Traces consist of a root trace with nested observations (spans, generations, events), all stored in ClickHouse. **Linking evals to traces** works through the `executionTraceId` mechanism: when an LLM-as-a-judge evaluator runs, it creates a new *internal* trace (environment `langfuse-llm-as-a-judge` or `langfuse-code-eval`) in the same project, passing `traceSinkParams` with `targetProjectId`, `traceId`, and `metadata` (`evalExecutionDeps.ts` lines 302–312). This internal trace records the judge's input prompt, output score, model, latency, and reasoning. The score event's `executionTraceId` field links the score back to this internal trace. **Online vs offline evals**: "online" (trace-upsert-triggered) runs eval rules immediately when traces arrive; "offline" evaluations are triggered via the UI batch-action dialog or the public API with a timestamp range for historical data. **Feedback and annotation** is a separate feature: Langfuse has `scores` with `source: "ANNOTATION"` as distinct from `source: "EVAL"`, both stored in the same scores table but tagged differently. The scores repository (`scores.ts`) handles both types. **The eval log** (`eval-log.tsx`) provides a UI table of all eval job executions with status, score, comment, and links to the target trace. Each evaluation also writes its own internal trace with environment labels that are excluded from triggering further evaluations (the `isInternalEvalEnvironment` guard in `isEvalTargetEnvironmentAllowed.ts` prevents infinite eval loops).

> **Editor's note.** Correction: the Python and JS tracing SDKs are not in this repository; `packages/shared/src/server/llm/` is the server's own LLM-calling code.

Citations: [worker/src/features/evaluation/evalExecutionDeps.ts:245-325](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionDeps.ts#L245-L325) · [packages/shared/src/server/otel/OtelIngestionProcessor.ts:1-50](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/otel/OtelIngestionProcessor.ts#L1-L50) · [worker/src/features/evaluation/isEvalTargetEnvironmentAllowed.ts:1-30](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/isEvalTargetEnvironmentAllowed.ts#L1-L30) · [packages/shared/src/server/repositories/scores.ts:1-30](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/repositories/scores.ts#L1-L30) · [web/src/features/evals/components/eval-log.tsx:1-60](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/components/eval-log.tsx#L1-L60)

### promptfoo/promptfoo (answered)

OpenTelemetry integration via src/tracing/otelSdk.ts and src/tracing/otelConfig.ts with OTLP HTTP export. Provider calls wrapped with withTracedProviderCall/withTestCaseSpan/withGraderSpan (src/tracing/targetTracer.ts). W3C traceparent headers generated and validated (evaluator.ts L1567-1583). Trace store (src/tracing/store.ts) manages SQLite-backed tracesTable and spansTable with retry-based stability detection (loadTraceData polls until span count stabilizes, src/assertions/index.ts L184-227). External trace providers: Braintrust, Langfuse, Tempo via TraceProvider interface (src/tracing/providers/index.ts L21-39). Traces linked to eval rows via traceId/evaluationId. No production traffic capture SDK exists; traces are generated only during eval execution.


Citations: [src/tracing/store.ts:1-100](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/tracing/store.ts#L1-L100) · [src/tracing/providers/index.ts:21-39](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/tracing/providers/index.ts#L21-L39) · [src/assertions/index.ts:183-227](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L183-L227) · [src/evaluator.ts:1558-1598](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1558-L1598)

### comet-ml/opik (answered)

Opik has deep observability built around its `@opik.track` decorator (`decorator/tracker.py:11-99` and `decorator/base_track_decorator.py:41-80`).

**SDK instrumentation**: The `@track` decorator wraps any function (LLM calls, internal steps, entire tasks) to automatically create spans and traces. It captures inputs/outputs, start/end times, tags, metadata, project name, and error info. The `OpikTrackDecorator` has `_start_span_inputs_preprocessor` and `_end_span_inputs_preprocessor` hooks. Spans and traces have IDs, can be nested, and support `update_current_span()`/`update_current_trace()` for mid-execution metadata. `@track` can be applied transparently to any callable; the evaluation system auto-wraps task functions that are not already tracked (`engine.py:279-281`).

**Context managers**: `start_as_current_span()` and `start_as_current_trace()` provide explicit context-manager access for fine-grained manual tracing.

**LLM provider integrations**: Opik has dedicated integration modules under `integrations/` covering: OpenAI, Anthropic, LangChain, LlamaIndex, Haystack, CrewAI, DSPy, Bedrock, Mistral, Ollama, Groq, Cerebras, AISuite, ADK, AgentSpec, Harbor, SageMaker, and Guardrails. Each logs the provider's request/response as Opik spans.

**OpenTelemetry**: The `integrations/otel/` module provides `OpikSpanProcessor` — an OpenTelemetry span processor that converts OTel spans into Opik spans, enabling integration with any OTel-instrumented code.

**Online vs offline evals**: The default path connects to the Opik backend to log experiments and feedback scores. For offline use, `record_traces_locally()` (`api_objects/local_recording.py`) captures traces to an in-memory handle that can be inspected post-hoc without a backend.

**Feedback and annotation**: `TracesAnnotationQueue` and `ThreadsAnnotationQueue` (`api_objects/annotation_queue.py`, linked from `__init__.py:2-4`) support human annotation workflows. Feedback scores are the primary mechanism for recording evaluation results on traces.

**Guardrails tracing**: The `guardrails/` module has its own `GuardrailsTrackDecorator` that logs guardrail validation results as special spans.


Citations: [sdks/python/src/opik/decorator/base_track_decorator.py:41-80](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/decorator/base_track_decorator.py#L41-L80) · [sdks/python/src/opik/integrations/otel/__init__.py:1-8](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/integrations/otel/__init__.py#L1-L8) · [sdks/python/src/opik/__init__.py:2-6](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/__init__.py#L2-L6)

### openai/evals (answered)

**SDK instrumentation and OpenTelemetry.** This repository does not implement OpenTelemetry or any SDK-level instrumentation. There is no tracing, no span export, and no OpenTelemetry dependency. Evals run as isolated batch processes, not as instrumented services.

**Online vs offline evals.** All evals are offline batch workloads. The `oaieval` CLI runs a single eval process that loads samples, queries a model, and records results. There is no online evaluation mode, no streaming pipeline, and no live traffic evaluation.

**What is captured.** During execution, `RecorderBase` stores typed events in memory and flushes them periodically: `match` (correct/incorrect comparisons), `sampling` (prompts and model outputs), `embedding`, `cond_logp`, `pick_option`, `function_call`, `metrics`, `error`, and `extra` (`evals/record.py:44-70`). Each event carries `run_id`, `sample_id`, `type`, `data`, `created_by`, and `created_at`. Token usage from sampling events is extracted after the run and folded into the final result dict (`evals/cli/oaieval.py:269-294`).

**Feedback and annotation.** There is no human feedback or annotation system. The `record_extra` method (`evals/record.py:259-260`) could be used as a generic escape hatch to log arbitrary data during an eval, but no tooling exists to collect, view, or manage annotations.

**Linked production data.** There is no mechanism to import or link production traces to evaluation runs. The Snowflake `Recorder` stores runs and events in relational tables with a `run_id` foreign key (`evals/record.py:492-511`), but this is only for the evaluation run results, not production traffic.


Citations: [evals/record.py:44-90](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L44-L90) · [evals/record.py:260-270](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L260-L270) · [evals/record.py:468-512](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/record.py#L468-L512) · [evals/cli/oaieval.py:269-295](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/cli/oaieval.py#L269-L295)

### confident-ai/deepeval (answered)

DeepEval has a comprehensive tracing subsystem that captures LLM calls, agent steps, retrievals, and tool invocations as spans within traces. The `TraceManager` (`deepeval/tracing/tracing.py:1-80`) manages a tree of spans of types `BaseSpan`, `LlmSpan`, `AgentSpan`, `RetrieverSpan`, and `ToolSpan` (`deepeval/tracing/types.py`). **SDK instrumentation** is automatic: `patch_openai_client()` and `patch_anthropic_client()` monkey-patch the SDK client methods (`chat.completions.create`, etc.) to wrap each call in a span (`deepeval/tracing/patchers.py:16-60`). The `@observe` decorator or `with trace(...)` context manager mark arbitrary Python functions as traceable spans. **OpenTelemetry integration** is two-way: `ConfidentSpanExporter` implements `SpanExporter` to push deepeval spans into the OTel ecosystem, and `ContextAwareSpanProcessor` routes incoming OTel spans either through the REST-based trace manager path (when inside a deepeval trace context) or directly to Confident AI's OTLP endpoint (`deepeval/tracing/otel/exporter.py:1-60`, `deepeval/tracing/otel/context_aware_processor.py:1-60`). Supported frameworks (LangChain, CrewAI, LlamaIndex, OpenAI Agents, PydanticAI, Google ADK, Mastra) get automatic OTel integration (`deepeval/tracing/integrations.py:6-22`). **Online evals**: traces collected during a `trace()` session are automatically linked to evaluation when `assert_test()` is called from within the context. **Offline evals**: `evaluate_trace()`, `evaluate_span()`, and `evaluate_thread()` in `deepeval/tracing/offline_evals/` can retrospectively evaluate a stored trace against a metric collection (`deepeval/tracing/offline_evals/trace.py:6-36`). Trace data (input, output, timestamps, metadata) is serialized as `BaseApiSpan` Pydantic models and uploaded to Confident AI for dashboards. Feedback and labeling happen through the Confident AI platform rather than the open-source library itself.


Citations: [deepeval/tracing/tracing.py:1-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/tracing.py#L1-L80) · [deepeval/tracing/patchers.py:16-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/patchers.py#L16-L60) · [deepeval/tracing/otel/exporter.py:1-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/exporter.py#L1-L60) · [deepeval/tracing/otel/context_aware_processor.py:1-60](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/otel/context_aware_processor.py#L1-L60) · [deepeval/tracing/integrations.py:1-22](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/integrations.py#L1-L22) · [deepeval/tracing/offline_evals/trace.py:6-36](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/tracing/offline_evals/trace.py#L6-L36)

### vibrantlabsai/ragas (answered)

Ragas provides built-in tracing via a callback-based architecture and first-party integrations with two tracing platforms.

**Built-in tracing**: The `RagasTracer` (callbacks.py:80-121) is a `BaseCallbackHandler` that records every evaluation run as a tree of `ChainRun` nodes. Each evaluation creates a tree: evaluation → row → metric → prompt, with parent-child relationships tracked via `run_id`/`parent_run_id`. This captures inputs, outputs, and metadata at every level. Traces are stored in memory during execution and attached to `EvaluationResult.ragas_traces` for post-hoc analysis. `parse_run_traces()` traverses the tree to produce flat per-row lists of metric scores and prompt I/O (callbacks.py:134-173).

**Tracing integrations** (integrations/tracing/): `LangfuseTrace` and `MLflowTrace` provide explicit platform-level observability. The `@observe()` decorator wraps evaluation functions so every metric computation and LLM call is recorded in the tracing backend. `sync_trace()` returns a trace object with `.get_url()` for deep linking into the observability platform (integrations/tracing/__init__.py:1-78). These are optional dependencies — imported lazily.

**LLM observability integrations**: `HeliconeConfig` proxies LLM calls through Helicone for request logging. `Langsmith` and `Opik` integrations are available (integrations/langsmith.py, integrations/opik.py) for tracking LLM calls and evaluation results. Framework integrations (LangChain, LlamaIndex, Griptape, LangGraph) bridge those ecosystems' own observability.

**Online vs offline evals**: No distinction exists — all evaluations are offline/batch by design. There is no production traffic capturing, SDK instrumentation for live applications, or OpenTelemetry integration. The project integrates with no APM/vendor for online monitoring of deployed LLM applications.

**Feedback and annotation**: The `PromptAnnotation`/`SampleAnnotation`/`MetricAnnotation` classes (dataset_schema.py:555-837) support human-in-the-loop annotation for metric training: users can edit prompt outputs, mark samples as accepted/rejected, and use this data to optimize instruction prompts or few-shot demonstrations via `MetricWithLLM.train()`. This is a training feedback loop, not production feedback collection.


Citations: [src/ragas/integrations/tracing/__init__.py:1-78](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/integrations/tracing/__init__.py#L1-L78) · [src/ragas/dataset_schema.py:555-660](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L555-L660) · [src/ragas/metrics/base.py:191-306](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L191-L306) · [src/ragas/integrations/helicone.py:1-30](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/integrations/helicone.py#L1-L30)

### EleutherAI/lm-evaluation-harness (answered)

The harness does **not** have built-in OpenTelemetry instrumentation, SDK tracing, or automated production trace capture. Evaluation observability is limited to: **Logging** — the `eval_logger` (standard Python logging) throughout the codebase, plus environment info via `add_env_info()` and tokenizer info via `add_tokenizer_info()` in `lm_eval/evaluator.py:421-422`. **WandbLogger** at `lm_eval/loggers/wandb_logger.py:24-60` integrates with Weights & Biases for experiment tracking, logging aggregated metrics and config. **TrackioLogger** at `lm_eval/loggers/trackio_logger.py` provides a lightweight local-first alternative with per-sample `Trace` objects that capture prompt/response pairs as conversational traces with gold targets and metric values attached as metadata. The `_sample_to_trace()` helper at `lm_eval/loggers/trackio_logger.py:15-54` converts eval samples into `trackio.Trace` objects with the natural prompt/response for each output type (loglikelihood, multiple_choice, generate_until). **Result files** are JSON with full config, git hash, environment, tokenizer info, task versions, and per-task hashes for reproducibility (`result_schema.py` defines the schema). **Online vs offline evals**: The harness is fundamentally offline — tasks are loaded from datasets, inference is run, metrics are computed. No production data ingestion or online eval triggers exist. **Feedback/annotation**: Not present — there is no annotation UI, no feedback API, and no mechanism to incorporate human feedback into eval results within the framework itself.


Citations: [lm_eval/loggers/trackio_logger.py:15-80](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/trackio_logger.py#L15-L80) · [lm_eval/loggers/wandb_logger.py:24-70](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/wandb_logger.py#L24-L70) · [lm_eval/loggers/evaluation_tracker.py:230-300](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/loggers/evaluation_tracker.py#L230-L300) · [lm_eval/evaluator.py:420-425](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/evaluator.py#L420-L425)

### Arize-ai/phoenix (answered)

Observability is Phoenix's primary domain, built on **OpenTelemetry** with **OpenInference semantic conventions**. The `phoenix-otel` package (`packages/phoenix-otel/src/phoenix/otel/otel.py`) provides a `register()` function that creates an OpenTelemetry `TracerProvider` configured to export spans to the Phoenix collector via OTLP (gRPC or HTTP/protobuf). Spans follow the `openinference.semconv.trace` schema with `SpanAttributes` for LLM-specific fields like `INPUT_VALUE`, `OUTPUT_VALUE`, `LLM_MODEL_NAME`, `LLM_TOKEN_COUNT_PROMPT`, `LLM_TOKEN_COUNT_COMPLETION`, `LLM_TOKEN_COUNT_TOTAL`, and `OPENINFERENCE_SPAN_KIND` (`src/phoenix/trace/otel.py:1-57`). The **trace module** (`src/phoenix/trace/`) handles span ingestion via OTLP protobuf decoding, attribute flattening/unflattening (`attributes.py`), conversion between protobuf and the internal `Span` schema (`schemas.py`). SDK instrumentation is documented for Python, JavaScript, and Java. **Online vs offline evals**: Evaluation results are linked to traces via `span_id` in the `Score.metadata` (`evaluators.py:273-291`). The online-evals runner (`evals/pxi/online_evals/run.py`) discovers spans, fetches full traces, runs LLM judges, and writes `SpanAnnotationData` back to the Phoenix server via `client.spans.log_span_annotations()`. The Phoenix server stores these annotations and serves them via GraphQL and REST APIs. **Trace-level tracing**: The `@trace` decorator (`tracing.py:96-332`) wraps evaluator and LLM calls with OpenInference spans, recording input/output values and span metadata. It uses `OITracer` from `openinference.instrumentation` and auto-injects `trace_id` from the span context into evaluator calls. The `Score._add_trace_id_to_scores()` method embeds the trace ID into each score's metadata for traceability. **Feedback and annotation**: The `to_annotation_dataframe()` utility in `utils.py:410-491` formats evaluation results as annotations (span_id, score, label, explanation, annotation_name, annotator_kind) for logging back to Phoenix.


Citations: [packages/phoenix-otel/src/phoenix/otel/otel.py:65-120](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-otel/src/phoenix/otel/otel.py#L65-L120) · [src/phoenix/trace/otel.py:1-57](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/trace/otel.py#L1-L57) · [packages/phoenix-evals/src/phoenix/evals/tracing.py:96-332](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/tracing.py#L96-L332) · [evals/pxi/online_evals/run.py:196-210](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L196-L210) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:273-291](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L273-L291) · [packages/phoenix-evals/src/phoenix/evals/utils.py:410-491](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L410-L491)

### NVIDIA/garak (not applicable)

garak has no SDK instrumentation, OpenTelemetry integration, online/offline evaluation distinction, production trace capture, or annotation/feedback collection system. It is purely an offline evaluation framework: it runs probes/detectors against a model endpoint and writes results to a JSONL report file. The closest feature is the `hitlog.jsonl` which captures individual failure details at eval time, but this is an output artifact, not a production observability integration.



### UKGovernmentBEIS/inspect_ai (answered)

**SDK instrumentation** is provided through the `events` module and the hooks system, not through OpenTelemetry.

**Event system** (`src/inspect_ai/event/`). During sample execution, a timeline of typed events is recorded: `ModelEvent`, `ToolEvent`, `ScoreEvent`, `SandboxEvent`, `StateEvent`, `StepEvent`, `SubtaskEvent`, `InputEvent`, `ApprovalEvent`, `ErrorEvent`, `StoreEvent`, `SpanEvent`, and more. These are organized into an `EventTree` modeling the hierarchical sample run. The `Timeline` builds a linearized, filterable view.

**Hooks** (`src/inspect_ai/hooks/`). Lifecycle hooks fire at `eval_set_start/end`, `run_start/end`, `task_start/end`, `sample_start/init/attempt_start/attempt_end/end/scoring`, and `before_model_generate`. Hooks are passed to `eval()`.

**Trace logging** (`src/inspect_ai/_util/trace.py`). A custom `TRACE` log level captures HTTP requests, model calls, and long-running actions. The `inspect trace` CLI (`src/inspect_ai/_cli/trace.py`) lists, dumps, filters traces by action type, and surfaces anomalies.

**No OpenTelemetry.** There is no `opentelemetry` dependency or exporter. The tracing system is file-based JSONL.

**Online vs offline evals.** "Online" evals run via `eval()` and write to log files in real-time. "Offline" analysis reads logged `EvalLog` objects via `read_eval_log()`, recomputes metrics via `recompute_metrics()`, or edits scores via `edit_score()`.

**Feedback and annotation.** The `precomputed_scores` scorer attaches externally computed scores by sample ID. `edit_score()` modifies logged scores with provenance tracking. There is no built-in annotation UI.


Citations: [src/inspect_ai/event/__init__.py:1-91](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/event/__init__.py#L1-L91) · [src/inspect_ai/hooks/_hooks.py:1-61](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/hooks/_hooks.py#L1-L61) · [src/inspect_ai/_util/trace.py:1-120](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_util/trace.py#L1-L120) · [src/inspect_ai/_cli/trace.py:1-50](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/_cli/trace.py#L1-L50) · [src/inspect_ai/log/_log.py:95-180](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/log/_log.py#L95-L180)
