LLMs Technical Reviews

How are traces or production data captured and linked to evaluations?

SDK instrumentation; OpenTelemetry; online vs offline evals; feedback and annotation; if absent, say so.

Verdict

Langfuse and Opik are the only projects that score production traces as they arrive. Phoenix has the cleanest OpenTelemetry ingest, but you script online scoring yourself. The other seven tools trace only their own eval runs, or nothing.

Trace platforms. Langfuse accepts OTLP and its own SDK events. Evaluation rules pick traces with a deterministic SHA-256 sample, so the same trace is always in or out. Each judge call is itself a trace, linked from the score by executionTraceId. Human annotations share the score table with source: ANNOTATION. Note that the SDKs are in separate repositories. Opik captures spans with @opik.track, an OTel span processor or its OTLP endpoint. A sampler using SecureRandom sends traces through Redis streams to LLM-judge or Python-metric scorers. Experiment items are ordinary traces, so offline and online scores sit side by side. Phoenix receives OTLP over gRPC or HTTP with OpenInference attributes. Scores, labels and code checks all become span annotations, and log_span_annotations attaches offline results to production spans. Its online runner in evals/pxi is internal and not shipped.

Libraries that trace evaluations. DeepEval has @observe and OpenAI and Anthropic patchers, plus an OTel processor that forwards spans to its vendor’s hosted platform. Online evaluation of production traffic is a feature of that platform. promptfoo starts a local OTLP receiver during a run so that assertions can check the spans your app emitted. It can also read traces from Langfuse, Braintrust or Tempo, but it never captures production traffic. Ragas records an evaluation → row → metric → prompt callback tree and has Langfuse and MLflow helpers. It has no OTel support.

Offline logs only. Inspect records typed events per sample and exposes lifecycle hooks, but has no OTel. lm-evaluation-harness logs to W&B or Trackio. OpenAI Evals records prompt and output events to JSONL. For garak the question is not applicable. It is a scanner with a report file.

Pick: Langfuse for rule-based online scoring with auditable judges. Pick: Opik when you also want a Python metric library on the same traces. Pick: Phoenix for standard OTLP ingest and annotation-driven review.

Per-project answers

langfuse/langfuse

answered

Langfuse is an observability platform, so traces are captured via a first-party Python/JS/TS SDK (langfuse pip package, JS SDK in packages/shared/src/server/llm/), and it also ingests OpenTelemetry data through the OTel ingestion pipeline (OtelIngestionProcessor.ts). Traces consist of a root trace with nested observations (spans, generations, events), all stored in ClickHouse. Linking evals to traces works through the executionTraceId mechanism: when an LLM-as-a-judge evaluator runs, it creates a new internal trace (environment langfuse-llm-as-a-judge or langfuse-code-eval) in the same project, passing traceSinkParams with targetProjectId, traceId, and metadata (evalExecutionDeps.ts lines 302–312). This internal trace records the judge's input prompt, output score, model, latency, and reasoning. The score event's executionTraceId field links the score back to this internal trace. Online vs offline evals: "online" (trace-upsert-triggered) runs eval rules immediately when traces arrive; "offline" evaluations are triggered via the UI batch-action dialog or the public API with a timestamp range for historical data. Feedback and annotation is a separate feature: Langfuse has scores with source: "ANNOTATION" as distinct from source: "EVAL", both stored in the same scores table but tagged differently. The scores repository (scores.ts) handles both types. The eval log (eval-log.tsx) provides a UI table of all eval job executions with status, score, comment, and links to the target trace. Each evaluation also writes its own internal trace with environment labels that are excluded from triggering further evaluations (the isInternalEvalEnvironment guard in isEvalTargetEnvironmentAllowed.ts prevents infinite eval loops).

Editor's note. Correction: the Python and JS tracing SDKs are not in this repository; packages/shared/src/server/llm/ is the server's own LLM-calling code.

promptfoo/promptfoo

answered

OpenTelemetry integration via src/tracing/otelSdk.ts and src/tracing/otelConfig.ts with OTLP HTTP export. Provider calls wrapped with withTracedProviderCall/withTestCaseSpan/withGraderSpan (src/tracing/targetTracer.ts). W3C traceparent headers generated and validated (evaluator.ts L1567-1583). Trace store (src/tracing/store.ts) manages SQLite-backed tracesTable and spansTable with retry-based stability detection (loadTraceData polls until span count stabilizes, src/assertions/index.ts L184-227). External trace providers: Braintrust, Langfuse, Tempo via TraceProvider interface (src/tracing/providers/index.ts L21-39). Traces linked to eval rows via traceId/evaluationId. No production traffic capture SDK exists; traces are generated only during eval execution.

comet-ml/opik

answered

Opik has deep observability built around its @opik.track decorator (decorator/tracker.py:11-99 and decorator/base_track_decorator.py:41-80).

SDK instrumentation: The @track decorator wraps any function (LLM calls, internal steps, entire tasks) to automatically create spans and traces. It captures inputs/outputs, start/end times, tags, metadata, project name, and error info. The OpikTrackDecorator has _start_span_inputs_preprocessor and _end_span_inputs_preprocessor hooks. Spans and traces have IDs, can be nested, and support update_current_span()/update_current_trace() for mid-execution metadata. @track can be applied transparently to any callable; the evaluation system auto-wraps task functions that are not already tracked (engine.py:279-281).

Context managers: start_as_current_span() and start_as_current_trace() provide explicit context-manager access for fine-grained manual tracing.

LLM provider integrations: Opik has dedicated integration modules under integrations/ covering: OpenAI, Anthropic, LangChain, LlamaIndex, Haystack, CrewAI, DSPy, Bedrock, Mistral, Ollama, Groq, Cerebras, AISuite, ADK, AgentSpec, Harbor, SageMaker, and Guardrails. Each logs the provider's request/response as Opik spans.

OpenTelemetry: The integrations/otel/ module provides OpikSpanProcessor — an OpenTelemetry span processor that converts OTel spans into Opik spans, enabling integration with any OTel-instrumented code.

Online vs offline evals: The default path connects to the Opik backend to log experiments and feedback scores. For offline use, record_traces_locally() (api_objects/local_recording.py) captures traces to an in-memory handle that can be inspected post-hoc without a backend.

Feedback and annotation: TracesAnnotationQueue and ThreadsAnnotationQueue (api_objects/annotation_queue.py, linked from __init__.py:2-4) support human annotation workflows. Feedback scores are the primary mechanism for recording evaluation results on traces.

Guardrails tracing: The guardrails/ module has its own GuardrailsTrackDecorator that logs guardrail validation results as special spans.

openai/evals

answered

SDK instrumentation and OpenTelemetry. This repository does not implement OpenTelemetry or any SDK-level instrumentation. There is no tracing, no span export, and no OpenTelemetry dependency. Evals run as isolated batch processes, not as instrumented services.

Online vs offline evals. All evals are offline batch workloads. The oaieval CLI runs a single eval process that loads samples, queries a model, and records results. There is no online evaluation mode, no streaming pipeline, and no live traffic evaluation.

What is captured. During execution, RecorderBase stores typed events in memory and flushes them periodically: match (correct/incorrect comparisons), sampling (prompts and model outputs), embedding, cond_logp, pick_option, function_call, metrics, error, and extra (evals/record.py:44-70). Each event carries run_id, sample_id, type, data, created_by, and created_at. Token usage from sampling events is extracted after the run and folded into the final result dict (evals/cli/oaieval.py:269-294).

Feedback and annotation. There is no human feedback or annotation system. The record_extra method (evals/record.py:259-260) could be used as a generic escape hatch to log arbitrary data during an eval, but no tooling exists to collect, view, or manage annotations.

Linked production data. There is no mechanism to import or link production traces to evaluation runs. The Snowflake Recorder stores runs and events in relational tables with a run_id foreign key (evals/record.py:492-511), but this is only for the evaluation run results, not production traffic.

confident-ai/deepeval

answered

DeepEval has a comprehensive tracing subsystem that captures LLM calls, agent steps, retrievals, and tool invocations as spans within traces. The TraceManager (deepeval/tracing/tracing.py:1-80) manages a tree of spans of types BaseSpan, LlmSpan, AgentSpan, RetrieverSpan, and ToolSpan (deepeval/tracing/types.py). SDK instrumentation is automatic: patch_openai_client() and patch_anthropic_client() monkey-patch the SDK client methods (chat.completions.create, etc.) to wrap each call in a span (deepeval/tracing/patchers.py:16-60). The @observe decorator or with trace(...) context manager mark arbitrary Python functions as traceable spans. OpenTelemetry integration is two-way: ConfidentSpanExporter implements SpanExporter to push deepeval spans into the OTel ecosystem, and ContextAwareSpanProcessor routes incoming OTel spans either through the REST-based trace manager path (when inside a deepeval trace context) or directly to Confident AI's OTLP endpoint (deepeval/tracing/otel/exporter.py:1-60, deepeval/tracing/otel/context_aware_processor.py:1-60). Supported frameworks (LangChain, CrewAI, LlamaIndex, OpenAI Agents, PydanticAI, Google ADK, Mastra) get automatic OTel integration (deepeval/tracing/integrations.py:6-22). Online evals: traces collected during a trace() session are automatically linked to evaluation when assert_test() is called from within the context. Offline evals: evaluate_trace(), evaluate_span(), and evaluate_thread() in deepeval/tracing/offline_evals/ can retrospectively evaluate a stored trace against a metric collection (deepeval/tracing/offline_evals/trace.py:6-36). Trace data (input, output, timestamps, metadata) is serialized as BaseApiSpan Pydantic models and uploaded to Confident AI for dashboards. Feedback and labeling happen through the Confident AI platform rather than the open-source library itself.

vibrantlabsai/ragas

answered

Ragas provides built-in tracing via a callback-based architecture and first-party integrations with two tracing platforms.

Built-in tracing: The RagasTracer (callbacks.py:80-121) is a BaseCallbackHandler that records every evaluation run as a tree of ChainRun nodes. Each evaluation creates a tree: evaluation → row → metric → prompt, with parent-child relationships tracked via run_id/parent_run_id. This captures inputs, outputs, and metadata at every level. Traces are stored in memory during execution and attached to EvaluationResult.ragas_traces for post-hoc analysis. parse_run_traces() traverses the tree to produce flat per-row lists of metric scores and prompt I/O (callbacks.py:134-173).

Tracing integrations (integrations/tracing/): LangfuseTrace and MLflowTrace provide explicit platform-level observability. The @observe() decorator wraps evaluation functions so every metric computation and LLM call is recorded in the tracing backend. sync_trace() returns a trace object with .get_url() for deep linking into the observability platform (integrations/tracing/init.py:1-78). These are optional dependencies — imported lazily.

LLM observability integrations: HeliconeConfig proxies LLM calls through Helicone for request logging. Langsmith and Opik integrations are available (integrations/langsmith.py, integrations/opik.py) for tracking LLM calls and evaluation results. Framework integrations (LangChain, LlamaIndex, Griptape, LangGraph) bridge those ecosystems' own observability.

Online vs offline evals: No distinction exists — all evaluations are offline/batch by design. There is no production traffic capturing, SDK instrumentation for live applications, or OpenTelemetry integration. The project integrates with no APM/vendor for online monitoring of deployed LLM applications.

Feedback and annotation: The PromptAnnotation/SampleAnnotation/MetricAnnotation classes (dataset_schema.py:555-837) support human-in-the-loop annotation for metric training: users can edit prompt outputs, mark samples as accepted/rejected, and use this data to optimize instruction prompts or few-shot demonstrations via MetricWithLLM.train(). This is a training feedback loop, not production feedback collection.

EleutherAI/lm-evaluation-harness

answered

The harness does not have built-in OpenTelemetry instrumentation, SDK tracing, or automated production trace capture. Evaluation observability is limited to: Logging — the eval_logger (standard Python logging) throughout the codebase, plus environment info via add_env_info() and tokenizer info via add_tokenizer_info() in lm_eval/evaluator.py:421-422. WandbLogger at lm_eval/loggers/wandb_logger.py:24-60 integrates with Weights & Biases for experiment tracking, logging aggregated metrics and config. TrackioLogger at lm_eval/loggers/trackio_logger.py provides a lightweight local-first alternative with per-sample Trace objects that capture prompt/response pairs as conversational traces with gold targets and metric values attached as metadata. The _sample_to_trace() helper at lm_eval/loggers/trackio_logger.py:15-54 converts eval samples into trackio.Trace objects with the natural prompt/response for each output type (loglikelihood, multiple_choice, generate_until). Result files are JSON with full config, git hash, environment, tokenizer info, task versions, and per-task hashes for reproducibility (result_schema.py defines the schema). Online vs offline evals: The harness is fundamentally offline — tasks are loaded from datasets, inference is run, metrics are computed. No production data ingestion or online eval triggers exist. Feedback/annotation: Not present — there is no annotation UI, no feedback API, and no mechanism to incorporate human feedback into eval results within the framework itself.

Arize-ai/phoenix

answered

Observability is Phoenix's primary domain, built on OpenTelemetry with OpenInference semantic conventions. The phoenix-otel package (packages/phoenix-otel/src/phoenix/otel/otel.py) provides a register() function that creates an OpenTelemetry TracerProvider configured to export spans to the Phoenix collector via OTLP (gRPC or HTTP/protobuf). Spans follow the openinference.semconv.trace schema with SpanAttributes for LLM-specific fields like INPUT_VALUE, OUTPUT_VALUE, LLM_MODEL_NAME, LLM_TOKEN_COUNT_PROMPT, LLM_TOKEN_COUNT_COMPLETION, LLM_TOKEN_COUNT_TOTAL, and OPENINFERENCE_SPAN_KIND (src/phoenix/trace/otel.py:1-57). The trace module (src/phoenix/trace/) handles span ingestion via OTLP protobuf decoding, attribute flattening/unflattening (attributes.py), conversion between protobuf and the internal Span schema (schemas.py). SDK instrumentation is documented for Python, JavaScript, and Java. Online vs offline evals: Evaluation results are linked to traces via span_id in the Score.metadata (evaluators.py:273-291). The online-evals runner (evals/pxi/online_evals/run.py) discovers spans, fetches full traces, runs LLM judges, and writes SpanAnnotationData back to the Phoenix server via client.spans.log_span_annotations(). The Phoenix server stores these annotations and serves them via GraphQL and REST APIs. Trace-level tracing: The @trace decorator (tracing.py:96-332) wraps evaluator and LLM calls with OpenInference spans, recording input/output values and span metadata. It uses OITracer from openinference.instrumentation and auto-injects trace_id from the span context into evaluator calls. The Score._add_trace_id_to_scores() method embeds the trace ID into each score's metadata for traceability. Feedback and annotation: The to_annotation_dataframe() utility in utils.py:410-491 formats evaluation results as annotations (span_id, score, label, explanation, annotation_name, annotator_kind) for logging back to Phoenix.

UKGovernmentBEIS/inspect_ai

answered

SDK instrumentation is provided through the events module and the hooks system, not through OpenTelemetry.

Event system (src/inspect_ai/event/). During sample execution, a timeline of typed events is recorded: ModelEvent, ToolEvent, ScoreEvent, SandboxEvent, StateEvent, StepEvent, SubtaskEvent, InputEvent, ApprovalEvent, ErrorEvent, StoreEvent, SpanEvent, and more. These are organized into an EventTree modeling the hierarchical sample run. The Timeline builds a linearized, filterable view.

Hooks (src/inspect_ai/hooks/). Lifecycle hooks fire at eval_set_start/end, run_start/end, task_start/end, sample_start/init/attempt_start/attempt_end/end/scoring, and before_model_generate. Hooks are passed to eval().

Trace logging (src/inspect_ai/_util/trace.py). A custom TRACE log level captures HTTP requests, model calls, and long-running actions. The inspect trace CLI (src/inspect_ai/_cli/trace.py) lists, dumps, filters traces by action type, and surfaces anomalies.

No OpenTelemetry. There is no opentelemetry dependency or exporter. The tracing system is file-based JSONL.

Online vs offline evals. "Online" evals run via eval() and write to log files in real-time. "Offline" analysis reads logged EvalLog objects via read_eval_log(), recomputes metrics via recompute_metrics(), or edits scores via edit_score().

Feedback and annotation. The precomputed_scores scorer attaches externally computed scores by sample ID. edit_score() modifies logged scores with provenance tracking. There is no built-in annotation UI.

NVIDIA/garak

not applicable

garak has no SDK instrumentation, OpenTelemetry integration, online/offline evaluation distinction, production trace capture, or annotation/feedback collection system. It is purely an offline evaluation framework: it runs probes/detectors against a model endpoint and writes results to a JSONL report file. The closest feature is the hitlog.jsonl which captures individual failure details at eval time, but this is an output artifact, not a production observability integration.

← How are evals executed and reported? · Does it support red-teaming or safety testing, and how? →