# langfuse/langfuse

> Self-hostable LLM tracing and eval platform: Next.js API, BullMQ worker, ClickHouse traces, and queued LLM-judge and code evaluators.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/langfuse/langfuse (reviewed at commit `1a21a4229c98dd0dbba2916631ca5fe3e05fd504`, 2026-10-06)
- Stars: 35452 · Language: TypeScript · License: MIT (ee/ folders under a commercial licence)
- Canonical page: https://llms-technical-reviews.com/p/langfuse/

## Overview

Langfuse is a server platform for LLM observability and evaluation. Your application sends traces (through the Langfuse SDKs, which live in separate repositories, or through any OpenTelemetry exporter). The platform stores them, shows them in a web UI, and can score them automatically with evaluators that run in a background worker. Datasets, prompt experiments, annotation queues and prompt management sit on top of the same trace and score store.

This repository is the server side only: a Next.js app in `web/` (UI, tRPC routes and the public REST and OTLP APIs), a Node worker in `worker/` that consumes BullMQ queues, shared code in `packages/shared/` (Prisma schema, ClickHouse repositories, eval execution), and a small Rust AI gateway in `ai-gateway/` that proxies OpenAI and Anthropic calls and logs them as generations. Evaluation is not a library you call from a test. It is a rule engine: you define an evaluator once, attach it to a filter and a sampling rate, and the worker scores matching traces, observations or dataset-run items as they arrive.

**Licence.** The repository is open core. The root [LICENSE](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/LICENSE#L1-L7) puts everything under MIT except `ee/`, `web/src/ee/` and `worker/src/ee/`, which fall under the commercial [ee/LICENSE](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/ee/LICENSE#L1-L11). Those folders hold SSO configuration, audit-log viewer, admin API, UI customisation, billing, usage metering and data retention. The tracing, ingestion and evaluation cores described below are all outside them. A `packages/shared/src/server/ee/` folder (ingestion masking, licence check) also exists, but the root LICENSE does not list that path. The plan table grants self-hosted enterprise RBAC, audit logs, data retention, protected prompt labels, the admin API and UI customisation. The `oss` plan has no limit on evaluators or annotation queues ([entitlements.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/entitlements/constants/entitlements.ts#L138-L181)).

## Architecture

```mermaid
flowchart LR
  SDK["SDKs / OTel exporters"] --> API["web: /api/public/ingestion, /otel/v1/traces"]
  API --> S3["Blob store (event JSON)"]
  API --> IQ["IngestionQueue / OtelIngestionQueue"]
  IQ --> ING["worker: IngestionService.mergeAndWrite"]
  ING --> CH["ClickHouse: traces, observations, scores"]
  ING --> TUQ["TraceUpsert queue"]
  IQ --> OBS["scheduleObservationEvals"]
  TUQ --> CREATE["createEvalJobs"]
  CREATE --> PG["Postgres: job_executions"]
  CREATE --> EXQ["EvalExecution queue"]
  OBS --> OEQ["LLM-judge / code-eval queues"]
  EXQ --> RUN["runLLMAsJudgeEvaluation"]
  OEQ --> RUN2["processObservationEval"]
  RUN --> DONE["completeEvalExecution"]
  RUN2 --> DONE
  DONE --> IQ
```

| Component | Path | Role |
|---|---|---|
| Public ingestion API | `web/src/pages/api/public/ingestion.ts` | Validates batches, rate-limits, hands off to `processEventBatch` |
| OTLP endpoint | `web/src/pages/api/public/otel/v1/traces/index.ts` | Accepts OTLP JSON or protobuf (optionally gzipped), publishes to the OTel queue |
| Batch processor | `packages/shared/src/server/ingestion/processEventBatch.ts` | Uploads each event to S3-compatible storage, then enqueues one job per entity |
| Ingestion worker | `worker/src/queues/ingestionQueue.ts`, `worker/src/services/IngestionService/` | Re-reads events from storage, merges them with the stored row, writes to ClickHouse |
| Eval scheduler | `worker/src/features/evaluation/evalService.ts` | Matches rules to traces/dataset items, dedups, samples, enqueues executions |
| Observation evals | `worker/src/features/evaluation/observationEval/` | Schedules and runs evaluators on single observations (spans, generations) |
| Evaluator runtimes | `packages/shared/src/server/evals/` | LLM-judge prompt compilation, code-eval dispatchers, decision-model mapping |
| Queue wiring | `worker/src/app.ts`, `worker/src/queues/evalQueue.ts` | Registers BullMQ processors and retry/redirect logic |
| Enterprise code | `ee/`, `web/src/ee/`, `worker/src/ee/` | Commercially licensed periphery (SSO, audit logs, billing, retention) |

## How a request flows

Take an OTel span from your app, scored by an LLM judge:

1. **Accept.** The OTLP route reads the body (with a size cap and gunzip), checks that ingestion is not suspended, and calls `processOtelIngestion` ([traces/index.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/pages/api/public/otel/v1/traces/index.ts#L34-L133)). That function uploads the raw resource spans to blob storage and enqueues an `OtelIngestionJob`.
2. **Convert.** The worker's OTel processor turns spans into Langfuse ingestion events with `OtelIngestionProcessor.processToIngestionEvents`. Observations go straight to `IngestionService`. Trace-level events take another pass through `processEventBatch` ([otelIngestionQueue.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/queues/otelIngestionQueue.ts#L536-L560), [L760-L785](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/queues/otelIngestionQueue.ts#L760-L785)).
3. **Store first, process later.** `processEventBatch` writes each event to the blob store, aborts if any upload failed, applies trace sampling, and adds one `IngestionJob` per entity to a sharded queue with a short delay ([processEventBatch.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/ingestion/processEventBatch.ts#L317-L401)). The blob store is the source of truth for replays.
4. **Merge and write.** The ingestion worker lists or downloads every stored event for that entity, sorts by upload time, and calls `IngestionService.mergeAndWrite` ([ingestionQueue.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/queues/ingestionQueue.ts#L160-L200)). For a trace, the service reads the current ClickHouse row, merges, queues the insert, and upserts the session in Postgres ([IngestionService](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/services/IngestionService/index.ts#L840-L952)).
5. **Trigger evals.** Unless a Redis cache says the project has no trace-based evaluators, the trace goes onto `TraceUpsertQueue` (same file). For observations, `processOtelEvents` calls `scheduleObservationEvals` directly ([processOtelEvents.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/otel-ingestion/processOtelEvents.ts#L103-L146)).
6. **Create jobs.** `createEvalJobs` loads active `evaluationRule` rows for trace and dataset targets, skips traces from internal `langfuse-*` environments to avoid eval-on-eval loops, and checks each rule's filter ([evalService.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L302-L380)). For a match it derives a deterministic job id, applies the sampling rate, inserts a `PENDING` `job_execution` with `skipDuplicates`, and enqueues it on `EvalExecutionQueue` with the rule's delay ([L763-L838](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L763-L838)).
7. **Execute.** `evaluate` reloads the job, resolves the rule and template, extracts variables from the trace (and dataset item, at its `validFrom` version), re-checks the environment and calls `executeLLMAsJudgeEvaluation` ([L1432-L1563](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L1432-L1563)). `runLLMAsJudgeEvaluation` validates the output definition, resolves the model connection and calls the LLM with a structured-output schema. The call itself is traced into the same project under the `langfuse-llm-as-a-judge` environment ([L1087-L1151](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L1087-L1151)).
8. **Persist the score.** `completeEvalExecution` uploads each score as an event file, enqueues it on the normal ingestion queue, and marks the job `COMPLETED` with the score id and execution trace id ([evalCompletion.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalCompletion.ts#L23-L98)). Scores travel the same path as user-submitted scores.

## Key components

### Evaluator types

The Prisma enum has `LLM_AS_JUDGE`, `CODE`, `DECISION_MODEL` and `FACET`. They do not run on the same paths. Trace- and dataset-level rules only load LLM-judge evaluators, and `evaluate` only runs the LLM-judge branch. Code and decision-model evaluators run in `processObservationEval`, which switches on the template type ([observationEvalProcessor.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/observationEval/observationEvalProcessor.ts#L267-L297)). Facets are rejected as evaluators ([L436-L499](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/observationEval/observationEvalProcessor.ts#L436-L499)). LLM-judge output can be numeric, boolean or categorical, validated by Zod and normalised into one or more scores.

### Code evaluators

`resolveConfiguredCodeEvalDispatcher` picks `aws-lambda` (separate Python and Node functions) or `insecure-local`. Outside development and test, it returns `null` unless `LANGFUSE_CODE_EVAL_DISPATCHER` is set ([codeEvalDispatchers.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/codeEvalDispatchers.ts#L13-L43)). The local dispatcher strips TypeScript types and runs the code in a `node:vm` context inside the worker process, TypeScript only ([localCodeEvalDispatcher.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/localCodeEvalDispatcher.ts#L15-L60)). `node:vm` is not a security boundary. A self-hoster who wants Python evaluators or untrusted code needs the Lambda route.

### Decision models

A decision-model evaluator asks typed questions (choice, score, boolean) of a connection whose adapter is `typesafe`. It returns probabilities and confidence that land in score metadata. Any other connection type blocks the evaluator ([runDecisionModelEvaluation.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/decisionModel/runDecisionModelEvaluation.ts#L105-L141), [types.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/llm/types.ts#L247-L262)).

### Sampling, dedup and auto-blocking

Sampling hashes the target id with SHA-256 under a fixed domain prefix and compares the result to the rate, so the same trace is always in or out ([deterministicSampling.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/deterministicSampling.ts#L1-L30)). Job ids derive from (project, rule, trace, dataset item, observation), and the Postgres primary key lets only one of two racing producers enqueue. When a judge call fails with an auth, billing or bad-config error, the evaluator is paused for every rule that uses it rather than retried forever ([evalService.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L941-L1060)).

### Prompt experiments

`experimentServiceClickhouse.ts` runs a stored prompt over a dataset. For each item it creates a dataset-run-item event with a deterministic trace id through the same `processEventBatch` path, calls the model, and schedules observation evals on the result ([experimentServiceClickhouse.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/experiments/experimentServiceClickhouse.ts#L74-L115)). Experiments driven from your own code go through the SDKs and arrive as ordinary traces linked to dataset items.

## Extending it

- **Evaluator templates.** LLM-judge templates are prompts with `{{variables}}` plus an output definition. Variable mappings point at trace, observation or dataset-item fields. A managed catalogue ships starter templates.
- **Code evaluators.** Write a TypeScript or Python function that returns `{scores: [...]}`. Python requires the Lambda dispatcher.
- **Model connections.** Judges use per-project LLM connections. The adapters are OpenAI, Anthropic, Azure, Bedrock, Vertex AI, Google AI Studio and TypeSafe.
- **Ingestion.** Anything that speaks OTLP over HTTP can send traces. The public REST API also accepts scores, so external eval harnesses can write results next to the built-in ones.

## Running it

`docker-compose.yml` runs six services: `langfuse-web`, `langfuse-worker`, ClickHouse, Postgres, Redis and MinIO (S3-compatible storage). All are required. Postgres holds configuration, rules and job executions. ClickHouse holds traces, observations and scores. Redis backs BullMQ. Blob storage holds every raw event. Without an enterprise licence key, a self-hosted instance falls back to the `oss` plan ([getPlan.ts](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/entitlements/server/getPlan.ts#L63-L69)). That plan includes tracing, datasets, evaluators and experiments.

## Strengths and caveats

- **Strength: eval as infrastructure.** Rules score live production traffic and experiment runs on the same queue path, with deterministic sampling, dedup, delays and per-evaluator pausing. This is closer to a production scoring service than to a test runner.
- **Strength: auditable judges.** Every judge call is itself a trace in your project, linked from the score by `executionTraceId`, so you can debug the evaluator with the tool you use for your app.
- **Strength: durable ingestion.** Raw events are written to blob storage before processing, and later events merge into earlier rows, which tolerates out-of-order SDK flushes.
- **Caveat: heavy to self-host.** Six services, two databases, and an ingestion path that depends on object storage. This is not something you embed in CI.
- **Caveat: no built-in metric library.** There are no BLEU, ROUGE or embedding metrics. Scores come from LLM judges, your code, decision models, humans or the API.
- **Caveat: evaluator types are uneven.** Code and decision-model evaluators only run in the observation/experiment path. Code evals are off in production until a dispatcher is configured.
- **Caveat: no red-teaming.** Safety is covered by a few judge templates (prompt-injection detection, rule adherence), not by attack generation.
- **Caveat: SDKs live elsewhere.** The Python and JS tracing SDKs are not in this repository, so this review covers their server contract, not their internals.

*Sources: code at 1a21a42, deepwiki-open wiki (10 pages), verified Q&A.*

## How langfuse/langfuse answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Langfuse provides **three evaluator types** that produce scores, plus deterministic sampling and blocking. **LLM-as-a-judge** (`evalService.ts` lines 941–1233) — the primary mechanism — constructs a prompt from a template, calls a configurable LLM (OpenAI, Anthropic, etc.) with Zod-validated structured output, and normalizes the response into **NUMERIC**, **BOOLEAN**, or **CATEGORICAL** scores (`outputDefinition.ts` lines 1–428). The output definition is a persisted schema (`PersistedEvalOutputDefinitionSchema`) supporting numeric ranges (`minValue`/`maxValue`), boolean verdicts, and categorical choices (single or multi-match). The `toNormalizedScores` helper (`evalService.ts` lines 1235–1270) converts the LLM response — e.g. a numeric 0–1, a boolean true/false, or a string array — into `CodeEvalScoreWithName[]` records with a shared comment (the model's reasoning). Each score carries a `dataType`, `value`, and optionally `configId` and `metadata`. **Code-based evaluators** (`codeEvalExecution.ts`) dispatch user-supplied code (JS/Python) to either an AWS Lambda dispatcher or a local process, returning scores in a `{scores: [...]}` JSON contract — NO built-in heuristic or statistical metrics exist outside these three types. **Decision-model evaluators** (`decisionModelEvaluatorExecution.ts`) call a separate "TypeSafe" API for choice/score/boolean questions and map responses into scores. **Deterministic sampling** (`deterministicSampling.ts` lines 1–31) controls what fraction of traces are scored, using a SHA‑256 over target ID and sampling rate. **Per-trace vs per-observation scoring** is handled via `targetObject`: `TRACE` scores the whole trace at the trace level, `DATASET` ties scores to dataset items, and `EVENT` (observation-level) schedules per-observation evaluations (`observationEval/`). All scores are persisted with `source: "EVAL"` and share an `EvalScoreWritePayload` structure (`evalScoreEvent.ts`).

> **Editor's note.** Correction: code and decision-model evaluators only run in the observation-level and experiment path (`processObservationEval`); trace- and dataset-level rules load only LLM-as-judge evaluators. The `insecure-local` code dispatcher runs TypeScript only, in a `node:vm` context inside the worker, and production has no code-eval dispatcher unless `LANGFUSE_CODE_EVAL_DISPATCHER` is set.

Citations: [worker/src/features/evaluation/evalService.ts:941-1270](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L941-L1270) · [packages/shared/src/server/evals/codeEvalExecution.ts:210-320](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/codeEvalExecution.ts#L210-L320) · [packages/shared/src/server/evals/decisionModelEvaluatorExecution.ts:252-307](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/decisionModelEvaluatorExecution.ts#L252-L307) · [worker/src/features/evaluation/evalScoreEvent.ts:1-68](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalScoreEvent.ts#L1-L68)

### How is LLM-as-a-judge implemented? (answered)

LLM-as-a-judge is implemented through a chain of modules converging on `runLLMAsJudgeEvaluation` in `evalService.ts` (lines 941–1233). **Judge prompts and rubrics** are stored as `EvalTemplate` records with a `prompt` (template string using `{{variable}}` syntax) and `promptMessages` (an array of system/user/assistant messages). The prompt is compiled by substituting extracted variables — trace fields like `input`, `output`, or dataset-item columns — via `compileEvalPrompt` in `llmEvaluatorExecution.ts` (lines 40–57). **Structured output** is enforced through Zod schemas: the `outputDefinition` stored on the template (legacy `{score, reasoning}` strings or v2 `{dataType, score, reasoning}` objects) is compiled into a Zod schema via `compilePersistedEvalOutputDefinition` (`outputDefinition.ts` lines 362–375). The schema is serialized to JSON Schema and passed to the AI SDK, which constraints the LLM's response shape. Results are validated with `validateEvalOutputResult` — numeric ranges, allowed categorical values, duplicates detection, and boolean type strictness are all checked (test file lines 541–679). **Judge model choice** is per-evaluator: each template stores `provider`, `model`, and `modelParams` which are resolved through `fetchModelConfig` (`evalExecutionDeps.ts` lines 327–337). The system supports OpenAI, Anthropic, Google, AWS Bedrock, and any OpenAI-compatible endpoint. **No built-in multi-sample/consensus** — each eval job produces a single LLM call. **No explicit calibration** exists, but **bias controls** are available through the prompt template itself (the user authors the rubric). **Auto-blocking** (`classifyEvaluatorLlmError.ts`) is key: when the judge LLM fails (auth, billing, endpoint unreachable), the evaluator is automatically paused at the database level so all rules using it stop until the user intervenes (lines 60–84). **Internal tracing** records each judge execution as an `LLMJudge`-environment trace in the user's project for debugging.


Citations: [worker/src/features/evaluation/evalService.ts:941-1233](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L941-L1233) · [packages/shared/src/server/evals/llmEvaluatorExecution.ts:40-87](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/llmEvaluatorExecution.ts#L40-L87) · [packages/shared/src/features/evals/outputDefinition.ts:286-427](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/features/evals/outputDefinition.ts#L286-L427) · [worker/src/features/evaluation/evalExecutionDeps.ts:245-325](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionDeps.ts#L245-L325) · [packages/shared/src/server/evals/classifyEvaluatorLlmError.ts:60-163](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/classifyEvaluatorLlmError.ts#L60-L163) · [worker/src/features/evaluation/executeLLMAsJudgeEvaluation.test.ts:541-723](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/executeLLMAsJudgeEvaluation.test.ts#L541-L723)

### How are test datasets and cases defined, generated and versioned? (answered)

Datasets are stored as `Dataset` and `DatasetItem` records in PostgreSQL, with ClickHouse for analytics queries. A dataset has an `inputSchema` and `expectedOutputSchema` (optional JSON Schema) for validation (`dataset-items.ts` lines 62–72). **Dataset items** carry `input`, `expectedOutput`, `metadata`, plus a `sourceTraceId`/`sourceObservationId` when created from traced data. Items are **versioned** via a temporal model: `validFrom`/`validTo` columns enable point-in-time snapshots — fetching an item at a specific date retrieves the version valid then (code in `evalService.ts` `extractVariablesFromTracingData` at line 1635: `...(datasetItemValidFrom ? { validFrom: datasetItemValidFrom } : { validTo: null })`). **Dataset runs** (`dataset-runs.ts`) execute a prompt experiment against dataset items: the `experimentServiceClickhouse.ts` takes a prompt template, replaces variables with each dataset item's `input`, calls the LLM, and creates `DatasetRunItem` records linking the experiment's output traces back to the original dataset items (lines 74–78). **No synthetic data generation** DSL exists — items are created via the UI, the public API (`POST /api/public/dataset-items`), or batch CSV/JSON upload. **Golden sets** are just datasets used as eval benchmarks: you create a dataset, map its columns to evaluator variables in an evaluation rule, and the rule scores every new trace that matches (or you run a batch evaluation). Eval jobs can target `TRACE` (live scoring) or `DATASET` (scoring dataset-run items). **Benchmark task registration** is handled through the evaluation rules system (`evaluationRule` Prisma model) — rules connect evaluators (judge templates) to dataset items or traces via filters, sampling rates, and variable mappings.


Citations: [packages/shared/src/server/repositories/dataset-items.ts:1-80](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/repositories/dataset-items.ts#L1-L80) · [packages/shared/src/domain/dataset-items.ts:1-38](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/domain/dataset-items.ts#L1-L38) · [worker/src/features/evaluation/evalService.ts:1565-1708](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L1565-L1708) · [worker/src/features/experiments/experimentServiceClickhouse.ts:48-80](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/experiments/experimentServiceClickhouse.ts#L48-L80)

### How are evals executed and reported? (answered)

Eval execution is a **job-queue pipeline** built on BullMQ (Redis). **Scheduling** (`evalService.ts` `createEvalJobs` lines 302–913): when a trace is upserted or a dataset run item created, the system fetches active evaluation rules for the project, filters them against the trace (filter state and sampling), deduplicates via deterministic job IDs, and enqueues `EvalExecutionQueue` jobs. **Three parallel queues** exist for trace-level and observation-level evals, routed by `EvalExecutionQueue.getInstance({shardingKey})` (line 841). **Parallelism** is handled by BullMQ worker concurrency; no built-in batching per model provider. **Caching** (`deterministicSampling.ts`) uses SHA‑256 hashing of the target ID to produce a deterministic `[0,1)` value for sampling decisions. A trace-cache optimization (`evalService.ts` lines 389–442) fetches trace data once for multiple configs. **Execution** calls into `runLLMAsJudgeEvaluation` (or code/decision-model executors), which calls the LLM, validates the structured output, and persists scores. Results flow through `completeEvalExecution` (`evalCompletion.ts` lines 23–98): scores are uploaded to S3 as JSON event files, queued for ClickHouse ingestion via `IngestionQueue`, and the `JobExecution` row is updated to `COMPLETED` with the primary score ID and execution trace ID. **Error handling** (`evalExecutionMetrics.ts`) classifies outcomes into `success`, `platform_error`, `upstream_error`, `customer_error`, or `cancelled`. LLM errors are categorized and can auto-block the evaluator. **CI integration** is user-defined (no built-in harness); users call the public API or run evaluations ad-hoc via the UI's "Run evaluation" batch dialog. **Results** appear as scores on traces/observations in the Langfuse UI, filterable and searchable via the Eval Logs table (`eval-log.tsx`). **Dashboards and regression** are user-built on top of the score data — the platform stores all scores with timestamps, trace IDs, and metadata, enabling external charting or the built-in analytics views.

> **Editor's note.** Correction: `deterministicSampling.ts` is sampling, not caching: it hashes a domain prefix plus the target id with SHA-256 and compares the result to the rule's rate. Trace-level `evaluate` only executes LLM-as-judge templates; code and decision-model evaluators run through the observation-eval queues.

Citations: [worker/src/features/evaluation/evalService.ts:302-913](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L302-L913) · [worker/src/features/evaluation/evalCompletion.ts:23-98](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalCompletion.ts#L23-L98) · [worker/src/features/evaluation/evalExecutionMetrics.ts:1-141](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionMetrics.ts#L1-L141) · [web/src/features/evals/components/eval-log.tsx:1-60](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/components/eval-log.tsx#L1-L60)

### How are traces or production data captured and linked to evaluations? (answered)

Langfuse **is** an observability platform, so traces are captured via a first-party Python/JS/TS SDK (`langfuse` pip package, JS SDK in `packages/shared/src/server/llm/`), and it also ingests OpenTelemetry data through the OTel ingestion pipeline (`OtelIngestionProcessor.ts`). Traces consist of a root trace with nested observations (spans, generations, events), all stored in ClickHouse. **Linking evals to traces** works through the `executionTraceId` mechanism: when an LLM-as-a-judge evaluator runs, it creates a new *internal* trace (environment `langfuse-llm-as-a-judge` or `langfuse-code-eval`) in the same project, passing `traceSinkParams` with `targetProjectId`, `traceId`, and `metadata` (`evalExecutionDeps.ts` lines 302–312). This internal trace records the judge's input prompt, output score, model, latency, and reasoning. The score event's `executionTraceId` field links the score back to this internal trace. **Online vs offline evals**: "online" (trace-upsert-triggered) runs eval rules immediately when traces arrive; "offline" evaluations are triggered via the UI batch-action dialog or the public API with a timestamp range for historical data. **Feedback and annotation** is a separate feature: Langfuse has `scores` with `source: "ANNOTATION"` as distinct from `source: "EVAL"`, both stored in the same scores table but tagged differently. The scores repository (`scores.ts`) handles both types. **The eval log** (`eval-log.tsx`) provides a UI table of all eval job executions with status, score, comment, and links to the target trace. Each evaluation also writes its own internal trace with environment labels that are excluded from triggering further evaluations (the `isInternalEvalEnvironment` guard in `isEvalTargetEnvironmentAllowed.ts` prevents infinite eval loops).

> **Editor's note.** Correction: the Python and JS tracing SDKs are not in this repository; `packages/shared/src/server/llm/` is the server's own LLM-calling code.

Citations: [worker/src/features/evaluation/evalExecutionDeps.ts:245-325](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionDeps.ts#L245-L325) · [packages/shared/src/server/otel/OtelIngestionProcessor.ts:1-50](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/otel/OtelIngestionProcessor.ts#L1-L50) · [worker/src/features/evaluation/isEvalTargetEnvironmentAllowed.ts:1-30](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/isEvalTargetEnvironmentAllowed.ts#L1-L30) · [packages/shared/src/server/repositories/scores.ts:1-30](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/repositories/scores.ts#L1-L30) · [web/src/features/evals/components/eval-log.tsx:1-60](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/components/eval-log.tsx#L1-L60)

### Does it support red-teaming or safety testing, and how? (answered)

Langfuse does NOT have a dedicated red-teaming framework (no adversarial-probe runners, attack plugins, or automated jailbreak testing pipelines). Instead, it provides **evaluator templates in the `safety` category** of its managed templates catalog (`managedTemplatesCatalog.ts` lines 53–58, 1193–1297) that users can use for manual or automated safety testing. These include: **Detect Prompt Injection** — an LLM-as-a-judge template that classifies whether an input contains a credible manipulation attempt (override instructions, reveal system prompt, bypass safeguards), returning a boolean verdict (lines 1242–1297). **Check Rule Adherence** — a boolean evaluator that checks if an output follows a defined policy or instruction set (lines 1193–1241). The **safety category** description (lines 54–57) states it "Monitors policy adherence, privacy leakage, and adversarial prompts." These are **template starters** — users must adapt the prompt and deploy an evaluation rule that runs on their traces. There are no built-in safety-specific tools, no adversarial dataset generators, no automated red-team loops, and no vulnerability-reporting workflow (the repo's `SECURITY.md` at root only links to a generic security policy). Safety testing must be authored by the user: create a dataset of adversarial inputs, write an evaluator template (or use the managed one), and run batch evaluations against it. The platform stores all results and can track regressions over time via the scores API.


Citations: [web/src/features/evals/v2/constants/managedTemplatesCatalog.ts:53-58](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/v2/constants/managedTemplatesCatalog.ts#L53-L58) · [web/src/features/evals/v2/constants/managedTemplatesCatalog.ts:1193-1297](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/v2/constants/managedTemplatesCatalog.ts#L1193-L1297) · [web/src/features/evals/v2/constants/managedTemplatesCatalog.ts:1242-1297](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/web/src/features/evals/v2/constants/managedTemplatesCatalog.ts#L1242-L1297)
