LLMs Technical Reviews

langfuse/langfuse

Self-hostable LLM tracing and eval platform: Next.js API, BullMQ worker, ClickHouse traces, and queued LLM-judge and code evaluators.

GitHub ↗★ 35kTypeScriptMIT (ee/ folders under a commercial licence)commit 1a21a42 · 2026-10-06homepage ↗

Overview

Langfuse is a server platform for LLM observability and evaluation. Your application sends traces (through the Langfuse SDKs, which live in separate repositories, or through any OpenTelemetry exporter). The platform stores them, shows them in a web UI, and can score them automatically with evaluators that run in a background worker. Datasets, prompt experiments, annotation queues and prompt management sit on top of the same trace and score store.

This repository is the server side only: a Next.js app in web/ (UI, tRPC routes and the public REST and OTLP APIs), a Node worker in worker/ that consumes BullMQ queues, shared code in packages/shared/ (Prisma schema, ClickHouse repositories, eval execution), and a small Rust AI gateway in ai-gateway/ that proxies OpenAI and Anthropic calls and logs them as generations. Evaluation is not a library you call from a test. It is a rule engine: you define an evaluator once, attach it to a filter and a sampling rate, and the worker scores matching traces, observations or dataset-run items as they arrive.

Licence. The repository is open core. The root LICENSE puts everything under MIT except ee/, web/src/ee/ and worker/src/ee/, which fall under the commercial ee/LICENSE. Those folders hold SSO configuration, audit-log viewer, admin API, UI customisation, billing, usage metering and data retention. The tracing, ingestion and evaluation cores described below are all outside them. A packages/shared/src/server/ee/ folder (ingestion masking, licence check) also exists, but the root LICENSE does not list that path. The plan table grants self-hosted enterprise RBAC, audit logs, data retention, protected prompt labels, the admin API and UI customisation. The oss plan has no limit on evaluators or annotation queues (entitlements.ts).

Architecture

flowchart LR
  SDK["SDKs / OTel exporters"] --> API["web: /api/public/ingestion, /otel/v1/traces"]
  API --> S3["Blob store (event JSON)"]
  API --> IQ["IngestionQueue / OtelIngestionQueue"]
  IQ --> ING["worker: IngestionService.mergeAndWrite"]
  ING --> CH["ClickHouse: traces, observations, scores"]
  ING --> TUQ["TraceUpsert queue"]
  IQ --> OBS["scheduleObservationEvals"]
  TUQ --> CREATE["createEvalJobs"]
  CREATE --> PG["Postgres: job_executions"]
  CREATE --> EXQ["EvalExecution queue"]
  OBS --> OEQ["LLM-judge / code-eval queues"]
  EXQ --> RUN["runLLMAsJudgeEvaluation"]
  OEQ --> RUN2["processObservationEval"]
  RUN --> DONE["completeEvalExecution"]
  RUN2 --> DONE
  DONE --> IQ
Component Path Role
Public ingestion API web/src/pages/api/public/ingestion.ts Validates batches, rate-limits, hands off to processEventBatch
OTLP endpoint web/src/pages/api/public/otel/v1/traces/index.ts Accepts OTLP JSON or protobuf (optionally gzipped), publishes to the OTel queue
Batch processor packages/shared/src/server/ingestion/processEventBatch.ts Uploads each event to S3-compatible storage, then enqueues one job per entity
Ingestion worker worker/src/queues/ingestionQueue.ts, worker/src/services/IngestionService/ Re-reads events from storage, merges them with the stored row, writes to ClickHouse
Eval scheduler worker/src/features/evaluation/evalService.ts Matches rules to traces/dataset items, dedups, samples, enqueues executions
Observation evals worker/src/features/evaluation/observationEval/ Schedules and runs evaluators on single observations (spans, generations)
Evaluator runtimes packages/shared/src/server/evals/ LLM-judge prompt compilation, code-eval dispatchers, decision-model mapping
Queue wiring worker/src/app.ts, worker/src/queues/evalQueue.ts Registers BullMQ processors and retry/redirect logic
Enterprise code ee/, web/src/ee/, worker/src/ee/ Commercially licensed periphery (SSO, audit logs, billing, retention)

How a request flows

Take an OTel span from your app, scored by an LLM judge:

  1. Accept. The OTLP route reads the body (with a size cap and gunzip), checks that ingestion is not suspended, and calls processOtelIngestion (traces/index.ts). That function uploads the raw resource spans to blob storage and enqueues an OtelIngestionJob.
  2. Convert. The worker’s OTel processor turns spans into Langfuse ingestion events with OtelIngestionProcessor.processToIngestionEvents. Observations go straight to IngestionService. Trace-level events take another pass through processEventBatch (otelIngestionQueue.ts, L760-L785).
  3. Store first, process later. processEventBatch writes each event to the blob store, aborts if any upload failed, applies trace sampling, and adds one IngestionJob per entity to a sharded queue with a short delay (processEventBatch.ts). The blob store is the source of truth for replays.
  4. Merge and write. The ingestion worker lists or downloads every stored event for that entity, sorts by upload time, and calls IngestionService.mergeAndWrite (ingestionQueue.ts). For a trace, the service reads the current ClickHouse row, merges, queues the insert, and upserts the session in Postgres (IngestionService).
  5. Trigger evals. Unless a Redis cache says the project has no trace-based evaluators, the trace goes onto TraceUpsertQueue (same file). For observations, processOtelEvents calls scheduleObservationEvals directly (processOtelEvents.ts).
  6. Create jobs. createEvalJobs loads active evaluationRule rows for trace and dataset targets, skips traces from internal langfuse-* environments to avoid eval-on-eval loops, and checks each rule’s filter (evalService.ts). For a match it derives a deterministic job id, applies the sampling rate, inserts a PENDING job_execution with skipDuplicates, and enqueues it on EvalExecutionQueue with the rule’s delay (L763-L838).
  7. Execute. evaluate reloads the job, resolves the rule and template, extracts variables from the trace (and dataset item, at its validFrom version), re-checks the environment and calls executeLLMAsJudgeEvaluation (L1432-L1563). runLLMAsJudgeEvaluation validates the output definition, resolves the model connection and calls the LLM with a structured-output schema. The call itself is traced into the same project under the langfuse-llm-as-a-judge environment (L1087-L1151).
  8. Persist the score. completeEvalExecution uploads each score as an event file, enqueues it on the normal ingestion queue, and marks the job COMPLETED with the score id and execution trace id (evalCompletion.ts). Scores travel the same path as user-submitted scores.

Key components

Evaluator types

The Prisma enum has LLM_AS_JUDGE, CODE, DECISION_MODEL and FACET. They do not run on the same paths. Trace- and dataset-level rules only load LLM-judge evaluators, and evaluate only runs the LLM-judge branch. Code and decision-model evaluators run in processObservationEval, which switches on the template type (observationEvalProcessor.ts). Facets are rejected as evaluators (L436-L499). LLM-judge output can be numeric, boolean or categorical, validated by Zod and normalised into one or more scores.

Code evaluators

resolveConfiguredCodeEvalDispatcher picks aws-lambda (separate Python and Node functions) or insecure-local. Outside development and test, it returns null unless LANGFUSE_CODE_EVAL_DISPATCHER is set (codeEvalDispatchers.ts). The local dispatcher strips TypeScript types and runs the code in a node:vm context inside the worker process, TypeScript only (localCodeEvalDispatcher.ts). node:vm is not a security boundary. A self-hoster who wants Python evaluators or untrusted code needs the Lambda route.

Decision models

A decision-model evaluator asks typed questions (choice, score, boolean) of a connection whose adapter is typesafe. It returns probabilities and confidence that land in score metadata. Any other connection type blocks the evaluator (runDecisionModelEvaluation.ts, types.ts).

Sampling, dedup and auto-blocking

Sampling hashes the target id with SHA-256 under a fixed domain prefix and compares the result to the rate, so the same trace is always in or out (deterministicSampling.ts). Job ids derive from (project, rule, trace, dataset item, observation), and the Postgres primary key lets only one of two racing producers enqueue. When a judge call fails with an auth, billing or bad-config error, the evaluator is paused for every rule that uses it rather than retried forever (evalService.ts).

Prompt experiments

experimentServiceClickhouse.ts runs a stored prompt over a dataset. For each item it creates a dataset-run-item event with a deterministic trace id through the same processEventBatch path, calls the model, and schedules observation evals on the result (experimentServiceClickhouse.ts). Experiments driven from your own code go through the SDKs and arrive as ordinary traces linked to dataset items.

Extending it

  • Evaluator templates. LLM-judge templates are prompts with {{variables}} plus an output definition. Variable mappings point at trace, observation or dataset-item fields. A managed catalogue ships starter templates.
  • Code evaluators. Write a TypeScript or Python function that returns {scores: [...]}. Python requires the Lambda dispatcher.
  • Model connections. Judges use per-project LLM connections. The adapters are OpenAI, Anthropic, Azure, Bedrock, Vertex AI, Google AI Studio and TypeSafe.
  • Ingestion. Anything that speaks OTLP over HTTP can send traces. The public REST API also accepts scores, so external eval harnesses can write results next to the built-in ones.

Running it

docker-compose.yml runs six services: langfuse-web, langfuse-worker, ClickHouse, Postgres, Redis and MinIO (S3-compatible storage). All are required. Postgres holds configuration, rules and job executions. ClickHouse holds traces, observations and scores. Redis backs BullMQ. Blob storage holds every raw event. Without an enterprise licence key, a self-hosted instance falls back to the oss plan (getPlan.ts). That plan includes tracing, datasets, evaluators and experiments.

Strengths and caveats

  • Strength: eval as infrastructure. Rules score live production traffic and experiment runs on the same queue path, with deterministic sampling, dedup, delays and per-evaluator pausing. This is closer to a production scoring service than to a test runner.
  • Strength: auditable judges. Every judge call is itself a trace in your project, linked from the score by executionTraceId, so you can debug the evaluator with the tool you use for your app.
  • Strength: durable ingestion. Raw events are written to blob storage before processing, and later events merge into earlier rows, which tolerates out-of-order SDK flushes.
  • Caveat: heavy to self-host. Six services, two databases, and an ingestion path that depends on object storage. This is not something you embed in CI.
  • Caveat: no built-in metric library. There are no BLEU, ROUGE or embedding metrics. Scores come from LLM judges, your code, decision models, humans or the API.
  • Caveat: evaluator types are uneven. Code and decision-model evaluators only run in the observation/experiment path. Code evals are off in production until a dispatcher is configured.
  • Caveat: no red-teaming. Safety is covered by a few judge templates (prompt-injection detection, rule adherence), not by attack generation.
  • Caveat: SDKs live elsewhere. The Python and JS tracing SDKs are not in this repository, so this review covers their server contract, not their internals.

Sources: code at 1a21a42, deepwiki-open wiki (10 pages), verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit 1a21a42. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Langfuse provides three evaluator types that produce scores, plus deterministic sampling and blocking. LLM-as-a-judge (evalService.ts lines 941–1233) — the primary mechanism — constructs a prompt from a template, calls a configurable LLM (OpenAI, Anthropic, etc.) with Zod-validated structured output, and normalizes the response into NUMERIC, BOOLEAN, or CATEGORICAL scores (outputDefinition.ts lines 1–428). The output definition is a persisted schema (PersistedEvalOutputDefinitionSchema) supporting numeric ranges (minValue/maxValue), boolean verdicts, and categorical choices (single or multi-match). The toNormalizedScores helper (evalService.ts lines 1235–1270) converts the LLM response — e.g. a numeric 0–1, a boolean true/false, or a string array — into CodeEvalScoreWithName[] records with a shared comment (the model's reasoning). Each score carries a dataType, value, and optionally configId and metadata. Code-based evaluators (codeEvalExecution.ts) dispatch user-supplied code (JS/Python) to either an AWS Lambda dispatcher or a local process, returning scores in a {scores: [...]} JSON contract — NO built-in heuristic or statistical metrics exist outside these three types. Decision-model evaluators (decisionModelEvaluatorExecution.ts) call a separate "TypeSafe" API for choice/score/boolean questions and map responses into scores. Deterministic sampling (deterministicSampling.ts lines 1–31) controls what fraction of traces are scored, using a SHA‑256 over target ID and sampling rate. Per-trace vs per-observation scoring is handled via targetObject: TRACE scores the whole trace at the trace level, DATASET ties scores to dataset items, and EVENT (observation-level) schedules per-observation evaluations (observationEval/). All scores are persisted with source: "EVAL" and share an EvalScoreWritePayload structure (evalScoreEvent.ts).

Editor's note. Correction: code and decision-model evaluators only run in the observation-level and experiment path (processObservationEval); trace- and dataset-level rules load only LLM-as-judge evaluators. The insecure-local code dispatcher runs TypeScript only, in a node:vm context inside the worker, and production has no code-eval dispatcher unless LANGFUSE_CODE_EVAL_DISPATCHER is set.

How is LLM-as-a-judge implemented?

answered

LLM-as-a-judge is implemented through a chain of modules converging on runLLMAsJudgeEvaluation in evalService.ts (lines 941–1233). Judge prompts and rubrics are stored as EvalTemplate records with a prompt (template string using {{variable}} syntax) and promptMessages (an array of system/user/assistant messages). The prompt is compiled by substituting extracted variables — trace fields like input, output, or dataset-item columns — via compileEvalPrompt in llmEvaluatorExecution.ts (lines 40–57). Structured output is enforced through Zod schemas: the outputDefinition stored on the template (legacy {score, reasoning} strings or v2 {dataType, score, reasoning} objects) is compiled into a Zod schema via compilePersistedEvalOutputDefinition (outputDefinition.ts lines 362–375). The schema is serialized to JSON Schema and passed to the AI SDK, which constraints the LLM's response shape. Results are validated with validateEvalOutputResult — numeric ranges, allowed categorical values, duplicates detection, and boolean type strictness are all checked (test file lines 541–679). Judge model choice is per-evaluator: each template stores provider, model, and modelParams which are resolved through fetchModelConfig (evalExecutionDeps.ts lines 327–337). The system supports OpenAI, Anthropic, Google, AWS Bedrock, and any OpenAI-compatible endpoint. No built-in multi-sample/consensus — each eval job produces a single LLM call. No explicit calibration exists, but bias controls are available through the prompt template itself (the user authors the rubric). Auto-blocking (classifyEvaluatorLlmError.ts) is key: when the judge LLM fails (auth, billing, endpoint unreachable), the evaluator is automatically paused at the database level so all rules using it stop until the user intervenes (lines 60–84). Internal tracing records each judge execution as an LLMJudge-environment trace in the user's project for debugging.

How are test datasets and cases defined, generated and versioned?

answered

Datasets are stored as Dataset and DatasetItem records in PostgreSQL, with ClickHouse for analytics queries. A dataset has an inputSchema and expectedOutputSchema (optional JSON Schema) for validation (dataset-items.ts lines 62–72). Dataset items carry input, expectedOutput, metadata, plus a sourceTraceId/sourceObservationId when created from traced data. Items are versioned via a temporal model: validFrom/validTo columns enable point-in-time snapshots — fetching an item at a specific date retrieves the version valid then (code in evalService.ts extractVariablesFromTracingData at line 1635: ...(datasetItemValidFrom ? { validFrom: datasetItemValidFrom } : { validTo: null })). Dataset runs (dataset-runs.ts) execute a prompt experiment against dataset items: the experimentServiceClickhouse.ts takes a prompt template, replaces variables with each dataset item's input, calls the LLM, and creates DatasetRunItem records linking the experiment's output traces back to the original dataset items (lines 74–78). No synthetic data generation DSL exists — items are created via the UI, the public API (POST /api/public/dataset-items), or batch CSV/JSON upload. Golden sets are just datasets used as eval benchmarks: you create a dataset, map its columns to evaluator variables in an evaluation rule, and the rule scores every new trace that matches (or you run a batch evaluation). Eval jobs can target TRACE (live scoring) or DATASET (scoring dataset-run items). Benchmark task registration is handled through the evaluation rules system (evaluationRule Prisma model) — rules connect evaluators (judge templates) to dataset items or traces via filters, sampling rates, and variable mappings.

How are evals executed and reported?

answered

Eval execution is a job-queue pipeline built on BullMQ (Redis). Scheduling (evalService.ts createEvalJobs lines 302–913): when a trace is upserted or a dataset run item created, the system fetches active evaluation rules for the project, filters them against the trace (filter state and sampling), deduplicates via deterministic job IDs, and enqueues EvalExecutionQueue jobs. Three parallel queues exist for trace-level and observation-level evals, routed by EvalExecutionQueue.getInstance({shardingKey}) (line 841). Parallelism is handled by BullMQ worker concurrency; no built-in batching per model provider. Caching (deterministicSampling.ts) uses SHA‑256 hashing of the target ID to produce a deterministic [0,1) value for sampling decisions. A trace-cache optimization (evalService.ts lines 389–442) fetches trace data once for multiple configs. Execution calls into runLLMAsJudgeEvaluation (or code/decision-model executors), which calls the LLM, validates the structured output, and persists scores. Results flow through completeEvalExecution (evalCompletion.ts lines 23–98): scores are uploaded to S3 as JSON event files, queued for ClickHouse ingestion via IngestionQueue, and the JobExecution row is updated to COMPLETED with the primary score ID and execution trace ID. Error handling (evalExecutionMetrics.ts) classifies outcomes into success, platform_error, upstream_error, customer_error, or cancelled. LLM errors are categorized and can auto-block the evaluator. CI integration is user-defined (no built-in harness); users call the public API or run evaluations ad-hoc via the UI's "Run evaluation" batch dialog. Results appear as scores on traces/observations in the Langfuse UI, filterable and searchable via the Eval Logs table (eval-log.tsx). Dashboards and regression are user-built on top of the score data — the platform stores all scores with timestamps, trace IDs, and metadata, enabling external charting or the built-in analytics views.

Editor's note. Correction: deterministicSampling.ts is sampling, not caching: it hashes a domain prefix plus the target id with SHA-256 and compares the result to the rule's rate. Trace-level evaluate only executes LLM-as-judge templates; code and decision-model evaluators run through the observation-eval queues.

How are traces or production data captured and linked to evaluations?

answered

Langfuse is an observability platform, so traces are captured via a first-party Python/JS/TS SDK (langfuse pip package, JS SDK in packages/shared/src/server/llm/), and it also ingests OpenTelemetry data through the OTel ingestion pipeline (OtelIngestionProcessor.ts). Traces consist of a root trace with nested observations (spans, generations, events), all stored in ClickHouse. Linking evals to traces works through the executionTraceId mechanism: when an LLM-as-a-judge evaluator runs, it creates a new internal trace (environment langfuse-llm-as-a-judge or langfuse-code-eval) in the same project, passing traceSinkParams with targetProjectId, traceId, and metadata (evalExecutionDeps.ts lines 302–312). This internal trace records the judge's input prompt, output score, model, latency, and reasoning. The score event's executionTraceId field links the score back to this internal trace. Online vs offline evals: "online" (trace-upsert-triggered) runs eval rules immediately when traces arrive; "offline" evaluations are triggered via the UI batch-action dialog or the public API with a timestamp range for historical data. Feedback and annotation is a separate feature: Langfuse has scores with source: "ANNOTATION" as distinct from source: "EVAL", both stored in the same scores table but tagged differently. The scores repository (scores.ts) handles both types. The eval log (eval-log.tsx) provides a UI table of all eval job executions with status, score, comment, and links to the target trace. Each evaluation also writes its own internal trace with environment labels that are excluded from triggering further evaluations (the isInternalEvalEnvironment guard in isEvalTargetEnvironmentAllowed.ts prevents infinite eval loops).

Editor's note. Correction: the Python and JS tracing SDKs are not in this repository; packages/shared/src/server/llm/ is the server's own LLM-calling code.

Does it support red-teaming or safety testing, and how?

answered

Langfuse does NOT have a dedicated red-teaming framework (no adversarial-probe runners, attack plugins, or automated jailbreak testing pipelines). Instead, it provides evaluator templates in the safety category of its managed templates catalog (managedTemplatesCatalog.ts lines 53–58, 1193–1297) that users can use for manual or automated safety testing. These include: Detect Prompt Injection — an LLM-as-a-judge template that classifies whether an input contains a credible manipulation attempt (override instructions, reveal system prompt, bypass safeguards), returning a boolean verdict (lines 1242–1297). Check Rule Adherence — a boolean evaluator that checks if an output follows a defined policy or instruction set (lines 1193–1241). The safety category description (lines 54–57) states it "Monitors policy adherence, privacy leakage, and adversarial prompts." These are template starters — users must adapt the prompt and deploy an evaluation rule that runs on their traces. There are no built-in safety-specific tools, no adversarial dataset generators, no automated red-team loops, and no vulnerability-reporting workflow (the repo's SECURITY.md at root only links to a generic security policy). Safety testing must be authored by the user: create a dataset of adversarial inputs, write an evaluator template (or use the managed one), and run batch evaluations against it. The platform stores all results and can track regressions over time via the scores API.