LLMs Technical Reviews

Arize-ai/phoenix

Self-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.

GitHub ↗★ 12kPythonElastic-2.0commit bbdce4f · 2026-10-06homepage ↗

Overview

Arize Phoenix evaluates LLM applications, and it starts from traces rather than benchmarks. You instrument your app with OpenTelemetry using the OpenInference semantic conventions, and Phoenix receives the spans over OTLP and stores them in SQLite or PostgreSQL. A React UI shows them as traces, sessions and projects. Evaluation is built on that store. Spans can be turned into dataset examples. Datasets are versioned. An experiment runs your task function, or a stored prompt, over a dataset version, and evaluators score each run. All results (task runs, evaluation scores, span annotations) are written back to the same database, so you can compare experiments side by side.

The code is split into four Python packages. The server and UI live in src/phoenix/ (arize-phoenix). phoenix-otel is a thin OTel setup helper. phoenix-client is the REST client that runs experiments. phoenix-evals is a standalone evaluator library: LLM-judge classifiers and code metrics over a pandas DataFrame. TypeScript equivalents live under js/packages/.

Phoenix answers different questions from the other tools in this category. lm-evaluation-harness compares base models on public benchmarks, and garak attacks a model endpoint. Phoenix asks what your app did on real or curated inputs and whether a change made it better. It has no attack generation and no academic benchmark suite. Note the licence: the root LICENSE is the Elastic License 2.0 (LICENSE), which is source-available, not OSI open source. It restricts offering Phoenix as a managed service.

Architecture

flowchart LR
  APP["Your app + OpenInference instrumentors"] --> OTEL["phoenix.otel.register()"]
  OTEL --> GRPC["gRPC OTLP Servicer"]
  OTEL --> HTTP["POST /v1/traces"]
  GRPC --> BULK["BulkInserter queue"]
  HTTP --> BULK
  BULK --> DB["SQLite / PostgreSQL"]
  DB --> UI["React UI + GraphQL"]
  UI --> DS["Datasets + versions"]
  CLIENT["phoenix-client run_experiment"] --> DS
  CLIENT --> EXP["Experiments, runs, evals"]
  RUNNER["ExperimentRunner daemon"] --> EXP
  EVALS["phoenix-evals evaluators"] --> CLIENT
  EXP --> DB
Component Path Role
OTel helper packages/phoenix-otel/ register() builds a TracerProvider with an OTLP exporter pointed at Phoenix
Ingest src/phoenix/server/grpc_server.py, server/api/routers/v1/traces.py OTLP/gRPC and OTLP/HTTP protobuf endpoints that decode spans and enqueue them
Bulk inserter src/phoenix/db/bulk_inserter.py Batches span and annotation inserts, computes span cost
Data model src/phoenix/db/models.py Traces, spans, annotations, datasets, versions, splits, experiments, prompts, users
Server src/phoenix/server/app.py FastAPI + Strawberry GraphQL app, daemons, auth
Experiment runner src/phoenix/server/daemons/experiment_runner.py Background execution of UI-launched prompt experiments and evaluators
Server evaluators src/phoenix/server/api/evaluators.py Built-in code evaluators (contains, exact match, regex, Levenshtein, JSON distance) and LLM evaluators
Client packages/phoenix-client/ REST client: spans, datasets, prompts, experiments.run_experiment
Evals library packages/phoenix-evals/ ClassificationEvaluator, prebuilt judges, evaluate_dataframe, provider adapters
UI js/app/ Trace explorer, dataset and experiment views, playground

How a request flows

Two paths matter: getting a trace in, and running an experiment.

Trace ingestion

  1. Instrument. phoenix.otel.register(project_name=..., auto_instrument=True) creates an OTel TracerProvider and exporter using PHOENIX_COLLECTOR_ENDPOINT, inferring gRPC or HTTP from the URL (otel.py). OpenInference instrumentors for OpenAI, LangChain, LlamaIndex and others then emit spans with LLM attributes.
  2. Receive. The gRPC Servicer.Export reads the project name from resource attributes, decodes each span in a thread pool and enqueues it (grpc_server.py). The HTTP route accepts only application/x-protobuf, handles gzip/deflate, and does the insert in a background task (traces.py).
  3. Persist. BulkInserter._bulk_insert drains the queue in bounded transactions. _insert_spans writes each span in a nested transaction, so one bad span doesn’t sink the batch, and then computes a SpanCost row from token counts (bulk_inserter.py). The app wires the inserter’s enqueue_span into the gRPC server at startup (app.py).

Experiment from code

  1. Dataset. Examples are created from spans in the UI, uploaded from a DataFrame or CSV, or added through the client. Each change creates a DatasetVersion, and every example has DatasetExampleRevision rows linked to versions, plus optional splits (models.py).
  2. Run. client.experiments.run_experiment(dataset=..., task=fn, evaluators=[...]) validates the task signature, creates an Experiment pinned to dataset.version_id, and builds one test case per (example x repetition) (experiments/__init__.py). Arguments are bound by name: input, expected, metadata, example.
  3. Trace the task. Each task call runs inside an OTel span sent to the experiment’s own project, so you can open the trace behind any failing row.
  4. Evaluate. evaluate_experiment runs each evaluator on each run. Evaluators can return a bool, float, str, (score, explanation) tuple or a dict, and results are posted as experiment evaluations.
  5. Compare. The client prints URLs for the dataset’s experiments page, where runs from different experiments are lined up per example.

Key components

phoenix-evals

ClassificationEvaluator renders a prompt template, calls llm.generate_classification, rejects labels outside the allowed set, and maps the label to a score with a label_score_map (evaluators.py). Prebuilt judges cover correctness, faithfulness, hallucination, document and retrieval relevance, toxicity, refusal, PII and tool selection or invocation. The OpenAI adapter tries native JSON-schema structured output first, falls back to tool calling only on a BadRequestError, and caches whichever worked per model (adapter.py). evaluate_dataframe runs evaluators over every row and adds score columns (evaluators.py). The async executor adapts concurrency to error rates. Each judgement is one call, with no built-in multi-judge consensus or position-bias control.

Server-side experiments

Experiments can also start from the UI. A prompt version plus model config (ExperimentPromptTask) is run over a dataset by the ExperimentRunner daemon. It schedules work across experiments, prioritises evals over retries over new tasks, and handles rate limits and cancellation (experiment_runner.py; constructed in app.py). Built-in code evaluators are registered with @register_builtin_evaluator (evaluators.py, L641-L655).

Annotations as the join point

Human labels, LLM-judge scores and code checks all end up as span, trace, document or session annotations, with an annotator kind. You can attach an offline evaluate_dataframe result to production spans with log_span_annotations_dataframe, and then filter traces in the UI by score.

Phoenix’s own eval suite

evals/pxi/ holds the team’s evals for Phoenix’s built-in assistant: YAML datasets, a CI gate and an online runner that annotates recent production spans (README.md). It isn’t shipped in the package. It’s a worked example of an online-eval loop you would build yourself on the client API.

Extending it

  • Custom evaluators: any function works in run_experiment. In phoenix-evals, @create_evaluator or create_classifier wrap a function or a rubric with labels.
  • Judge providers: adapters exist for OpenAI, Anthropic, Google, LiteLLM and LangChain, chosen via LLM(provider=..., model=...).
  • Instrumentation: any OTel exporter works, because the ingest is plain OTLP. OpenInference attributes are what light up LLM-specific views.
  • APIs: REST (/v1/...), GraphQL and an MCP server expose spans, datasets, prompts and experiments.

Running it

  • pip install arize-phoenix then phoenix serve (main.py). The UI and OTLP/HTTP default to port 6006, OTLP/gRPC to 4317. Docker, Helm and Kustomize manifests are in the repo.
  • Storage defaults to sqlite:///<working_dir>/phoenix.db (working dir ~/.phoenix). Set PHOENIX_SQL_DATABASE_URL or the Postgres variables for PostgreSQL (config.py).
  • App side: pip install arize-phoenix-otel plus the OpenInference instrumentors you need. Evals: pip install arize-phoenix-evals and a judge provider key.

Strengths and caveats

  • Strength: traces and evals share one store. Datasets come from spans, experiment tasks are traced, and scores land as annotations, so moving from a failing score to the exact trace is one click.
  • Strength: real dataset versioning. Experiments pin a version, and examples carry revisions and splits. That makes regression comparisons meaningful.
  • Strength: standard ingest. OTLP over gRPC or HTTP, so you’re not locked into a proprietary SDK.
  • Caveat: licence. ELv2, not Apache or MIT. Fine for internal use, restrictive for anyone hosting it for others.
  • Caveat: a server to run. Experiments and annotations need a live Phoenix with a database. phoenix-evals alone is only a DataFrame scorer.
  • Caveat: no red-teaming or benchmark suite. Toxicity and PII judges classify outputs after the fact. Adversarial input generation must come from elsewhere.
  • Caveat: CI is DIY. There is no packaged pass/fail gate; the in-repo evals/pxi gate shows how to build one.

Sources: code at bbdce4f, verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit bbdce4f. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Phoenix provides three categories of evaluators: code/heuristic, LLM-based, and statistical. Code evaluators include exact_match (string equality check) at packages/phoenix-evals/src/phoenix/evals/metrics/exact_match.py:37, MatchesRegex (regex matching) and PrecisionRecallFScore (macro/micro/weighted precision, recall, F-beta with binary one-vs-rest) at packages/phoenix-evals/src/phoenix/evals/metrics/precision_recall.py:66-424. LLM-based evaluators subclass ClassificationEvaluator and include CorrectnessEvaluator, FaithfulnessEvaluator, HallucinationEvaluator, CompletenessEvaluator, ConcisenessEvaluator, DocumentRelevanceEvaluator, RetrievalRelevanceEvaluator, ToxicityEvaluator, RefusalEvaluator, PiiDetectionEvaluator, UserFrictionEvaluator, and three tool-oriented evaluators (ToolInvocationEvaluator, ToolSelectionEvaluator, ToolResponseHandlingEvaluator) — all in packages/phoenix-evals/src/phoenix/evals/metrics/. Each LLM evaluator loads a generated prompt config (e.g. _correctness_classification_evaluator_config.py defines a rubric string, model choices like {"correct": 1.0, "incorrect": 0.0}, and optimization_direction). The base Evaluator class (packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506) handles input remapping via input_mapping (field-name or JSONPath-based), Pydantic schema validation, and returns Score objects. The Score dataclass (evaluators.py:158-271) carries name, score (float/int), label (string), explanation, metadata, kind ("human"|"llm"|"code"), and direction ("maximize"|"minimize"|"neutral"). Per-trace vs per-task: Evaluators operate on a single EvalInput dict (one row = one trace or span); evaluate_dataframe() and async_evaluate_dataframe() apply an evaluator list over every row and add score columns back to the DataFrame. Custom evaluators are created via the @create_evaluator(name=..., kind=...) decorator which wraps any sync/async function into an Evaluator instance with auto-generated Pydantic input schema from function signature, and _convert_to_score handles returns of type Score, bool, float, str, dict, or tuple.

How is LLM-as-a-judge implemented?

answered

LLM-as-a-judge is implemented through the LLM wrapper class (packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440) which delegates to provider-specific adapters registered via a singleton ProviderRegistry and AdapterRegistry (registries.py:39-146). Supported adapters include OpenAI, Anthropic, Google Gemini, LangChain, and LiteLLM — each implementing BaseLLMAdapter with generate_text, generate_object, and async variants (types.py:22-108). The ClassificationEvaluator (evaluators.py:601-825) renders a prompt template, calls llm.generate_classification(), which builds a JSON schema via generate_classification_schema() (wrapper.py:451-524) and calls generate_object() with that schema. The schema is a structured JSON object requiring a label string (constrained by enum or oneOf for valid choices) and optionally an explanation string. Structured output vs tool calling: On OpenAI, the adapter first tries native structured output (response_format: {"type": "json_schema", ...}) and falls back to tool calling on BadRequestError, caching the preferred method per model (openai/adapter.py:130-240). No built-in multi-sample / consensus — each evaluation is a single LLM call. Judge model choice is fully configurable: pass any LLM(provider=..., model=...) to the evaluator constructor. The online-evals runner (evals/pxi/online_evals/judge.py:1-55) reads PHOENIX_AGENTS_EVALS_PROVIDER and PHOENIX_AGENTS_EVALS_MODEL env vars (defaulting to OpenAI/gpt-5.5). The Harbor evaluation verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) uses PHOENIX_EVAL_JUDGE_PROVIDER and PHOENIX_EVAL_JUDGE_MODEL env vars (defaulting to OpenAI/gpt-5-nano). Bias controls: Each evaluator's rubric is stored in the generated config files (e.g. _faithfulness_classification_evaluator_config.py has detailed grading criteria). The RateLimiter (rate_limiters.py) uses an AdaptiveTokenBucket that backs off on rate-limit errors but does not implement calibration or answer-order randomization.

Editor's note. Correction: the PHOENIX_AGENTS_EVALS_* and PHOENIX_EVAL_JUDGE_* defaults belong to the internal evals/ scripts, not the product. In phoenix-evals the judge is whatever LLM(provider=..., model=...) the user passes.

How are test datasets and cases defined, generated and versioned?

answered

Test datasets are defined primarily as pandas DataFrames in Python code. Phoenix provides a download_benchmark_dataset(task, dataset_name) function (packages/phoenix-evals/src/phoenix/evals/utils.py:17-36) that fetches zipped JSONL files from Google Cloud Storage (storage.googleapis.com/arize-phoenix-assets/evals/) and loads them as DataFrames. There is no built-in dataset versioning, golden set management, or registry — datasets are code-managed. For experiment-style evaluations within the Phoenix app, datasets are curated from production spans saved into the Phoenix project (via the UI or API), then used in experiments. The PXI eval harness (evals/pxi/) defines datasets as YAML files (evals/pxi/datasets/). Each YAML file (e.g. product_knowledge.yaml:1-700) declares a dataset_name, description, a list of evaluators to run (e.g. assistant_text_substrings_match, correct_tools_called), and examples. Each example has an id, splits (e.g. [regression, dev]), an input with an ordered messages list, expected output constraints (e.g. assistant_text.contains_all, tools.forbidden), and metadata including human annotation agreement scores. Examples can be marked adversarial: true for wrong-premise tests. The evals/harbor/ directory contains a Harbor-based eval framework with agents/, tasks/, environments/, and verifiers/ — the llm_judge verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) defines a matches_reference LLM judge that compares agent replies to reference answers with grading notes. Synthetic data generation is described in a cookbook tutorial (docs/phoenix/cookbook/tracing/generating-synthetic-datasets-for-llm-evaluators-and-agents.mdx) — this is a documented pattern, not built-in tooling. There is no DSL, version-controlled dataset format, or benchmark task registry beyond the download URL convention.

Editor's note. Correction: Phoenix does version datasets. The server has DatasetVersion, DatasetExampleRevision and DatasetSplit tables (src/phoenix/db/models.py), and experiments are pinned to a dataset_version_id, so datasets are not merely code-managed.

How are evals executed and reported?

answered

Evaluation execution uses a producer-consumer pattern with two executor classes in packages/phoenix-evals/src/phoenix/evals/executors.py. SyncExecutor (executors.py:464-573) is a synchronous loop with configurable max_retries (default 10), exit_on_error (default True), tqdm progress bar, and SIGINT handling. AsyncExecutor (executors.py:170-462) is an async producer-consumer with configurable concurrency (default 3), retries with requeue into a priority queue, per-task timeout (default 60s), and a dynamic concurrency controller using an AIMD (Additive Increase/Multiplicative Decrease) algorithm (executors.py:73-167) that adapts target concurrency based on error rates — automatically scaling down on failures and back up during success windows, with a collapse mode that drops concurrency to 1 on multiple errors within a window. The high-level evaluate_dataframe() (evaluators.py:1408-1555) and async_evaluate_dataframe() (evaluators.py:1558-1735) accept a DataFrame and list of evaluators, create the task list as a Cartesian product of (row × evaluator), execute, and add {score_name}_score and {evaluator_name}_execution_details columns. The get_executor_on_sync_context() helper (executors.py:576-645) automatically selects between sync and async execution based on the current thread and event loop. Online (production) evals: The PXI online-eval runner (evals/pxi/online_evals/run.py:445-483) is a CLI script that discovers candidate spans via client.spans.get_spans(), deduplicates against existing annotations, applies deterministic per-trace sampling (_sampled() using SHA256 of artifact ID), fetches full trace spans, evaluates with up to 8 concurrent evaluations, and writes SpanAnnotationData back to the Phoenix server in batches of 100. The runner produces a RunSummary with discovered/evaluated/sampled-out/already-annotated/error counts and supports GitHub Step Summary output. CI integration: no built-in CI integration — users run evals as scripts or notebooks and can compare results manually. Result storage: scores are either columns in a DataFrame (offline) or annotations persisted to the Phoenix server via client.spans.log_span_annotations(). The to_annotation_dataframe() utility (utils.py:410-491) reformats DataFrame scores into annotation format for logging.

Editor's note. Correction: evals/pxi (online runner, CI gate) is Phoenix's internal eval suite for its own assistant and is not shipped in the package. User-facing execution is client.experiments.run_experiment or the server-side ExperimentRunner daemon, with results compared per example in the UI.

How are traces or production data captured and linked to evaluations?

answered

Observability is Phoenix's primary domain, built on OpenTelemetry with OpenInference semantic conventions. The phoenix-otel package (packages/phoenix-otel/src/phoenix/otel/otel.py) provides a register() function that creates an OpenTelemetry TracerProvider configured to export spans to the Phoenix collector via OTLP (gRPC or HTTP/protobuf). Spans follow the openinference.semconv.trace schema with SpanAttributes for LLM-specific fields like INPUT_VALUE, OUTPUT_VALUE, LLM_MODEL_NAME, LLM_TOKEN_COUNT_PROMPT, LLM_TOKEN_COUNT_COMPLETION, LLM_TOKEN_COUNT_TOTAL, and OPENINFERENCE_SPAN_KIND (src/phoenix/trace/otel.py:1-57). The trace module (src/phoenix/trace/) handles span ingestion via OTLP protobuf decoding, attribute flattening/unflattening (attributes.py), conversion between protobuf and the internal Span schema (schemas.py). SDK instrumentation is documented for Python, JavaScript, and Java. Online vs offline evals: Evaluation results are linked to traces via span_id in the Score.metadata (evaluators.py:273-291). The online-evals runner (evals/pxi/online_evals/run.py) discovers spans, fetches full traces, runs LLM judges, and writes SpanAnnotationData back to the Phoenix server via client.spans.log_span_annotations(). The Phoenix server stores these annotations and serves them via GraphQL and REST APIs. Trace-level tracing: The @trace decorator (tracing.py:96-332) wraps evaluator and LLM calls with OpenInference spans, recording input/output values and span metadata. It uses OITracer from openinference.instrumentation and auto-injects trace_id from the span context into evaluator calls. The Score._add_trace_id_to_scores() method embeds the trace ID into each score's metadata for traceability. Feedback and annotation: The to_annotation_dataframe() utility in utils.py:410-491 formats evaluation results as annotations (span_id, score, label, explanation, annotation_name, annotator_kind) for logging back to Phoenix.

Does it support red-teaming or safety testing, and how?

insufficient evidence

Phoenix does not provide built-in red-teaming or adversarial safety testing tooling. There is no dedicated module for adversarial probes, attack plugins, jailbreak prompt generation, prompt injection testing suites, or automated vulnerability reporting. The ToxicityEvaluator and PiiDetectionEvaluator are content-safety evaluators that can flag toxic or PII-containing outputs, but these are after-the-fact classifiers, not adversarial generators. The PXI eval dataset (evals/pxi/datasets/product_knowledge.yaml) includes examples tagged adversarial: true — these are test cases where the user prompt presents a wrong premise to verify the agent corrects it, but this is a manual convention in the dataset YAML, not an automated adversarial generation system. Tutorials on jailbreak/prompt-injection defense and realtime guardrails exist as educational cookbook patterns using Phoenix's observability to monitor guardrail effectiveness, not as built-in tooling. There is no attack fuzzer, no adversarial probe optimizer, and no vulnerability reporting pipeline at the reviewed commit.