# Arize-ai/phoenix

> Self-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/Arize-ai/phoenix (reviewed at commit `bbdce4f193ef85b8fcdad1df7109f18e93d44971`, 2026-10-06)
- Stars: 11735 · Language: Python · License: Elastic-2.0
- Canonical page: https://llms-technical-reviews.com/p/phoenix/

## Overview

Arize Phoenix evaluates LLM **applications**, and it starts from traces rather than benchmarks. You instrument your app with OpenTelemetry using the OpenInference semantic conventions, and Phoenix receives the spans over OTLP and stores them in SQLite or PostgreSQL. A React UI shows them as traces, sessions and projects. Evaluation is built on that store. Spans can be turned into dataset examples. Datasets are versioned. An experiment runs your task function, or a stored prompt, over a dataset version, and evaluators score each run. All results (task runs, evaluation scores, span annotations) are written back to the same database, so you can compare experiments side by side.

The code is split into four Python packages. The server and UI live in `src/phoenix/` (`arize-phoenix`). `phoenix-otel` is a thin OTel setup helper. `phoenix-client` is the REST client that runs experiments. `phoenix-evals` is a standalone evaluator library: LLM-judge classifiers and code metrics over a pandas DataFrame. TypeScript equivalents live under `js/packages/`.

Phoenix answers different questions from the other tools in this category. lm-evaluation-harness compares base models on public benchmarks, and garak attacks a model endpoint. Phoenix asks what your app did on real or curated inputs and whether a change made it better. It has no attack generation and no academic benchmark suite. Note the licence: the root `LICENSE` is the Elastic License 2.0 ([LICENSE](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/LICENSE#L1-L5)), which is source-available, not OSI open source. It restricts offering Phoenix as a managed service.

## Architecture

```mermaid
flowchart LR
  APP["Your app + OpenInference instrumentors"] --> OTEL["phoenix.otel.register()"]
  OTEL --> GRPC["gRPC OTLP Servicer"]
  OTEL --> HTTP["POST /v1/traces"]
  GRPC --> BULK["BulkInserter queue"]
  HTTP --> BULK
  BULK --> DB["SQLite / PostgreSQL"]
  DB --> UI["React UI + GraphQL"]
  UI --> DS["Datasets + versions"]
  CLIENT["phoenix-client run_experiment"] --> DS
  CLIENT --> EXP["Experiments, runs, evals"]
  RUNNER["ExperimentRunner daemon"] --> EXP
  EVALS["phoenix-evals evaluators"] --> CLIENT
  EXP --> DB
```

| Component | Path | Role |
|---|---|---|
| OTel helper | `packages/phoenix-otel/` | `register()` builds a TracerProvider with an OTLP exporter pointed at Phoenix |
| Ingest | `src/phoenix/server/grpc_server.py`, `server/api/routers/v1/traces.py` | OTLP/gRPC and OTLP/HTTP protobuf endpoints that decode spans and enqueue them |
| Bulk inserter | `src/phoenix/db/bulk_inserter.py` | Batches span and annotation inserts, computes span cost |
| Data model | `src/phoenix/db/models.py` | Traces, spans, annotations, datasets, versions, splits, experiments, prompts, users |
| Server | `src/phoenix/server/app.py` | FastAPI + Strawberry GraphQL app, daemons, auth |
| Experiment runner | `src/phoenix/server/daemons/experiment_runner.py` | Background execution of UI-launched prompt experiments and evaluators |
| Server evaluators | `src/phoenix/server/api/evaluators.py` | Built-in code evaluators (contains, exact match, regex, Levenshtein, JSON distance) and LLM evaluators |
| Client | `packages/phoenix-client/` | REST client: spans, datasets, prompts, `experiments.run_experiment` |
| Evals library | `packages/phoenix-evals/` | `ClassificationEvaluator`, prebuilt judges, `evaluate_dataframe`, provider adapters |
| UI | `js/app/` | Trace explorer, dataset and experiment views, playground |

## How a request flows

Two paths matter: getting a trace in, and running an experiment.

**Trace ingestion**

1. **Instrument.** `phoenix.otel.register(project_name=..., auto_instrument=True)` creates an OTel `TracerProvider` and exporter using `PHOENIX_COLLECTOR_ENDPOINT`, inferring gRPC or HTTP from the URL ([otel.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-otel/src/phoenix/otel/otel.py#L65-L105)). OpenInference instrumentors for OpenAI, LangChain, LlamaIndex and others then emit spans with LLM attributes.
2. **Receive.** The gRPC `Servicer.Export` reads the project name from resource attributes, decodes each span in a thread pool and enqueues it ([grpc_server.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/grpc_server.py#L33-L52)). The HTTP route accepts only `application/x-protobuf`, handles gzip/deflate, and does the insert in a background task ([traces.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/api/routers/v1/traces.py#L440-L511)).
3. **Persist.** `BulkInserter._bulk_insert` drains the queue in bounded transactions. `_insert_spans` writes each span in a nested transaction, so one bad span doesn't sink the batch, and then computes a `SpanCost` row from token counts ([bulk_inserter.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/db/bulk_inserter.py#L121-L212)). The app wires the inserter's `enqueue_span` into the gRPC server at startup ([app.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/app.py#L675-L700)).

**Experiment from code**

1. **Dataset.** Examples are created from spans in the UI, uploaded from a DataFrame or CSV, or added through the client. Each change creates a `DatasetVersion`, and every example has `DatasetExampleRevision` rows linked to versions, plus optional splits ([models.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/db/models.py#L1549-L1660)).
2. **Run.** `client.experiments.run_experiment(dataset=..., task=fn, evaluators=[...])` validates the task signature, creates an `Experiment` pinned to `dataset.version_id`, and builds one test case per (example x repetition) ([experiments/\_\_init\_\_.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-client/src/phoenix/client/resources/experiments/__init__.py#L744-L1010)). Arguments are bound by name: `input`, `expected`, `metadata`, `example`.
3. **Trace the task.** Each task call runs inside an OTel span sent to the experiment's own project, so you can open the trace behind any failing row.
4. **Evaluate.** `evaluate_experiment` runs each evaluator on each run. Evaluators can return a `bool`, `float`, `str`, `(score, explanation)` tuple or a dict, and results are posted as experiment evaluations.
5. **Compare.** The client prints URLs for the dataset's experiments page, where runs from different experiments are lined up per example.

## Key components

### phoenix-evals

`ClassificationEvaluator` renders a prompt template, calls `llm.generate_classification`, rejects labels outside the allowed set, and maps the label to a score with a `label_score_map` ([evaluators.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L743-L783)). Prebuilt judges cover correctness, faithfulness, hallucination, document and retrieval relevance, toxicity, refusal, PII and tool selection or invocation. The OpenAI adapter tries native JSON-schema structured output first, falls back to tool calling only on a `BadRequestError`, and caches whichever worked per model ([adapter.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/adapters/openai/adapter.py#L130-L175)). `evaluate_dataframe` runs evaluators over every row and adds score columns ([evaluators.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L1409-L1430)). The async executor adapts concurrency to error rates. Each judgement is one call, with no built-in multi-judge consensus or position-bias control.

### Server-side experiments

Experiments can also start from the UI. A prompt version plus model config (`ExperimentPromptTask`) is run over a dataset by the `ExperimentRunner` daemon. It schedules work across experiments, prioritises evals over retries over new tasks, and handles rate limits and cancellation ([experiment_runner.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/daemons/experiment_runner.py#L1-L40); constructed in [app.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/app.py#L1069-L1075)). Built-in code evaluators are registered with `@register_builtin_evaluator` ([evaluators.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/api/evaluators.py#L547-L560), [L641-L655](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/api/evaluators.py#L641-L655)).

### Annotations as the join point

Human labels, LLM-judge scores and code checks all end up as span, trace, document or session annotations, with an annotator kind. You can attach an offline `evaluate_dataframe` result to production spans with `log_span_annotations_dataframe`, and then filter traces in the UI by score.

### Phoenix's own eval suite

`evals/pxi/` holds the team's evals for Phoenix's built-in assistant: YAML datasets, a CI gate and an online runner that annotates recent production spans ([README.md](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/README.md#L1-L30)). It isn't shipped in the package. It's a worked example of an online-eval loop you would build yourself on the client API.

## Extending it

- **Custom evaluators:** any function works in `run_experiment`. In phoenix-evals, `@create_evaluator` or `create_classifier` wrap a function or a rubric with labels.
- **Judge providers:** adapters exist for OpenAI, Anthropic, Google, LiteLLM and LangChain, chosen via `LLM(provider=..., model=...)`.
- **Instrumentation:** any OTel exporter works, because the ingest is plain OTLP. OpenInference attributes are what light up LLM-specific views.
- **APIs:** REST (`/v1/...`), GraphQL and an MCP server expose spans, datasets, prompts and experiments.

## Running it

- `pip install arize-phoenix` then `phoenix serve` ([main.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/server/main.py#L17-L30)). The UI and OTLP/HTTP default to port 6006, OTLP/gRPC to 4317. Docker, Helm and Kustomize manifests are in the repo.
- Storage defaults to `sqlite:///<working_dir>/phoenix.db` (working dir `~/.phoenix`). Set `PHOENIX_SQL_DATABASE_URL` or the Postgres variables for PostgreSQL ([config.py](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/config.py#L3362-L3381)).
- App side: `pip install arize-phoenix-otel` plus the OpenInference instrumentors you need. Evals: `pip install arize-phoenix-evals` and a judge provider key.

## Strengths and caveats

- **Strength: traces and evals share one store.** Datasets come from spans, experiment tasks are traced, and scores land as annotations, so moving from a failing score to the exact trace is one click.
- **Strength: real dataset versioning.** Experiments pin a version, and examples carry revisions and splits. That makes regression comparisons meaningful.
- **Strength: standard ingest.** OTLP over gRPC or HTTP, so you're not locked into a proprietary SDK.
- **Caveat: licence.** ELv2, not Apache or MIT. Fine for internal use, restrictive for anyone hosting it for others.
- **Caveat: a server to run.** Experiments and annotations need a live Phoenix with a database. phoenix-evals alone is only a DataFrame scorer.
- **Caveat: no red-teaming or benchmark suite.** Toxicity and PII judges classify outputs after the fact. Adversarial input generation must come from elsewhere.
- **Caveat: CI is DIY.** There is no packaged pass/fail gate; the in-repo `evals/pxi` gate shows how to build one.

*Sources: code at bbdce4f, verified Q&A.*

## How Arize-ai/phoenix answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Phoenix provides three categories of evaluators: code/heuristic, LLM-based, and statistical. **Code evaluators** include `exact_match` (string equality check) at `packages/phoenix-evals/src/phoenix/evals/metrics/exact_match.py:37`, `MatchesRegex` (regex matching) and `PrecisionRecallFScore` (macro/micro/weighted precision, recall, F-beta with binary one-vs-rest) at `packages/phoenix-evals/src/phoenix/evals/metrics/precision_recall.py:66-424`. **LLM-based evaluators** subclass `ClassificationEvaluator` and include `CorrectnessEvaluator`, `FaithfulnessEvaluator`, `HallucinationEvaluator`, `CompletenessEvaluator`, `ConcisenessEvaluator`, `DocumentRelevanceEvaluator`, `RetrievalRelevanceEvaluator`, `ToxicityEvaluator`, `RefusalEvaluator`, `PiiDetectionEvaluator`, `UserFrictionEvaluator`, and three tool-oriented evaluators (`ToolInvocationEvaluator`, `ToolSelectionEvaluator`, `ToolResponseHandlingEvaluator`) — all in `packages/phoenix-evals/src/phoenix/evals/metrics/`. Each LLM evaluator loads a generated prompt config (e.g. `_correctness_classification_evaluator_config.py` defines a rubric string, model `choices` like `{"correct": 1.0, "incorrect": 0.0}`, and `optimization_direction`). The base `Evaluator` class (`packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506`) handles input remapping via `input_mapping` (field-name or JSONPath-based), Pydantic schema validation, and returns `Score` objects. The `Score` dataclass (`evaluators.py:158-271`) carries `name`, `score` (float/int), `label` (string), `explanation`, `metadata`, `kind` ("human"|"llm"|"code"), and `direction` ("maximize"|"minimize"|"neutral"). **Per-trace vs per-task**: Evaluators operate on a single `EvalInput` dict (one row = one trace or span); `evaluate_dataframe()` and `async_evaluate_dataframe()` apply an evaluator list over every row and add score columns back to the DataFrame. **Custom evaluators** are created via the `@create_evaluator(name=..., kind=...)` decorator which wraps any sync/async function into an `Evaluator` instance with auto-generated Pydantic input schema from function signature, and `_convert_to_score` handles returns of type `Score`, `bool`, `float`, `str`, `dict`, or `tuple`.


Citations: [packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L303-L506) · [packages/phoenix-evals/src/phoenix/evals/metrics/__init__.py:1-37](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/metrics/__init__.py#L1-L37) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:828-1128](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L828-L1128)

### How is LLM-as-a-judge implemented? (answered)

LLM-as-a-judge is implemented through the `LLM` wrapper class (`packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440`) which delegates to provider-specific **adapters** registered via a singleton `ProviderRegistry` and `AdapterRegistry` (`registries.py:39-146`). Supported adapters include OpenAI, Anthropic, Google Gemini, LangChain, and LiteLLM — each implementing `BaseLLMAdapter` with `generate_text`, `generate_object`, and async variants (`types.py:22-108`). The `ClassificationEvaluator` (`evaluators.py:601-825`) renders a prompt template, calls `llm.generate_classification()`, which builds a JSON schema via `generate_classification_schema()` (`wrapper.py:451-524`) and calls `generate_object()` with that schema. The schema is a structured JSON object requiring a `label` string (constrained by `enum` or `oneOf` for valid choices) and optionally an `explanation` string. **Structured output vs tool calling**: On OpenAI, the adapter first tries native structured output (`response_format: {"type": "json_schema", ...}`) and falls back to tool calling on `BadRequestError`, caching the preferred method per model (`openai/adapter.py:130-240`). **No built-in multi-sample / consensus** — each evaluation is a single LLM call. **Judge model choice** is fully configurable: pass any `LLM(provider=..., model=...)` to the evaluator constructor. The online-evals runner (`evals/pxi/online_evals/judge.py:1-55`) reads `PHOENIX_AGENTS_EVALS_PROVIDER` and `PHOENIX_AGENTS_EVALS_MODEL` env vars (defaulting to OpenAI/gpt-5.5). The Harbor evaluation verifier (`evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44`) uses `PHOENIX_EVAL_JUDGE_PROVIDER` and `PHOENIX_EVAL_JUDGE_MODEL` env vars (defaulting to OpenAI/gpt-5-nano). **Bias controls**: Each evaluator's rubric is stored in the generated config files (e.g. `_faithfulness_classification_evaluator_config.py` has detailed grading criteria). The `RateLimiter` (`rate_limiters.py`) uses an `AdaptiveTokenBucket` that backs off on rate-limit errors but does not implement calibration or answer-order randomization.

> **Editor's note.** Correction: the PHOENIX_AGENTS_EVALS_* and PHOENIX_EVAL_JUDGE_* defaults belong to the internal evals/ scripts, not the product. In phoenix-evals the judge is whatever LLM(provider=..., model=...) the user passes.

Citations: [packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py#L97-L440) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:601-825](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L601-L825) · [packages/phoenix-evals/src/phoenix/evals/llm/adapters/openai/adapter.py:130-240](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/adapters/openai/adapter.py#L130-L240) · [packages/phoenix-evals/src/phoenix/evals/llm/registries.py:39-146](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/registries.py#L39-L146)

### How are test datasets and cases defined, generated and versioned? (answered)

Test datasets are defined primarily as **pandas DataFrames** in Python code. Phoenix provides a `download_benchmark_dataset(task, dataset_name)` function (`packages/phoenix-evals/src/phoenix/evals/utils.py:17-36`) that fetches zipped JSONL files from Google Cloud Storage (`storage.googleapis.com/arize-phoenix-assets/evals/`) and loads them as DataFrames. There is no built-in dataset versioning, golden set management, or registry — datasets are code-managed. For **experiment-style evaluations** within the Phoenix app, datasets are curated from production spans saved into the Phoenix project (via the UI or API), then used in experiments. The **PXI eval harness** (`evals/pxi/`) defines datasets as YAML files (`evals/pxi/datasets/`). Each YAML file (e.g. `product_knowledge.yaml:1-700`) declares a `dataset_name`, `description`, a list of evaluators to run (e.g. `assistant_text_substrings_match`, `correct_tools_called`), and `examples`. Each example has an `id`, `splits` (e.g. `[regression, dev]`), an `input` with an ordered `messages` list, `expected` output constraints (e.g. `assistant_text.contains_all`, `tools.forbidden`), and `metadata` including human annotation agreement scores. Examples can be marked `adversarial: true` for wrong-premise tests. The `evals/harbor/` directory contains a Harbor-based eval framework with `agents/`, `tasks/`, `environments/`, and `verifiers/` — the llm_judge verifier (`evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44`) defines a `matches_reference` LLM judge that compares agent replies to reference answers with grading notes. **Synthetic data generation** is described in a cookbook tutorial (`docs/phoenix/cookbook/tracing/generating-synthetic-datasets-for-llm-evaluators-and-agents.mdx`) — this is a documented pattern, not built-in tooling. There is no DSL, version-controlled dataset format, or benchmark task registry beyond the download URL convention.

> **Editor's note.** Correction: Phoenix does version datasets. The server has DatasetVersion, DatasetExampleRevision and DatasetSplit tables (src/phoenix/db/models.py), and experiments are pinned to a dataset_version_id, so datasets are not merely code-managed.

Citations: [packages/phoenix-evals/src/phoenix/evals/utils.py:17-36](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L17-L36)

### How are evals executed and reported? (answered)

Evaluation execution uses a **producer-consumer pattern** with two executor classes in `packages/phoenix-evals/src/phoenix/evals/executors.py`. `SyncExecutor` (`executors.py:464-573`) is a synchronous loop with configurable `max_retries` (default 10), `exit_on_error` (default True), tqdm progress bar, and SIGINT handling. `AsyncExecutor` (`executors.py:170-462`) is an async producer-consumer with configurable `concurrency` (default 3), retries with requeue into a priority queue, per-task `timeout` (default 60s), and a **dynamic concurrency controller** using an AIMD (Additive Increase/Multiplicative Decrease) algorithm (`executors.py:73-167`) that adapts target concurrency based on error rates — automatically scaling down on failures and back up during success windows, with a collapse mode that drops concurrency to 1 on multiple errors within a window. The high-level `evaluate_dataframe()` (`evaluators.py:1408-1555`) and `async_evaluate_dataframe()` (`evaluators.py:1558-1735`) accept a DataFrame and list of evaluators, create the task list as a Cartesian product of (row × evaluator), execute, and add `{score_name}_score` and `{evaluator_name}_execution_details` columns. The `get_executor_on_sync_context()` helper (`executors.py:576-645`) automatically selects between sync and async execution based on the current thread and event loop. **Online (production) evals**: The PXI online-eval runner (`evals/pxi/online_evals/run.py:445-483`) is a CLI script that discovers candidate spans via `client.spans.get_spans()`, deduplicates against existing annotations, applies deterministic per-trace sampling (`_sampled()` using SHA256 of artifact ID), fetches full trace spans, evaluates with up to 8 concurrent evaluations, and writes `SpanAnnotationData` back to the Phoenix server in batches of 100. The runner produces a `RunSummary` with discovered/evaluated/sampled-out/already-annotated/error counts and supports GitHub Step Summary output. **CI integration**: no built-in CI integration — users run evals as scripts or notebooks and can compare results manually. **Result storage**: scores are either columns in a DataFrame (offline) or annotations persisted to the Phoenix server via `client.spans.log_span_annotations()`. The `to_annotation_dataframe()` utility (`utils.py:410-491`) reformats DataFrame scores into annotation format for logging.

> **Editor's note.** Correction: evals/pxi (online runner, CI gate) is Phoenix's internal eval suite for its own assistant and is not shipped in the package. User-facing execution is client.experiments.run_experiment or the server-side ExperimentRunner daemon, with results compared per example in the UI.

Citations: [packages/phoenix-evals/src/phoenix/evals/executors.py:170-462](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/executors.py#L170-L462) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:1408-1735](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L1408-L1735) · [evals/pxi/online_evals/run.py:39-120](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L39-L120) · [evals/pxi/online_evals/run.py:210-342](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L210-L342) · [packages/phoenix-evals/src/phoenix/evals/executors.py:73-167](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/executors.py#L73-L167) · [packages/phoenix-evals/src/phoenix/evals/utils.py:410-491](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L410-L491)

### How are traces or production data captured and linked to evaluations? (answered)

Observability is Phoenix's primary domain, built on **OpenTelemetry** with **OpenInference semantic conventions**. The `phoenix-otel` package (`packages/phoenix-otel/src/phoenix/otel/otel.py`) provides a `register()` function that creates an OpenTelemetry `TracerProvider` configured to export spans to the Phoenix collector via OTLP (gRPC or HTTP/protobuf). Spans follow the `openinference.semconv.trace` schema with `SpanAttributes` for LLM-specific fields like `INPUT_VALUE`, `OUTPUT_VALUE`, `LLM_MODEL_NAME`, `LLM_TOKEN_COUNT_PROMPT`, `LLM_TOKEN_COUNT_COMPLETION`, `LLM_TOKEN_COUNT_TOTAL`, and `OPENINFERENCE_SPAN_KIND` (`src/phoenix/trace/otel.py:1-57`). The **trace module** (`src/phoenix/trace/`) handles span ingestion via OTLP protobuf decoding, attribute flattening/unflattening (`attributes.py`), conversion between protobuf and the internal `Span` schema (`schemas.py`). SDK instrumentation is documented for Python, JavaScript, and Java. **Online vs offline evals**: Evaluation results are linked to traces via `span_id` in the `Score.metadata` (`evaluators.py:273-291`). The online-evals runner (`evals/pxi/online_evals/run.py`) discovers spans, fetches full traces, runs LLM judges, and writes `SpanAnnotationData` back to the Phoenix server via `client.spans.log_span_annotations()`. The Phoenix server stores these annotations and serves them via GraphQL and REST APIs. **Trace-level tracing**: The `@trace` decorator (`tracing.py:96-332`) wraps evaluator and LLM calls with OpenInference spans, recording input/output values and span metadata. It uses `OITracer` from `openinference.instrumentation` and auto-injects `trace_id` from the span context into evaluator calls. The `Score._add_trace_id_to_scores()` method embeds the trace ID into each score's metadata for traceability. **Feedback and annotation**: The `to_annotation_dataframe()` utility in `utils.py:410-491` formats evaluation results as annotations (span_id, score, label, explanation, annotation_name, annotator_kind) for logging back to Phoenix.


Citations: [packages/phoenix-otel/src/phoenix/otel/otel.py:65-120](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-otel/src/phoenix/otel/otel.py#L65-L120) · [src/phoenix/trace/otel.py:1-57](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/src/phoenix/trace/otel.py#L1-L57) · [packages/phoenix-evals/src/phoenix/evals/tracing.py:96-332](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/tracing.py#L96-L332) · [evals/pxi/online_evals/run.py:196-210](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/evals/pxi/online_evals/run.py#L196-L210) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:273-291](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L273-L291) · [packages/phoenix-evals/src/phoenix/evals/utils.py:410-491](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L410-L491)

### Does it support red-teaming or safety testing, and how? (insufficient evidence)

Phoenix does not provide built-in red-teaming or adversarial safety testing tooling. There is no dedicated module for adversarial probes, attack plugins, jailbreak prompt generation, prompt injection testing suites, or automated vulnerability reporting. The `ToxicityEvaluator` and `PiiDetectionEvaluator` are content-safety evaluators that can flag toxic or PII-containing outputs, but these are after-the-fact classifiers, not adversarial generators. The PXI eval dataset (`evals/pxi/datasets/product_knowledge.yaml`) includes examples tagged `adversarial: true` — these are test cases where the user prompt presents a wrong premise to verify the agent corrects it, but this is a manual convention in the dataset YAML, not an automated adversarial generation system. Tutorials on jailbreak/prompt-injection defense and realtime guardrails exist as educational cookbook patterns using Phoenix's observability to monitor guardrail effectiveness, not as built-in tooling. There is no attack fuzzer, no adversarial probe optimizer, and no vulnerability reporting pipeline at the reviewed commit.


