Arize-ai/phoenix
Self-hosted OTLP trace store and UI for LLM apps, with versioned datasets, experiments and LLM-judge evaluators on top.
Overview
Arize Phoenix evaluates LLM applications, and it starts from traces rather than benchmarks. You instrument your app with OpenTelemetry using the OpenInference semantic conventions, and Phoenix receives the spans over OTLP and stores them in SQLite or PostgreSQL. A React UI shows them as traces, sessions and projects. Evaluation is built on that store. Spans can be turned into dataset examples. Datasets are versioned. An experiment runs your task function, or a stored prompt, over a dataset version, and evaluators score each run. All results (task runs, evaluation scores, span annotations) are written back to the same database, so you can compare experiments side by side.
The code is split into four Python packages. The server and UI live in src/phoenix/ (arize-phoenix). phoenix-otel is a thin OTel setup helper. phoenix-client is the REST client that runs experiments. phoenix-evals is a standalone evaluator library: LLM-judge classifiers and code metrics over a pandas DataFrame. TypeScript equivalents live under js/packages/.
Phoenix answers different questions from the other tools in this category. lm-evaluation-harness compares base models on public benchmarks, and garak attacks a model endpoint. Phoenix asks what your app did on real or curated inputs and whether a change made it better. It has no attack generation and no academic benchmark suite. Note the licence: the root LICENSE is the Elastic License 2.0 (LICENSE), which is source-available, not OSI open source. It restricts offering Phoenix as a managed service.
Architecture
flowchart LR
APP["Your app + OpenInference instrumentors"] --> OTEL["phoenix.otel.register()"]
OTEL --> GRPC["gRPC OTLP Servicer"]
OTEL --> HTTP["POST /v1/traces"]
GRPC --> BULK["BulkInserter queue"]
HTTP --> BULK
BULK --> DB["SQLite / PostgreSQL"]
DB --> UI["React UI + GraphQL"]
UI --> DS["Datasets + versions"]
CLIENT["phoenix-client run_experiment"] --> DS
CLIENT --> EXP["Experiments, runs, evals"]
RUNNER["ExperimentRunner daemon"] --> EXP
EVALS["phoenix-evals evaluators"] --> CLIENT
EXP --> DB
| Component | Path | Role |
|---|---|---|
| OTel helper | packages/phoenix-otel/ |
register() builds a TracerProvider with an OTLP exporter pointed at Phoenix |
| Ingest | src/phoenix/server/grpc_server.py, server/api/routers/v1/traces.py |
OTLP/gRPC and OTLP/HTTP protobuf endpoints that decode spans and enqueue them |
| Bulk inserter | src/phoenix/db/bulk_inserter.py |
Batches span and annotation inserts, computes span cost |
| Data model | src/phoenix/db/models.py |
Traces, spans, annotations, datasets, versions, splits, experiments, prompts, users |
| Server | src/phoenix/server/app.py |
FastAPI + Strawberry GraphQL app, daemons, auth |
| Experiment runner | src/phoenix/server/daemons/experiment_runner.py |
Background execution of UI-launched prompt experiments and evaluators |
| Server evaluators | src/phoenix/server/api/evaluators.py |
Built-in code evaluators (contains, exact match, regex, Levenshtein, JSON distance) and LLM evaluators |
| Client | packages/phoenix-client/ |
REST client: spans, datasets, prompts, experiments.run_experiment |
| Evals library | packages/phoenix-evals/ |
ClassificationEvaluator, prebuilt judges, evaluate_dataframe, provider adapters |
| UI | js/app/ |
Trace explorer, dataset and experiment views, playground |
How a request flows
Two paths matter: getting a trace in, and running an experiment.
Trace ingestion
- Instrument.
phoenix.otel.register(project_name=..., auto_instrument=True)creates an OTelTracerProviderand exporter usingPHOENIX_COLLECTOR_ENDPOINT, inferring gRPC or HTTP from the URL (otel.py). OpenInference instrumentors for OpenAI, LangChain, LlamaIndex and others then emit spans with LLM attributes. - Receive. The gRPC
Servicer.Exportreads the project name from resource attributes, decodes each span in a thread pool and enqueues it (grpc_server.py). The HTTP route accepts onlyapplication/x-protobuf, handles gzip/deflate, and does the insert in a background task (traces.py). - Persist.
BulkInserter._bulk_insertdrains the queue in bounded transactions._insert_spanswrites each span in a nested transaction, so one bad span doesn’t sink the batch, and then computes aSpanCostrow from token counts (bulk_inserter.py). The app wires the inserter’senqueue_spaninto the gRPC server at startup (app.py).
Experiment from code
- Dataset. Examples are created from spans in the UI, uploaded from a DataFrame or CSV, or added through the client. Each change creates a
DatasetVersion, and every example hasDatasetExampleRevisionrows linked to versions, plus optional splits (models.py). - Run.
client.experiments.run_experiment(dataset=..., task=fn, evaluators=[...])validates the task signature, creates anExperimentpinned todataset.version_id, and builds one test case per (example x repetition) (experiments/__init__.py). Arguments are bound by name:input,expected,metadata,example. - Trace the task. Each task call runs inside an OTel span sent to the experiment’s own project, so you can open the trace behind any failing row.
- Evaluate.
evaluate_experimentruns each evaluator on each run. Evaluators can return abool,float,str,(score, explanation)tuple or a dict, and results are posted as experiment evaluations. - Compare. The client prints URLs for the dataset’s experiments page, where runs from different experiments are lined up per example.
Key components
phoenix-evals
ClassificationEvaluator renders a prompt template, calls llm.generate_classification, rejects labels outside the allowed set, and maps the label to a score with a label_score_map (evaluators.py). Prebuilt judges cover correctness, faithfulness, hallucination, document and retrieval relevance, toxicity, refusal, PII and tool selection or invocation. The OpenAI adapter tries native JSON-schema structured output first, falls back to tool calling only on a BadRequestError, and caches whichever worked per model (adapter.py). evaluate_dataframe runs evaluators over every row and adds score columns (evaluators.py). The async executor adapts concurrency to error rates. Each judgement is one call, with no built-in multi-judge consensus or position-bias control.
Server-side experiments
Experiments can also start from the UI. A prompt version plus model config (ExperimentPromptTask) is run over a dataset by the ExperimentRunner daemon. It schedules work across experiments, prioritises evals over retries over new tasks, and handles rate limits and cancellation (experiment_runner.py; constructed in app.py). Built-in code evaluators are registered with @register_builtin_evaluator (evaluators.py, L641-L655).
Annotations as the join point
Human labels, LLM-judge scores and code checks all end up as span, trace, document or session annotations, with an annotator kind. You can attach an offline evaluate_dataframe result to production spans with log_span_annotations_dataframe, and then filter traces in the UI by score.
Phoenix’s own eval suite
evals/pxi/ holds the team’s evals for Phoenix’s built-in assistant: YAML datasets, a CI gate and an online runner that annotates recent production spans (README.md). It isn’t shipped in the package. It’s a worked example of an online-eval loop you would build yourself on the client API.
Extending it
- Custom evaluators: any function works in
run_experiment. In phoenix-evals,@create_evaluatororcreate_classifierwrap a function or a rubric with labels. - Judge providers: adapters exist for OpenAI, Anthropic, Google, LiteLLM and LangChain, chosen via
LLM(provider=..., model=...). - Instrumentation: any OTel exporter works, because the ingest is plain OTLP. OpenInference attributes are what light up LLM-specific views.
- APIs: REST (
/v1/...), GraphQL and an MCP server expose spans, datasets, prompts and experiments.
Running it
pip install arize-phoenixthenphoenix serve(main.py). The UI and OTLP/HTTP default to port 6006, OTLP/gRPC to 4317. Docker, Helm and Kustomize manifests are in the repo.- Storage defaults to
sqlite:///<working_dir>/phoenix.db(working dir~/.phoenix). SetPHOENIX_SQL_DATABASE_URLor the Postgres variables for PostgreSQL (config.py). - App side:
pip install arize-phoenix-otelplus the OpenInference instrumentors you need. Evals:pip install arize-phoenix-evalsand a judge provider key.
Strengths and caveats
- Strength: traces and evals share one store. Datasets come from spans, experiment tasks are traced, and scores land as annotations, so moving from a failing score to the exact trace is one click.
- Strength: real dataset versioning. Experiments pin a version, and examples carry revisions and splits. That makes regression comparisons meaningful.
- Strength: standard ingest. OTLP over gRPC or HTTP, so you’re not locked into a proprietary SDK.
- Caveat: licence. ELv2, not Apache or MIT. Fine for internal use, restrictive for anyone hosting it for others.
- Caveat: a server to run. Experiments and annotations need a live Phoenix with a database. phoenix-evals alone is only a DataFrame scorer.
- Caveat: no red-teaming or benchmark suite. Toxicity and PII judges classify outputs after the fact. Adversarial input generation must come from elsewhere.
- Caveat: CI is DIY. There is no packaged pass/fail gate; the in-repo
evals/pxigate shows how to build one.
Sources: code at bbdce4f, verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit bbdce4f. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredPhoenix provides three categories of evaluators: code/heuristic, LLM-based, and statistical. Code evaluators include exact_match (string equality check) at packages/phoenix-evals/src/phoenix/evals/metrics/exact_match.py:37, MatchesRegex (regex matching) and PrecisionRecallFScore (macro/micro/weighted precision, recall, F-beta with binary one-vs-rest) at packages/phoenix-evals/src/phoenix/evals/metrics/precision_recall.py:66-424. LLM-based evaluators subclass ClassificationEvaluator and include CorrectnessEvaluator, FaithfulnessEvaluator, HallucinationEvaluator, CompletenessEvaluator, ConcisenessEvaluator, DocumentRelevanceEvaluator, RetrievalRelevanceEvaluator, ToxicityEvaluator, RefusalEvaluator, PiiDetectionEvaluator, UserFrictionEvaluator, and three tool-oriented evaluators (ToolInvocationEvaluator, ToolSelectionEvaluator, ToolResponseHandlingEvaluator) — all in packages/phoenix-evals/src/phoenix/evals/metrics/. Each LLM evaluator loads a generated prompt config (e.g. _correctness_classification_evaluator_config.py defines a rubric string, model choices like {"correct": 1.0, "incorrect": 0.0}, and optimization_direction). The base Evaluator class (packages/phoenix-evals/src/phoenix/evals/evaluators.py:303-506) handles input remapping via input_mapping (field-name or JSONPath-based), Pydantic schema validation, and returns Score objects. The Score dataclass (evaluators.py:158-271) carries name, score (float/int), label (string), explanation, metadata, kind ("human"|"llm"|"code"), and direction ("maximize"|"minimize"|"neutral"). Per-trace vs per-task: Evaluators operate on a single EvalInput dict (one row = one trace or span); evaluate_dataframe() and async_evaluate_dataframe() apply an evaluator list over every row and add score columns back to the DataFrame. Custom evaluators are created via the @create_evaluator(name=..., kind=...) decorator which wraps any sync/async function into an Evaluator instance with auto-generated Pydantic input schema from function signature, and _convert_to_score handles returns of type Score, bool, float, str, dict, or tuple.
How is LLM-as-a-judge implemented?
answeredLLM-as-a-judge is implemented through the LLM wrapper class (packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440) which delegates to provider-specific adapters registered via a singleton ProviderRegistry and AdapterRegistry (registries.py:39-146). Supported adapters include OpenAI, Anthropic, Google Gemini, LangChain, and LiteLLM — each implementing BaseLLMAdapter with generate_text, generate_object, and async variants (types.py:22-108). The ClassificationEvaluator (evaluators.py:601-825) renders a prompt template, calls llm.generate_classification(), which builds a JSON schema via generate_classification_schema() (wrapper.py:451-524) and calls generate_object() with that schema. The schema is a structured JSON object requiring a label string (constrained by enum or oneOf for valid choices) and optionally an explanation string. Structured output vs tool calling: On OpenAI, the adapter first tries native structured output (response_format: {"type": "json_schema", ...}) and falls back to tool calling on BadRequestError, caching the preferred method per model (openai/adapter.py:130-240). No built-in multi-sample / consensus — each evaluation is a single LLM call. Judge model choice is fully configurable: pass any LLM(provider=..., model=...) to the evaluator constructor. The online-evals runner (evals/pxi/online_evals/judge.py:1-55) reads PHOENIX_AGENTS_EVALS_PROVIDER and PHOENIX_AGENTS_EVALS_MODEL env vars (defaulting to OpenAI/gpt-5.5). The Harbor evaluation verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) uses PHOENIX_EVAL_JUDGE_PROVIDER and PHOENIX_EVAL_JUDGE_MODEL env vars (defaulting to OpenAI/gpt-5-nano). Bias controls: Each evaluator's rubric is stored in the generated config files (e.g. _faithfulness_classification_evaluator_config.py has detailed grading criteria). The RateLimiter (rate_limiters.py) uses an AdaptiveTokenBucket that backs off on rate-limit errors but does not implement calibration or answer-order randomization.
How are test datasets and cases defined, generated and versioned?
answeredTest datasets are defined primarily as pandas DataFrames in Python code. Phoenix provides a download_benchmark_dataset(task, dataset_name) function (packages/phoenix-evals/src/phoenix/evals/utils.py:17-36) that fetches zipped JSONL files from Google Cloud Storage (storage.googleapis.com/arize-phoenix-assets/evals/) and loads them as DataFrames. There is no built-in dataset versioning, golden set management, or registry — datasets are code-managed. For experiment-style evaluations within the Phoenix app, datasets are curated from production spans saved into the Phoenix project (via the UI or API), then used in experiments. The PXI eval harness (evals/pxi/) defines datasets as YAML files (evals/pxi/datasets/). Each YAML file (e.g. product_knowledge.yaml:1-700) declares a dataset_name, description, a list of evaluators to run (e.g. assistant_text_substrings_match, correct_tools_called), and examples. Each example has an id, splits (e.g. [regression, dev]), an input with an ordered messages list, expected output constraints (e.g. assistant_text.contains_all, tools.forbidden), and metadata including human annotation agreement scores. Examples can be marked adversarial: true for wrong-premise tests. The evals/harbor/ directory contains a Harbor-based eval framework with agents/, tasks/, environments/, and verifiers/ — the llm_judge verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) defines a matches_reference LLM judge that compares agent replies to reference answers with grading notes. Synthetic data generation is described in a cookbook tutorial (docs/phoenix/cookbook/tracing/generating-synthetic-datasets-for-llm-evaluators-and-agents.mdx) — this is a documented pattern, not built-in tooling. There is no DSL, version-controlled dataset format, or benchmark task registry beyond the download URL convention.
How are evals executed and reported?
answeredEvaluation execution uses a producer-consumer pattern with two executor classes in packages/phoenix-evals/src/phoenix/evals/executors.py. SyncExecutor (executors.py:464-573) is a synchronous loop with configurable max_retries (default 10), exit_on_error (default True), tqdm progress bar, and SIGINT handling. AsyncExecutor (executors.py:170-462) is an async producer-consumer with configurable concurrency (default 3), retries with requeue into a priority queue, per-task timeout (default 60s), and a dynamic concurrency controller using an AIMD (Additive Increase/Multiplicative Decrease) algorithm (executors.py:73-167) that adapts target concurrency based on error rates — automatically scaling down on failures and back up during success windows, with a collapse mode that drops concurrency to 1 on multiple errors within a window. The high-level evaluate_dataframe() (evaluators.py:1408-1555) and async_evaluate_dataframe() (evaluators.py:1558-1735) accept a DataFrame and list of evaluators, create the task list as a Cartesian product of (row × evaluator), execute, and add {score_name}_score and {evaluator_name}_execution_details columns. The get_executor_on_sync_context() helper (executors.py:576-645) automatically selects between sync and async execution based on the current thread and event loop. Online (production) evals: The PXI online-eval runner (evals/pxi/online_evals/run.py:445-483) is a CLI script that discovers candidate spans via client.spans.get_spans(), deduplicates against existing annotations, applies deterministic per-trace sampling (_sampled() using SHA256 of artifact ID), fetches full trace spans, evaluates with up to 8 concurrent evaluations, and writes SpanAnnotationData back to the Phoenix server in batches of 100. The runner produces a RunSummary with discovered/evaluated/sampled-out/already-annotated/error counts and supports GitHub Step Summary output. CI integration: no built-in CI integration — users run evals as scripts or notebooks and can compare results manually. Result storage: scores are either columns in a DataFrame (offline) or annotations persisted to the Phoenix server via client.spans.log_span_annotations(). The to_annotation_dataframe() utility (utils.py:410-491) reformats DataFrame scores into annotation format for logging.
How are traces or production data captured and linked to evaluations?
answeredObservability is Phoenix's primary domain, built on OpenTelemetry with OpenInference semantic conventions. The phoenix-otel package (packages/phoenix-otel/src/phoenix/otel/otel.py) provides a register() function that creates an OpenTelemetry TracerProvider configured to export spans to the Phoenix collector via OTLP (gRPC or HTTP/protobuf). Spans follow the openinference.semconv.trace schema with SpanAttributes for LLM-specific fields like INPUT_VALUE, OUTPUT_VALUE, LLM_MODEL_NAME, LLM_TOKEN_COUNT_PROMPT, LLM_TOKEN_COUNT_COMPLETION, LLM_TOKEN_COUNT_TOTAL, and OPENINFERENCE_SPAN_KIND (src/phoenix/trace/otel.py:1-57). The trace module (src/phoenix/trace/) handles span ingestion via OTLP protobuf decoding, attribute flattening/unflattening (attributes.py), conversion between protobuf and the internal Span schema (schemas.py). SDK instrumentation is documented for Python, JavaScript, and Java. Online vs offline evals: Evaluation results are linked to traces via span_id in the Score.metadata (evaluators.py:273-291). The online-evals runner (evals/pxi/online_evals/run.py) discovers spans, fetches full traces, runs LLM judges, and writes SpanAnnotationData back to the Phoenix server via client.spans.log_span_annotations(). The Phoenix server stores these annotations and serves them via GraphQL and REST APIs. Trace-level tracing: The @trace decorator (tracing.py:96-332) wraps evaluator and LLM calls with OpenInference spans, recording input/output values and span metadata. It uses OITracer from openinference.instrumentation and auto-injects trace_id from the span context into evaluator calls. The Score._add_trace_id_to_scores() method embeds the trace ID into each score's metadata for traceability. Feedback and annotation: The to_annotation_dataframe() utility in utils.py:410-491 formats evaluation results as annotations (span_id, score, label, explanation, annotation_name, annotator_kind) for logging back to Phoenix.
Does it support red-teaming or safety testing, and how?
insufficient evidencePhoenix does not provide built-in red-teaming or adversarial safety testing tooling. There is no dedicated module for adversarial probes, attack plugins, jailbreak prompt generation, prompt injection testing suites, or automated vulnerability reporting. The ToxicityEvaluator and PiiDetectionEvaluator are content-safety evaluators that can flag toxic or PII-containing outputs, but these are after-the-fact classifiers, not adversarial generators. The PXI eval dataset (evals/pxi/datasets/product_knowledge.yaml) includes examples tagged adversarial: true — these are test cases where the user prompt presents a wrong premise to verify the agent corrects it, but this is a manual convention in the dataset YAML, not an automated adversarial generation system. Tutorials on jailbreak/prompt-injection defense and realtime guardrails exist as educational cookbook patterns using Phoenix's observability to monitor guardrail effectiveness, not as built-in tooling. There is no attack fuzzer, no adversarial probe optimizer, and no vulnerability reporting pipeline at the reviewed commit.