# How are test datasets and cases defined, generated and versioned?

> LLM evals and testing — a good answer covers: File formats/DSL; synthetic data generation; versioning; golden sets; benchmark task registries.

Canonical page: https://llms-technical-reviews.com/evals/q/datasets/

## Verdict

[Phoenix](/p/phoenix/), [Opik](/p/opik/) and [Langfuse](/p/langfuse/) are the only projects that version test data themselves. [DeepEval](/p/deepeval/) and [Ragas](/p/ragas/) are best at generating test data. [lm-evaluation-harness](/p/lm-evaluation-harness/) and [OpenAI Evals](/p/openai-evals/) are benchmark registries.

**Versioned datasets on a server.** Phoenix stores `DatasetVersion`, per-example revisions and splits, and each experiment is pinned to a dataset version. Langfuse versions items in time with `validFrom`/`validTo`. A dataset-level eval reads each item at its `validFrom` version, and items can point back to the trace they came from. Opik has dataset versions and test suites whose items carry their own evaluator configs and pass policies (`runs_per_item`, `pass_threshold`). None of the three generates synthetic data.

**Datasets in code, with generators.** DeepEval's `Synthesizer` rewrites seed inputs with evolutions (reasoning, multi-context, hypothetical and others). It also wraps about 15 standard benchmarks, but dataset versions live on the vendor's hosted platform. Ragas's `TestsetGenerator` builds a knowledge graph from your documents, invents personas and writes single-hop and multi-hop questions. Its `version_experiment` commits your working tree to a git branch. [promptfoo](/p/promptfoo/) reads tests from YAML, CSV, XLSX, Google Sheets or Hugging Face, and expands array variables into every combination. Versioning is git plus a config snapshot per run. [Inspect](/p/inspect_ai/) `Sample`s can carry sandbox files and setup. Versioning is left to the source, such as a Hugging Face `revision`.

**Benchmark and attack registries.** lm-evaluation-harness has about 220 benchmark folders and close to 14,000 task YAML files, each with a `metadata.version`. OpenAI Evals has 463 eval YAML files named `<base>.<split>.v<N>`, and its JSONL samples are stored in Git LFS. [garak](/p/garak/) ships attack payloads as JSON under `garak/data/`. A copy in the user data directory overrides them. It has no versioning and no generation.

Pick: Phoenix or Opik when experiments must be pinned to a dataset version.
Pick: Ragas or DeepEval to generate questions from your own documents.
Pick: lm-evaluation-harness for public benchmarks with versioned task configs.

## Per-project answers

### langfuse/langfuse (answered)

Datasets are stored as `Dataset` and `DatasetItem` records in PostgreSQL, with ClickHouse for analytics queries. A dataset has an `inputSchema` and `expectedOutputSchema` (optional JSON Schema) for validation (`dataset-items.ts` lines 62–72). **Dataset items** carry `input`, `expectedOutput`, `metadata`, plus a `sourceTraceId`/`sourceObservationId` when created from traced data. Items are **versioned** via a temporal model: `validFrom`/`validTo` columns enable point-in-time snapshots — fetching an item at a specific date retrieves the version valid then (code in `evalService.ts` `extractVariablesFromTracingData` at line 1635: `...(datasetItemValidFrom ? { validFrom: datasetItemValidFrom } : { validTo: null })`). **Dataset runs** (`dataset-runs.ts`) execute a prompt experiment against dataset items: the `experimentServiceClickhouse.ts` takes a prompt template, replaces variables with each dataset item's `input`, calls the LLM, and creates `DatasetRunItem` records linking the experiment's output traces back to the original dataset items (lines 74–78). **No synthetic data generation** DSL exists — items are created via the UI, the public API (`POST /api/public/dataset-items`), or batch CSV/JSON upload. **Golden sets** are just datasets used as eval benchmarks: you create a dataset, map its columns to evaluator variables in an evaluation rule, and the rule scores every new trace that matches (or you run a batch evaluation). Eval jobs can target `TRACE` (live scoring) or `DATASET` (scoring dataset-run items). **Benchmark task registration** is handled through the evaluation rules system (`evaluationRule` Prisma model) — rules connect evaluators (judge templates) to dataset items or traces via filters, sampling rates, and variable mappings.


Citations: [packages/shared/src/server/repositories/dataset-items.ts:1-80](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/repositories/dataset-items.ts#L1-L80) · [packages/shared/src/domain/dataset-items.ts:1-38](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/domain/dataset-items.ts#L1-L38) · [worker/src/features/evaluation/evalService.ts:1565-1708](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L1565-L1708) · [worker/src/features/experiments/experimentServiceClickhouse.ts:48-80](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/experiments/experimentServiceClickhouse.ts#L48-L80)

### promptfoo/promptfoo (answered)

Test cases defined in tests[] of promptfooconfig.yaml. readStandaloneTestsFile (src/util/testCaseReader.ts L113) supports HuggingFace (huggingface://datasets/), Google Sheets (CSV), Azure Blob (az://), CSV, JSON/YAML, XLSX, SharePoint, and JavaScript/Python files. HuggingFace datasets fetched with pagination and concurrent page requests. generateVarCombinations (evaluator.ts L1981-2028) expands array-vars into Cartesian test rows. Scenarios (evaluator.ts L2441-2498) merge scenario config with per-test overrides. defaultTest provides golden-set base assertions. No built-in dataset versioning beyond git; evalsTable stores each runs config snapshot for reproducibility.


Citations: [src/util/testCaseReader.ts:92-119](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/util/testCaseReader.ts#L92-L119) · [src/integrations/huggingfaceDatasets.ts:1-50](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/integrations/huggingfaceDatasets.ts#L1-L50) · [src/evaluator.ts:1981-2028](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1981-L2028) · [src/evaluator.ts:2441-2498](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L2441-L2498) · [src/database/tables.ts:58-77](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/database/tables.ts#L58-L77)

### comet-ml/opik (answered)

Opik treats datasets as first-class objects managed through a backend API with versioning built in.

**File formats/DSL**: Datasets are collections of items (dicts with arbitrary JSON-serializable content). Created via the SDK (`client.create_dataset(name)` → `dataset.insert(items)`) or the UI. Items can contain arbitrary data fields, tags, metadata, and trace/span references. The `DatasetItem` model (`api_objects/dataset/dataset_item.py:40-65`) has an `EvaluatorItem` list — each item can carry per-item evaluator configs (type + config dict, currently only 'llm_judge' supported) and per-item execution policies (runs_per_item, pass_threshold). Datasets export to pandas via `to_pandas()`, to JSON via `to_json()`, and stream items as chunks or individually through `stream_items()` / `get_items()`.

**Versioning**: `Dataset` objects have a `get_version_info()` returning `DatasetVersionPublic` (id, version_name, etc.). `DatasetVersion` provides a read-only snapshot at a specific version. Test suites (`api_objects/dataset/test_suite/test_suite.py`) are a wrapper around datasets: a `TestSuite` uses a `Dataset` internally and `TestSuiteVersion` wraps a `DatasetVersion`. The `TestSuite` API exposes `get_version_view()` for pinning to a specific version. 

**Synthetic data generation**: No built-in synthetic data generators exist in the SDK, but `SimulatedUser` (`simulation/simulated_user.py:10-100`) generates synthetic user messages via an LLM based on a persona prompt — used in multi-turn simulation but not for direct dataset creation.

**Benchmark task registry**: No static benchmark registry is included. Tasks are callables (`LLMTask = Callable[[Dict[str, Any]], Any]`) passed at evaluation time. Opik integrates with Ragas via `RagasMetricWrapper` but does not ship standard benchmark datasets.

**Dataset filtering**: Items can be filtered at retrieval time via OQL filter strings supporting fields like `data.field_name`, `tags`, `created_at` etc. Filters also support `dataset_filter_string` in `evaluate()` calls.


Citations: [sdks/python/src/opik/api_objects/dataset/dataset_item.py:9-65](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/dataset_item.py#L9-L65) · [sdks/python/src/opik/api_objects/dataset/dataset.py:100-300](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/dataset.py#L100-L300) · [sdks/python/src/opik/api_objects/dataset/test_suite/test_suite.py:84-120](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/api_objects/dataset/test_suite/test_suite.py#L84-L120) · [sdks/python/src/opik/simulation/simulated_user.py:10-100](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/simulation/simulated_user.py#L10-L100)

### openai/evals (answered)

**File formats and DSL.** Test data uses the JSONL format with standard fields `"input"` (a chat-message list or plain text) and `"ideal"` (a string or list of accepted answers). For example, `test_fuzzy_match/samples.jsonl` entries contain OpenAI-format chat arrays with example few-shot messages and an `"ideal"` answer list (`evals/registry/data/test_fuzzy_match/samples.jsonl:1-3`). Eval configurations are YAML files in `evals/registry/evals/` — each file names a base spec (e.g. `ab:` with `metrics: [accuracy]`) and one or more versioned split entries (e.g. `ab.dev.v0:`) linking to a class, `samples_jsonl` path, `eval_type`, and `modelgraded_spec` (`evals/registry/evals/ab.yaml:1-11`). The `Registry` class loads these YAML files from `evals/registry/evals/`, `completion_fns/`, `solvers/`, `modelgraded/`, and `eval_sets/` directories (`evals/registry.py:103-331`).

**Synthetic data generation.** Custom generators live in `evals/registry/data/*/` — e.g. `simple_physics_engine/samples_generator.py`, `poker_analysis/poker_analysis_sample_generator.py`, `mazes/nxn_maze_eval_generator.py`, and `solve-for-variable/tools/main.py`. These emit JSONL files consumed by the evals.

**External dataset integration.** The `MultipleChoice` eval class loads from HuggingFace datasets (HellaSwag, Hendrycks MMLU) via `datasets.load_dataset()` using `hf://` URLs (`evals/elsuite/multiple_choice.py:20-48`). The `Lambada` eval similarly loads `EleutherAI/lambada_openai` from HuggingFace (`evals/elsuite/lambada.py:42-44`).

**Versioning.** Versioning follows a `{base_eval}.{split}.v{N}` convention — e.g. `ab.dev.v0`, `prompt-injection.dev.v0`, `human-safety.test.v0`. Splits (dev, test, etc.) are freeform strings. The `registry_path` parameter supports loading multiple registry directories, and `~/.evals` is a secondary path.

**Benchmark registries.** There are 463 eval YAML files in `evals/registry/evals/` and 472 data directories in `evals/registry/data/`. Eval sets like `test-all` and `test-basic` list multiple eval names to run together (`evals/registry/eval_sets/test-all.yaml:1-21`).


Citations: [evals/elsuite/multiple_choice.py:20-48](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/multiple_choice.py#L20-L48) · [evals/data.py:47-108](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/data.py#L47-L108) · [evals/registry/data/test_fuzzy_match/samples.jsonl:1-3](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/data/test_fuzzy_match/samples.jsonl#L1-L3)

### confident-ai/deepeval (answered)

**`EvaluationDataset`** (`deepeval/dataset/dataset.py:90-250`) is the core data container, holding either `Golden` objects (single-turn) or `ConversationalGolden` objects (multi-turn). Each `Golden` wraps `input`, `actual_output`, `expected_output`, `context`, `retrieval_context`, `tools_called`, `expected_tools`, `additional_metadata`, and optional `persona`/`scenario`/`expected_outcome` fields (`deepeval/dataset/golden.py:1-180`). **File formats**: `add_test_cases_from_csv_file()` parses CSV with configurable column names and list delimiters (`deepeval/dataset/dataset.py:266-340`). JSON/JSONL loading is supported through `add_test_cases_from_json_file()`. Column mappings support `input`, `actual_output`, `expected_output`, `context`, `retrieval_context` and tool call fields. **Synthetic generation**: The `Synthesizer` class (`deepeval/synthesizer/synthesizer.py:1-80`) produces `SyntheticData` / `ConversationalScenario` from source documents via evolution techniques — `Reasoning`, `Multi-context`, `Concretizing`, `Constrained`, `Comparative`, `Hypothetical`, `In-Breadth` (`deepeval/synthesizer/schema.py:14-20`). These use LLM-driven evolution templates to re-write seed inputs, and `FiltrationConfig` / `EvolutionConfig` / `StylingConfig` control quality and diversity. Conversational goldens support `Persona` (demographics, speaking style, interruption behavior) and `BackgroundNoiseSettings` for voice simulations (`deepeval/dataset/golden.py:78-150`). **Versioning**: Datasets have `_alias`, `_id`, and `_version` fields, and the `DatasetVersion` / `CreateDatasetVersionHttpResponse` API types enable versioned snapshots on Confident AI (`deepeval/dataset/api.py:49-61`). **Benchmark task registries**: The `benchmarks/` package contains ~15 standard benchmarks (MMLU, HellaSwag, BIG-Bench-Hard, GSM8K, DROP, SQuAD, HumanEval, etc.) each as a `DeepEvalBaseBenchmark` subclass that loads HuggingFace datasets and evaluates via `load_benchmark_dataset()` → `evaluate()` (`deepeval/benchmarks/base_benchmark.py:16-32`).


Citations: [deepeval/dataset/dataset.py:90-250](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/dataset.py#L90-L250) · [deepeval/dataset/golden.py:1-180](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/golden.py#L1-L180) · [deepeval/synthesizer/synthesizer.py:1-80](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/synthesizer/synthesizer.py#L1-L80) · [deepeval/synthesizer/schema.py:14-20](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/synthesizer/schema.py#L14-L20) · [deepeval/dataset/api.py:49-61](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/dataset/api.py#L49-L61) · [deepeval/benchmarks/base_benchmark.py:16-32](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/benchmarks/base_benchmark.py#L16-L32)

### vibrantlabsai/ragas (answered)

Test datasets are structured around Pydantic sample models and a flexible backend storage layer.

**Data model**: `SingleTurnSample` holds `user_input`, `retrieved_contexts`, `response`, `reference`, `rubrics`, and metadata fields (persona_name, query_style, query_length). `MultiTurnSample` holds a list of `HumanMessage|AIMessage|ToolMessage` objects with conversation validation (ToolMessage must follow an AIMessage that called tools). Both extend `BaseSample` with `to_dict()`/`get_features()` (dataset_schema.py:29-180).

**File formats/DSL**: `EvaluationDataset` wraps lists of samples and supports conversion to/from Hugging Face datasets, pandas DataFrames, CSV, and JSONL files (.to_csv/.to_jsonl/.from_jsonl) (dataset_schema.py:186-406). The `DataTable` class (dataset.py) provides a list-like interface with pluggable `BaseBackend` persistence: `LocalCSVBackend` (per-row CSV), `LocalJSONLBackend`, `GDriveBackend` (Google Drive), and `InMemoryBackend`. Backends are resolved by name via a registry (e.g. "local/csv", "gdrive") (backends/base.py, backends/registry.py, backends/local_csv.py).

**Synthetic data generation**: Test data is generated via `BaseSynthesizer` subclasses that operate on a `KnowledgeGraph` of nodes (chunks of documents) and edges (relationships). `SingleHopQuerySynthesizer` generates simple lookup queries: it samples nodes, terms, personas, query styles, and lengths, then calls an LLM prompt (`QueryAnswerGenerationPrompt`) to produce a question+reference answer pair (testset/synthesizers/single_hop/base.py:30-47). `MultiHopQuerySynthesizer` generates questions requiring 2-3 steps of reasoning over related nodes. The `Persona` system enables persona-based query generation (e.g. "expert", "novice"). Query-level `transforms` (splitting, filtering, relationship building) prepare documents before synthesis.

**Versioning**: `version_experiment()` (experiment.py:21-100) snapshots the current codebase state to git: it creates a commit and a git branch named `ragas/{experiment_name}`, returning the commit hash. This links evaluation results to a specific code version.

**Golden sets and registries**: No benchmark task registry or golden test set management exists in the codebase. Golden/holdout data lives in user-managed files loaded via `DataTable.load()` from backends. The `Dataset` class provides `train_test_split()` for splitting data into training and testing sets for metric alignment.


Citations: [src/ragas/dataset_schema.py:29-180](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L29-L180) · [src/ragas/dataset_schema.py:186-300](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L186-L300) · [src/ragas/dataset.py:31-120](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset.py#L31-L120) · [src/ragas/testset/synthesizers/single_hop/base.py:1-80](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/testset/synthesizers/single_hop/base.py#L1-L80) · [src/ragas/experiment.py:21-100](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L21-L100) · [src/ragas/backends/local_csv.py:1-80](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/backends/local_csv.py#L1-L80)

### EleutherAI/lm-evaluation-harness (answered)

Tasks are defined as **YAML config files** in `lm_eval/tasks/`, organized into subdirectories by benchmark (arc, mmlu, hellaswag, etc.). Each YAML declares `task` name, `dataset_path` (HuggingFace dataset identifier), `dataset_name` (subset), split mappings (`training_split`, `validation_split`, `test_split`), `output_type` (`loglikelihood`, `multiple_choice`, `loglikelihood_rolling`, `generate_until`), prompt templates (`doc_to_text`, `doc_to_target`, `doc_to_choice` with Jinja2-style `{{}}` interpolation), and `metric_list`. Example: `lm_eval/tasks/arc/arc_easy.yaml` with `dataset_path: allenai/ai2_arc`. **Custom dataset functions** can be specified via `process_docs` as an inline Python callable (serialized in output). **Task discovery** uses `TaskManager` (`lm_eval/tasks/manager.py`) and `TaskIndex` which scans all YAML files in the tasks directory (and optional `include_path`), building a registry of tasks, groups, and tags. **Versioning** is per-task via `metadata.version` in each YAML (e.g. `version: 1.0`), stored in `task.VERSION`. The `TaskConfig` dataclass at `lm_eval/config/task.py:82-168` defines all configuration fields. **Groups** are defined in YAML too (e.g. `lm_eval/tasks/leaderboard/leaderboard.yaml`) with `group` name, `task` list, and `aggregate_metric_list` for hierarchical aggregation. Tags allow task selection by category. **Synthetic data generation** is not a built-in feature; tasks load from HuggingFace datasets hub or local scripts. The `bear` and `toxigen` tasks use published datasets. The `model_written_evals/` directory contains tasks from the "advanced AI risk" benchmark (EleutherAI/advanced_ai_risk dataset), auto-generated by `_generate_configs.py` scripts. There is no formal golden set mechanism beyond the per-task split definitions.


Citations: [lm_eval/tasks/arc/arc_easy.yaml:1-20](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/arc/arc_easy.yaml#L1-L20) · [lm_eval/tasks/manager.py:37-100](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/manager.py#L37-L100) · [lm_eval/config/task.py:82-170](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/config/task.py#L82-L170) · [lm_eval/tasks/leaderboard/leaderboard.yaml:1-22](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/tasks/leaderboard/leaderboard.yaml#L1-L22)

### Arize-ai/phoenix (answered)

Test datasets are defined primarily as **pandas DataFrames** in Python code. Phoenix provides a `download_benchmark_dataset(task, dataset_name)` function (`packages/phoenix-evals/src/phoenix/evals/utils.py:17-36`) that fetches zipped JSONL files from Google Cloud Storage (`storage.googleapis.com/arize-phoenix-assets/evals/`) and loads them as DataFrames. There is no built-in dataset versioning, golden set management, or registry — datasets are code-managed. For **experiment-style evaluations** within the Phoenix app, datasets are curated from production spans saved into the Phoenix project (via the UI or API), then used in experiments. The **PXI eval harness** (`evals/pxi/`) defines datasets as YAML files (`evals/pxi/datasets/`). Each YAML file (e.g. `product_knowledge.yaml:1-700`) declares a `dataset_name`, `description`, a list of evaluators to run (e.g. `assistant_text_substrings_match`, `correct_tools_called`), and `examples`. Each example has an `id`, `splits` (e.g. `[regression, dev]`), an `input` with an ordered `messages` list, `expected` output constraints (e.g. `assistant_text.contains_all`, `tools.forbidden`), and `metadata` including human annotation agreement scores. Examples can be marked `adversarial: true` for wrong-premise tests. The `evals/harbor/` directory contains a Harbor-based eval framework with `agents/`, `tasks/`, `environments/`, and `verifiers/` — the llm_judge verifier (`evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44`) defines a `matches_reference` LLM judge that compares agent replies to reference answers with grading notes. **Synthetic data generation** is described in a cookbook tutorial (`docs/phoenix/cookbook/tracing/generating-synthetic-datasets-for-llm-evaluators-and-agents.mdx`) — this is a documented pattern, not built-in tooling. There is no DSL, version-controlled dataset format, or benchmark task registry beyond the download URL convention.

> **Editor's note.** Correction: Phoenix does version datasets. The server has DatasetVersion, DatasetExampleRevision and DatasetSplit tables (src/phoenix/db/models.py), and experiments are pinned to a dataset_version_id, so datasets are not merely code-managed.

Citations: [packages/phoenix-evals/src/phoenix/evals/utils.py:17-36](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/utils.py#L17-L36)

### NVIDIA/garak (answered)

Test data is stored as JSON files under `garak/data/`, organized by attack technique: `payloads/harmful_behaviors.json`, `dan/`, `harmbench/`, `autodan/`, `gcg/`, `beast/`, `donotanswer/`, etc. Each payload JSON includes `garak_payload_name`, `payload_types`, `intent` (intent taxonomy code), `detector_name`, and a `payloads` array of prompt strings (e.g. `data/payloads/harmful_behaviors.json:1-24`).

Data loading goes through the `LocalDataPath` resolver (`data/__init__.py:29-123`) which searches user data dir (`~/.local/share/garak/data/`) first, then the bundled package dir — enabling user overrides. The `DANProbeMeta` metaclass (`probes/dan.py:28-60`) auto-detects a `prompt_file` attribute from the class name, opens it via `data_path / prompt_file`, and loads JSON arrays into `self.prompts`.

The intent taxonomy in `data/cas/trait_typology.json` defines structured intents (e.g. `S006instructions`, `T009ignore`). `IntentProbe` (`probes/base.py:846-953`) loads intents from `intentservice`, maps them to stubs via `data/cas/intent_stubs/`, and builds prompts dynamically. Probes have tier ratings (OF_CONCERN=1 through UNLISTED=4) via `probes/_tier.py`. Config files are in `resources/garak.core.yaml` (base), `~/.config/garak/garak.site.yaml` (user), and `configs/fast.json` / `configs/bag.yaml` (bundled run templates). No formal dataset versioning — prompts are static JSON files committed to the repo; `plugin_cache.json` tracks plugin metadata timestamps for cache invalidation. No synthetic data generation is built in.


Citations: [garak/data/payloads/harmful_behaviors.json:1-24](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/data/payloads/harmful_behaviors.json#L1-L24) · [garak/data/__init__.py:29-123](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/data/__init__.py#L29-L123) · [garak/probes/dan.py:28-60](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/probes/dan.py#L28-L60) · [garak/probes/base.py:846-953](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/probes/base.py#L846-L953)

### UKGovernmentBEIS/inspect_ai (answered)

**Dataset definition.** A `Dataset` is a `Sequence[Sample]` (`src/inspect_ai/dataset/_dataset.py:144-221`). Each `Sample` has `input`, `target`, `choices`, `id`, `metadata`, `sandbox`, `files`, `setup`, and `checkpoint` fields. The input can be a string or a list of `ChatMessage` objects.

**File formats and sources** (`src/inspect_ai/dataset/_sources/`):
- **hf_dataset** — loads from Hugging Face datasets with retry logic for transient errors, field mapping via `FieldSpec`, and support for splits/shuffling
- **json_dataset** — JSON array of sample dicts
- **csv_dataset** — CSV file with configurable field mapping
- **file_dataset** — generic file (auto-detects JSON/CSV/JSONL by extension)
- **example_dataset** — inline sample generation via a function

**Synthetic data generation** is done in user code by writing Python functions that return `MemoryDataset(list[Sample])` or by implementing a `SampleSource` protocol. There is no built-in synthetic data DSL beyond the `Sample` constructor.

**Versioning.** Inspect does not version datasets natively — versioning is the data provider's responsibility (e.g. Hugging Face dataset revisions pinned by `revision=` in `hf_dataset()`). The `eval_set` system provides log-based versioning through eval set IDs and manifest files.

**Benchmark task registry.** Tasks are registered with `@task` decorator and can be discovered via `inspect list` or `list_tasks()`. The `Task` object binds a `Dataset`, a `Solver`/`Plan`, and optional `Scorer`s, metrics, and config.

**Golden sets.** There is no built-in "golden set" concept — tasks specify their own dataset at construction time. The `SampleSource` protocol allows dynamic sample feeding. The `precomputed_scores` scorer attaches externally computed labels to existing logs by sample ID.


Citations: [src/inspect_ai/dataset/__init__.py:3-27](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/__init__.py#L3-L27) · [src/inspect_ai/dataset/_dataset.py:29-121](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_dataset.py#L29-L121) · [src/inspect_ai/dataset/_dataset.py:144-299](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_dataset.py#L144-L299) · [src/inspect_ai/dataset/_sources/hf.py:1-100](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/dataset/_sources/hf.py#L1-L100)
