How are test datasets and cases defined, generated and versioned?
File formats/DSL; synthetic data generation; versioning; golden sets; benchmark task registries.
Verdict
Phoenix, Opik and Langfuse are the only projects that version test data themselves. DeepEval and Ragas are best at generating test data. lm-evaluation-harness and OpenAI Evals are benchmark registries.
Versioned datasets on a server. Phoenix stores DatasetVersion, per-example revisions and splits, and each experiment is pinned to a dataset version. Langfuse versions items in time with validFrom/validTo. A dataset-level eval reads each item at its validFrom version, and items can point back to the trace they came from. Opik has dataset versions and test suites whose items carry their own evaluator configs and pass policies (runs_per_item, pass_threshold). None of the three generates synthetic data.
Datasets in code, with generators. DeepEval’s Synthesizer rewrites seed inputs with evolutions (reasoning, multi-context, hypothetical and others). It also wraps about 15 standard benchmarks, but dataset versions live on the vendor’s hosted platform. Ragas’s TestsetGenerator builds a knowledge graph from your documents, invents personas and writes single-hop and multi-hop questions. Its version_experiment commits your working tree to a git branch. promptfoo reads tests from YAML, CSV, XLSX, Google Sheets or Hugging Face, and expands array variables into every combination. Versioning is git plus a config snapshot per run. Inspect Samples can carry sandbox files and setup. Versioning is left to the source, such as a Hugging Face revision.
Benchmark and attack registries. lm-evaluation-harness has about 220 benchmark folders and close to 14,000 task YAML files, each with a metadata.version. OpenAI Evals has 463 eval YAML files named <base>.<split>.v<N>, and its JSONL samples are stored in Git LFS. garak ships attack payloads as JSON under garak/data/. A copy in the user data directory overrides them. It has no versioning and no generation.
Pick: Phoenix or Opik when experiments must be pinned to a dataset version. Pick: Ragas or DeepEval to generate questions from your own documents. Pick: lm-evaluation-harness for public benchmarks with versioned task configs.
Per-project answers
langfuse/langfuse
answeredDatasets are stored as Dataset and DatasetItem records in PostgreSQL, with ClickHouse for analytics queries. A dataset has an inputSchema and expectedOutputSchema (optional JSON Schema) for validation (dataset-items.ts lines 62–72). Dataset items carry input, expectedOutput, metadata, plus a sourceTraceId/sourceObservationId when created from traced data. Items are versioned via a temporal model: validFrom/validTo columns enable point-in-time snapshots — fetching an item at a specific date retrieves the version valid then (code in evalService.ts extractVariablesFromTracingData at line 1635: ...(datasetItemValidFrom ? { validFrom: datasetItemValidFrom } : { validTo: null })). Dataset runs (dataset-runs.ts) execute a prompt experiment against dataset items: the experimentServiceClickhouse.ts takes a prompt template, replaces variables with each dataset item's input, calls the LLM, and creates DatasetRunItem records linking the experiment's output traces back to the original dataset items (lines 74–78). No synthetic data generation DSL exists — items are created via the UI, the public API (POST /api/public/dataset-items), or batch CSV/JSON upload. Golden sets are just datasets used as eval benchmarks: you create a dataset, map its columns to evaluator variables in an evaluation rule, and the rule scores every new trace that matches (or you run a batch evaluation). Eval jobs can target TRACE (live scoring) or DATASET (scoring dataset-run items). Benchmark task registration is handled through the evaluation rules system (evaluationRule Prisma model) — rules connect evaluators (judge templates) to dataset items or traces via filters, sampling rates, and variable mappings.
promptfoo/promptfoo
answeredTest cases defined in tests[] of promptfooconfig.yaml. readStandaloneTestsFile (src/util/testCaseReader.ts L113) supports HuggingFace (huggingface://datasets/), Google Sheets (CSV), Azure Blob (az://), CSV, JSON/YAML, XLSX, SharePoint, and JavaScript/Python files. HuggingFace datasets fetched with pagination and concurrent page requests. generateVarCombinations (evaluator.ts L1981-2028) expands array-vars into Cartesian test rows. Scenarios (evaluator.ts L2441-2498) merge scenario config with per-test overrides. defaultTest provides golden-set base assertions. No built-in dataset versioning beyond git; evalsTable stores each runs config snapshot for reproducibility.
comet-ml/opik
answeredOpik treats datasets as first-class objects managed through a backend API with versioning built in.
File formats/DSL: Datasets are collections of items (dicts with arbitrary JSON-serializable content). Created via the SDK (client.create_dataset(name) → dataset.insert(items)) or the UI. Items can contain arbitrary data fields, tags, metadata, and trace/span references. The DatasetItem model (api_objects/dataset/dataset_item.py:40-65) has an EvaluatorItem list — each item can carry per-item evaluator configs (type + config dict, currently only 'llm_judge' supported) and per-item execution policies (runs_per_item, pass_threshold). Datasets export to pandas via to_pandas(), to JSON via to_json(), and stream items as chunks or individually through stream_items() / get_items().
Versioning: Dataset objects have a get_version_info() returning DatasetVersionPublic (id, version_name, etc.). DatasetVersion provides a read-only snapshot at a specific version. Test suites (api_objects/dataset/test_suite/test_suite.py) are a wrapper around datasets: a TestSuite uses a Dataset internally and TestSuiteVersion wraps a DatasetVersion. The TestSuite API exposes get_version_view() for pinning to a specific version.
Synthetic data generation: No built-in synthetic data generators exist in the SDK, but SimulatedUser (simulation/simulated_user.py:10-100) generates synthetic user messages via an LLM based on a persona prompt — used in multi-turn simulation but not for direct dataset creation.
Benchmark task registry: No static benchmark registry is included. Tasks are callables (LLMTask = Callable[[Dict[str, Any]], Any]) passed at evaluation time. Opik integrates with Ragas via RagasMetricWrapper but does not ship standard benchmark datasets.
Dataset filtering: Items can be filtered at retrieval time via OQL filter strings supporting fields like data.field_name, tags, created_at etc. Filters also support dataset_filter_string in evaluate() calls.
openai/evals
answeredFile formats and DSL. Test data uses the JSONL format with standard fields "input" (a chat-message list or plain text) and "ideal" (a string or list of accepted answers). For example, test_fuzzy_match/samples.jsonl entries contain OpenAI-format chat arrays with example few-shot messages and an "ideal" answer list (evals/registry/data/test_fuzzy_match/samples.jsonl:1-3). Eval configurations are YAML files in evals/registry/evals/ — each file names a base spec (e.g. ab: with metrics: [accuracy]) and one or more versioned split entries (e.g. ab.dev.v0:) linking to a class, samples_jsonl path, eval_type, and modelgraded_spec (evals/registry/evals/ab.yaml:1-11). The Registry class loads these YAML files from evals/registry/evals/, completion_fns/, solvers/, modelgraded/, and eval_sets/ directories (evals/registry.py:103-331).
Synthetic data generation. Custom generators live in evals/registry/data/*/ — e.g. simple_physics_engine/samples_generator.py, poker_analysis/poker_analysis_sample_generator.py, mazes/nxn_maze_eval_generator.py, and solve-for-variable/tools/main.py. These emit JSONL files consumed by the evals.
External dataset integration. The MultipleChoice eval class loads from HuggingFace datasets (HellaSwag, Hendrycks MMLU) via datasets.load_dataset() using hf:// URLs (evals/elsuite/multiple_choice.py:20-48). The Lambada eval similarly loads EleutherAI/lambada_openai from HuggingFace (evals/elsuite/lambada.py:42-44).
Versioning. Versioning follows a {base_eval}.{split}.v{N} convention — e.g. ab.dev.v0, prompt-injection.dev.v0, human-safety.test.v0. Splits (dev, test, etc.) are freeform strings. The registry_path parameter supports loading multiple registry directories, and ~/.evals is a secondary path.
Benchmark registries. There are 463 eval YAML files in evals/registry/evals/ and 472 data directories in evals/registry/data/. Eval sets like test-all and test-basic list multiple eval names to run together (evals/registry/eval_sets/test-all.yaml:1-21).
confident-ai/deepeval
answeredEvaluationDataset (deepeval/dataset/dataset.py:90-250) is the core data container, holding either Golden objects (single-turn) or ConversationalGolden objects (multi-turn). Each Golden wraps input, actual_output, expected_output, context, retrieval_context, tools_called, expected_tools, additional_metadata, and optional persona/scenario/expected_outcome fields (deepeval/dataset/golden.py:1-180). File formats: add_test_cases_from_csv_file() parses CSV with configurable column names and list delimiters (deepeval/dataset/dataset.py:266-340). JSON/JSONL loading is supported through add_test_cases_from_json_file(). Column mappings support input, actual_output, expected_output, context, retrieval_context and tool call fields. Synthetic generation: The Synthesizer class (deepeval/synthesizer/synthesizer.py:1-80) produces SyntheticData / ConversationalScenario from source documents via evolution techniques — Reasoning, Multi-context, Concretizing, Constrained, Comparative, Hypothetical, In-Breadth (deepeval/synthesizer/schema.py:14-20). These use LLM-driven evolution templates to re-write seed inputs, and FiltrationConfig / EvolutionConfig / StylingConfig control quality and diversity. Conversational goldens support Persona (demographics, speaking style, interruption behavior) and BackgroundNoiseSettings for voice simulations (deepeval/dataset/golden.py:78-150). Versioning: Datasets have _alias, _id, and _version fields, and the DatasetVersion / CreateDatasetVersionHttpResponse API types enable versioned snapshots on Confident AI (deepeval/dataset/api.py:49-61). Benchmark task registries: The benchmarks/ package contains ~15 standard benchmarks (MMLU, HellaSwag, BIG-Bench-Hard, GSM8K, DROP, SQuAD, HumanEval, etc.) each as a DeepEvalBaseBenchmark subclass that loads HuggingFace datasets and evaluates via load_benchmark_dataset() → evaluate() (deepeval/benchmarks/base_benchmark.py:16-32).
vibrantlabsai/ragas
answeredTest datasets are structured around Pydantic sample models and a flexible backend storage layer.
Data model: SingleTurnSample holds user_input, retrieved_contexts, response, reference, rubrics, and metadata fields (persona_name, query_style, query_length). MultiTurnSample holds a list of HumanMessage|AIMessage|ToolMessage objects with conversation validation (ToolMessage must follow an AIMessage that called tools). Both extend BaseSample with to_dict()/get_features() (dataset_schema.py:29-180).
File formats/DSL: EvaluationDataset wraps lists of samples and supports conversion to/from Hugging Face datasets, pandas DataFrames, CSV, and JSONL files (.to_csv/.to_jsonl/.from_jsonl) (dataset_schema.py:186-406). The DataTable class (dataset.py) provides a list-like interface with pluggable BaseBackend persistence: LocalCSVBackend (per-row CSV), LocalJSONLBackend, GDriveBackend (Google Drive), and InMemoryBackend. Backends are resolved by name via a registry (e.g. "local/csv", "gdrive") (backends/base.py, backends/registry.py, backends/local_csv.py).
Synthetic data generation: Test data is generated via BaseSynthesizer subclasses that operate on a KnowledgeGraph of nodes (chunks of documents) and edges (relationships). SingleHopQuerySynthesizer generates simple lookup queries: it samples nodes, terms, personas, query styles, and lengths, then calls an LLM prompt (QueryAnswerGenerationPrompt) to produce a question+reference answer pair (testset/synthesizers/single_hop/base.py:30-47). MultiHopQuerySynthesizer generates questions requiring 2-3 steps of reasoning over related nodes. The Persona system enables persona-based query generation (e.g. "expert", "novice"). Query-level transforms (splitting, filtering, relationship building) prepare documents before synthesis.
Versioning: version_experiment() (experiment.py:21-100) snapshots the current codebase state to git: it creates a commit and a git branch named ragas/{experiment_name}, returning the commit hash. This links evaluation results to a specific code version.
Golden sets and registries: No benchmark task registry or golden test set management exists in the codebase. Golden/holdout data lives in user-managed files loaded via DataTable.load() from backends. The Dataset class provides train_test_split() for splitting data into training and testing sets for metric alignment.
EleutherAI/lm-evaluation-harness
answeredTasks are defined as YAML config files in lm_eval/tasks/, organized into subdirectories by benchmark (arc, mmlu, hellaswag, etc.). Each YAML declares task name, dataset_path (HuggingFace dataset identifier), dataset_name (subset), split mappings (training_split, validation_split, test_split), output_type (loglikelihood, multiple_choice, loglikelihood_rolling, generate_until), prompt templates (doc_to_text, doc_to_target, doc_to_choice with Jinja2-style {{}} interpolation), and metric_list. Example: lm_eval/tasks/arc/arc_easy.yaml with dataset_path: allenai/ai2_arc. Custom dataset functions can be specified via process_docs as an inline Python callable (serialized in output). Task discovery uses TaskManager (lm_eval/tasks/manager.py) and TaskIndex which scans all YAML files in the tasks directory (and optional include_path), building a registry of tasks, groups, and tags. Versioning is per-task via metadata.version in each YAML (e.g. version: 1.0), stored in task.VERSION. The TaskConfig dataclass at lm_eval/config/task.py:82-168 defines all configuration fields. Groups are defined in YAML too (e.g. lm_eval/tasks/leaderboard/leaderboard.yaml) with group name, task list, and aggregate_metric_list for hierarchical aggregation. Tags allow task selection by category. Synthetic data generation is not a built-in feature; tasks load from HuggingFace datasets hub or local scripts. The bear and toxigen tasks use published datasets. The model_written_evals/ directory contains tasks from the "advanced AI risk" benchmark (EleutherAI/advanced_ai_risk dataset), auto-generated by _generate_configs.py scripts. There is no formal golden set mechanism beyond the per-task split definitions.
Arize-ai/phoenix
answeredTest datasets are defined primarily as pandas DataFrames in Python code. Phoenix provides a download_benchmark_dataset(task, dataset_name) function (packages/phoenix-evals/src/phoenix/evals/utils.py:17-36) that fetches zipped JSONL files from Google Cloud Storage (storage.googleapis.com/arize-phoenix-assets/evals/) and loads them as DataFrames. There is no built-in dataset versioning, golden set management, or registry — datasets are code-managed. For experiment-style evaluations within the Phoenix app, datasets are curated from production spans saved into the Phoenix project (via the UI or API), then used in experiments. The PXI eval harness (evals/pxi/) defines datasets as YAML files (evals/pxi/datasets/). Each YAML file (e.g. product_knowledge.yaml:1-700) declares a dataset_name, description, a list of evaluators to run (e.g. assistant_text_substrings_match, correct_tools_called), and examples. Each example has an id, splits (e.g. [regression, dev]), an input with an ordered messages list, expected output constraints (e.g. assistant_text.contains_all, tools.forbidden), and metadata including human annotation agreement scores. Examples can be marked adversarial: true for wrong-premise tests. The evals/harbor/ directory contains a Harbor-based eval framework with agents/, tasks/, environments/, and verifiers/ — the llm_judge verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) defines a matches_reference LLM judge that compares agent replies to reference answers with grading notes. Synthetic data generation is described in a cookbook tutorial (docs/phoenix/cookbook/tracing/generating-synthetic-datasets-for-llm-evaluators-and-agents.mdx) — this is a documented pattern, not built-in tooling. There is no DSL, version-controlled dataset format, or benchmark task registry beyond the download URL convention.
NVIDIA/garak
answeredTest data is stored as JSON files under garak/data/, organized by attack technique: payloads/harmful_behaviors.json, dan/, harmbench/, autodan/, gcg/, beast/, donotanswer/, etc. Each payload JSON includes garak_payload_name, payload_types, intent (intent taxonomy code), detector_name, and a payloads array of prompt strings (e.g. data/payloads/harmful_behaviors.json:1-24).
Data loading goes through the LocalDataPath resolver (data/__init__.py:29-123) which searches user data dir (~/.local/share/garak/data/) first, then the bundled package dir — enabling user overrides. The DANProbeMeta metaclass (probes/dan.py:28-60) auto-detects a prompt_file attribute from the class name, opens it via data_path / prompt_file, and loads JSON arrays into self.prompts.
The intent taxonomy in data/cas/trait_typology.json defines structured intents (e.g. S006instructions, T009ignore). IntentProbe (probes/base.py:846-953) loads intents from intentservice, maps them to stubs via data/cas/intent_stubs/, and builds prompts dynamically. Probes have tier ratings (OF_CONCERN=1 through UNLISTED=4) via probes/_tier.py. Config files are in resources/garak.core.yaml (base), ~/.config/garak/garak.site.yaml (user), and configs/fast.json / configs/bag.yaml (bundled run templates). No formal dataset versioning — prompts are static JSON files committed to the repo; plugin_cache.json tracks plugin metadata timestamps for cache invalidation. No synthetic data generation is built in.
UKGovernmentBEIS/inspect_ai
answeredDataset definition. A Dataset is a Sequence[Sample] (src/inspect_ai/dataset/_dataset.py:144-221). Each Sample has input, target, choices, id, metadata, sandbox, files, setup, and checkpoint fields. The input can be a string or a list of ChatMessage objects.
File formats and sources (src/inspect_ai/dataset/_sources/):
- hf_dataset — loads from Hugging Face datasets with retry logic for transient errors, field mapping via
FieldSpec, and support for splits/shuffling - json_dataset — JSON array of sample dicts
- csv_dataset — CSV file with configurable field mapping
- file_dataset — generic file (auto-detects JSON/CSV/JSONL by extension)
- example_dataset — inline sample generation via a function
Synthetic data generation is done in user code by writing Python functions that return MemoryDataset(list[Sample]) or by implementing a SampleSource protocol. There is no built-in synthetic data DSL beyond the Sample constructor.
Versioning. Inspect does not version datasets natively — versioning is the data provider's responsibility (e.g. Hugging Face dataset revisions pinned by revision= in hf_dataset()). The eval_set system provides log-based versioning through eval set IDs and manifest files.
Benchmark task registry. Tasks are registered with @task decorator and can be discovered via inspect list or list_tasks(). The Task object binds a Dataset, a Solver/Plan, and optional Scorers, metrics, and config.
Golden sets. There is no built-in "golden set" concept — tasks specify their own dataset at construction time. The SampleSource protocol allows dynamic sample feeding. The precomputed_scores scorer attaches externally computed labels to existing logs by sample ID.
← How is LLM-as-a-judge implemented? · How are evals executed and reported? →