# vibrantlabsai/ragas

> Python library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/vibrantlabsai/ragas (reviewed at commit `298b68274234c060deacab3cf5fb52aa3a20e885`, 2026-02-24)
- Stars: 15950 · Language: Python · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/ragas/

## Overview

Ragas is a Python library for scoring LLM applications, best known for its retrieval-augmented generation (RAG) metrics: faithfulness, answer relevancy, context precision and context recall. Most metrics are short LLM-as-a-judge pipelines. A metric breaks a response into statements, asks a judge model for structured verdicts, and returns a ratio. A smaller set is purely lexical or embedding-based (exact match, BLEU, ROUGE, chrF, string similarity, semantic similarity), and one faithfulness variant swaps the judge for Vectara's HHEM classifier.

At the pinned commit the codebase is midway through a migration, and that shapes how you use it. The **legacy API** is `evaluate()` over an `EvaluationDataset`, with metric objects such as `faithfulness` that implement `single_turn_ascore`. It still works but warns that it is deprecated in favour of `@experiment`. The **new API** is `ragas.metrics.collections`: standalone metrics called as `await metric.ascore(user_input=..., response=..., ...)` that return a `MetricResult`. You orchestrate them yourself inside an `@experiment` function, and results go to a pluggable storage backend. The two metric families have different base classes and are not interchangeable.

The second major feature is test-set generation. `TestsetGenerator` builds a knowledge graph from your documents through extractors, splitters and relationship builders, invents personas, and synthesises single-hop and multi-hop questions with reference answers. Ragas also includes metric alignment and prompt optimisation (a genetic optimiser and a DSPy adapter) trained on human annotations.

## Architecture

```mermaid
flowchart LR
  DS["EvaluationDataset / Dataset"] --> EV["evaluate() (legacy)"]
  DS --> EXP["@experiment arun"]
  EV --> EXE["Executor (async jobs)"]
  EXE --> MET["Metric.single_turn_ascore"]
  MET --> PP["PydanticPrompt"]
  PP --> LLM["llm_factory: Instructor LLM"]
  EXP --> UF["your function + collections metrics"]
  UF --> LLM
  EXP --> BE["Backends: CSV / JSONL / GDrive"]
  EV --> RES["EvaluationResult + RagasTracer"]
  KG["KnowledgeGraph + transforms"] --> TG["TestsetGenerator"]
  TG --> DS
```

| Component | Path | Role |
|---|---|---|
| Legacy runner | `src/ragas/evaluation.py` | `evaluate`/`aevaluate`: defaults, LLM injection, job submission, result assembly |
| Executor | `src/ragas/executor.py` | Async job queue with `max_workers`, batching, cancel, NaN on error |
| Experiments | `src/ragas/experiment.py` | `@experiment` decorator, `arun`, git-based `version_experiment` |
| Legacy metrics | `src/ragas/metrics/_*.py`, `metrics/base.py` | `Metric`, `MetricWithLLM`, `SingleTurnMetric`, `MultiTurnMetric` |
| New metrics | `src/ragas/metrics/collections/` | `BaseMetric.ascore(**kwargs) -> MetricResult`, one folder per metric |
| Custom metrics | `metrics/discrete.py`, `numeric.py`, `ranking.py`, `base.py` | `@discrete_metric` / `@numeric_metric` / `@ranking_metric`, `SimpleLLMMetric` |
| Prompts | `src/ragas/prompt/pydantic_prompt.py` | Typed instruction + few-shot prompt, output parsing, fix-format retry |
| LLMs | `src/ragas/llms/` | `llm_factory` (Instructor/LiteLLM), LangChain and LlamaIndex wrappers |
| Data | `dataset_schema.py`, `dataset.py`, `backends/` | Samples, `EvaluationDataset`, `DataTable` with local CSV/JSONL, in-memory, Google Drive |
| Testset | `src/ragas/testset/` | Knowledge graph, transforms, personas, query synthesizers |
| Integrations | `src/ragas/integrations/` | LangChain, LlamaIndex, LangGraph, Langfuse/MLflow tracing, Opik, Helicone, Bedrock, ... |
| CLI | `src/ragas/cli.py` | `quickstart` templates, `evals` with baseline gate tables |

## How a request flows

Legacy path, `evaluate(dataset, metrics=[faithfulness])`:

1. **Validate and default.** `aevaluate` emits a `DeprecationWarning`, converts a Hugging Face `Dataset` into an `EvaluationDataset`, checks required columns, and falls back to four default metrics when none are given ([evaluation.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L59-L155)).
2. **Inject models.** Any `MetricWithLLM` without an LLM gets `llm_factory("gpt-4o-mini", client=OpenAI())`, embeddings are inferred from that LLM's provider, and every metric's `init(run_config)` runs ([evaluation.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L163-L200)).
3. **Submit jobs.** An `Executor` and a `RagasTracer` callback are created. For every row and metric, `metric.single_turn_ascore(sample, row_callbacks)` is submitted as one job with the run timeout ([evaluation.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L202-L278)).
4. **Run concurrently.** Jobs run under `RunConfig.max_workers` (16 by default, 180 s timeout, 10 retries) ([run_config.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/run_config.py#L51-L54)). A failed job logs the exception and returns `NaN` unless `raise_exceptions=True` ([executor.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/executor.py#L64-L138)).
5. **Score one row.** `single_turn_ascore` keeps only the metric's required columns, opens a callback group, awaits `_single_turn_ascore` under `asyncio.wait_for`, and sends an analytics event ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L460-L513)). Faithfulness generates statements, then NLI verdicts, through `PydanticPrompt.generate`.
6. **Call the judge.** `PydanticPrompt.generate_multiple` renders the instruction, JSON schema, examples and input into one string. It dispatches to a LangChain LLM, an Instructor LLM (structured output via `response_model`) or a `BaseRagasLLM`, then validates the output into the Pydantic model ([pydantic_prompt.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/prompt/pydantic_prompt.py#L188-L320)).
7. **Assemble.** Results are mapped back to `{metric_name: score}` per row and wrapped in an `EvaluationResult` that carries the tracer's run tree and optional cost data ([evaluation.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L284-L300)).

New path: decorate an async function with `@experiment()`, call collections metrics inside it, and run `await fn.arun(dataset, name=...)`. `arun` schedules one task per dataset row, appends each returned row to an `Experiment` table, prints a warning for any row that raised, and saves through the backend ([experiment.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L116-L232)).

## Key components

### Collections metrics

A collections metric declares its components in `__init__` (for example `llm: InstructorBaseRagasLLM`), and the base class rejects legacy LangChain-style wrappers ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/collections/base.py#L13-L59)). `Faithfulness` is a clear example. It makes one call that turns the answer into atomic statements and one NLI call over the joined contexts, and it returns faithful/total, or `NaN` when no statements come back ([metric.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/collections/faithfulness/metric.py#L88-L163)). These metrics are `SimpleBaseMetric` subclasses with `score`/`ascore`/`batch_score` ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L689-L752)), not `Metric`, so `evaluate()` rejects them.

### LLM layer

`llm_factory(model, provider, client, adapter, cache, mode)` wraps a pre-built client with Instructor (`from_openai`, `from_anthropic`, `from_genai`, `from_litellm`, ...) or LiteLLM and returns an object with `generate`/`agenerate(prompt, response_model)` ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L606-L640), [L1117-L1163](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L1117-L1163)). Passing `cache=DiskCacheBackend()` memoises generations on disk. Legacy `BaseRagasLLM` subclasses get the same wrapper on `generate_text` ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L55-L66)).

### Rubric judges

`AspectCritic` turns a natural-language definition into a Yes/No instruction, and predefined critics cover harmfulness, maliciousness, coherence and similar aspects. It has a `strictness` parameter meant for majority voting, but at this commit both the single-turn and multi-turn paths call `generate` once and vote over a one-element list ([_aspect_critic.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_aspect_critic.py#L143-L212)). Treat it as a single-sample judge. `SimpleCriteriaScore` follows the same pattern.

### Tracing and results

`RagasTracer` is a LangChain-style callback handler that records an evaluation → row → metric → prompt tree of `ChainRun`s with inputs and outputs ([callbacks.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/callbacks.py#L80-L121)). This is what lets you inspect why a row scored low. Langfuse and MLflow helpers under `integrations/tracing/` and LangSmith/Opik/Helicone integrations push to external tools. There is no OpenTelemetry or production-traffic capture.

### Testset generation

`TestsetGenerator.generate(testset_size, query_distribution, num_personas)` uses the default single-hop and multi-hop synthesizers when no distribution is given, generates personas from the knowledge graph, and runs scenario and sample generation through the same `Executor` ([generate.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/testset/synthesizers/generate.py#L414-L520)). Entry points exist for LangChain documents, LlamaIndex documents and raw chunks.

## Extending it

- **Quick custom metric.** Decorate a function with `@discrete_metric(name=..., allowed_values=[...])`, `@numeric_metric` or `@ranking_metric`, or use `SimpleLLMMetric` with your own prompt and response model. It can be saved, loaded and aligned against human labels.
- **Full metric.** In the collections style, subclass `BaseMetric` and implement `ascore`. In the legacy style, subclass `MetricWithLLM` + `SingleTurnMetric` and implement `_single_turn_ascore`.
- **Prompts.** Subclass `PydanticPrompt[Input, Output]` with an `instruction` and `examples`. Prompts can be translated with `adapt()` and saved to JSON.
- **Storage.** Implement `BaseBackend` and register it, or use `local/csv`, `local/jsonl`, `inmemory` or `gdrive`.
- **Versioning.** `version_experiment(name)` stages tracked changes, commits them, and creates a `ragas/<name>` branch, so a result can be tied to code. It writes to your git repository, so call it deliberately ([experiment.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L21-L100)).

## Running it

- `pip install ragas` (Python 3.9+). Extras add GitPython (`git`), Langfuse and MLflow (`tracing`), Google Drive, DSPy and others.
- Create an LLM explicitly: `llm = llm_factory("gpt-4o-mini", client=AsyncOpenAI())`. Use an async client with collections metrics, because `agenerate` raises on a sync client. The client argument is mandatory: "text-only mode" was removed, and `llm_factory(model)` without a client raises a `ValueError` with migration instructions ([base.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L690-L699)).
- `ragas quickstart` scaffolds example projects. `ragas evals` prints current-versus-baseline tables with a pass/fail gate that tolerates regressions under 0.01 ([cli.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/cli.py#L54-L98)).
- Usage analytics are sent by default, and setting `RAGAS_DO_NOT_TRACK=true` turns them off.

## Strengths and caveats

- **Strength: the reference RAG metrics.** Faithfulness, context precision/recall, noise sensitivity and context entity recall are well-defined, small pipelines whose prompts are readable Pydantic classes.
- **Strength: structured judge I/O.** Every judge call has typed input and output models, enforced by Instructor or validated with a fix-format retry, so parsing failures are rare and explicit.
- **Strength: testset generation** from your own corpus with personas and multi-hop questions is more complete than in most eval libraries.
- **Caveat: two APIs at once.** Legacy metrics with `evaluate()` and collections metrics with `@experiment` coexist, with different base classes, call signatures and LLM wrappers. Docs and examples mix them.
- **Caveat: errors become NaN.** Both runners swallow per-row failures by default (`NaN` in `evaluate`, a printed warning and a dropped row in `arun`), so check counts, not just means. `arun` also starts every row at once with no concurrency limit.
- **Caveat: `strictness` voting is inert** in `AspectCritic` and `SimpleCriteriaScore` at this commit.
- **Caveat: CLI `evals` is half-wired.** It looks for a project object with `get_dataset`/`get_experiment` and an experiment with `run_async`, and the code notes that the Project class is not implemented yet ([cli.py](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/cli.py#L370-L457)). Prefer the Python API for CI gates.
- **Caveat: offline only.** No OpenTelemetry, no online scoring of production traffic and no red-teaming tooling.

*Sources: code at 298b682, deepwiki-open wiki (12 pages), verified Q&A.*

## How vibrantlabsai/ragas answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Ragas provides ~30+ built-in metrics across three implementation categories.

**Heuristic/statistical metrics** (no LLM call): `ExactMatch` (string equality), `StringPresence` (substring match), `NonLLMStringSimilarity` (Levenshtein/Hamming/Jaro/Jaro-Winkler via rapidfuzz), `BleuScore`, `RougeScore`, `ChrfScore`, and `DataCompyScore`. These are pure Python computations implementing SingleTurnMetric. Example: `ExactMatch` returns `float(sample.reference == sample.response)` (metrics/_string.py:29-35).

**LLM-based metrics**: `Faithfulness` (statement decomposition → NLI verdicts), `AnswerRelevancy`, `AnswerCorrectness`, `FactualCorrectness`, `ContextPrecision` (multiple variants: LLM, NonLLM, ID-based), `ContextRecall`, `ContextEntityRecall`, `NoiseSensitivity`, `SummarizationScore`, `TopicAdherenceScore`, `ToolCallAccuracy`, `ToolCallF1`, `MultiModalFaithfulness`, `MultiModalRelevance`, `SQLSemanticEquivalence`, `GoalAccuracy` (agent). These extend `MetricWithLLM` and implement `_single_turn_ascore()` which calls a `PydanticPrompt.generate()` against an LLM. Example: Faithfulness breaks the response into atomic statements using `StatementGeneratorPrompt`, then judges each against retrieved contexts via `NLIStatementPrompt` (metrics/_faithfulness.py:152-214).

**Hybrid/embedding-based**: `FaithfulnesswithHHEM` replaces the LLM NLI stage with a HuggingFace transformer (`vectara/hallucination_evaluation_model`) (metrics/_faithfulness.py:217-273). `SemanticSimilarity` uses embeddings. `NonLLM variants` of ContextPrecision/ContextRecall use embedding cosine similarity instead of LLM calls.

**Custom metrics**: Three decorator-based primitives: `@discrete_metric` (categorical output, e.g. 'pass'/'fail'), `@numeric_metric` (continuous float), `@ranking_metric` (ranked output). The `SimpleLLMMetric` class accepts a custom prompt string and Pydantic response model for ad-hoc LLM-as-a-judge metrics with save/load/reproducibility (metrics/base.py:846-935). `AspectCritic` and `SimpleCriteriaScore` take user-defined rubric strings for binary/discrete scoring with automatic strictness-based majority voting.

**Per-trace vs per-task**: All metrics implement `SingleTurnMetric` (single QA pair) or `MultiTurnMetric` (conversation). The `Metric` base class declares `required_columns` per type, e.g. Faithfulness needs `{user_input, response, retrieved_contexts}` for single-turn. Collection metrics (e.g. `AnswerCorrectness` composite) internally orchestrate multiple sub-metrics per trace.


Citations: [src/ragas/metrics/__init__.py:1-99](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/__init__.py#L1-L99) · [src/ragas/metrics/base.py:74-94](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L74-L94) · [src/ragas/metrics/_string.py:20-99](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_string.py#L20-L99) · [src/ragas/metrics/_faithfulness.py:134-215](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_faithfulness.py#L134-L215) · [src/ragas/metrics/base.py:688-935](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L688-L935) · [src/ragas/metrics/discrete.py:1-50](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/discrete.py#L1-L50)

### How is LLM-as-a-judge implemented? (answered)

LLM-as-a-judge is a core primitive in Ragas, built on `PydanticPrompt` with structured output.

**Judge prompts and rubrics**: The `PydanticPrompt` class (prompt/pydantic_prompt.py:82-350) is a generic template that takes an `instruction` string and `input_model`/`output_model` as Pydantic type parameters. The instruction describes the judgement criteria. At inference time, `to_string()` renders the instruction, output schema (auto-generated JSON Schema from `output_model.model_json_schema()`), few-shot examples, and the actual input into one prompt sent to the LLM. For instance, `NLIStatementPrompt` instructs "return verdict as 1 if the statement can be directly inferred..." with a `StatementFaithfulnessAnswer` output model (metrics/_faithfulness.py:73-131).

**Structured output**: `PydanticPrompt.generate()` calls the LLM with `response_model=self.output_model`, enforcing structured JSON via the `Instructor` library (llms/base.py:606-748). The `llm_factory()` wraps provider clients with instructor patching (OpenAI → `instructor.from_openai`, Anthropic → `instructor.from_anthropic`, etc.) and passes the Pydantic model as `response_model`. Returns are validated immediately into the output Pydantic model, with automatic retry (`retries_left: int = 3`) via `RagasOutputParser.parse_output_string()` which on parse failure calls `FixOutputFormat` to ask the LLM to fix malformed JSON (prompt/pydantic_prompt.py:525-558).

**Multi-sample/consensus**: The `Ensember` class (metrics/base.py:641-686) implements majority voting over multiple LLM outputs for the same input. `AspectCritic` and `SimpleCriteriaScore` have a `strictness` parameter — when >1, they call the LLM N times and aggregate via `Counter(verdicts).most_common(1)[0][0]` (metrics/_aspect_critic.py:155-165, metrics/_simple_criteria.py:152-162).

**Judge model choice**: The judge LLM is set per-metric via `metric.llm`, or inherited from the `evaluate()`/`@experiment` call-level llm parameter. `llm_factory()` supports OpenAI, Anthropic, Google/Gemini, Groq, Mistral, Azure, LiteLLM (100+ providers), and more. It auto-detects the correct instructor patching strategy. The default fallback when no LLM is provided is `gpt-4o-mini` via `OpenAI()` (evaluation.py:176-179).

**Calibration/bias controls**: No dedicated calibration or bias-mitigation module exists. The project provides an `Ensember` for consensus voting but no systematic bias measurement, calibration curves, or position-bias detection. The `AspectCritic` with predefined aspects (harmfulness, maliciousness, etc.) can be used as a safety check, but bias controls are absent.

> **Editor's note.** Correction: `strictness` is inert at this commit. `AspectCritic` and `SimpleCriteriaScore` call `prompt.generate` once and vote over a one-element list (src/ragas/metrics/_aspect_critic.py L173-L212), so no multi-sample consensus happens.

Citations: [src/ragas/prompt/pydantic_prompt.py:82-350](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/prompt/pydantic_prompt.py#L82-L350) · [src/ragas/llms/base.py:606-748](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L606-L748) · [src/ragas/prompt/pydantic_prompt.py:525-558](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/prompt/pydantic_prompt.py#L525-L558) · [src/ragas/metrics/base.py:641-686](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L641-L686) · [src/ragas/metrics/_aspect_critic.py:75-165](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_aspect_critic.py#L75-L165) · [src/ragas/evaluation.py:170-180](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L170-L180)

### How are test datasets and cases defined, generated and versioned? (answered)

Test datasets are structured around Pydantic sample models and a flexible backend storage layer.

**Data model**: `SingleTurnSample` holds `user_input`, `retrieved_contexts`, `response`, `reference`, `rubrics`, and metadata fields (persona_name, query_style, query_length). `MultiTurnSample` holds a list of `HumanMessage|AIMessage|ToolMessage` objects with conversation validation (ToolMessage must follow an AIMessage that called tools). Both extend `BaseSample` with `to_dict()`/`get_features()` (dataset_schema.py:29-180).

**File formats/DSL**: `EvaluationDataset` wraps lists of samples and supports conversion to/from Hugging Face datasets, pandas DataFrames, CSV, and JSONL files (.to_csv/.to_jsonl/.from_jsonl) (dataset_schema.py:186-406). The `DataTable` class (dataset.py) provides a list-like interface with pluggable `BaseBackend` persistence: `LocalCSVBackend` (per-row CSV), `LocalJSONLBackend`, `GDriveBackend` (Google Drive), and `InMemoryBackend`. Backends are resolved by name via a registry (e.g. "local/csv", "gdrive") (backends/base.py, backends/registry.py, backends/local_csv.py).

**Synthetic data generation**: Test data is generated via `BaseSynthesizer` subclasses that operate on a `KnowledgeGraph` of nodes (chunks of documents) and edges (relationships). `SingleHopQuerySynthesizer` generates simple lookup queries: it samples nodes, terms, personas, query styles, and lengths, then calls an LLM prompt (`QueryAnswerGenerationPrompt`) to produce a question+reference answer pair (testset/synthesizers/single_hop/base.py:30-47). `MultiHopQuerySynthesizer` generates questions requiring 2-3 steps of reasoning over related nodes. The `Persona` system enables persona-based query generation (e.g. "expert", "novice"). Query-level `transforms` (splitting, filtering, relationship building) prepare documents before synthesis.

**Versioning**: `version_experiment()` (experiment.py:21-100) snapshots the current codebase state to git: it creates a commit and a git branch named `ragas/{experiment_name}`, returning the commit hash. This links evaluation results to a specific code version.

**Golden sets and registries**: No benchmark task registry or golden test set management exists in the codebase. Golden/holdout data lives in user-managed files loaded via `DataTable.load()` from backends. The `Dataset` class provides `train_test_split()` for splitting data into training and testing sets for metric alignment.


Citations: [src/ragas/dataset_schema.py:29-180](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L29-L180) · [src/ragas/dataset_schema.py:186-300](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L186-L300) · [src/ragas/dataset.py:31-120](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset.py#L31-L120) · [src/ragas/testset/synthesizers/single_hop/base.py:1-80](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/testset/synthesizers/single_hop/base.py#L1-L80) · [src/ragas/experiment.py:21-100](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L21-L100) · [src/ragas/backends/local_csv.py:1-80](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/backends/local_csv.py#L1-L80)

### How are evals executed and reported? (answered)

Evaluation execution is centered on the `Executor` class and the newer `@experiment` decorator pattern.

**Runners**: The `Executor` (executor.py:18-216) manages async job queues. Jobs are submitted via `.submit(callable, *args, name=...)` and executed concurrently via `as_completed()` with a configurable `max_workers` limit from `RunConfig`. It wraps each callable with error handling (returns `np.nan` on failure unless `raise_exceptions=True`). Both synchronous `.results()` and async `.aresults()` paths are supported. In the `evaluate()`/`aevaluate()` flow (now deprecated in favor of `@experiment`), one `Executor` is created per evaluation run. Each metric's `single_turn_ascore()` or `multi_turn_ascore()` is submitted as a separate job per dataset row (evaluation.py:244-276).

**Parallelism and caching**: Concurrency is controlled by `RunConfig.max_workers` (default unlimited). Batching is optional via `batch_size` — batches process as nested progress bars. LLM responses are cacheable via `CacheInterface`: `DiskCacheBackend` (diskcache) stores results keyed by SHA-256 of the function + arguments. The `cacher` decorator wraps `generate_text`/`agenerate_text` (llms/base.py:60-66). This provides ~60x speedup for repeated evaluations with identical inputs.

**CI integration**: No native CI integration — tests use standard `pytest` via `make test`. The CLI (`ragas` command via typer) can run experiments from the command line, making it scriptable in CI pipelines.

**Result storage**: Results are returned as `EvaluationResult` (dataset_schema.py:411-552), which exposes per-metric scores as dict-of-lists, mean-aggregated scores via `_repr_dict`, cost tracking via `CostCallbackHandler` (total tokens, total cost), and full run traces via `RagasTracer` (callback tree of evaluation → rows → metrics → prompts). Conversion to pandas is available via `.to_pandas()`. The `Experiment`/`@experiment` decorator pattern (experiment.py:103-232) saves results to a backend automatically.

**Comparison, regression, dashboards**: The CLI (cli.py:54-99) implements baseline comparison: it renders Rich tables showing current, baseline, delta values and pass/fail gates with small-regression tolerance (±0.01). No web dashboard exists. The CLI supports `ragas evaluate` for running experiments and `ragas compare` for baseline comparisons with categorical frequency tables.

> **Editor's note.** Correction: `RunConfig.max_workers` defaults to 16, not unlimited. The CLI has no `evaluate` or `compare` commands, only `evals` (baseline gate tables via `--baseline`), `quickstart` and `hello_world`, and `evals` depends on a Project class the code marks as not implemented.

Citations: [src/ragas/executor.py:18-216](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/executor.py#L18-L216) · [src/ragas/evaluation.py:59-345](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L59-L345) · [src/ragas/experiment.py:103-232](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/experiment.py#L103-L232) · [src/ragas/cli.py:54-100](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/cli.py#L54-L100) · [src/ragas/dataset_schema.py:411-552](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L411-L552)

### How are traces or production data captured and linked to evaluations? (answered)

Ragas provides built-in tracing via a callback-based architecture and first-party integrations with two tracing platforms.

**Built-in tracing**: The `RagasTracer` (callbacks.py:80-121) is a `BaseCallbackHandler` that records every evaluation run as a tree of `ChainRun` nodes. Each evaluation creates a tree: evaluation → row → metric → prompt, with parent-child relationships tracked via `run_id`/`parent_run_id`. This captures inputs, outputs, and metadata at every level. Traces are stored in memory during execution and attached to `EvaluationResult.ragas_traces` for post-hoc analysis. `parse_run_traces()` traverses the tree to produce flat per-row lists of metric scores and prompt I/O (callbacks.py:134-173).

**Tracing integrations** (integrations/tracing/): `LangfuseTrace` and `MLflowTrace` provide explicit platform-level observability. The `@observe()` decorator wraps evaluation functions so every metric computation and LLM call is recorded in the tracing backend. `sync_trace()` returns a trace object with `.get_url()` for deep linking into the observability platform (integrations/tracing/__init__.py:1-78). These are optional dependencies — imported lazily.

**LLM observability integrations**: `HeliconeConfig` proxies LLM calls through Helicone for request logging. `Langsmith` and `Opik` integrations are available (integrations/langsmith.py, integrations/opik.py) for tracking LLM calls and evaluation results. Framework integrations (LangChain, LlamaIndex, Griptape, LangGraph) bridge those ecosystems' own observability.

**Online vs offline evals**: No distinction exists — all evaluations are offline/batch by design. There is no production traffic capturing, SDK instrumentation for live applications, or OpenTelemetry integration. The project integrates with no APM/vendor for online monitoring of deployed LLM applications.

**Feedback and annotation**: The `PromptAnnotation`/`SampleAnnotation`/`MetricAnnotation` classes (dataset_schema.py:555-837) support human-in-the-loop annotation for metric training: users can edit prompt outputs, mark samples as accepted/rejected, and use this data to optimize instruction prompts or few-shot demonstrations via `MetricWithLLM.train()`. This is a training feedback loop, not production feedback collection.


Citations: [src/ragas/integrations/tracing/__init__.py:1-78](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/integrations/tracing/__init__.py#L1-L78) · [src/ragas/dataset_schema.py:555-660](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/dataset_schema.py#L555-L660) · [src/ragas/metrics/base.py:191-306](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L191-L306) · [src/ragas/integrations/helicone.py:1-30](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/integrations/helicone.py#L1-L30)

### Does it support red-teaming or safety testing, and how? (answered)

Ragas does **not** have a dedicated red-teaming or safety testing framework. There are no adversarial probes, attack plugins, jailbreak tests, prompt injection test suites, or vulnerability reporting capabilities.

**What it does instead**: The `AspectCritic` metric (metrics/_aspect_critic.py:75-242) evaluates outputs against custom criteria defined as natural-language strings. Ragas ships several predefined aspects including `harmfulness` ("Does the submission cause or have the potential to cause harm...") and `maliciousness` ("Is the submission intended to harm, deceive, or exploit users?") (metrics/_aspect_critic.py:215-221). These are LLM-as-a-judge binary verdicts, not adversarial probing tools — they check whether a *given output* is harmful, rather than attempting to *provoke* harmful outputs.

**What is absent**: No adversarial input generation (no probe synthesis, no attack chaining, no red-team scenario specification). No jailbreak/prompt injection detection metrics — while an `AspectCritic` with a custom definition could be written to judge a response for injection success, it provides no tooling to generate injection attempts. There is no vulnerability reporting mechanism.

**Tangential reference**: A comment in `llms/adapters/__init__.py:80` references "HARM_CATEGORY_JAILBREAK" in the context of Google Gemini safety settings causing issues with the Instructor library — this is an upstream workaround note, not a ragas feature.


Citations: [src/ragas/metrics/_aspect_critic.py:75-242](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_aspect_critic.py#L75-L242) · [src/ragas/metrics/_aspect_critic.py:215-235](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_aspect_critic.py#L215-L235) · [src/ragas/llms/adapters/__init__.py:78-82](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/adapters/__init__.py#L78-L82)
