vibrantlabsai/ragas
Python library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.
Overview
Ragas is a Python library for scoring LLM applications, best known for its retrieval-augmented generation (RAG) metrics: faithfulness, answer relevancy, context precision and context recall. Most metrics are short LLM-as-a-judge pipelines. A metric breaks a response into statements, asks a judge model for structured verdicts, and returns a ratio. A smaller set is purely lexical or embedding-based (exact match, BLEU, ROUGE, chrF, string similarity, semantic similarity), and one faithfulness variant swaps the judge for Vectara’s HHEM classifier.
At the pinned commit the codebase is midway through a migration, and that shapes how you use it. The legacy API is evaluate() over an EvaluationDataset, with metric objects such as faithfulness that implement single_turn_ascore. It still works but warns that it is deprecated in favour of @experiment. The new API is ragas.metrics.collections: standalone metrics called as await metric.ascore(user_input=..., response=..., ...) that return a MetricResult. You orchestrate them yourself inside an @experiment function, and results go to a pluggable storage backend. The two metric families have different base classes and are not interchangeable.
The second major feature is test-set generation. TestsetGenerator builds a knowledge graph from your documents through extractors, splitters and relationship builders, invents personas, and synthesises single-hop and multi-hop questions with reference answers. Ragas also includes metric alignment and prompt optimisation (a genetic optimiser and a DSPy adapter) trained on human annotations.
Architecture
flowchart LR
DS["EvaluationDataset / Dataset"] --> EV["evaluate() (legacy)"]
DS --> EXP["@experiment arun"]
EV --> EXE["Executor (async jobs)"]
EXE --> MET["Metric.single_turn_ascore"]
MET --> PP["PydanticPrompt"]
PP --> LLM["llm_factory: Instructor LLM"]
EXP --> UF["your function + collections metrics"]
UF --> LLM
EXP --> BE["Backends: CSV / JSONL / GDrive"]
EV --> RES["EvaluationResult + RagasTracer"]
KG["KnowledgeGraph + transforms"] --> TG["TestsetGenerator"]
TG --> DS
| Component | Path | Role |
|---|---|---|
| Legacy runner | src/ragas/evaluation.py |
evaluate/aevaluate: defaults, LLM injection, job submission, result assembly |
| Executor | src/ragas/executor.py |
Async job queue with max_workers, batching, cancel, NaN on error |
| Experiments | src/ragas/experiment.py |
@experiment decorator, arun, git-based version_experiment |
| Legacy metrics | src/ragas/metrics/_*.py, metrics/base.py |
Metric, MetricWithLLM, SingleTurnMetric, MultiTurnMetric |
| New metrics | src/ragas/metrics/collections/ |
BaseMetric.ascore(**kwargs) -> MetricResult, one folder per metric |
| Custom metrics | metrics/discrete.py, numeric.py, ranking.py, base.py |
@discrete_metric / @numeric_metric / @ranking_metric, SimpleLLMMetric |
| Prompts | src/ragas/prompt/pydantic_prompt.py |
Typed instruction + few-shot prompt, output parsing, fix-format retry |
| LLMs | src/ragas/llms/ |
llm_factory (Instructor/LiteLLM), LangChain and LlamaIndex wrappers |
| Data | dataset_schema.py, dataset.py, backends/ |
Samples, EvaluationDataset, DataTable with local CSV/JSONL, in-memory, Google Drive |
| Testset | src/ragas/testset/ |
Knowledge graph, transforms, personas, query synthesizers |
| Integrations | src/ragas/integrations/ |
LangChain, LlamaIndex, LangGraph, Langfuse/MLflow tracing, Opik, Helicone, Bedrock, … |
| CLI | src/ragas/cli.py |
quickstart templates, evals with baseline gate tables |
How a request flows
Legacy path, evaluate(dataset, metrics=[faithfulness]):
- Validate and default.
aevaluateemits aDeprecationWarning, converts a Hugging FaceDatasetinto anEvaluationDataset, checks required columns, and falls back to four default metrics when none are given (evaluation.py). - Inject models. Any
MetricWithLLMwithout an LLM getsllm_factory("gpt-4o-mini", client=OpenAI()), embeddings are inferred from that LLM’s provider, and every metric’sinit(run_config)runs (evaluation.py). - Submit jobs. An
Executorand aRagasTracercallback are created. For every row and metric,metric.single_turn_ascore(sample, row_callbacks)is submitted as one job with the run timeout (evaluation.py). - Run concurrently. Jobs run under
RunConfig.max_workers(16 by default, 180 s timeout, 10 retries) (run_config.py). A failed job logs the exception and returnsNaNunlessraise_exceptions=True(executor.py). - Score one row.
single_turn_ascorekeeps only the metric’s required columns, opens a callback group, awaits_single_turn_ascoreunderasyncio.wait_for, and sends an analytics event (base.py). Faithfulness generates statements, then NLI verdicts, throughPydanticPrompt.generate. - Call the judge.
PydanticPrompt.generate_multiplerenders the instruction, JSON schema, examples and input into one string. It dispatches to a LangChain LLM, an Instructor LLM (structured output viaresponse_model) or aBaseRagasLLM, then validates the output into the Pydantic model (pydantic_prompt.py). - Assemble. Results are mapped back to
{metric_name: score}per row and wrapped in anEvaluationResultthat carries the tracer’s run tree and optional cost data (evaluation.py).
New path: decorate an async function with @experiment(), call collections metrics inside it, and run await fn.arun(dataset, name=...). arun schedules one task per dataset row, appends each returned row to an Experiment table, prints a warning for any row that raised, and saves through the backend (experiment.py).
Key components
Collections metrics
A collections metric declares its components in __init__ (for example llm: InstructorBaseRagasLLM), and the base class rejects legacy LangChain-style wrappers (base.py). Faithfulness is a clear example. It makes one call that turns the answer into atomic statements and one NLI call over the joined contexts, and it returns faithful/total, or NaN when no statements come back (metric.py). These metrics are SimpleBaseMetric subclasses with score/ascore/batch_score (base.py), not Metric, so evaluate() rejects them.
LLM layer
llm_factory(model, provider, client, adapter, cache, mode) wraps a pre-built client with Instructor (from_openai, from_anthropic, from_genai, from_litellm, …) or LiteLLM and returns an object with generate/agenerate(prompt, response_model) (base.py, L1117-L1163). Passing cache=DiskCacheBackend() memoises generations on disk. Legacy BaseRagasLLM subclasses get the same wrapper on generate_text (base.py).
Rubric judges
AspectCritic turns a natural-language definition into a Yes/No instruction, and predefined critics cover harmfulness, maliciousness, coherence and similar aspects. It has a strictness parameter meant for majority voting, but at this commit both the single-turn and multi-turn paths call generate once and vote over a one-element list (_aspect_critic.py). Treat it as a single-sample judge. SimpleCriteriaScore follows the same pattern.
Tracing and results
RagasTracer is a LangChain-style callback handler that records an evaluation → row → metric → prompt tree of ChainRuns with inputs and outputs (callbacks.py). This is what lets you inspect why a row scored low. Langfuse and MLflow helpers under integrations/tracing/ and LangSmith/Opik/Helicone integrations push to external tools. There is no OpenTelemetry or production-traffic capture.
Testset generation
TestsetGenerator.generate(testset_size, query_distribution, num_personas) uses the default single-hop and multi-hop synthesizers when no distribution is given, generates personas from the knowledge graph, and runs scenario and sample generation through the same Executor (generate.py). Entry points exist for LangChain documents, LlamaIndex documents and raw chunks.
Extending it
- Quick custom metric. Decorate a function with
@discrete_metric(name=..., allowed_values=[...]),@numeric_metricor@ranking_metric, or useSimpleLLMMetricwith your own prompt and response model. It can be saved, loaded and aligned against human labels. - Full metric. In the collections style, subclass
BaseMetricand implementascore. In the legacy style, subclassMetricWithLLM+SingleTurnMetricand implement_single_turn_ascore. - Prompts. Subclass
PydanticPrompt[Input, Output]with aninstructionandexamples. Prompts can be translated withadapt()and saved to JSON. - Storage. Implement
BaseBackendand register it, or uselocal/csv,local/jsonl,inmemoryorgdrive. - Versioning.
version_experiment(name)stages tracked changes, commits them, and creates aragas/<name>branch, so a result can be tied to code. It writes to your git repository, so call it deliberately (experiment.py).
Running it
pip install ragas(Python 3.9+). Extras add GitPython (git), Langfuse and MLflow (tracing), Google Drive, DSPy and others.- Create an LLM explicitly:
llm = llm_factory("gpt-4o-mini", client=AsyncOpenAI()). Use an async client with collections metrics, becauseagenerateraises on a sync client. The client argument is mandatory: “text-only mode” was removed, andllm_factory(model)without a client raises aValueErrorwith migration instructions (base.py). ragas quickstartscaffolds example projects.ragas evalsprints current-versus-baseline tables with a pass/fail gate that tolerates regressions under 0.01 (cli.py).- Usage analytics are sent by default, and setting
RAGAS_DO_NOT_TRACK=trueturns them off.
Strengths and caveats
- Strength: the reference RAG metrics. Faithfulness, context precision/recall, noise sensitivity and context entity recall are well-defined, small pipelines whose prompts are readable Pydantic classes.
- Strength: structured judge I/O. Every judge call has typed input and output models, enforced by Instructor or validated with a fix-format retry, so parsing failures are rare and explicit.
- Strength: testset generation from your own corpus with personas and multi-hop questions is more complete than in most eval libraries.
- Caveat: two APIs at once. Legacy metrics with
evaluate()and collections metrics with@experimentcoexist, with different base classes, call signatures and LLM wrappers. Docs and examples mix them. - Caveat: errors become NaN. Both runners swallow per-row failures by default (
NaNinevaluate, a printed warning and a dropped row inarun), so check counts, not just means.arunalso starts every row at once with no concurrency limit. - Caveat:
strictnessvoting is inert inAspectCriticandSimpleCriteriaScoreat this commit. - Caveat: CLI
evalsis half-wired. It looks for a project object withget_dataset/get_experimentand an experiment withrun_async, and the code notes that the Project class is not implemented yet (cli.py). Prefer the Python API for CI gates. - Caveat: offline only. No OpenTelemetry, no online scoring of production traffic and no red-teaming tooling.
Sources: code at 298b682, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit 298b682. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredRagas provides ~30+ built-in metrics across three implementation categories.
Heuristic/statistical metrics (no LLM call): ExactMatch (string equality), StringPresence (substring match), NonLLMStringSimilarity (Levenshtein/Hamming/Jaro/Jaro-Winkler via rapidfuzz), BleuScore, RougeScore, ChrfScore, and DataCompyScore. These are pure Python computations implementing SingleTurnMetric. Example: ExactMatch returns float(sample.reference == sample.response) (metrics/_string.py:29-35).
LLM-based metrics: Faithfulness (statement decomposition → NLI verdicts), AnswerRelevancy, AnswerCorrectness, FactualCorrectness, ContextPrecision (multiple variants: LLM, NonLLM, ID-based), ContextRecall, ContextEntityRecall, NoiseSensitivity, SummarizationScore, TopicAdherenceScore, ToolCallAccuracy, ToolCallF1, MultiModalFaithfulness, MultiModalRelevance, SQLSemanticEquivalence, GoalAccuracy (agent). These extend MetricWithLLM and implement _single_turn_ascore() which calls a PydanticPrompt.generate() against an LLM. Example: Faithfulness breaks the response into atomic statements using StatementGeneratorPrompt, then judges each against retrieved contexts via NLIStatementPrompt (metrics/_faithfulness.py:152-214).
Hybrid/embedding-based: FaithfulnesswithHHEM replaces the LLM NLI stage with a HuggingFace transformer (vectara/hallucination_evaluation_model) (metrics/_faithfulness.py:217-273). SemanticSimilarity uses embeddings. NonLLM variants of ContextPrecision/ContextRecall use embedding cosine similarity instead of LLM calls.
Custom metrics: Three decorator-based primitives: @discrete_metric (categorical output, e.g. 'pass'/'fail'), @numeric_metric (continuous float), @ranking_metric (ranked output). The SimpleLLMMetric class accepts a custom prompt string and Pydantic response model for ad-hoc LLM-as-a-judge metrics with save/load/reproducibility (metrics/base.py:846-935). AspectCritic and SimpleCriteriaScore take user-defined rubric strings for binary/discrete scoring with automatic strictness-based majority voting.
Per-trace vs per-task: All metrics implement SingleTurnMetric (single QA pair) or MultiTurnMetric (conversation). The Metric base class declares required_columns per type, e.g. Faithfulness needs {user_input, response, retrieved_contexts} for single-turn. Collection metrics (e.g. AnswerCorrectness composite) internally orchestrate multiple sub-metrics per trace.
How is LLM-as-a-judge implemented?
answeredLLM-as-a-judge is a core primitive in Ragas, built on PydanticPrompt with structured output.
Judge prompts and rubrics: The PydanticPrompt class (prompt/pydantic_prompt.py:82-350) is a generic template that takes an instruction string and input_model/output_model as Pydantic type parameters. The instruction describes the judgement criteria. At inference time, to_string() renders the instruction, output schema (auto-generated JSON Schema from output_model.model_json_schema()), few-shot examples, and the actual input into one prompt sent to the LLM. For instance, NLIStatementPrompt instructs "return verdict as 1 if the statement can be directly inferred..." with a StatementFaithfulnessAnswer output model (metrics/_faithfulness.py:73-131).
Structured output: PydanticPrompt.generate() calls the LLM with response_model=self.output_model, enforcing structured JSON via the Instructor library (llms/base.py:606-748). The llm_factory() wraps provider clients with instructor patching (OpenAI → instructor.from_openai, Anthropic → instructor.from_anthropic, etc.) and passes the Pydantic model as response_model. Returns are validated immediately into the output Pydantic model, with automatic retry (retries_left: int = 3) via RagasOutputParser.parse_output_string() which on parse failure calls FixOutputFormat to ask the LLM to fix malformed JSON (prompt/pydantic_prompt.py:525-558).
Multi-sample/consensus: The Ensember class (metrics/base.py:641-686) implements majority voting over multiple LLM outputs for the same input. AspectCritic and SimpleCriteriaScore have a strictness parameter — when >1, they call the LLM N times and aggregate via Counter(verdicts).most_common(1)[0][0] (metrics/_aspect_critic.py:155-165, metrics/_simple_criteria.py:152-162).
Judge model choice: The judge LLM is set per-metric via metric.llm, or inherited from the evaluate()/@experiment call-level llm parameter. llm_factory() supports OpenAI, Anthropic, Google/Gemini, Groq, Mistral, Azure, LiteLLM (100+ providers), and more. It auto-detects the correct instructor patching strategy. The default fallback when no LLM is provided is gpt-4o-mini via OpenAI() (evaluation.py:176-179).
Calibration/bias controls: No dedicated calibration or bias-mitigation module exists. The project provides an Ensember for consensus voting but no systematic bias measurement, calibration curves, or position-bias detection. The AspectCritic with predefined aspects (harmfulness, maliciousness, etc.) can be used as a safety check, but bias controls are absent.
strictness is inert at this commit. AspectCritic and SimpleCriteriaScore call prompt.generate once and vote over a one-element list (src/ragas/metrics/_aspect_critic.py L173-L212), so no multi-sample consensus happens.How are test datasets and cases defined, generated and versioned?
answeredTest datasets are structured around Pydantic sample models and a flexible backend storage layer.
Data model: SingleTurnSample holds user_input, retrieved_contexts, response, reference, rubrics, and metadata fields (persona_name, query_style, query_length). MultiTurnSample holds a list of HumanMessage|AIMessage|ToolMessage objects with conversation validation (ToolMessage must follow an AIMessage that called tools). Both extend BaseSample with to_dict()/get_features() (dataset_schema.py:29-180).
File formats/DSL: EvaluationDataset wraps lists of samples and supports conversion to/from Hugging Face datasets, pandas DataFrames, CSV, and JSONL files (.to_csv/.to_jsonl/.from_jsonl) (dataset_schema.py:186-406). The DataTable class (dataset.py) provides a list-like interface with pluggable BaseBackend persistence: LocalCSVBackend (per-row CSV), LocalJSONLBackend, GDriveBackend (Google Drive), and InMemoryBackend. Backends are resolved by name via a registry (e.g. "local/csv", "gdrive") (backends/base.py, backends/registry.py, backends/local_csv.py).
Synthetic data generation: Test data is generated via BaseSynthesizer subclasses that operate on a KnowledgeGraph of nodes (chunks of documents) and edges (relationships). SingleHopQuerySynthesizer generates simple lookup queries: it samples nodes, terms, personas, query styles, and lengths, then calls an LLM prompt (QueryAnswerGenerationPrompt) to produce a question+reference answer pair (testset/synthesizers/single_hop/base.py:30-47). MultiHopQuerySynthesizer generates questions requiring 2-3 steps of reasoning over related nodes. The Persona system enables persona-based query generation (e.g. "expert", "novice"). Query-level transforms (splitting, filtering, relationship building) prepare documents before synthesis.
Versioning: version_experiment() (experiment.py:21-100) snapshots the current codebase state to git: it creates a commit and a git branch named ragas/{experiment_name}, returning the commit hash. This links evaluation results to a specific code version.
Golden sets and registries: No benchmark task registry or golden test set management exists in the codebase. Golden/holdout data lives in user-managed files loaded via DataTable.load() from backends. The Dataset class provides train_test_split() for splitting data into training and testing sets for metric alignment.
How are evals executed and reported?
answeredEvaluation execution is centered on the Executor class and the newer @experiment decorator pattern.
Runners: The Executor (executor.py:18-216) manages async job queues. Jobs are submitted via .submit(callable, *args, name=...) and executed concurrently via as_completed() with a configurable max_workers limit from RunConfig. It wraps each callable with error handling (returns np.nan on failure unless raise_exceptions=True). Both synchronous .results() and async .aresults() paths are supported. In the evaluate()/aevaluate() flow (now deprecated in favor of @experiment), one Executor is created per evaluation run. Each metric's single_turn_ascore() or multi_turn_ascore() is submitted as a separate job per dataset row (evaluation.py:244-276).
Parallelism and caching: Concurrency is controlled by RunConfig.max_workers (default unlimited). Batching is optional via batch_size — batches process as nested progress bars. LLM responses are cacheable via CacheInterface: DiskCacheBackend (diskcache) stores results keyed by SHA-256 of the function + arguments. The cacher decorator wraps generate_text/agenerate_text (llms/base.py:60-66). This provides ~60x speedup for repeated evaluations with identical inputs.
CI integration: No native CI integration — tests use standard pytest via make test. The CLI (ragas command via typer) can run experiments from the command line, making it scriptable in CI pipelines.
Result storage: Results are returned as EvaluationResult (dataset_schema.py:411-552), which exposes per-metric scores as dict-of-lists, mean-aggregated scores via _repr_dict, cost tracking via CostCallbackHandler (total tokens, total cost), and full run traces via RagasTracer (callback tree of evaluation → rows → metrics → prompts). Conversion to pandas is available via .to_pandas(). The Experiment/@experiment decorator pattern (experiment.py:103-232) saves results to a backend automatically.
Comparison, regression, dashboards: The CLI (cli.py:54-99) implements baseline comparison: it renders Rich tables showing current, baseline, delta values and pass/fail gates with small-regression tolerance (±0.01). No web dashboard exists. The CLI supports ragas evaluate for running experiments and ragas compare for baseline comparisons with categorical frequency tables.
RunConfig.max_workers defaults to 16, not unlimited. The CLI has no evaluate or compare commands, only evals (baseline gate tables via --baseline), quickstart and hello_world, and evals depends on a Project class the code marks as not implemented.How are traces or production data captured and linked to evaluations?
answeredRagas provides built-in tracing via a callback-based architecture and first-party integrations with two tracing platforms.
Built-in tracing: The RagasTracer (callbacks.py:80-121) is a BaseCallbackHandler that records every evaluation run as a tree of ChainRun nodes. Each evaluation creates a tree: evaluation → row → metric → prompt, with parent-child relationships tracked via run_id/parent_run_id. This captures inputs, outputs, and metadata at every level. Traces are stored in memory during execution and attached to EvaluationResult.ragas_traces for post-hoc analysis. parse_run_traces() traverses the tree to produce flat per-row lists of metric scores and prompt I/O (callbacks.py:134-173).
Tracing integrations (integrations/tracing/): LangfuseTrace and MLflowTrace provide explicit platform-level observability. The @observe() decorator wraps evaluation functions so every metric computation and LLM call is recorded in the tracing backend. sync_trace() returns a trace object with .get_url() for deep linking into the observability platform (integrations/tracing/init.py:1-78). These are optional dependencies — imported lazily.
LLM observability integrations: HeliconeConfig proxies LLM calls through Helicone for request logging. Langsmith and Opik integrations are available (integrations/langsmith.py, integrations/opik.py) for tracking LLM calls and evaluation results. Framework integrations (LangChain, LlamaIndex, Griptape, LangGraph) bridge those ecosystems' own observability.
Online vs offline evals: No distinction exists — all evaluations are offline/batch by design. There is no production traffic capturing, SDK instrumentation for live applications, or OpenTelemetry integration. The project integrates with no APM/vendor for online monitoring of deployed LLM applications.
Feedback and annotation: The PromptAnnotation/SampleAnnotation/MetricAnnotation classes (dataset_schema.py:555-837) support human-in-the-loop annotation for metric training: users can edit prompt outputs, mark samples as accepted/rejected, and use this data to optimize instruction prompts or few-shot demonstrations via MetricWithLLM.train(). This is a training feedback loop, not production feedback collection.
Does it support red-teaming or safety testing, and how?
answeredRagas does not have a dedicated red-teaming or safety testing framework. There are no adversarial probes, attack plugins, jailbreak tests, prompt injection test suites, or vulnerability reporting capabilities.
What it does instead: The AspectCritic metric (metrics/_aspect_critic.py:75-242) evaluates outputs against custom criteria defined as natural-language strings. Ragas ships several predefined aspects including harmfulness ("Does the submission cause or have the potential to cause harm...") and maliciousness ("Is the submission intended to harm, deceive, or exploit users?") (metrics/_aspect_critic.py:215-221). These are LLM-as-a-judge binary verdicts, not adversarial probing tools — they check whether a given output is harmful, rather than attempting to provoke harmful outputs.
What is absent: No adversarial input generation (no probe synthesis, no attack chaining, no red-team scenario specification). No jailbreak/prompt injection detection metrics — while an AspectCritic with a custom definition could be written to judge a response for injection success, it provides no tooling to generate injection attempts. There is no vulnerability reporting mechanism.
Tangential reference: A comment in llms/adapters/__init__.py:80 references "HARM_CATEGORY_JAILBREAK" in the context of Google Gemini safety settings causing issues with the Instructor library — this is an upstream workaround note, not a ragas feature.