LLMs Technical Reviews

vibrantlabsai/ragas

Python library of RAG and agent metrics (faithfulness, context precision/recall, ...) with knowledge-graph testset generation.

GitHub ↗★ 16kPythonApache-2.0commit 298b682 · 2026-02-24homepage ↗

Overview

Ragas is a Python library for scoring LLM applications, best known for its retrieval-augmented generation (RAG) metrics: faithfulness, answer relevancy, context precision and context recall. Most metrics are short LLM-as-a-judge pipelines. A metric breaks a response into statements, asks a judge model for structured verdicts, and returns a ratio. A smaller set is purely lexical or embedding-based (exact match, BLEU, ROUGE, chrF, string similarity, semantic similarity), and one faithfulness variant swaps the judge for Vectara’s HHEM classifier.

At the pinned commit the codebase is midway through a migration, and that shapes how you use it. The legacy API is evaluate() over an EvaluationDataset, with metric objects such as faithfulness that implement single_turn_ascore. It still works but warns that it is deprecated in favour of @experiment. The new API is ragas.metrics.collections: standalone metrics called as await metric.ascore(user_input=..., response=..., ...) that return a MetricResult. You orchestrate them yourself inside an @experiment function, and results go to a pluggable storage backend. The two metric families have different base classes and are not interchangeable.

The second major feature is test-set generation. TestsetGenerator builds a knowledge graph from your documents through extractors, splitters and relationship builders, invents personas, and synthesises single-hop and multi-hop questions with reference answers. Ragas also includes metric alignment and prompt optimisation (a genetic optimiser and a DSPy adapter) trained on human annotations.

Architecture

flowchart LR
  DS["EvaluationDataset / Dataset"] --> EV["evaluate() (legacy)"]
  DS --> EXP["@experiment arun"]
  EV --> EXE["Executor (async jobs)"]
  EXE --> MET["Metric.single_turn_ascore"]
  MET --> PP["PydanticPrompt"]
  PP --> LLM["llm_factory: Instructor LLM"]
  EXP --> UF["your function + collections metrics"]
  UF --> LLM
  EXP --> BE["Backends: CSV / JSONL / GDrive"]
  EV --> RES["EvaluationResult + RagasTracer"]
  KG["KnowledgeGraph + transforms"] --> TG["TestsetGenerator"]
  TG --> DS
Component Path Role
Legacy runner src/ragas/evaluation.py evaluate/aevaluate: defaults, LLM injection, job submission, result assembly
Executor src/ragas/executor.py Async job queue with max_workers, batching, cancel, NaN on error
Experiments src/ragas/experiment.py @experiment decorator, arun, git-based version_experiment
Legacy metrics src/ragas/metrics/_*.py, metrics/base.py Metric, MetricWithLLM, SingleTurnMetric, MultiTurnMetric
New metrics src/ragas/metrics/collections/ BaseMetric.ascore(**kwargs) -> MetricResult, one folder per metric
Custom metrics metrics/discrete.py, numeric.py, ranking.py, base.py @discrete_metric / @numeric_metric / @ranking_metric, SimpleLLMMetric
Prompts src/ragas/prompt/pydantic_prompt.py Typed instruction + few-shot prompt, output parsing, fix-format retry
LLMs src/ragas/llms/ llm_factory (Instructor/LiteLLM), LangChain and LlamaIndex wrappers
Data dataset_schema.py, dataset.py, backends/ Samples, EvaluationDataset, DataTable with local CSV/JSONL, in-memory, Google Drive
Testset src/ragas/testset/ Knowledge graph, transforms, personas, query synthesizers
Integrations src/ragas/integrations/ LangChain, LlamaIndex, LangGraph, Langfuse/MLflow tracing, Opik, Helicone, Bedrock, …
CLI src/ragas/cli.py quickstart templates, evals with baseline gate tables

How a request flows

Legacy path, evaluate(dataset, metrics=[faithfulness]):

  1. Validate and default. aevaluate emits a DeprecationWarning, converts a Hugging Face Dataset into an EvaluationDataset, checks required columns, and falls back to four default metrics when none are given (evaluation.py).
  2. Inject models. Any MetricWithLLM without an LLM gets llm_factory("gpt-4o-mini", client=OpenAI()), embeddings are inferred from that LLM’s provider, and every metric’s init(run_config) runs (evaluation.py).
  3. Submit jobs. An Executor and a RagasTracer callback are created. For every row and metric, metric.single_turn_ascore(sample, row_callbacks) is submitted as one job with the run timeout (evaluation.py).
  4. Run concurrently. Jobs run under RunConfig.max_workers (16 by default, 180 s timeout, 10 retries) (run_config.py). A failed job logs the exception and returns NaN unless raise_exceptions=True (executor.py).
  5. Score one row. single_turn_ascore keeps only the metric’s required columns, opens a callback group, awaits _single_turn_ascore under asyncio.wait_for, and sends an analytics event (base.py). Faithfulness generates statements, then NLI verdicts, through PydanticPrompt.generate.
  6. Call the judge. PydanticPrompt.generate_multiple renders the instruction, JSON schema, examples and input into one string. It dispatches to a LangChain LLM, an Instructor LLM (structured output via response_model) or a BaseRagasLLM, then validates the output into the Pydantic model (pydantic_prompt.py).
  7. Assemble. Results are mapped back to {metric_name: score} per row and wrapped in an EvaluationResult that carries the tracer’s run tree and optional cost data (evaluation.py).

New path: decorate an async function with @experiment(), call collections metrics inside it, and run await fn.arun(dataset, name=...). arun schedules one task per dataset row, appends each returned row to an Experiment table, prints a warning for any row that raised, and saves through the backend (experiment.py).

Key components

Collections metrics

A collections metric declares its components in __init__ (for example llm: InstructorBaseRagasLLM), and the base class rejects legacy LangChain-style wrappers (base.py). Faithfulness is a clear example. It makes one call that turns the answer into atomic statements and one NLI call over the joined contexts, and it returns faithful/total, or NaN when no statements come back (metric.py). These metrics are SimpleBaseMetric subclasses with score/ascore/batch_score (base.py), not Metric, so evaluate() rejects them.

LLM layer

llm_factory(model, provider, client, adapter, cache, mode) wraps a pre-built client with Instructor (from_openai, from_anthropic, from_genai, from_litellm, …) or LiteLLM and returns an object with generate/agenerate(prompt, response_model) (base.py, L1117-L1163). Passing cache=DiskCacheBackend() memoises generations on disk. Legacy BaseRagasLLM subclasses get the same wrapper on generate_text (base.py).

Rubric judges

AspectCritic turns a natural-language definition into a Yes/No instruction, and predefined critics cover harmfulness, maliciousness, coherence and similar aspects. It has a strictness parameter meant for majority voting, but at this commit both the single-turn and multi-turn paths call generate once and vote over a one-element list (_aspect_critic.py). Treat it as a single-sample judge. SimpleCriteriaScore follows the same pattern.

Tracing and results

RagasTracer is a LangChain-style callback handler that records an evaluation → row → metric → prompt tree of ChainRuns with inputs and outputs (callbacks.py). This is what lets you inspect why a row scored low. Langfuse and MLflow helpers under integrations/tracing/ and LangSmith/Opik/Helicone integrations push to external tools. There is no OpenTelemetry or production-traffic capture.

Testset generation

TestsetGenerator.generate(testset_size, query_distribution, num_personas) uses the default single-hop and multi-hop synthesizers when no distribution is given, generates personas from the knowledge graph, and runs scenario and sample generation through the same Executor (generate.py). Entry points exist for LangChain documents, LlamaIndex documents and raw chunks.

Extending it

  • Quick custom metric. Decorate a function with @discrete_metric(name=..., allowed_values=[...]), @numeric_metric or @ranking_metric, or use SimpleLLMMetric with your own prompt and response model. It can be saved, loaded and aligned against human labels.
  • Full metric. In the collections style, subclass BaseMetric and implement ascore. In the legacy style, subclass MetricWithLLM + SingleTurnMetric and implement _single_turn_ascore.
  • Prompts. Subclass PydanticPrompt[Input, Output] with an instruction and examples. Prompts can be translated with adapt() and saved to JSON.
  • Storage. Implement BaseBackend and register it, or use local/csv, local/jsonl, inmemory or gdrive.
  • Versioning. version_experiment(name) stages tracked changes, commits them, and creates a ragas/<name> branch, so a result can be tied to code. It writes to your git repository, so call it deliberately (experiment.py).

Running it

  • pip install ragas (Python 3.9+). Extras add GitPython (git), Langfuse and MLflow (tracing), Google Drive, DSPy and others.
  • Create an LLM explicitly: llm = llm_factory("gpt-4o-mini", client=AsyncOpenAI()). Use an async client with collections metrics, because agenerate raises on a sync client. The client argument is mandatory: “text-only mode” was removed, and llm_factory(model) without a client raises a ValueError with migration instructions (base.py).
  • ragas quickstart scaffolds example projects. ragas evals prints current-versus-baseline tables with a pass/fail gate that tolerates regressions under 0.01 (cli.py).
  • Usage analytics are sent by default, and setting RAGAS_DO_NOT_TRACK=true turns them off.

Strengths and caveats

  • Strength: the reference RAG metrics. Faithfulness, context precision/recall, noise sensitivity and context entity recall are well-defined, small pipelines whose prompts are readable Pydantic classes.
  • Strength: structured judge I/O. Every judge call has typed input and output models, enforced by Instructor or validated with a fix-format retry, so parsing failures are rare and explicit.
  • Strength: testset generation from your own corpus with personas and multi-hop questions is more complete than in most eval libraries.
  • Caveat: two APIs at once. Legacy metrics with evaluate() and collections metrics with @experiment coexist, with different base classes, call signatures and LLM wrappers. Docs and examples mix them.
  • Caveat: errors become NaN. Both runners swallow per-row failures by default (NaN in evaluate, a printed warning and a dropped row in arun), so check counts, not just means. arun also starts every row at once with no concurrency limit.
  • Caveat: strictness voting is inert in AspectCritic and SimpleCriteriaScore at this commit.
  • Caveat: CLI evals is half-wired. It looks for a project object with get_dataset/get_experiment and an experiment with run_async, and the code notes that the Project class is not implemented yet (cli.py). Prefer the Python API for CI gates.
  • Caveat: offline only. No OpenTelemetry, no online scoring of production traffic and no red-teaming tooling.

Sources: code at 298b682, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit 298b682. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Ragas provides ~30+ built-in metrics across three implementation categories.

Heuristic/statistical metrics (no LLM call): ExactMatch (string equality), StringPresence (substring match), NonLLMStringSimilarity (Levenshtein/Hamming/Jaro/Jaro-Winkler via rapidfuzz), BleuScore, RougeScore, ChrfScore, and DataCompyScore. These are pure Python computations implementing SingleTurnMetric. Example: ExactMatch returns float(sample.reference == sample.response) (metrics/_string.py:29-35).

LLM-based metrics: Faithfulness (statement decomposition → NLI verdicts), AnswerRelevancy, AnswerCorrectness, FactualCorrectness, ContextPrecision (multiple variants: LLM, NonLLM, ID-based), ContextRecall, ContextEntityRecall, NoiseSensitivity, SummarizationScore, TopicAdherenceScore, ToolCallAccuracy, ToolCallF1, MultiModalFaithfulness, MultiModalRelevance, SQLSemanticEquivalence, GoalAccuracy (agent). These extend MetricWithLLM and implement _single_turn_ascore() which calls a PydanticPrompt.generate() against an LLM. Example: Faithfulness breaks the response into atomic statements using StatementGeneratorPrompt, then judges each against retrieved contexts via NLIStatementPrompt (metrics/_faithfulness.py:152-214).

Hybrid/embedding-based: FaithfulnesswithHHEM replaces the LLM NLI stage with a HuggingFace transformer (vectara/hallucination_evaluation_model) (metrics/_faithfulness.py:217-273). SemanticSimilarity uses embeddings. NonLLM variants of ContextPrecision/ContextRecall use embedding cosine similarity instead of LLM calls.

Custom metrics: Three decorator-based primitives: @discrete_metric (categorical output, e.g. 'pass'/'fail'), @numeric_metric (continuous float), @ranking_metric (ranked output). The SimpleLLMMetric class accepts a custom prompt string and Pydantic response model for ad-hoc LLM-as-a-judge metrics with save/load/reproducibility (metrics/base.py:846-935). AspectCritic and SimpleCriteriaScore take user-defined rubric strings for binary/discrete scoring with automatic strictness-based majority voting.

Per-trace vs per-task: All metrics implement SingleTurnMetric (single QA pair) or MultiTurnMetric (conversation). The Metric base class declares required_columns per type, e.g. Faithfulness needs {user_input, response, retrieved_contexts} for single-turn. Collection metrics (e.g. AnswerCorrectness composite) internally orchestrate multiple sub-metrics per trace.

How is LLM-as-a-judge implemented?

answered

LLM-as-a-judge is a core primitive in Ragas, built on PydanticPrompt with structured output.

Judge prompts and rubrics: The PydanticPrompt class (prompt/pydantic_prompt.py:82-350) is a generic template that takes an instruction string and input_model/output_model as Pydantic type parameters. The instruction describes the judgement criteria. At inference time, to_string() renders the instruction, output schema (auto-generated JSON Schema from output_model.model_json_schema()), few-shot examples, and the actual input into one prompt sent to the LLM. For instance, NLIStatementPrompt instructs "return verdict as 1 if the statement can be directly inferred..." with a StatementFaithfulnessAnswer output model (metrics/_faithfulness.py:73-131).

Structured output: PydanticPrompt.generate() calls the LLM with response_model=self.output_model, enforcing structured JSON via the Instructor library (llms/base.py:606-748). The llm_factory() wraps provider clients with instructor patching (OpenAI → instructor.from_openai, Anthropic → instructor.from_anthropic, etc.) and passes the Pydantic model as response_model. Returns are validated immediately into the output Pydantic model, with automatic retry (retries_left: int = 3) via RagasOutputParser.parse_output_string() which on parse failure calls FixOutputFormat to ask the LLM to fix malformed JSON (prompt/pydantic_prompt.py:525-558).

Multi-sample/consensus: The Ensember class (metrics/base.py:641-686) implements majority voting over multiple LLM outputs for the same input. AspectCritic and SimpleCriteriaScore have a strictness parameter — when >1, they call the LLM N times and aggregate via Counter(verdicts).most_common(1)[0][0] (metrics/_aspect_critic.py:155-165, metrics/_simple_criteria.py:152-162).

Judge model choice: The judge LLM is set per-metric via metric.llm, or inherited from the evaluate()/@experiment call-level llm parameter. llm_factory() supports OpenAI, Anthropic, Google/Gemini, Groq, Mistral, Azure, LiteLLM (100+ providers), and more. It auto-detects the correct instructor patching strategy. The default fallback when no LLM is provided is gpt-4o-mini via OpenAI() (evaluation.py:176-179).

Calibration/bias controls: No dedicated calibration or bias-mitigation module exists. The project provides an Ensember for consensus voting but no systematic bias measurement, calibration curves, or position-bias detection. The AspectCritic with predefined aspects (harmfulness, maliciousness, etc.) can be used as a safety check, but bias controls are absent.

Editor's note. Correction: strictness is inert at this commit. AspectCritic and SimpleCriteriaScore call prompt.generate once and vote over a one-element list (src/ragas/metrics/_aspect_critic.py L173-L212), so no multi-sample consensus happens.

How are test datasets and cases defined, generated and versioned?

answered

Test datasets are structured around Pydantic sample models and a flexible backend storage layer.

Data model: SingleTurnSample holds user_input, retrieved_contexts, response, reference, rubrics, and metadata fields (persona_name, query_style, query_length). MultiTurnSample holds a list of HumanMessage|AIMessage|ToolMessage objects with conversation validation (ToolMessage must follow an AIMessage that called tools). Both extend BaseSample with to_dict()/get_features() (dataset_schema.py:29-180).

File formats/DSL: EvaluationDataset wraps lists of samples and supports conversion to/from Hugging Face datasets, pandas DataFrames, CSV, and JSONL files (.to_csv/.to_jsonl/.from_jsonl) (dataset_schema.py:186-406). The DataTable class (dataset.py) provides a list-like interface with pluggable BaseBackend persistence: LocalCSVBackend (per-row CSV), LocalJSONLBackend, GDriveBackend (Google Drive), and InMemoryBackend. Backends are resolved by name via a registry (e.g. "local/csv", "gdrive") (backends/base.py, backends/registry.py, backends/local_csv.py).

Synthetic data generation: Test data is generated via BaseSynthesizer subclasses that operate on a KnowledgeGraph of nodes (chunks of documents) and edges (relationships). SingleHopQuerySynthesizer generates simple lookup queries: it samples nodes, terms, personas, query styles, and lengths, then calls an LLM prompt (QueryAnswerGenerationPrompt) to produce a question+reference answer pair (testset/synthesizers/single_hop/base.py:30-47). MultiHopQuerySynthesizer generates questions requiring 2-3 steps of reasoning over related nodes. The Persona system enables persona-based query generation (e.g. "expert", "novice"). Query-level transforms (splitting, filtering, relationship building) prepare documents before synthesis.

Versioning: version_experiment() (experiment.py:21-100) snapshots the current codebase state to git: it creates a commit and a git branch named ragas/{experiment_name}, returning the commit hash. This links evaluation results to a specific code version.

Golden sets and registries: No benchmark task registry or golden test set management exists in the codebase. Golden/holdout data lives in user-managed files loaded via DataTable.load() from backends. The Dataset class provides train_test_split() for splitting data into training and testing sets for metric alignment.

How are evals executed and reported?

answered

Evaluation execution is centered on the Executor class and the newer @experiment decorator pattern.

Runners: The Executor (executor.py:18-216) manages async job queues. Jobs are submitted via .submit(callable, *args, name=...) and executed concurrently via as_completed() with a configurable max_workers limit from RunConfig. It wraps each callable with error handling (returns np.nan on failure unless raise_exceptions=True). Both synchronous .results() and async .aresults() paths are supported. In the evaluate()/aevaluate() flow (now deprecated in favor of @experiment), one Executor is created per evaluation run. Each metric's single_turn_ascore() or multi_turn_ascore() is submitted as a separate job per dataset row (evaluation.py:244-276).

Parallelism and caching: Concurrency is controlled by RunConfig.max_workers (default unlimited). Batching is optional via batch_size — batches process as nested progress bars. LLM responses are cacheable via CacheInterface: DiskCacheBackend (diskcache) stores results keyed by SHA-256 of the function + arguments. The cacher decorator wraps generate_text/agenerate_text (llms/base.py:60-66). This provides ~60x speedup for repeated evaluations with identical inputs.

CI integration: No native CI integration — tests use standard pytest via make test. The CLI (ragas command via typer) can run experiments from the command line, making it scriptable in CI pipelines.

Result storage: Results are returned as EvaluationResult (dataset_schema.py:411-552), which exposes per-metric scores as dict-of-lists, mean-aggregated scores via _repr_dict, cost tracking via CostCallbackHandler (total tokens, total cost), and full run traces via RagasTracer (callback tree of evaluation → rows → metrics → prompts). Conversion to pandas is available via .to_pandas(). The Experiment/@experiment decorator pattern (experiment.py:103-232) saves results to a backend automatically.

Comparison, regression, dashboards: The CLI (cli.py:54-99) implements baseline comparison: it renders Rich tables showing current, baseline, delta values and pass/fail gates with small-regression tolerance (±0.01). No web dashboard exists. The CLI supports ragas evaluate for running experiments and ragas compare for baseline comparisons with categorical frequency tables.

Editor's note. Correction: RunConfig.max_workers defaults to 16, not unlimited. The CLI has no evaluate or compare commands, only evals (baseline gate tables via --baseline), quickstart and hello_world, and evals depends on a Project class the code marks as not implemented.

How are traces or production data captured and linked to evaluations?

answered

Ragas provides built-in tracing via a callback-based architecture and first-party integrations with two tracing platforms.

Built-in tracing: The RagasTracer (callbacks.py:80-121) is a BaseCallbackHandler that records every evaluation run as a tree of ChainRun nodes. Each evaluation creates a tree: evaluation → row → metric → prompt, with parent-child relationships tracked via run_id/parent_run_id. This captures inputs, outputs, and metadata at every level. Traces are stored in memory during execution and attached to EvaluationResult.ragas_traces for post-hoc analysis. parse_run_traces() traverses the tree to produce flat per-row lists of metric scores and prompt I/O (callbacks.py:134-173).

Tracing integrations (integrations/tracing/): LangfuseTrace and MLflowTrace provide explicit platform-level observability. The @observe() decorator wraps evaluation functions so every metric computation and LLM call is recorded in the tracing backend. sync_trace() returns a trace object with .get_url() for deep linking into the observability platform (integrations/tracing/init.py:1-78). These are optional dependencies — imported lazily.

LLM observability integrations: HeliconeConfig proxies LLM calls through Helicone for request logging. Langsmith and Opik integrations are available (integrations/langsmith.py, integrations/opik.py) for tracking LLM calls and evaluation results. Framework integrations (LangChain, LlamaIndex, Griptape, LangGraph) bridge those ecosystems' own observability.

Online vs offline evals: No distinction exists — all evaluations are offline/batch by design. There is no production traffic capturing, SDK instrumentation for live applications, or OpenTelemetry integration. The project integrates with no APM/vendor for online monitoring of deployed LLM applications.

Feedback and annotation: The PromptAnnotation/SampleAnnotation/MetricAnnotation classes (dataset_schema.py:555-837) support human-in-the-loop annotation for metric training: users can edit prompt outputs, mark samples as accepted/rejected, and use this data to optimize instruction prompts or few-shot demonstrations via MetricWithLLM.train(). This is a training feedback loop, not production feedback collection.

Does it support red-teaming or safety testing, and how?

answered

Ragas does not have a dedicated red-teaming or safety testing framework. There are no adversarial probes, attack plugins, jailbreak tests, prompt injection test suites, or vulnerability reporting capabilities.

What it does instead: The AspectCritic metric (metrics/_aspect_critic.py:75-242) evaluates outputs against custom criteria defined as natural-language strings. Ragas ships several predefined aspects including harmfulness ("Does the submission cause or have the potential to cause harm...") and maliciousness ("Is the submission intended to harm, deceive, or exploit users?") (metrics/_aspect_critic.py:215-221). These are LLM-as-a-judge binary verdicts, not adversarial probing tools — they check whether a given output is harmful, rather than attempting to provoke harmful outputs.

What is absent: No adversarial input generation (no probe synthesis, no attack chaining, no red-team scenario specification). No jailbreak/prompt injection detection metrics — while an AspectCritic with a custom definition could be written to judge a response for injection success, it provides no tooling to generate injection attempts. There is no vulnerability reporting mechanism.

Tangential reference: A comment in llms/adapters/__init__.py:80 references "HARM_CATEGORY_JAILBREAK" in the context of Google Gemini safety settings causing issues with the Instructor library — this is an upstream workaround note, not a ragas feature.