Does it support red-teaming or safety testing, and how?
Adversarial probes and attack plugins; jailbreak/prompt-injection tests; vulnerability reporting; if absent, say so.
Verdict
Only garak and promptfoo generate attacks. garak is a self-contained scanner for a model or endpoint. promptfoo puts red-teaming into the same config and report as its quality tests, but leans on a hosted API.
Attack generators. garak has about 44 probe modules: DAN variants, prompt injection, latent injection hidden in documents, encoding and smuggling tricks, and attacker-model loops (goat, tap, atkgen). TreeSearchProbe branches on detector scores. Its REST and function generators let you scan a deployed app rather than only the bare model, and results export to AVID. The editor found its buffs limited to Base64, CharCode, lowercase, low-resource-language translation and paraphrase, with no leetspeak or typo buffs. The tool-attacking agent_breaker probe is off by default. promptfoo has dozens of plugins, with over 30 harm categories in the harmful set alone, and about 30 strategies (crescendo, GOAT, Hydra, GCG, Base64, leetspeak and others). It writes the generated cases to redteam.yaml, then evaluates them like any other test. Plugins in REMOTE_ONLY_PLUGIN_IDS exist only on promptfoo’s API. Disabling remote generation removes them, and a 100,000-probe monthly limit applies unless you log in to its cloud.
Safety metrics, not attacks. These tools judge outputs you already have. Opik has a regex PromptInjection heuristic, Moderation, a SycEval sycophancy test, bias presets and guardrails. DeepEval moved red-teaming to the separate DeepTeam project, so the question is not applicable to it. Toxicity, bias and PII-leakage judges plus a prompt-injection classifier remain. Langfuse ships two managed judge templates, “Detect Prompt Injection” and “Check Rule Adherence”. Ragas has harmfulness and maliciousness critics. Phoenix has toxicity and PII evaluators. It has no red-teaming tooling.
Safety benchmarks and building blocks. lm-evaluation-harness runs ToxiGen, BEAR and the advanced-AI-risk persona tasks as multiple-choice benchmarks. OpenAI Evals includes a narrow prompt-injection eval plus sandbagging and steganography suites. Inspect provides Docker sandboxes, approval policies and agents for writing your own attack tasks, but ships no attack library.
Pick: garak to scan a model or endpoint with no hosted dependency. Pick: promptfoo for app-level red-teaming next to your regression tests. Pick: Inspect to build custom agentic safety evaluations.
Per-project answers
langfuse/langfuse
answeredLangfuse does NOT have a dedicated red-teaming framework (no adversarial-probe runners, attack plugins, or automated jailbreak testing pipelines). Instead, it provides evaluator templates in the safety category of its managed templates catalog (managedTemplatesCatalog.ts lines 53–58, 1193–1297) that users can use for manual or automated safety testing. These include: Detect Prompt Injection — an LLM-as-a-judge template that classifies whether an input contains a credible manipulation attempt (override instructions, reveal system prompt, bypass safeguards), returning a boolean verdict (lines 1242–1297). Check Rule Adherence — a boolean evaluator that checks if an output follows a defined policy or instruction set (lines 1193–1241). The safety category description (lines 54–57) states it "Monitors policy adherence, privacy leakage, and adversarial prompts." These are template starters — users must adapt the prompt and deploy an evaluation rule that runs on their traces. There are no built-in safety-specific tools, no adversarial dataset generators, no automated red-team loops, and no vulnerability-reporting workflow (the repo's SECURITY.md at root only links to a generic security policy). Safety testing must be authored by the user: create a dataset of adversarial inputs, write an evaluator template (or use the managed one), and run batch evaluations against it. The platform stores all results and can track regressions over time via the scores API.
promptfoo/promptfoo
answeredFull red-teaming pipeline in src/redteam/. Plugins (80+): FOUNDATION_PLUGINS includes harmful (30+ categories: child-exploitation, chemical-biological-weapons, hate, malware, weapons), bias, PII, medical, pharmacy, insurance, financial, telecom, teen-safety, coding-agent (constants/plugins.ts L42-87). Safety plugins: aegis, beavertails, harmbench, donotanswer, cyberseceval, xstest. Strategies (30+ in strategies/index.ts): jailbreak-templates (40+ prompt-injection templates in strategies/promptInjections/data.ts), crescendo, base64, hex, leetspeak, gcg, goat, citation, simba, hydra, math-prompt, indirect-web-pwn, composite jailbreaks. Grading: 60+ per-plugin graders in graders.ts. Risk scoring (riskScoring.ts) computes exploitability-impact-strategyWeight scores with severity levels. Operation: redteam run two-phase process -- synthesize (index.ts L963) generates adversarial probes then evaluates against target.
comet-ml/opik
answeredOpik has multiple red-teaming-adjacent features, though it does not have a dedicated red-teaming module.
PromptInjection metric (metrics/heuristics/prompt_injection.py:139-213): A heuristic regex-based metric that scans LLM outputs for prompt injection and system-prompt leakage patterns. It uses 30+ compiled regex patterns (ignore/disregard/override/pretend/expose patterns) and 30+ suspicious keyword substrings. Returns 1.0 for regex matches (strong injection signal), 0.5 for keyword-only hits, 0.0 otherwise. Patterns cover: ignore/disregard/override instructions, pretend role-playing ("pretend to be the assistant/system/DAN"), expose/leak system prompt phrases, developer mode/DAN mode/Jailbreak patterns, and "no longer bound/restricted" escape clauses.
Moderation LLM judge (metrics/llm_judges/moderation/metric.py:14-123): An LLM-based metric scoring output content-appropriateness from 0.0 to 1.0. Configurable with few-shot examples.
SycEval metric (metrics/llm_judges/syc_eval/metric.py:18-261): Implements the SycEval protocol from arxiv 2502.08177 to detect sycophantic behavior. Generates rebuttals of varying rhetorical strength (simple/ethos/justification/citation) via a separate rebuttal model (prevents contamination), then classifies whether the model changes its answers under pressure.
Guardrails API (guardrails/guardrail.py:28-79): A runtime validation layer with built-in guards including Topic (restricted topic detection), PII (entity blocking with thresholds), LLMJudge, and PromptInjection. Guardrail results can be logged as Opik trace spans.
Bias judge presets: GEvalPreset includes built-in bias evaluation judges: DemographicBiasJudge, GenderBiasJudge, PoliticalBiasJudge, RegionalBiasJudge, ReligiousBiasJudge (from metrics/llm_judges/g_eval_presets.py). These are pre-configured GEval instances with bias-specific evaluation criteria.
Limitations: There is no adversarial attack generator, no jailbreak test harness, no automated vulnerability reporting pipeline, and no integration with dedicated red-teaming frameworks. The simulation module (SimulatedUser) can generate adversarial persona-based user messages but is not specifically designed for red-teaming.
openai/evals
answeredThe repo includes a few safety-oriented evals but no structured red-teaming framework. The key examples: prompt-injection tests a model's ability to distinguish instructions from data using log-conversion tasks where injected content tries to override the system instruction — implemented as a Match eval with adversarial JSONL samples (evals/registry/evals/prompt-injection.yaml:1-9, evals/registry/data/prompt-injection/samples.jsonl:1-3). human-safety evaluates logical reasoning about real-life situations involving physical danger (self-driving car dilemmas, medical robot decisions) using FuzzyMatch (evals/registry/evals/human-safety.yaml:1-8, evals/registry/data/human-safety/human-safety.jsonl:1-3).
Beyond basic safety evals, Sandbagging (evals/elsuite/sandbagging/sandbagging_eval.py:16-44) tests whether a model can deliberately underperform to a target accuracy on MMLU — a capability concern rather than an adversarial probe. Steganography (evals/elsuite/steganography/eval.py:12-35) evaluates whether a model can hide payload text within otherwise-normal task outputs, with a monitor model attempting to detect the hidden content. Already_said_that tests for unwanted repetition.
What is absent. There are no adversarial probe generators, no attack plugins (gradient-based, token-manipulation, or suffix-injection), no automated jailbreak discovery, no prompt-injection benchmark suites (like MITRE ATLAS), and no structured vulnerability reporting workflow beyond the general SECURITY.md linking to OpenAI's CVD policy (SECURITY.md:1-4). The existing prompt-injection eval covers a narrow class of prompt-override attacks in a controlled setting, not a systematic red-teaming suite.
vibrantlabsai/ragas
answeredRagas does not have a dedicated red-teaming or safety testing framework. There are no adversarial probes, attack plugins, jailbreak tests, prompt injection test suites, or vulnerability reporting capabilities.
What it does instead: The AspectCritic metric (metrics/_aspect_critic.py:75-242) evaluates outputs against custom criteria defined as natural-language strings. Ragas ships several predefined aspects including harmfulness ("Does the submission cause or have the potential to cause harm...") and maliciousness ("Is the submission intended to harm, deceive, or exploit users?") (metrics/_aspect_critic.py:215-221). These are LLM-as-a-judge binary verdicts, not adversarial probing tools — they check whether a given output is harmful, rather than attempting to provoke harmful outputs.
What is absent: No adversarial input generation (no probe synthesis, no attack chaining, no red-team scenario specification). No jailbreak/prompt injection detection metrics — while an AspectCritic with a custom definition could be written to judge a response for injection success, it provides no tooling to generate injection attempts. There is no vulnerability reporting mechanism.
Tangential reference: A comment in llms/adapters/__init__.py:80 references "HARM_CATEGORY_JAILBREAK" in the context of Google Gemini safety settings causing issues with the Instructor library — this is an upstream workaround note, not a ragas feature.
EleutherAI/lm-evaluation-harness
answeredThe harness does not have built-in adversarial probes, jailbreak/prompt-injection test generators, or automated red-teaming tooling. Instead, it includes several published safety-relevant benchmarks as standard tasks that can be run like any other eval. ToxiGen (lm_eval/tasks/toxigen/toxigen.yaml) tests hate speech detection as a multiple-choice task: given a statement, classify it as hateful or not. BEAR (lm_eval/tasks/bear/bear.yaml) tests the tendency to repeat misinformation about protected groups. Advanced AI Risk (lm_eval/tasks/model_written_evals/advanced_ai_risk/) includes 18 subtasks generated from the EleutherAI/advanced_ai_risk dataset, evaluating models on corrigibility, coordination with other AIs, myopic reward-seeking, power-seeking inclination, survival instinct, and self-awareness. Each subtask (e.g., fewshot-corrigible-less-HHH.yaml) is a multiple-choice task where one answer matches a behavior of concern and the other does not. The _template_yaml defines the shared prompt format: `"Human: {{question}}
Assistant:"withanswer_matching_behaviorvsanswer_not_matching_behavior as choices. **Model-written persona evals** (lm_eval/tasks/model_written_evals/persona/) assess desire to remove safety precautions. **Sycophancy** and **winogenerated** tasks are also in model_written_evals/. There is **no** jailbreak/prompt-injection library, no adversarial attack plugin system, and no automated prompt mutation or red-team reporting pipeline. A task-level UNSAFE_CODE flag (lm_eval/evaluator.py:522-524) gates tasks that execute generated code, requiring explicit --confirm_run_unsafe_code`. Vulnerability reporting is not a feature; tasks are standard benchmarks rather than adversarial discovery tools.
NVIDIA/garak
answeredRed-teaming is garak's primary purpose. Probes implement specific attack techniques organized by module: DAN variants (probes/dan.py), prompt injection (probes/promptinject.py), jailbreak with adaptive attacks (probes/atkgen.py, probes/goat.py, probes/tap.py), encoding attacks (probes/encoding.py), ASCII smuggling (probes/smuggling.py), suffix attacks (probes/suffix.py, probes/gcg.py), and 40+ more. Each probe sets a goal, intent (taxonomy code), MISP-format tags, and a tier rating. TreeSearchProbe (probes/base.py:483-688) explores attack surface breadth-first or depth-first, branching on detector score thresholds. IterativeProbe (probes/base.py:691-843) supports multi-turn conversational attacks.
Buffs transform prompts at runtime: encoding/rot13/base64, leetspeak, typos, and payload injection. They sit between probe prompt generation and model invocation.
Detectors specific to red-teaming include Jailbreak (detectors/judge.py:176-264) using JailbreakBench methodology, Refusal (detectors/judge.py:123-157), ModelAsJudge for custom goal evaluation, substring-based detectors for known-bad signatures (detectors/knownbadsignatures.py), HuggingFace toxicity classifiers (detectors/lmrc.py), and leakage detection (detectors/leakreplay.py).
The EvaluationJudge framework (resources/red_team/evaluation.py:64-144) provides reusable scoring: judge_score() rates 1-10, on_topic_score() does YES/NO classification. Modality matching (harnesses/base.py:293-301) ensures probes target compatible model input types. No built-in CVE or structured vulnerability reporting beyond AVID report export.
UKGovernmentBEIS/inspect_ai
answeredInspect provides infrastructure useful for red-teaming but has no built-in adversarial probes, attack plugins, jailbreak harnesses, or dedicated red-teaming tooling.
What exists:
- Scanner support (
src/inspect_ai/_eval/task/scan.py). Aneval_setcan attachScannerConfigobjects that run per-sample analysis via the optionalinspect_scoutpackage. Designed for post-hoc scanning of completed transcripts, not adversarial generation. - Review system (
src/inspect_ai/review/). TheReviewerprotocol intercepts tool call results mid-execution and cancontinue,terminate, orescalate. - Sandbox environments (
src/inspect_ai/util/_sandbox/). Docker-based sandboxes with diagnostics and egress controls provide isolation for running untrusted model outputs. - Model-level safety settings. Several model provider modules expose safety/abuse-detection parameters (Anthropic, OpenAI, Google, Grok) but as passthroughs, not a managed red-teaming feature.
What is absent:
- No built-in adversarial attack library (no prompt injection generators, no jailbreak test suites, no fuzzing tools)
- No dedicated red-team evaluation harness or report format
- No vulnerability reporting workflow or CVE tracking
- No built-in harmfulness classifiers or refusal detectors
- No automated red-teaming loop that generates increasingly adversarial inputs
How users do red-teaming today: By writing custom @scorer functions that test for specific failure modes, using the Tool system to build adversarial tool environments, and leveraging the Plan/Solver system to construct multi-turn probe sequences. The framework is flexible enough for ad-hoc red-teaming through these primitives, but provides no turnkey solution.
Arize-ai/phoenix
insufficient evidencePhoenix does not provide built-in red-teaming or adversarial safety testing tooling. There is no dedicated module for adversarial probes, attack plugins, jailbreak prompt generation, prompt injection testing suites, or automated vulnerability reporting. The ToxicityEvaluator and PiiDetectionEvaluator are content-safety evaluators that can flag toxic or PII-containing outputs, but these are after-the-fact classifiers, not adversarial generators. The PXI eval dataset (evals/pxi/datasets/product_knowledge.yaml) includes examples tagged adversarial: true — these are test cases where the user prompt presents a wrong premise to verify the agent corrects it, but this is a manual convention in the dataset YAML, not an automated adversarial generation system. Tutorials on jailbreak/prompt-injection defense and realtime guardrails exist as educational cookbook patterns using Phoenix's observability to monitor guardrail effectiveness, not as built-in tooling. There is no attack fuzzer, no adversarial probe optimizer, and no vulnerability reporting pipeline at the reviewed commit.
confident-ai/deepeval
not applicableRed-teaming was previously part of this repository but has been extracted into a separate project. A README at deepeval/red_teaming/README.md states: "The Red Teaming module is now in DeepTeam for deepeval-v3.0 onwards" and directs users to https://github.com/confident-ai/deepteam. The current codebase has no adversarial probes, jailbreak tests, prompt-injection detectors, or attack plugins. Instead, safety-related evaluation is handled through metrics like ToxicityMetric, BiasMetric, MisuseMetric, RoleViolationMetric, NonAdviceMetric, and PIILeakageMetric, which apply LLM judges to classify outputs as toxic, biased, non-compliant, etc., but these are evaluation metrics, not adversarial generation or vulnerability disclosure tooling.
← How are traces or production data captured and linked to evaluations?