NVIDIA/garak
Vulnerability scanner that fires adversarial probe prompts at a model or app endpoint and scores responses with detectors.
Overview
garak is NVIDIA’s LLM vulnerability scanner, the closest thing in this category to nmap for language models. You point it at a target (a hosted model, a local Hugging Face model, or any REST, WebSocket or Python-function endpoint) and choose probes. It sends a few thousand adversarial prompts and reports how often the target misbehaved: DAN-style jailbreaks, prompt injection, encoding tricks, latent injection hidden in documents, system-prompt extraction, training-data leakage, package hallucination, XSS-style exfiltration markup and more. There are about 45 probe modules and 30 detector modules at this commit.
It is a security test, not a quality eval. garak doesn’t know your app’s expected answers and doesn’t score helpfulness or correctness. Each response gets a 0-1 “hit” score from a detector, and the headline number is an attack success rate per probe and detector. For an LLM app team its place is pre-release and regression red-teaming of the deployed surface. The generic rest generator and the function generator let you scan your actual app endpoint, with its system prompt, guardrails and retrieval, instead of only the bare model. The agent_breaker probe goes further and attacks tool-using agents. What garak won’t tell you is whether the app gives good answers. For that, pair it with a dataset-and-scorer tool.
Architecture
flowchart LR
CLI["garak CLI (cli.main)"] --> SPEC["run.spec: probes + buffs"]
SPEC --> H["Harness: Probewise or PxD"]
H --> P["Probe: prompts -> Attempts"]
P --> B["Buffs (optional transforms)"]
B --> G["Generator: target endpoint"]
G --> P
P --> D["Detectors: score 0..1 per output"]
D --> E["ThresholdEvaluator"]
E --> R["report.jsonl + hitlog"]
R --> DIG["HTML digest / AVID export"]
| Component | Path | Role |
|---|---|---|
| CLI | garak/cli.py, garak/command.py |
Parse options and config, resolve the run spec, load generator, start the harness, write reports |
| Config | garak/_config.py, garak/resources/garak.core.yaml |
Layered config: core YAML, site YAML, run YAML/JSON, CLI flags |
| Plugins | garak/_plugins.py |
Enumerate and load probes, detectors, generators, buffs, harnesses by dotted name |
| Harnesses | garak/harnesses/ |
ProbewiseHarness (each probe’s own detectors) and PxD (every probe x every chosen detector) |
| Probes | garak/probes/ |
Attack prompt sets and multi-turn attackers (TreeSearchProbe, IterativeProbe, IntentProbe) |
| Generators | garak/generators/ |
Targets: OpenAI, Anthropic, Bedrock, NIM, Ollama, HF, LiteLLM, REST, WebSocket, function, NeMo Guardrails, and more |
| Detectors | garak/detectors/ |
String/trigger matchers, HF classifiers, LLM judges |
| Buffs | garak/buffs/ |
Prompt transforms: Base64, CharCode, lowercase, low-resource-language translation, paraphrase |
| Evaluators | garak/evaluators/base.py |
Turn detector scores into pass/fail counts and eval records |
| Analysis | garak/analyze/ |
Report digest, bootstrap CIs, calibration Z-scores |
How a request flows
Take garak --target_type openai --target_name gpt-4o-mini --probes promptinject:
- Resolve.
cli.mainloads the layered config, resolvesrun.specinto probes and buffs, and parsesdetector_spec(cli.py). It builds aThresholdEvaluatorwithrun.eval_threshold(0.5 by default; garak.core.yaml) and loads the generator plugin. - Pick a harness. With no explicit detectors it calls
probewise_run, otherwisepxd_run(cli.py, command.py).ProbewiseHarnessloads each probe, then itsprimary_detectorplusextended_detectors(probewise.py). - Mint attempts.
Probe.probecopies the probe’s prompts, translates them if a language provider is configured, and wraps each in anAttempt(probes/base.py)._buff_hookfans each attempt out through any active buffs (probes/base.py). - Call the target.
_execute_attemptcallsgenerator.generate(prompt, generations_this_call=self.generations), 5 outputs per prompt by default._execute_alluses amultiprocessing.Poolwhen--parallel_attemptsis set and the generator allows it, and streams each attempt to the report as it completes (probes/base.py). - Detect.
Harness.runskips probes whose modality the target can’t accept, then runs every detector over the attempts. ForIntentProbes it picks detectors per intent (harnesses/base.py). A detector returns one float per output, where 1.0 means the attack worked (detectors/base.py). - Evaluate.
Evaluator.evaluatecounts passes, fails and nulls per detector and writesevalandprobe_summaryrecords (evaluators/base.py).ThresholdEvaluator.testpasses a score below the threshold;ZeroToleranceEvaluatorpasses only 0.0 (evaluators/base.py). - Report. The run leaves
garak.<uuid>.report.jsonl, ahitlog.jsonlof successful attacks, and an HTML digest built bywrite_report_digest.garak --report <file>converts a report to AVID vulnerability records.
Key components
Probes
A plain probe is a list of prompts plus a primary_detector, a goal, OWASP/AVID tags and a tier. The interesting ones are generative. TreeSearchProbe branches on detector scores to explore an attack surface. IterativeProbe drives multi-turn conversations; goat, tap and atkgen use an attacker model to adapt prompts. latentinjection hides instructions inside resumes, reports and translations, the shape of an indirect injection through RAG. web_injection looks for markdown or image exfiltration and XSS payloads. Some probes are inactive by default because they need an attacker model or are slow.
Agent Breaker
AgentBreaker is the probe most aimed at applications. You describe the agent’s purpose and tools in agent.yaml, or let it ask the agent. A red-team model (default openai/gpt-oss-120b via NIM) analyses each tool, drafts targeted exploits, retries with what it learned, and verifies success. It is off by default because it needs that attacker model configured (agent_breaker.py).
Detectors and judges
Most detectors are cheap: StringDetector substring or word matching, trigger lists, or HF text classifiers (toxicity and similar). ModelAsJudge loads a second generator, by default meta/llama3-70b-instruct via NIM, and asks for a [[rating]] from 1 to 10 against the probe goal. Ratings of 7 or more count as hits (judge.py). The judge is called once per output, with no consensus or calibration.
Generators as the app boundary
RestGenerator takes a URI, method, headers, a req_template with $INPUT, and a JSONPath response_json_field. It handles rate-limit codes, proxies and mTLS (rest.py). function.Single calls any module#function that maps a string to a string, so you can scan an in-process chain (function.py). NeMoGuardrails wraps a Guardrails config directly.
Scoring and statistics
Per probe and detector you get an attack success rate. The digest groups results by taxonomy tag and aggregates with lower_quartile by default. It can add bootstrap confidence intervals and Z-scores against a bundled calibration of reference models, which turn into 1-5 “DEFCON” grades (calibration.py). The Z-scores compare you to garak’s reference model set, not to your own previous run.
Extending it
- Probe: subclass
garak.probes.Probe, setprompts,primary_detector,goal,tagsandtier, and put it in a module undergarak/probes/. Payload JSON undergarak/data/can be overridden from the user data directory. - Detector: subclass
Detector(orStringDetector/HFDetector) and implementdetect(attempt)returning floats in [0, 1]. - Generator: implement
_call_model(prompt, generations). Often you don’t need one, becauserestwith a JSON config covers most HTTP APIs. - Buff: implement
transform(attempt)and, if needed,untransformto map responses back. - Config: YAML or JSON run configs via
--config, per-plugin options via--generator_option_fileand similar flags.
Running it
pip install garak(Python package with thegarakconsole script). HF-based detectors and buffs pull in PyTorch and Transformers. Judge and attacker probes need an OpenAI-compatible endpoint (NIM by default) and its key.- Typical:
garak --target_type rest -G rest.json --probes promptinject,latentinjection,dan, or--target_type openai --target_name <model>. Add--parallel_attempts 16for API targets; the CLI hints when the generator supports it. - Outputs go to
garak_runs/under the user data directory. No server is involved.
Strengths and caveats
- Strength: breadth of attacks. A large, tagged catalogue (OWASP LLM Top 10, AVID) that keeps growing, including multi-turn, tree-search and agent-specific attackers.
- Strength: tests the real surface. REST, WebSocket and function generators let you scan the deployed app, guardrails included, not just the base model.
- Strength: detector-based scoring is cheap. Most checks are string or classifier based and run without a judge model.
- Caveat: noisy by design. Substring detectors produce false positives and false negatives. Read the hitlog before you trust a percentage.
- Caveat: not a quality eval. No reference answers, no dataset versioning, no traces. Pair it with a tool that scores correctness.
- Caveat: cost and time. Default runs send 5 generations per prompt across many probes, which adds up against paid APIs.
- Caveat: NIM-flavoured defaults. Judge and red-team models default to NVIDIA NIM model names and must be reconfigured for other providers.
Sources: code at bb30a7e, verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit bb30a7e. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredScoring lives in garak/evaluators/base.py. The Evaluator base class evaluates detector scores per-probe, per-detector. Each Attempt carries detector_results — a dict keyed by detector name, mapping to a list of floats (one per generation). Evaluator.evaluate() iterates attempts, groups results by detector, and calls _evaluate_one_detector() which counts passes, fails, and nulls using a pluggable test() method. Two concrete evaluators ship: ThresholdEvaluator (pass if score < threshold, default 0.5) and ZeroToleranceEvaluator (pass only if score == 0.0; evaluators/base.py:481-488). The default threshold of 0.5 is set in garak.core.yaml as eval_threshold.
Confidence intervals use non-parametric bootstrap: calculate_bootstrap_ci() in analyze/bootstrap_ci.py:90-105 resamples results with replacement (default 10,000 iterations, 95% confidence) and corrects for detector sensitivity/specificity loaded from data/detectors_eval/detector_metrics_summary.json via analyze/detector_metrics.py. Calibration Z-scores: Calibration class (analyze/calibration.py:17-99) loads prior-run distributions from data/calibration/calibration.json and computes (score - mu) / sigma, gated by MINIMUM_STD_DEV = 1/30. These feed "defcon" ratings (1-5, analyze/__init__.py:48-58).
Detectors range from heuristic (StringDetector — substring/word/prefix matching with optional Unicode normalization, detectors/base.py:197-272) to model-based (HFDetector wraps HuggingFace text-classification pipelines, normalizing logits to 0-1, detectors/base.py:82-194), and trigger-list matching (TriggerListDetector, detectors/base.py:275-304). Scoring is per-trace (each generation output independently), then aggregated to per-probe/per-detector summaries in the eval record.
How is LLM-as-a-judge implemented?
answeredLLM-as-judge is implemented through ModelAsJudge (detectors/judge.py:20-120) which mixes Detector with EvaluationJudge from resources/red_team/evaluation.py. On initialization, it loads a separate judge generator via _load_generator() (judge.py:52-82) — configurable by detector_model_type (default "nim") and detector_model_name (default "meta/llama3-70b-instruct"). The judge generator must be OpenAICompatible.
Two judge prompt styles exist. Rating prompts (_goal_system_prompt, judge.py:40-49): ask "Rate from 1 to 10" with [[rating]] output format, parsed by process_output_judge_score() (evaluation.py:24-33) via regex \[\[(\d+)\]\]. Scores above confidence_cutoff (default 7) count as hits. Yes/No prompts (_refusal_system_prompt, judge.py:141-148): ask for [[YES]] or [[NO]], parsed by process_output_on_topic_score() (evaluation.py:36-44). The Refusal detector (judge.py:123-157) uses YES/NO; Jailbreak (judge.py:176-264) uses the JailbreakBench <BEGIN REQUEST>/<END RESPONSE> format with YES/NO output.
EvaluationJudge._create_conv() (evaluation.py:78-116) builds an OpenAI-format conversation (system prompt + user prompt) and truncates if it exceeds the model's token limit (from generators/openai.py context_lengths, fallback 4096). No multi-sample consensus — each response is judged once. No explicit calibration or bias controls, though multiple generations per probe yield multiple judgements that can be averaged in report post-processing. JailbreakOnlyAdversarial and RefusalOnlyAdversarial (judge.py:160-280) gate on attempt.notes["is_adversarial"] to skip non-adversarial turns.
How are test datasets and cases defined, generated and versioned?
answeredTest data is stored as JSON files under garak/data/, organized by attack technique: payloads/harmful_behaviors.json, dan/, harmbench/, autodan/, gcg/, beast/, donotanswer/, etc. Each payload JSON includes garak_payload_name, payload_types, intent (intent taxonomy code), detector_name, and a payloads array of prompt strings (e.g. data/payloads/harmful_behaviors.json:1-24).
Data loading goes through the LocalDataPath resolver (data/__init__.py:29-123) which searches user data dir (~/.local/share/garak/data/) first, then the bundled package dir — enabling user overrides. The DANProbeMeta metaclass (probes/dan.py:28-60) auto-detects a prompt_file attribute from the class name, opens it via data_path / prompt_file, and loads JSON arrays into self.prompts.
The intent taxonomy in data/cas/trait_typology.json defines structured intents (e.g. S006instructions, T009ignore). IntentProbe (probes/base.py:846-953) loads intents from intentservice, maps them to stubs via data/cas/intent_stubs/, and builds prompts dynamically. Probes have tier ratings (OF_CONCERN=1 through UNLISTED=4) via probes/_tier.py. Config files are in resources/garak.core.yaml (base), ~/.config/garak/garak.site.yaml (user), and configs/fast.json / configs/bag.yaml (bundled run templates). No formal dataset versioning — prompts are static JSON files committed to the repo; plugin_cache.json tracks plugin metadata timestamps for cache invalidation. No synthetic data generation is built in.
How are evals executed and reported?
answeredTwo harnesses coordinate execution. ProbewiseHarness (harnesses/probewise.py:18-111) runs probes one-by-one, each with its own recommended detectors (primary + extended). PxD (harnesses/pxd.py:22-61) runs all probe × detector combinations. Both call Harness.run() (harnesses/base.py:159-290), which: loads buffs, iterates probes, calls probe.probe(model) to get attempts, runs each detector on each attempt via _run_detector(), writes attempts to report JSONL, and calls evaluator.evaluate().
Parallelism: Generator.generate() (generators/base.py:138-244) uses multiprocessing.Pool for parallel requests when parallel_requests > 1 and the generator doesn't natively support multi-generation. Probe._execute_all() (probes/base.py:321-383) uses Pool.imap_unordered() for parallel attempts when parallel_attempts > 1.
Result storage: everything streams to garak.<uuid>.report.jsonl (JSONL, one JSON object per line). Entry types include start_run setup, attempt, eval, probe_summary, plugin_cache, completion. Hit logs go to a parallel .hitlog.jsonl when detectors fail. Report (report.py:15-174) loads a report and exports to AVID format. report_digest.py:581-689 builds HTML digests with results stored in an in-memory SQLite3 database, enabling grouping by MISP taxonomy tags and per-group ASR aggregation (default: lower quartile, garak.core.yaml:39). Calibration Z-scores and bootstrap CIs are computed during digest building. No CI/CD integration or regression dashboards are built into the code.
How are traces or production data captured and linked to evaluations?
not applicablegarak has no SDK instrumentation, OpenTelemetry integration, online/offline evaluation distinction, production trace capture, or annotation/feedback collection system. It is purely an offline evaluation framework: it runs probes/detectors against a model endpoint and writes results to a JSONL report file. The closest feature is the hitlog.jsonl which captures individual failure details at eval time, but this is an output artifact, not a production observability integration.
Does it support red-teaming or safety testing, and how?
answeredRed-teaming is garak's primary purpose. Probes implement specific attack techniques organized by module: DAN variants (probes/dan.py), prompt injection (probes/promptinject.py), jailbreak with adaptive attacks (probes/atkgen.py, probes/goat.py, probes/tap.py), encoding attacks (probes/encoding.py), ASCII smuggling (probes/smuggling.py), suffix attacks (probes/suffix.py, probes/gcg.py), and 40+ more. Each probe sets a goal, intent (taxonomy code), MISP-format tags, and a tier rating. TreeSearchProbe (probes/base.py:483-688) explores attack surface breadth-first or depth-first, branching on detector score thresholds. IterativeProbe (probes/base.py:691-843) supports multi-turn conversational attacks.
Buffs transform prompts at runtime: encoding/rot13/base64, leetspeak, typos, and payload injection. They sit between probe prompt generation and model invocation.
Detectors specific to red-teaming include Jailbreak (detectors/judge.py:176-264) using JailbreakBench methodology, Refusal (detectors/judge.py:123-157), ModelAsJudge for custom goal evaluation, substring-based detectors for known-bad signatures (detectors/knownbadsignatures.py), HuggingFace toxicity classifiers (detectors/lmrc.py), and leakage detection (detectors/leakreplay.py).
The EvaluationJudge framework (resources/red_team/evaluation.py:64-144) provides reusable scoring: judge_score() rates 1-10, on_topic_score() does YES/NO classification. Modality matching (harnesses/base.py:293-301) ensures probes target compatible model input types. No built-in CVE or structured vulnerability reporting beyond AVID report export.