LLMs Technical Reviews

NVIDIA/garak

Vulnerability scanner that fires adversarial probe prompts at a model or app endpoint and scores responses with detectors.

GitHub ↗★ 9.5kPythonApache-2.0commit bb30a7e · 2026-10-02homepage ↗

Overview

garak is NVIDIA’s LLM vulnerability scanner, the closest thing in this category to nmap for language models. You point it at a target (a hosted model, a local Hugging Face model, or any REST, WebSocket or Python-function endpoint) and choose probes. It sends a few thousand adversarial prompts and reports how often the target misbehaved: DAN-style jailbreaks, prompt injection, encoding tricks, latent injection hidden in documents, system-prompt extraction, training-data leakage, package hallucination, XSS-style exfiltration markup and more. There are about 45 probe modules and 30 detector modules at this commit.

It is a security test, not a quality eval. garak doesn’t know your app’s expected answers and doesn’t score helpfulness or correctness. Each response gets a 0-1 “hit” score from a detector, and the headline number is an attack success rate per probe and detector. For an LLM app team its place is pre-release and regression red-teaming of the deployed surface. The generic rest generator and the function generator let you scan your actual app endpoint, with its system prompt, guardrails and retrieval, instead of only the bare model. The agent_breaker probe goes further and attacks tool-using agents. What garak won’t tell you is whether the app gives good answers. For that, pair it with a dataset-and-scorer tool.

Architecture

flowchart LR
  CLI["garak CLI (cli.main)"] --> SPEC["run.spec: probes + buffs"]
  SPEC --> H["Harness: Probewise or PxD"]
  H --> P["Probe: prompts -> Attempts"]
  P --> B["Buffs (optional transforms)"]
  B --> G["Generator: target endpoint"]
  G --> P
  P --> D["Detectors: score 0..1 per output"]
  D --> E["ThresholdEvaluator"]
  E --> R["report.jsonl + hitlog"]
  R --> DIG["HTML digest / AVID export"]
Component Path Role
CLI garak/cli.py, garak/command.py Parse options and config, resolve the run spec, load generator, start the harness, write reports
Config garak/_config.py, garak/resources/garak.core.yaml Layered config: core YAML, site YAML, run YAML/JSON, CLI flags
Plugins garak/_plugins.py Enumerate and load probes, detectors, generators, buffs, harnesses by dotted name
Harnesses garak/harnesses/ ProbewiseHarness (each probe’s own detectors) and PxD (every probe x every chosen detector)
Probes garak/probes/ Attack prompt sets and multi-turn attackers (TreeSearchProbe, IterativeProbe, IntentProbe)
Generators garak/generators/ Targets: OpenAI, Anthropic, Bedrock, NIM, Ollama, HF, LiteLLM, REST, WebSocket, function, NeMo Guardrails, and more
Detectors garak/detectors/ String/trigger matchers, HF classifiers, LLM judges
Buffs garak/buffs/ Prompt transforms: Base64, CharCode, lowercase, low-resource-language translation, paraphrase
Evaluators garak/evaluators/base.py Turn detector scores into pass/fail counts and eval records
Analysis garak/analyze/ Report digest, bootstrap CIs, calibration Z-scores

How a request flows

Take garak --target_type openai --target_name gpt-4o-mini --probes promptinject:

  1. Resolve. cli.main loads the layered config, resolves run.spec into probes and buffs, and parses detector_spec (cli.py). It builds a ThresholdEvaluator with run.eval_threshold (0.5 by default; garak.core.yaml) and loads the generator plugin.
  2. Pick a harness. With no explicit detectors it calls probewise_run, otherwise pxd_run (cli.py, command.py). ProbewiseHarness loads each probe, then its primary_detector plus extended_detectors (probewise.py).
  3. Mint attempts. Probe.probe copies the probe’s prompts, translates them if a language provider is configured, and wraps each in an Attempt (probes/base.py). _buff_hook fans each attempt out through any active buffs (probes/base.py).
  4. Call the target. _execute_attempt calls generator.generate(prompt, generations_this_call=self.generations), 5 outputs per prompt by default. _execute_all uses a multiprocessing.Pool when --parallel_attempts is set and the generator allows it, and streams each attempt to the report as it completes (probes/base.py).
  5. Detect. Harness.run skips probes whose modality the target can’t accept, then runs every detector over the attempts. For IntentProbes it picks detectors per intent (harnesses/base.py). A detector returns one float per output, where 1.0 means the attack worked (detectors/base.py).
  6. Evaluate. Evaluator.evaluate counts passes, fails and nulls per detector and writes eval and probe_summary records (evaluators/base.py). ThresholdEvaluator.test passes a score below the threshold; ZeroToleranceEvaluator passes only 0.0 (evaluators/base.py).
  7. Report. The run leaves garak.<uuid>.report.jsonl, a hitlog.jsonl of successful attacks, and an HTML digest built by write_report_digest. garak --report <file> converts a report to AVID vulnerability records.

Key components

Probes

A plain probe is a list of prompts plus a primary_detector, a goal, OWASP/AVID tags and a tier. The interesting ones are generative. TreeSearchProbe branches on detector scores to explore an attack surface. IterativeProbe drives multi-turn conversations; goat, tap and atkgen use an attacker model to adapt prompts. latentinjection hides instructions inside resumes, reports and translations, the shape of an indirect injection through RAG. web_injection looks for markdown or image exfiltration and XSS payloads. Some probes are inactive by default because they need an attacker model or are slow.

Agent Breaker

AgentBreaker is the probe most aimed at applications. You describe the agent’s purpose and tools in agent.yaml, or let it ask the agent. A red-team model (default openai/gpt-oss-120b via NIM) analyses each tool, drafts targeted exploits, retries with what it learned, and verifies success. It is off by default because it needs that attacker model configured (agent_breaker.py).

Detectors and judges

Most detectors are cheap: StringDetector substring or word matching, trigger lists, or HF text classifiers (toxicity and similar). ModelAsJudge loads a second generator, by default meta/llama3-70b-instruct via NIM, and asks for a [[rating]] from 1 to 10 against the probe goal. Ratings of 7 or more count as hits (judge.py). The judge is called once per output, with no consensus or calibration.

Generators as the app boundary

RestGenerator takes a URI, method, headers, a req_template with $INPUT, and a JSONPath response_json_field. It handles rate-limit codes, proxies and mTLS (rest.py). function.Single calls any module#function that maps a string to a string, so you can scan an in-process chain (function.py). NeMoGuardrails wraps a Guardrails config directly.

Scoring and statistics

Per probe and detector you get an attack success rate. The digest groups results by taxonomy tag and aggregates with lower_quartile by default. It can add bootstrap confidence intervals and Z-scores against a bundled calibration of reference models, which turn into 1-5 “DEFCON” grades (calibration.py). The Z-scores compare you to garak’s reference model set, not to your own previous run.

Extending it

  • Probe: subclass garak.probes.Probe, set prompts, primary_detector, goal, tags and tier, and put it in a module under garak/probes/. Payload JSON under garak/data/ can be overridden from the user data directory.
  • Detector: subclass Detector (or StringDetector/HFDetector) and implement detect(attempt) returning floats in [0, 1].
  • Generator: implement _call_model(prompt, generations). Often you don’t need one, because rest with a JSON config covers most HTTP APIs.
  • Buff: implement transform(attempt) and, if needed, untransform to map responses back.
  • Config: YAML or JSON run configs via --config, per-plugin options via --generator_option_file and similar flags.

Running it

  • pip install garak (Python package with the garak console script). HF-based detectors and buffs pull in PyTorch and Transformers. Judge and attacker probes need an OpenAI-compatible endpoint (NIM by default) and its key.
  • Typical: garak --target_type rest -G rest.json --probes promptinject,latentinjection,dan, or --target_type openai --target_name <model>. Add --parallel_attempts 16 for API targets; the CLI hints when the generator supports it.
  • Outputs go to garak_runs/ under the user data directory. No server is involved.

Strengths and caveats

  • Strength: breadth of attacks. A large, tagged catalogue (OWASP LLM Top 10, AVID) that keeps growing, including multi-turn, tree-search and agent-specific attackers.
  • Strength: tests the real surface. REST, WebSocket and function generators let you scan the deployed app, guardrails included, not just the base model.
  • Strength: detector-based scoring is cheap. Most checks are string or classifier based and run without a judge model.
  • Caveat: noisy by design. Substring detectors produce false positives and false negatives. Read the hitlog before you trust a percentage.
  • Caveat: not a quality eval. No reference answers, no dataset versioning, no traces. Pair it with a tool that scores correctness.
  • Caveat: cost and time. Default runs send 5 generations per prompt across many probes, which adds up against paid APIs.
  • Caveat: NIM-flavoured defaults. Judge and red-team models default to NVIDIA NIM model names and must be reconfigured for other providers.

Sources: code at bb30a7e, verified Q&A.

How it answers the LLM evals and testing questions

Each answer was drafted by a code-reading agent at commit bb30a7e. Its citations were checked mechanically. Compare with the other llm evals and testing →

Which evaluation metrics and scorers are provided, and how are they implemented?

answered

Scoring lives in garak/evaluators/base.py. The Evaluator base class evaluates detector scores per-probe, per-detector. Each Attempt carries detector_results — a dict keyed by detector name, mapping to a list of floats (one per generation). Evaluator.evaluate() iterates attempts, groups results by detector, and calls _evaluate_one_detector() which counts passes, fails, and nulls using a pluggable test() method. Two concrete evaluators ship: ThresholdEvaluator (pass if score < threshold, default 0.5) and ZeroToleranceEvaluator (pass only if score == 0.0; evaluators/base.py:481-488). The default threshold of 0.5 is set in garak.core.yaml as eval_threshold.

Confidence intervals use non-parametric bootstrap: calculate_bootstrap_ci() in analyze/bootstrap_ci.py:90-105 resamples results with replacement (default 10,000 iterations, 95% confidence) and corrects for detector sensitivity/specificity loaded from data/detectors_eval/detector_metrics_summary.json via analyze/detector_metrics.py. Calibration Z-scores: Calibration class (analyze/calibration.py:17-99) loads prior-run distributions from data/calibration/calibration.json and computes (score - mu) / sigma, gated by MINIMUM_STD_DEV = 1/30. These feed "defcon" ratings (1-5, analyze/__init__.py:48-58).

Detectors range from heuristic (StringDetector — substring/word/prefix matching with optional Unicode normalization, detectors/base.py:197-272) to model-based (HFDetector wraps HuggingFace text-classification pipelines, normalizing logits to 0-1, detectors/base.py:82-194), and trigger-list matching (TriggerListDetector, detectors/base.py:275-304). Scoring is per-trace (each generation output independently), then aggregated to per-probe/per-detector summaries in the eval record.

How is LLM-as-a-judge implemented?

answered

LLM-as-judge is implemented through ModelAsJudge (detectors/judge.py:20-120) which mixes Detector with EvaluationJudge from resources/red_team/evaluation.py. On initialization, it loads a separate judge generator via _load_generator() (judge.py:52-82) — configurable by detector_model_type (default "nim") and detector_model_name (default "meta/llama3-70b-instruct"). The judge generator must be OpenAICompatible.

Two judge prompt styles exist. Rating prompts (_goal_system_prompt, judge.py:40-49): ask "Rate from 1 to 10" with [[rating]] output format, parsed by process_output_judge_score() (evaluation.py:24-33) via regex \[\[(\d+)\]\]. Scores above confidence_cutoff (default 7) count as hits. Yes/No prompts (_refusal_system_prompt, judge.py:141-148): ask for [[YES]] or [[NO]], parsed by process_output_on_topic_score() (evaluation.py:36-44). The Refusal detector (judge.py:123-157) uses YES/NO; Jailbreak (judge.py:176-264) uses the JailbreakBench <BEGIN REQUEST>/<END RESPONSE> format with YES/NO output.

EvaluationJudge._create_conv() (evaluation.py:78-116) builds an OpenAI-format conversation (system prompt + user prompt) and truncates if it exceeds the model's token limit (from generators/openai.py context_lengths, fallback 4096). No multi-sample consensus — each response is judged once. No explicit calibration or bias controls, though multiple generations per probe yield multiple judgements that can be averaged in report post-processing. JailbreakOnlyAdversarial and RefusalOnlyAdversarial (judge.py:160-280) gate on attempt.notes["is_adversarial"] to skip non-adversarial turns.

How are test datasets and cases defined, generated and versioned?

answered

Test data is stored as JSON files under garak/data/, organized by attack technique: payloads/harmful_behaviors.json, dan/, harmbench/, autodan/, gcg/, beast/, donotanswer/, etc. Each payload JSON includes garak_payload_name, payload_types, intent (intent taxonomy code), detector_name, and a payloads array of prompt strings (e.g. data/payloads/harmful_behaviors.json:1-24).

Data loading goes through the LocalDataPath resolver (data/__init__.py:29-123) which searches user data dir (~/.local/share/garak/data/) first, then the bundled package dir — enabling user overrides. The DANProbeMeta metaclass (probes/dan.py:28-60) auto-detects a prompt_file attribute from the class name, opens it via data_path / prompt_file, and loads JSON arrays into self.prompts.

The intent taxonomy in data/cas/trait_typology.json defines structured intents (e.g. S006instructions, T009ignore). IntentProbe (probes/base.py:846-953) loads intents from intentservice, maps them to stubs via data/cas/intent_stubs/, and builds prompts dynamically. Probes have tier ratings (OF_CONCERN=1 through UNLISTED=4) via probes/_tier.py. Config files are in resources/garak.core.yaml (base), ~/.config/garak/garak.site.yaml (user), and configs/fast.json / configs/bag.yaml (bundled run templates). No formal dataset versioning — prompts are static JSON files committed to the repo; plugin_cache.json tracks plugin metadata timestamps for cache invalidation. No synthetic data generation is built in.

How are evals executed and reported?

answered

Two harnesses coordinate execution. ProbewiseHarness (harnesses/probewise.py:18-111) runs probes one-by-one, each with its own recommended detectors (primary + extended). PxD (harnesses/pxd.py:22-61) runs all probe × detector combinations. Both call Harness.run() (harnesses/base.py:159-290), which: loads buffs, iterates probes, calls probe.probe(model) to get attempts, runs each detector on each attempt via _run_detector(), writes attempts to report JSONL, and calls evaluator.evaluate().

Parallelism: Generator.generate() (generators/base.py:138-244) uses multiprocessing.Pool for parallel requests when parallel_requests > 1 and the generator doesn't natively support multi-generation. Probe._execute_all() (probes/base.py:321-383) uses Pool.imap_unordered() for parallel attempts when parallel_attempts > 1.

Result storage: everything streams to garak.<uuid>.report.jsonl (JSONL, one JSON object per line). Entry types include start_run setup, attempt, eval, probe_summary, plugin_cache, completion. Hit logs go to a parallel .hitlog.jsonl when detectors fail. Report (report.py:15-174) loads a report and exports to AVID format. report_digest.py:581-689 builds HTML digests with results stored in an in-memory SQLite3 database, enabling grouping by MISP taxonomy tags and per-group ASR aggregation (default: lower quartile, garak.core.yaml:39). Calibration Z-scores and bootstrap CIs are computed during digest building. No CI/CD integration or regression dashboards are built into the code.

How are traces or production data captured and linked to evaluations?

not applicable

garak has no SDK instrumentation, OpenTelemetry integration, online/offline evaluation distinction, production trace capture, or annotation/feedback collection system. It is purely an offline evaluation framework: it runs probes/detectors against a model endpoint and writes results to a JSONL report file. The closest feature is the hitlog.jsonl which captures individual failure details at eval time, but this is an output artifact, not a production observability integration.

Does it support red-teaming or safety testing, and how?

answered

Red-teaming is garak's primary purpose. Probes implement specific attack techniques organized by module: DAN variants (probes/dan.py), prompt injection (probes/promptinject.py), jailbreak with adaptive attacks (probes/atkgen.py, probes/goat.py, probes/tap.py), encoding attacks (probes/encoding.py), ASCII smuggling (probes/smuggling.py), suffix attacks (probes/suffix.py, probes/gcg.py), and 40+ more. Each probe sets a goal, intent (taxonomy code), MISP-format tags, and a tier rating. TreeSearchProbe (probes/base.py:483-688) explores attack surface breadth-first or depth-first, branching on detector score thresholds. IterativeProbe (probes/base.py:691-843) supports multi-turn conversational attacks.

Buffs transform prompts at runtime: encoding/rot13/base64, leetspeak, typos, and payload injection. They sit between probe prompt generation and model invocation.

Detectors specific to red-teaming include Jailbreak (detectors/judge.py:176-264) using JailbreakBench methodology, Refusal (detectors/judge.py:123-157), ModelAsJudge for custom goal evaluation, substring-based detectors for known-bad signatures (detectors/knownbadsignatures.py), HuggingFace toxicity classifiers (detectors/lmrc.py), and leakage detection (detectors/leakreplay.py).

The EvaluationJudge framework (resources/red_team/evaluation.py:64-144) provides reusable scoring: judge_score() rates 1-10, on_topic_score() does YES/NO classification. Modality matching (harnesses/base.py:293-301) ensures probes target compatible model input types. No built-in CVE or structured vulnerability reporting beyond AVID report export.

Editor's note. Correction: the shipped buffs are Base64, CharCode, Lowercase, LRLBuff (low-resource-language translation) and the PegasusT5/Fast paraphrasers in garak/buffs/. There are no leetspeak, typo or payload-injection buffs.