# promptfoo/promptfoo

> Local-first CLI and Node library that runs prompt x provider x test matrices, grades them with ~70 assertion types, and red-teams targets.

- Category: [LLM evals and testing](https://llms-technical-reviews.com/evals/)
- Repository: https://github.com/promptfoo/promptfoo (reviewed at commit `1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61`, 2026-10-06)
- Stars: 25766 · Language: TypeScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/promptfoo/

## Overview

promptfoo is a command-line tool and Node library for testing LLM prompts, providers and agents. You write a `promptfooconfig.yaml` with prompts, providers and tests. Each test has variables and a list of assertions. `promptfoo eval` expands this into a matrix (every prompt × every provider × every test, optionally repeated), calls each provider, grades every output, and stores the results in a local SQLite database. `promptfoo view` serves a React UI over that database.

The same engine runs red-team scans. `promptfoo redteam run` first generates adversarial test cases with plugins (what to attack) and strategies (how to wrap the attack), writes them to `redteam.yaml`, and then runs an ordinary eval over that file. Red-teaming is a test generator in front of the eval engine, not a separate system.

The whole repository is MIT-licensed. There is no `ee/` folder or enterprise-licensed code. The commercial boundary is a hosted service instead. Many red-team plugins, some grading and part of attack generation call promptfoo's API, and a free monthly probe limit applies when you are not logged in to promptfoo Cloud (details below).

## Architecture

```mermaid
flowchart LR
  CFG["promptfooconfig.yaml"] --> CLI["CLI: doEval"]
  RT["redteam generate"] --> YAML["redteam.yaml"]
  YAML --> CLI
  CLI --> EV["Evaluator (evaluator.ts)"]
  EV --> RUN["runEval: render + call"]
  RUN --> PROV["Providers (src/providers)"]
  PROV --> CACHE["Cache (memory + disk)"]
  RUN --> ASSERT["runAssertions"]
  ASSERT --> GRADE["llmGrading matchers"]
  ASSERT --> TRACE["Trace store"]
  OTLP["OTLP receiver /v1/traces"] --> TRACE
  EV --> DB["SQLite via Drizzle"]
  DB --> UI["Express server + React UI"]
  RT --> REMOTE["Remote generation API"]
```

| Component | Path | Role |
|---|---|---|
| CLI and commands | `src/main.ts`, `src/commands/`, `src/node/doEval.ts` | Parse flags, resolve configs, call `evaluate` |
| Evaluator | `src/evaluator.ts` | Builds the test matrix, schedules steps, runs comparisons, persists rows |
| Single step | `runEval` in `src/evaluator.ts` | Renders the prompt with Nunjucks, calls the provider, transforms, grades |
| Assertions | `src/assertions/` | `ASSERTION_HANDLERS` registry and `runAssertions` aggregation |
| Model graders | `src/matchers/llmGrading.ts` | `llm-rubric`, factuality, closed-QA, G-Eval and others |
| Providers | `src/providers/` | Over 100 provider modules, including generic HTTP, WebSocket, script and MCP targets |
| Cache | `src/cache.ts` | `cache-manager` over Keyv: memory plus a `KeyvFile` disk store |
| Storage | `src/database/tables.ts` | Drizzle schema for evals, results, prompts, datasets, traces, spans |
| Tracing | `src/tracing/` | OTel SDK setup, local OTLP receiver, trace store, external trace providers |
| Red team | `src/redteam/` | Plugins, strategies, graders, risk scoring, remote generation |
| Web UI | `src/server/`, `src/app/` | Express API and React app for results and red-team setup |

## How a request flows

For `promptfoo eval`:

1. **Resolve.** `doEval` loads and merges configs and builds a `TestSuite`. The library `evaluate` sets base path, env and config context, picks the Node runtime and constructs an `Evaluator` ([evaluator.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L5576-L5627)).
2. **Expand.** The evaluator merges `defaultTest` and scenarios into each test, expands array-valued vars into combinations (`generateVarCombinations`), and creates one `RunEvalOptions` per prompt × provider × test × repeat.
3. **Schedule.** Concurrency defaults to 4 (`-j` overrides it). It drops to 1 when a prompt uses the `_conversation` variable, a test uses `storeOutputAs`, or a browser provider keeps a persistent session ([evaluator.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L3224-L3259)). Concurrent steps run through `async.forEachOfLimit` and periodically flush prompt metrics to the store ([L4394-L4425](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L4394-L4425)).
4. **Run a step.** `runEval` wraps the step in a cache namespace keyed by repeat index. `runEvalInternal` then fills runtime vars, renders the prompt, optionally creates an agent workspace, and generates a W3C trace context when tracing is on ([L1618-L1700](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1618-L1700)). It then calls the provider, applies transforms and grades the response.
5. **Grade.** `runAssertions` expands `assert-set` groups and runs assertions with bounded concurrency (default 3). Each result is added to an `AssertionsResult` with its weight and metric name, and the final pass/score applies the test `threshold` or a custom scoring function ([assertions/index.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L755-L800), [L860-L911](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L860-L911)).
6. **Compare.** After all rows finish, `select-best` and `max-score` assertions grade outputs for the same test across providers. Rows with model-graded assertions can be deferred and graded in groups per grading provider.
7. **Persist.** Rows go to `eval_results` and the run summary to `evals` in SQLite. `evals.config` keeps a snapshot of the resolved config ([tables.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/database/tables.ts#L58-L100)). Optional JSON, CSV or HTML output files are written too.

## Key components

### Assertions

`ASSERTION_HANDLERS` maps about 70 base types to handlers ([assertions/index.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L229-L321)). They fall into four groups:
- deterministic checks: `equals`, `contains*`, `regex`, `is-json`, `is-sql`, `cost`, `latency`, `word-count`;
- reference metrics: `bleu`, `rouge-n`, `levenshtein`, `similar` (embeddings), `tool-call-f1`;
- model-graded checks: `llm-rubric`, `g-eval`, `factuality`, `answer-relevance`, `context-*`, `moderation`;
- trace and trajectory checks over recorded spans: `trace-span-count`, `trace-error-spans`, `trajectory:tool-sequence`, `trajectory:goal-success`.

Any type can be negated with a `not-` prefix. `javascript`, `python`, `ruby` and `webhook` cover custom logic.

### LLM grading

`matchesLlmRubric` renders a grading prompt, calls the grading provider and parses a JSON `{pass, score, reason}` ([llmGrading.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L181-L268)). The grader comes from `options.provider` (or `--grader`), otherwise a default grading provider. Inside a red-team config with no explicit grader, it prefers promptfoo's remote grading endpoint. Grading calls are not deduplicated. `ProviderGroupedCallQueue` only groups queued calls by provider id so one grader's calls run together ([providerCallQueue.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/scheduler/providerCallQueue.ts#L14-L57)).

### Cache

Provider responses go to a `cache-manager` cache with a memory store and a `KeyvFile` disk store at `~/.promptfoo/cache/cache.json`, with a 14-day default TTL that `PROMPTFOO_CACHE_TTL` can change ([cache.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/cache.ts#L43-L51), [L80-L105](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/cache.ts#L80-L105)). Repeats get separate namespaces, so `--repeat 3` makes three real calls.

### Tracing

With tracing enabled, promptfoo starts a local Express OTLP receiver. It accepts `/v1/traces` (JSON or protobuf) and `/v1/logs` and writes spans into SQLite `traces`/`spans` tables ([otlpReceiver.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/tracing/otlpReceiver.ts#L456-L515)). The target app receives a `traceparent` and exports its spans there. Trace assertions then poll until the span count settles. Langfuse, Braintrust and Tempo can also serve as trace sources. Tracing exists to grade agent runs, not to monitor production.

### Red team

`synthesize` in `src/redteam/index.ts` drives plugins (dozens, under `plugins/`) and strategies (about 30, under `strategies/`) to produce test cases ([index.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/index.ts#L963-L975)). `doRedteamRun` calls `doGenerateRedteam`, then `doEval` on the generated file ([shared.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/shared.ts#L95-L160)). Results are scored by per-plugin graders and summarised with severity-weighted risk scores.

How much runs locally is the main thing to know. `REMOTE_ONLY_PLUGIN_IDS` lists plugins that have no local implementation and always call the remote API: BOLA/BFLA, SSRF, indirect prompt injection, RAG poisoning, the MCP plugin, the coding-agent sets and all the medical, financial, pharmacy, insurance, e-commerce, telecom and real-estate packs ([plugins.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/constants/plugins.ts#L487-L534)). Other generation goes remote when you are logged in to Cloud or have no `OPENAI_API_KEY`. The default endpoint is `api.promptfoo.app` ([remoteGeneration.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/remoteGeneration.ts#L30-L43), [L164-L189](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/remoteGeneration.ts#L164-L189)). `PROMPTFOO_DISABLE_REMOTE_GENERATION` turns all of this off ([L73-L91](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/remoteGeneration.ts#L73-L91)), which also disables the remote-only plugins. Generation also checks a 100,000-probe monthly limit, counted from the local database, unless you are logged in to Cloud ([redteamProbeLimit.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/util/redteamProbeLimit.ts#L59-L78), [generate.ts](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/commands/generate.ts#L285-L302)).

## Extending it

- **Providers.** Use `file://provider.js` or `.py` custom providers, the generic `http` and `websocket` providers, or `exec:` scripts. Any `id:` string is resolved through `providerRegistry.ts`.
- **Assertions.** Use `javascript`, `python` and `ruby` assertions, either inline or from a file. They return a boolean, a score or a full `GradingResult`. `assertScoringFunction` replaces the default weighted aggregation.
- **Tests.** Tests can come from CSV, JSON, YAML, XLSX, Google Sheets, HuggingFace datasets or a JS/Python generator.
- **Hooks.** `extensions` run code before and after all tests or each test.
- **Red team.** Custom plugins load from a file (`plugins/custom.ts`). Custom strategies are `file://` JS modules.

## Running it

Run `npx promptfoo@latest init`, then `promptfoo eval` and `promptfoo view` (the UI defaults to port 15500). Node.js 22.22 or newer is required. All state lives under `~/.promptfoo` (SQLite database and cache). Results can be shared to promptfoo's hosted service, but this is optional. A Dockerfile and Helm chart serve the web UI for teams. For CI, run `promptfoo eval` with a non-zero exit on failures. Telemetry can be turned off with `PROMPTFOO_DISABLE_TELEMETRY`.

## Strengths and caveats

- **Strength: breadth of checks.** About 70 assertion types cover deterministic checks, reference metrics, LLM judges and trace/trajectory checks. Weights, thresholds and named metrics make mixed scorecards easy.
- **Strength: zero infrastructure.** A single Node process plus SQLite, with response caching. It is easy to put in CI and cheap to re-run.
- **Strength: the most complete red-team generator in this group.** Plugins, wrapping strategies, multi-turn attackers and per-plugin graders feed the same eval and report pipeline.
- **Caveat: red team is partly a client of a hosted API.** Many plugins only exist server-side, generation goes remote by default without an OpenAI key, and there is a free-tier probe cap. Air-gapped use means disabling remote generation and losing those plugins.
- **Caveat: offline by design.** It has no production trace ingestion, sampling or online scoring. Traces exist only for runs it starts itself.
- **Caveat: one very large file.** `src/evaluator.ts` is over 5,600 lines and holds scheduling, comparison, persistence and timeouts, which makes changes to the core risky.
- **Caveat: single-machine storage.** SQLite works for one user. Team history needs the hosted service or the self-hosted server image.

*Sources: code at 1df3eb8, deepwiki-open wiki (12 pages), verified Q&A.*

## How promptfoo/promptfoo answers the LLM evals and testing questions

### Which evaluation metrics and scorers are provided, and how are they implemented? (answered)

Built-in metrics registered via ASSERTION_HANDLERS map. Heuristic: equals, contains/contains-all/contains-any, regex, starts-with, is-json, is-sql, is-html, is-xml, is-refusal, finish-reason, cost. Statistical/distance: levenshtein, bleu, gleu, rouge-n, meteor, perplexity-score, similar (cosine/dot/euclidean), latency, word-count, tool-call-f1. Model-based: llm-rubric, model-graded-closedqa, factuality, g-eval, agent-rubric, search-rubric (src/matchers/llmGrading.ts). Trace-aware: trace-error-spans, trace-span-count, trace-span-duration, five trajectory types load spans from SQLite via loadTraceData. Custom: javascript/python/ruby file:// references return GradingResult. weight=0 forces metric-only pass. Custom per-test assertScoringFunction overrides default aggregation.


Citations: [src/assertions/index.ts:229-320](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L229-L320) · [src/assertions/index.ts:125-137](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L125-L137) · [src/assertions/index.ts:670-678](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L670-L678) · [src/evaluator.ts:2632-2642](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L2632-L2642)

### How is LLM-as-a-judge implemented? (answered)

Four LLM-as-a-judge matchers in src/matchers/llmGrading.ts. matchesLlmRubric (L181) uses a rubric prompt with runJsonGradingPrompt to parse JSON pass/score/reason. matchesFactuality (L323) scores against expert answer using categories A-E with configurable scoring. matchesClosedQa (L382) expects trailing Y/N. matchesGEval (L454) implements G-Eval: generates evaluation steps via LLM then scores 1-10. All use extractFirstJsonObject for structured parsing. Judge provider set via test.options.provider. ProviderGroupedCallQueue batches identical judge calls (evaluator.ts L538-550). Remote grading via doRemoteGrading when no local provider set. Inverse assertions flip pass/score via invertScore. g-eval supports multi-criteria arrays with averaged scoring (src/assertions/geval.ts L26-86).

> **Editor's note.** Correction: `ProviderGroupedCallQueue` does not batch or deduplicate identical judge calls; it groups queued grading calls by provider id. Remote grading for `llm-rubric` only applies inside a red-team config with no explicit grader or rubric prompt, not whenever no local provider is set.

Citations: [src/matchers/llmGrading.ts:181-267](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L181-L267) · [src/matchers/llmGrading.ts:454-612](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L454-L612) · [src/matchers/llmGrading.ts:382-442](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L382-L442) · [src/evaluator.ts:538-550](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L538-L550) · [src/assertions/geval.ts:9-116](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/geval.ts#L9-L116)

### How are test datasets and cases defined, generated and versioned? (answered)

Test cases defined in tests[] of promptfooconfig.yaml. readStandaloneTestsFile (src/util/testCaseReader.ts L113) supports HuggingFace (huggingface://datasets/), Google Sheets (CSV), Azure Blob (az://), CSV, JSON/YAML, XLSX, SharePoint, and JavaScript/Python files. HuggingFace datasets fetched with pagination and concurrent page requests. generateVarCombinations (evaluator.ts L1981-2028) expands array-vars into Cartesian test rows. Scenarios (evaluator.ts L2441-2498) merge scenario config with per-test overrides. defaultTest provides golden-set base assertions. No built-in dataset versioning beyond git; evalsTable stores each runs config snapshot for reproducibility.


Citations: [src/util/testCaseReader.ts:92-119](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/util/testCaseReader.ts#L92-L119) · [src/integrations/huggingfaceDatasets.ts:1-50](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/integrations/huggingfaceDatasets.ts#L1-L50) · [src/evaluator.ts:1981-2028](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1981-L2028) · [src/evaluator.ts:2441-2498](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L2441-L2498) · [src/database/tables.ts:58-77](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/database/tables.ts#L58-L77)

### How are evals executed and reported? (answered)

evaluate() (evaluator.ts L5576-5627) creates Evaluator instance that builds RunEvalOptions from provider-prompt-test-vars cross product. runEval (L1618) renders prompt via Nunjucks, calls provider, applies transforms, runs assertions. Parallelism via async.forEachOfLimit with configurable max-concurrency (default 4, max 20). Caching (src/cache.ts) uses two-level memory+disk store (KeyvFile) with 14-day TTL, namespace isolation via withCacheNamespace for repeat runs. CI: CIProgressReporter and cli-progress SingleBar. SQLite persistence via Drizzle: evalsTable (id, author, config snapshot, results JSON, prompts, isRedteam) and evalResultsTable. Also writes JSON/JSONL output files. Comparison assertions: select-best and max-score compare outputs across providers. Web dashboard (src/app/) displays history and side-by-side results.

> **Editor's note.** Correction: the default max concurrency is 4, but there is no maximum of 20; concurrency is forced to 1 when `_conversation`, `storeOutputAs` or a persistent browser session is used.

Citations: [src/evaluator.ts:5576-5627](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L5576-L5627) · [src/evaluator.ts:1618-1684](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1618-L1684) · [src/cache.ts:1-80](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/cache.ts#L1-L80) · [src/database/tables.ts:58-98](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/database/tables.ts#L58-L98) · [src/evaluator.ts:880-892](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L880-L892)

### How are traces or production data captured and linked to evaluations? (answered)

OpenTelemetry integration via src/tracing/otelSdk.ts and src/tracing/otelConfig.ts with OTLP HTTP export. Provider calls wrapped with withTracedProviderCall/withTestCaseSpan/withGraderSpan (src/tracing/targetTracer.ts). W3C traceparent headers generated and validated (evaluator.ts L1567-1583). Trace store (src/tracing/store.ts) manages SQLite-backed tracesTable and spansTable with retry-based stability detection (loadTraceData polls until span count stabilizes, src/assertions/index.ts L184-227). External trace providers: Braintrust, Langfuse, Tempo via TraceProvider interface (src/tracing/providers/index.ts L21-39). Traces linked to eval rows via traceId/evaluationId. No production traffic capture SDK exists; traces are generated only during eval execution.


Citations: [src/tracing/store.ts:1-100](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/tracing/store.ts#L1-L100) · [src/tracing/providers/index.ts:21-39](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/tracing/providers/index.ts#L21-L39) · [src/assertions/index.ts:183-227](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/index.ts#L183-L227) · [src/evaluator.ts:1558-1598](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L1558-L1598)

### Does it support red-teaming or safety testing, and how? (answered)

Full red-teaming pipeline in src/redteam/. Plugins (80+): FOUNDATION_PLUGINS includes harmful (30+ categories: child-exploitation, chemical-biological-weapons, hate, malware, weapons), bias, PII, medical, pharmacy, insurance, financial, telecom, teen-safety, coding-agent (constants/plugins.ts L42-87). Safety plugins: aegis, beavertails, harmbench, donotanswer, cyberseceval, xstest. Strategies (30+ in strategies/index.ts): jailbreak-templates (40+ prompt-injection templates in strategies/promptInjections/data.ts), crescendo, base64, hex, leetspeak, gcg, goat, citation, simba, hydra, math-prompt, indirect-web-pwn, composite jailbreaks. Grading: 60+ per-plugin graders in graders.ts. Risk scoring (riskScoring.ts) computes exploitability-impact-strategyWeight scores with severity levels. Operation: redteam run two-phase process -- synthesize (index.ts L963) generates adversarial probes then evaluates against target.


Citations: [src/redteam/constants/plugins.ts:42-87](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/constants/plugins.ts#L42-L87) · [src/redteam/strategies/promptInjections/data.ts:1-10](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/strategies/promptInjections/data.ts#L1-L10) · [src/redteam/graders.ts:1-100](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/graders.ts#L1-L100) · [src/redteam/riskScoring.ts:1-60](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/riskScoring.ts#L1-L60) · [src/redteam/index.ts:963-1000](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/index.ts#L963-L1000) · [src/redteam/strategies/index.ts:42-240](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/redteam/strategies/index.ts#L42-L240)
