promptfoo/promptfoo
Local-first CLI and Node library that runs prompt x provider x test matrices, grades them with ~70 assertion types, and red-teams targets.
Overview
promptfoo is a command-line tool and Node library for testing LLM prompts, providers and agents. You write a promptfooconfig.yaml with prompts, providers and tests. Each test has variables and a list of assertions. promptfoo eval expands this into a matrix (every prompt × every provider × every test, optionally repeated), calls each provider, grades every output, and stores the results in a local SQLite database. promptfoo view serves a React UI over that database.
The same engine runs red-team scans. promptfoo redteam run first generates adversarial test cases with plugins (what to attack) and strategies (how to wrap the attack), writes them to redteam.yaml, and then runs an ordinary eval over that file. Red-teaming is a test generator in front of the eval engine, not a separate system.
The whole repository is MIT-licensed. There is no ee/ folder or enterprise-licensed code. The commercial boundary is a hosted service instead. Many red-team plugins, some grading and part of attack generation call promptfoo’s API, and a free monthly probe limit applies when you are not logged in to promptfoo Cloud (details below).
Architecture
flowchart LR
CFG["promptfooconfig.yaml"] --> CLI["CLI: doEval"]
RT["redteam generate"] --> YAML["redteam.yaml"]
YAML --> CLI
CLI --> EV["Evaluator (evaluator.ts)"]
EV --> RUN["runEval: render + call"]
RUN --> PROV["Providers (src/providers)"]
PROV --> CACHE["Cache (memory + disk)"]
RUN --> ASSERT["runAssertions"]
ASSERT --> GRADE["llmGrading matchers"]
ASSERT --> TRACE["Trace store"]
OTLP["OTLP receiver /v1/traces"] --> TRACE
EV --> DB["SQLite via Drizzle"]
DB --> UI["Express server + React UI"]
RT --> REMOTE["Remote generation API"]
| Component | Path | Role |
|---|---|---|
| CLI and commands | src/main.ts, src/commands/, src/node/doEval.ts |
Parse flags, resolve configs, call evaluate |
| Evaluator | src/evaluator.ts |
Builds the test matrix, schedules steps, runs comparisons, persists rows |
| Single step | runEval in src/evaluator.ts |
Renders the prompt with Nunjucks, calls the provider, transforms, grades |
| Assertions | src/assertions/ |
ASSERTION_HANDLERS registry and runAssertions aggregation |
| Model graders | src/matchers/llmGrading.ts |
llm-rubric, factuality, closed-QA, G-Eval and others |
| Providers | src/providers/ |
Over 100 provider modules, including generic HTTP, WebSocket, script and MCP targets |
| Cache | src/cache.ts |
cache-manager over Keyv: memory plus a KeyvFile disk store |
| Storage | src/database/tables.ts |
Drizzle schema for evals, results, prompts, datasets, traces, spans |
| Tracing | src/tracing/ |
OTel SDK setup, local OTLP receiver, trace store, external trace providers |
| Red team | src/redteam/ |
Plugins, strategies, graders, risk scoring, remote generation |
| Web UI | src/server/, src/app/ |
Express API and React app for results and red-team setup |
How a request flows
For promptfoo eval:
- Resolve.
doEvalloads and merges configs and builds aTestSuite. The libraryevaluatesets base path, env and config context, picks the Node runtime and constructs anEvaluator(evaluator.ts). - Expand. The evaluator merges
defaultTestand scenarios into each test, expands array-valued vars into combinations (generateVarCombinations), and creates oneRunEvalOptionsper prompt × provider × test × repeat. - Schedule. Concurrency defaults to 4 (
-joverrides it). It drops to 1 when a prompt uses the_conversationvariable, a test usesstoreOutputAs, or a browser provider keeps a persistent session (evaluator.ts). Concurrent steps run throughasync.forEachOfLimitand periodically flush prompt metrics to the store (L4394-L4425). - Run a step.
runEvalwraps the step in a cache namespace keyed by repeat index.runEvalInternalthen fills runtime vars, renders the prompt, optionally creates an agent workspace, and generates a W3C trace context when tracing is on (L1618-L1700). It then calls the provider, applies transforms and grades the response. - Grade.
runAssertionsexpandsassert-setgroups and runs assertions with bounded concurrency (default 3). Each result is added to anAssertionsResultwith its weight and metric name, and the final pass/score applies the testthresholdor a custom scoring function (assertions/index.ts, L860-L911). - Compare. After all rows finish,
select-bestandmax-scoreassertions grade outputs for the same test across providers. Rows with model-graded assertions can be deferred and graded in groups per grading provider. - Persist. Rows go to
eval_resultsand the run summary toevalsin SQLite.evals.configkeeps a snapshot of the resolved config (tables.ts). Optional JSON, CSV or HTML output files are written too.
Key components
Assertions
ASSERTION_HANDLERS maps about 70 base types to handlers (assertions/index.ts). They fall into four groups:
- deterministic checks:
equals,contains*,regex,is-json,is-sql,cost,latency,word-count; - reference metrics:
bleu,rouge-n,levenshtein,similar(embeddings),tool-call-f1; - model-graded checks:
llm-rubric,g-eval,factuality,answer-relevance,context-*,moderation; - trace and trajectory checks over recorded spans:
trace-span-count,trace-error-spans,trajectory:tool-sequence,trajectory:goal-success.
Any type can be negated with a not- prefix. javascript, python, ruby and webhook cover custom logic.
LLM grading
matchesLlmRubric renders a grading prompt, calls the grading provider and parses a JSON {pass, score, reason} (llmGrading.ts). The grader comes from options.provider (or --grader), otherwise a default grading provider. Inside a red-team config with no explicit grader, it prefers promptfoo’s remote grading endpoint. Grading calls are not deduplicated. ProviderGroupedCallQueue only groups queued calls by provider id so one grader’s calls run together (providerCallQueue.ts).
Cache
Provider responses go to a cache-manager cache with a memory store and a KeyvFile disk store at ~/.promptfoo/cache/cache.json, with a 14-day default TTL that PROMPTFOO_CACHE_TTL can change (cache.ts, L80-L105). Repeats get separate namespaces, so --repeat 3 makes three real calls.
Tracing
With tracing enabled, promptfoo starts a local Express OTLP receiver. It accepts /v1/traces (JSON or protobuf) and /v1/logs and writes spans into SQLite traces/spans tables (otlpReceiver.ts). The target app receives a traceparent and exports its spans there. Trace assertions then poll until the span count settles. Langfuse, Braintrust and Tempo can also serve as trace sources. Tracing exists to grade agent runs, not to monitor production.
Red team
synthesize in src/redteam/index.ts drives plugins (dozens, under plugins/) and strategies (about 30, under strategies/) to produce test cases (index.ts). doRedteamRun calls doGenerateRedteam, then doEval on the generated file (shared.ts). Results are scored by per-plugin graders and summarised with severity-weighted risk scores.
How much runs locally is the main thing to know. REMOTE_ONLY_PLUGIN_IDS lists plugins that have no local implementation and always call the remote API: BOLA/BFLA, SSRF, indirect prompt injection, RAG poisoning, the MCP plugin, the coding-agent sets and all the medical, financial, pharmacy, insurance, e-commerce, telecom and real-estate packs (plugins.ts). Other generation goes remote when you are logged in to Cloud or have no OPENAI_API_KEY. The default endpoint is api.promptfoo.app (remoteGeneration.ts, L164-L189). PROMPTFOO_DISABLE_REMOTE_GENERATION turns all of this off (L73-L91), which also disables the remote-only plugins. Generation also checks a 100,000-probe monthly limit, counted from the local database, unless you are logged in to Cloud (redteamProbeLimit.ts, generate.ts).
Extending it
- Providers. Use
file://provider.jsor.pycustom providers, the generichttpandwebsocketproviders, orexec:scripts. Anyid:string is resolved throughproviderRegistry.ts. - Assertions. Use
javascript,pythonandrubyassertions, either inline or from a file. They return a boolean, a score or a fullGradingResult.assertScoringFunctionreplaces the default weighted aggregation. - Tests. Tests can come from CSV, JSON, YAML, XLSX, Google Sheets, HuggingFace datasets or a JS/Python generator.
- Hooks.
extensionsrun code before and after all tests or each test. - Red team. Custom plugins load from a file (
plugins/custom.ts). Custom strategies arefile://JS modules.
Running it
Run npx promptfoo@latest init, then promptfoo eval and promptfoo view (the UI defaults to port 15500). Node.js 22.22 or newer is required. All state lives under ~/.promptfoo (SQLite database and cache). Results can be shared to promptfoo’s hosted service, but this is optional. A Dockerfile and Helm chart serve the web UI for teams. For CI, run promptfoo eval with a non-zero exit on failures. Telemetry can be turned off with PROMPTFOO_DISABLE_TELEMETRY.
Strengths and caveats
- Strength: breadth of checks. About 70 assertion types cover deterministic checks, reference metrics, LLM judges and trace/trajectory checks. Weights, thresholds and named metrics make mixed scorecards easy.
- Strength: zero infrastructure. A single Node process plus SQLite, with response caching. It is easy to put in CI and cheap to re-run.
- Strength: the most complete red-team generator in this group. Plugins, wrapping strategies, multi-turn attackers and per-plugin graders feed the same eval and report pipeline.
- Caveat: red team is partly a client of a hosted API. Many plugins only exist server-side, generation goes remote by default without an OpenAI key, and there is a free-tier probe cap. Air-gapped use means disabling remote generation and losing those plugins.
- Caveat: offline by design. It has no production trace ingestion, sampling or online scoring. Traces exist only for runs it starts itself.
- Caveat: one very large file.
src/evaluator.tsis over 5,600 lines and holds scheduling, comparison, persistence and timeouts, which makes changes to the core risky. - Caveat: single-machine storage. SQLite works for one user. Team history needs the hosted service or the self-hosted server image.
Sources: code at 1df3eb8, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the LLM evals and testing questions
Each answer was drafted by a code-reading agent at commit 1df3eb8. Its citations were checked mechanically. Compare with the other llm evals and testing →
Which evaluation metrics and scorers are provided, and how are they implemented?
answeredBuilt-in metrics registered via ASSERTION_HANDLERS map. Heuristic: equals, contains/contains-all/contains-any, regex, starts-with, is-json, is-sql, is-html, is-xml, is-refusal, finish-reason, cost. Statistical/distance: levenshtein, bleu, gleu, rouge-n, meteor, perplexity-score, similar (cosine/dot/euclidean), latency, word-count, tool-call-f1. Model-based: llm-rubric, model-graded-closedqa, factuality, g-eval, agent-rubric, search-rubric (src/matchers/llmGrading.ts). Trace-aware: trace-error-spans, trace-span-count, trace-span-duration, five trajectory types load spans from SQLite via loadTraceData. Custom: javascript/python/ruby file:// references return GradingResult. weight=0 forces metric-only pass. Custom per-test assertScoringFunction overrides default aggregation.
How is LLM-as-a-judge implemented?
answeredFour LLM-as-a-judge matchers in src/matchers/llmGrading.ts. matchesLlmRubric (L181) uses a rubric prompt with runJsonGradingPrompt to parse JSON pass/score/reason. matchesFactuality (L323) scores against expert answer using categories A-E with configurable scoring. matchesClosedQa (L382) expects trailing Y/N. matchesGEval (L454) implements G-Eval: generates evaluation steps via LLM then scores 1-10. All use extractFirstJsonObject for structured parsing. Judge provider set via test.options.provider. ProviderGroupedCallQueue batches identical judge calls (evaluator.ts L538-550). Remote grading via doRemoteGrading when no local provider set. Inverse assertions flip pass/score via invertScore. g-eval supports multi-criteria arrays with averaged scoring (src/assertions/geval.ts L26-86).
ProviderGroupedCallQueue does not batch or deduplicate identical judge calls; it groups queued grading calls by provider id. Remote grading for llm-rubric only applies inside a red-team config with no explicit grader or rubric prompt, not whenever no local provider is set.How are test datasets and cases defined, generated and versioned?
answeredTest cases defined in tests[] of promptfooconfig.yaml. readStandaloneTestsFile (src/util/testCaseReader.ts L113) supports HuggingFace (huggingface://datasets/), Google Sheets (CSV), Azure Blob (az://), CSV, JSON/YAML, XLSX, SharePoint, and JavaScript/Python files. HuggingFace datasets fetched with pagination and concurrent page requests. generateVarCombinations (evaluator.ts L1981-2028) expands array-vars into Cartesian test rows. Scenarios (evaluator.ts L2441-2498) merge scenario config with per-test overrides. defaultTest provides golden-set base assertions. No built-in dataset versioning beyond git; evalsTable stores each runs config snapshot for reproducibility.
How are evals executed and reported?
answeredevaluate() (evaluator.ts L5576-5627) creates Evaluator instance that builds RunEvalOptions from provider-prompt-test-vars cross product. runEval (L1618) renders prompt via Nunjucks, calls provider, applies transforms, runs assertions. Parallelism via async.forEachOfLimit with configurable max-concurrency (default 4, max 20). Caching (src/cache.ts) uses two-level memory+disk store (KeyvFile) with 14-day TTL, namespace isolation via withCacheNamespace for repeat runs. CI: CIProgressReporter and cli-progress SingleBar. SQLite persistence via Drizzle: evalsTable (id, author, config snapshot, results JSON, prompts, isRedteam) and evalResultsTable. Also writes JSON/JSONL output files. Comparison assertions: select-best and max-score compare outputs across providers. Web dashboard (src/app/) displays history and side-by-side results.
_conversation, storeOutputAs or a persistent browser session is used.How are traces or production data captured and linked to evaluations?
answeredOpenTelemetry integration via src/tracing/otelSdk.ts and src/tracing/otelConfig.ts with OTLP HTTP export. Provider calls wrapped with withTracedProviderCall/withTestCaseSpan/withGraderSpan (src/tracing/targetTracer.ts). W3C traceparent headers generated and validated (evaluator.ts L1567-1583). Trace store (src/tracing/store.ts) manages SQLite-backed tracesTable and spansTable with retry-based stability detection (loadTraceData polls until span count stabilizes, src/assertions/index.ts L184-227). External trace providers: Braintrust, Langfuse, Tempo via TraceProvider interface (src/tracing/providers/index.ts L21-39). Traces linked to eval rows via traceId/evaluationId. No production traffic capture SDK exists; traces are generated only during eval execution.
Does it support red-teaming or safety testing, and how?
answeredFull red-teaming pipeline in src/redteam/. Plugins (80+): FOUNDATION_PLUGINS includes harmful (30+ categories: child-exploitation, chemical-biological-weapons, hate, malware, weapons), bias, PII, medical, pharmacy, insurance, financial, telecom, teen-safety, coding-agent (constants/plugins.ts L42-87). Safety plugins: aegis, beavertails, harmbench, donotanswer, cyberseceval, xstest. Strategies (30+ in strategies/index.ts): jailbreak-templates (40+ prompt-injection templates in strategies/promptInjections/data.ts), crescendo, base64, hex, leetspeak, gcg, goat, citation, simba, hydra, math-prompt, indirect-web-pwn, composite jailbreaks. Grading: 60+ per-plugin graders in graders.ts. Risk scoring (riskScoring.ts) computes exploitability-impact-strategyWeight scores with severity levels. Operation: redteam run two-phase process -- synthesize (index.ts L963) generates adversarial probes then evaluates against target.