firecrawl/firecrawl
Self-hostable scraping API that races fetch engines in a waterfall, converts pages to Markdown and runs LLM extraction.
Overview
Firecrawl is a web data API: you send it a URL and get back clean Markdown, HTML, links, screenshots or LLM-extracted JSON. You can also crawl a whole site, map its URLs, search the web, or run an “agent” job. The repository is a monorepo. The core is apps/api, a TypeScript/Express service with its own queue workers. Around it are SDKs for Python, JavaScript, Go, Rust, .NET, Ruby, PHP, Java and Elixir, a Playwright microservice, a Go HTML-to-Markdown service and a Rust native module.
The code is written for the hosted service first. Much of apps/api handles billing, credits, keyless access, zero-data-retention, threat protection and “safe mode”. The best scraping path, Fire Engine (a Chrome-over-CDP and TLS-client service with stealth proxies), is not in this repository. The API only calls it over HTTP when FIRE_ENGINE_BETA_URL is set (available.ts). A self-hosted Firecrawl therefore runs the bundled Playwright service plus a plain fetch fallback, and it loses actions, screenshots and stealth.
The interesting engineering is the scrape pipeline. Engines are ranked by the features each request needs, then raced in a staggered waterfall. A long chain of transformers then turns the winning result into whatever formats were asked for. The AGPL-3.0 licence covers the server. The SDKs are MIT.
Architecture
flowchart LR
C["SDK / HTTP client"] --> R["Express routers v1 / v2"]
R --> MW["auth, credits, blocklist middleware"]
MW --> SC["scrapeController (inline)"]
MW --> CC["crawlController"]
CC --> Q["NuQ queue (Postgres / FDB)"]
Q --> W["nuq-worker: processJob"]
SC --> P["scrapeURL pipeline"]
W --> P
P --> FL["buildFallbackList"]
FL --> E["engines: fire-engine, playwright, fetch, pdf, document..."]
P --> T["executeTransformers"]
T --> MD["Rust transformHtml + Go/turndown Markdown"]
T --> LLM["LLM extract (AI SDK)"]
W --> RD["Redis: crawl state, visited set"]
| Component | Path | Role |
|---|---|---|
| HTTP server | apps/api/src/index.ts, routes/v1.ts, routes/v2.ts |
Express app; mounts /v1, /v2 and middleware chains per endpoint |
| Controllers | apps/api/src/controllers/v2/ |
scrape, crawl, batch-scrape, map, search, extract, agent, parse and status endpoints |
| Queue (NuQ) | apps/api/src/services/worker/nuq.ts, apps/nuq-postgres/nuq.sql |
Postgres job table with FOR UPDATE SKIP LOCKED; RabbitMQ or LISTEN for completion; optional FoundationDB backend |
| Scrape worker | apps/api/src/services/worker/scrape-worker.ts |
processJob: runs a scrape, then for crawls discovers and enqueues links |
| scrapeURL | apps/api/src/scraper/scrapeURL/index.ts |
Feature flags, engine waterfall, retries on AddFeatureError |
| Engines | apps/api/src/scraper/scrapeURL/engines/ |
fire-engine client, playwright, fetch, pdf, document, image, index, wikipedia, x-twitter |
| Transformers | apps/api/src/scraper/scrapeURL/transformers/ |
HTML cleanup, Markdown, links, images, LLM extract, summary, diff |
| Crawler | apps/api/src/scraper/WebScraper/crawler.ts, lib/crawl-redis.ts |
Link filtering (Rust), robots.txt, sitemaps, Redis visited/lock sets |
| Native module | apps/api/native/ |
Rust (napi-rs): HTML transform, link filtering, Markdown post-processing |
| Playwright service | apps/playwright-service-ts/api.ts |
Self-host browser renderer behind a /scrape endpoint |
| SDKs | apps/*-sdk/ |
Thin REST clients for nine languages |
How a request flows
Take POST /v2/scrape with formats: ["markdown"]:
- Route and gate. The
/scraperoute runs auth (with keyless access allowed), a country check, a credit check and the scrape blocklist beforescrapeController(v2.ts). - Validate and lock. The controller parses the body with
scrapeRequestSchemaand resolves safe mode (scrape.ts). It then takes a per-team slot inteamConcurrencySemaphore, which gets two thirds of the request timeout to acquire (scrape.ts). - Run inline. A single scrape does not go through the queue. The controller builds a
NuQJobwithskipNuq: trueand callsprocessJobInternaldirectly, in the API process (scrape.ts). - Worker wrapper.
processJobsets the abort timer, checks for a cancelled crawl, and racesstartWebScraperPipelineagainst the remaining time (scrape-worker.ts).runWebScrapercallsscrapeURLonce, or up to three times for crawl pages (runWebScraper.ts). - Choose engines.
buildFallbackListfirst checks the Exchange data catalogue and the blocklist, then scores every configured engine. Each request flag has a priority. An engine qualifies when it covers at least half the summed priority. The list is then sorted by support score and a fixedqualitynumber (engines/index.ts, L955-L1070). - Waterfall.
scrapeURLLoopstarts the first engine. If it has not finished after its “max reasonable time” plus a configured delay, the next engine starts while the first keeps running. The first engine to return an acceptable result wins and the others are aborted (index.ts, L1100-L1125). - Judge the result.
scrapeURLLoopIterconverts the HTML to Markdown for a quality check. A non-empty body or a non-2xx status counts as success. A 401, 403 or 429 underproxy: "auto"throwsAddFeatureError(["stealthProxy"])(index.ts). The outer loop inscrapeURLcatches that, adds the flag and rebuilds the waterfall (index.ts). - Transform.
executeTransformersruns a fixed stack in order: HTML cleanup, Markdown, links, images, metadata, then LLM extract, summary, query, agent, diff and format coercion (transformers/index.ts). The document goes back in the HTTP response.
Key components
Engine waterfall
Engines are declared in one table with feature support and a quality score. Index and cache engines rank highest, then Fire Engine Chrome CDP (50), TLS client (10), Playwright (20) and fetch (5). Stealth variants have negative quality, so they run only when stealth is required (engines/index.ts, L405-L450). The list depends on configuration. Without Fire Engine and without Wikipedia or X credentials, it is playwright, fetch, pdf, document, and image when FirePDF is configured. The self-host Playwright engine declares actions: false and screenshot: false, so those features fail with ActionsNotSupportedError or come back with an “unsupported features” warning (index.ts).
Files take a detour. With Fire Engine available, a PDF or DOCX URL goes through a browser engine first. The file comes back via AddFeatureError with a prefetch, and the pdf or document engine then parses those bytes (engines/index.ts).
HTML to Markdown
htmlTransform calls transformHtml from the Rust module with includeTags, excludeTags and onlyMainContent, and falls back to cheerio on error (removeUnwantedElements.ts). parseMarkdown tries three converters in order: the Go HTTP service, the Go shared library over koffi FFI, then Turndown with GFM. Every path ends in Rust postProcessMarkdown (html-to-markdown.ts). If main-content extraction gives empty Markdown, deriveMarkdownFromHTML reruns it on the full page (transformers/index.ts).
LLM extraction
generateCompletions defaults to gpt-4o-mini with gpt-4.1-mini as the retry model. It does not chunk long pages. It trims the Markdown to 80% of the model’s input limit and adds a warning. The prompt tells the model to ignore instructions found in the content (llmExtract.ts). Providers come from the Vercel AI SDK: OpenAI, Ollama (the default when OLLAMA_BASE_URL is set), Anthropic, Groq, Google, OpenRouter, Fireworks, DeepInfra and Vertex (generic-ai.ts).
NuQ queue
Crawls, batch scrapes and async jobs go through NuQ, a custom queue in Postgres. Workers claim jobs with FOR UPDATE SKIP LOCKED, ordered by priority and age. Waiters hear about completion through a RabbitMQ reply queue or Postgres LISTEN (nuq.ts). The schema lives in apps/nuq-postgres/nuq.sql (nuq.sql). NUQ_BACKEND=fdb switches to FoundationDB (queue-jobs.ts). BullMQ on Redis is still used, but only for side queues such as billing, deep research, llms.txt and precrawl.
Crawler
crawlController fetches robots.txt, creates a NuQ crawl group with maxConcurrency and the user’s delay, saves the crawl in Redis and enqueues one kickoff job (crawl.ts). The code that would copy robots.txt Crawl-delay into the crawl is commented out, so politeness delay is whatever the caller sets. After each page, the worker extracts links and filters them through Rust filterLinks (include/exclude regexes, depth, robots, subdomains, backward crawling) (crawler.ts). It runs threat-protection checks and enqueues each link that wins lockURL (scrape-worker.ts). lockURL enforces the page limit and dedupes on a normalized URL, or on a permutation key when deduplicateSimilarURLs is on (crawl-redis.ts).
Extending it
- A new engine. Add a name to the
Engineunion, a handler, a max-reasonable-time function and a feature/quality entry inengines/index.ts. The waterfall picks it up from the scores. - A new output format. Add a transformer function and place it in
transformerStack. Order matters, because later transformers read fields set by earlier ones. - Models. Point
OPENAI_BASE_URLat any OpenAI-compatible endpoint, or setOLLAMA_BASE_URL. Other providers read their standard API-key variables. - Clients. The nine SDKs and the separate CLI and MCP projects all wrap the same REST API, so a new endpoint means a controller plus a route first.
Running it
- Docker Compose. The root
docker-compose.yamlruns the API (whose harness also starts the workers), the Playwright service, Redis, RabbitMQ and NuQ Postgres. Only port 3002 is published, and the default isUSE_DB_AUTHENTICATION=false, so the API is unauthenticated (docker-compose.yaml, harness.ts). - Optional services.
PROXY_SERVERfor the fetch and Playwright engines,SEARXNG_ENDPOINTfor search, an LLM key for JSON/summary formats, and Fire Engine or FirePDF URLs if you have access to them. - Playwright service. It runs headless Chromium with a random user agent, routes traffic through an SSRF-checking local proxy, and can block media. It has no stealth patches (api.ts).
Strengths and caveats
- Strength: a well-engineered fetch strategy. The hedged waterfall gives a fast engine a head start without letting a slow one block the request. Flag-based scoring keeps engines that cannot do what was asked out of the list.
- Strength: wide input coverage. HTML, PDFs, Office files and images go through one pipeline, with a single document model and many output formats.
- Strength: built for scale. The Postgres queue with
SKIP LOCKED, the Redis visited sets and the per-team semaphores are production designs, not demo code. - Caveat: self-host is a different product. Fire Engine (stealth, actions, screenshots, mobile, location) is a closed external service. Self-hosted, you get plain Playwright and
fetch. - Caveat: SaaS code everywhere. Billing, credit reservation, Exchange, safe mode and ZDR branches make the core paths long and harder to follow or fork.
- Caveat: robots.txt is only partly honoured. Disallow rules are enforced (unless
ignoreRobotsTxt), butCrawl-delayis ignored. - Caveat: LLM extraction truncates. Long pages are trimmed to fit the context window, not chunked and merged, so data near the end of a big page can be lost.
Sources: code at 8c84d8b, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (41 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 8c84d8b. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPlain HTTP vs Headless Browser: Pages are fetched through a waterfall of engines defined in scrapeURL/engines/index.ts:82-104 — the list includes fetch (plain HTTP via undici), playwright (a Playwright microservice), fire-engine;chrome-cdp (proprietary Chrome CDP service), and fire-engine;tlsclient (TLS-client engine). The fetch engine (engines/fetch/index.ts:86-229) uses undici.fetch with redirect-following and charset detection from HTTP headers and HTML <meta> tags. Headless browser rendering goes through either the Playwright microservice (engines/playwright/index.ts:8-53) which POSTs the URL plus wait/timeout/headers, or the Fire Engine's Chrome CDP (engines/fire-engine/index.ts:346-627) which sends a rich request with actions, wait, screenshot, mobile, geolocation, and proxy parameters, then polls for completion (performFireEngineScrape, lines 70-308).
JS Rendering: JS rendering is always-on for the Chrome CDP engine since it runs actual Chrome. The TLS Client engine (tlsclient) is a lighter non-JS engine for speed. The fetch engine does zero JS rendering.
Waiting Strategy: The waitFor option controls how long the browser waits before capturing; default is 0ms, but 2000ms for branding scrapes (fire-engine/index.ts:63). It's transformed into a wait action (lines 389-395). The overall scrape timeout defaults to 300s (scrapeURL/index.ts:966) and each engine has its own max-reasonable-time, 30s for Playwright (engines/playwright/index.ts:65-67), 15s for fetch (engines/fetch/index.ts:231-233).
Content Types Supported: The system handles HTML/web pages (all engines), PDFs (engines/pdf), Office documents (DOCX, XLSX, PPTX, ODT, CSV, EPUB, RTF via lib/document-formats.ts:1-32), images (PNG, JPEG, TIFF, GIF, WebP, AVIF via lib/image-formats.ts:1-34), plain text and JSON. PDFs can be processed by FirePDF for OCR. Binary files can be returned as rawBase64.
How is content extracted or converted?
answeredHTML→Markdown Conversion: The primary conversion is HTML to Markdown via lib/html-to-markdown.ts. It supports two backends: an HTTP-based Go HTML-to-Markdown service (html-to-markdown-client.ts) or a native Go shared library loaded via FFI (GoMarkdownConverter, lines 12-49). Both feed through Rust-based postProcessMarkdown from @mendable/firecrawl-rs (line 76). The converter is invoked by the deriveMarkdownFromHTML transformer (scrapeURL/transformers/index.ts:86-194) which applies htmlTransform for boilerplate removal and readability processing, then calls parseMarkdown.
Boilerplate Removal: The htmlTransform function in lib/removeUnwantedElements.ts strips navigation, ads, and non-content elements. Controlled by the onlyMainContent option (defaults to true). If the resulting markdown is empty, it falls back to full content extraction (line 166-192).
Selectors: Users can pass CSS selectors via includeTags/excludeTags options. The actions system supports click, wait, scroll, write, press, scrape, executeJavascript, and screenshot actions (fire-engine/index.ts:313-321).
Schema-based Extraction: JSON/extract schemas for structured data extraction. The scrapeURL/transformers/index.ts pipeline (lines 618-643) runs: performLLMExtract, performDeterministicJson, performQuery, performSummary, performAgent, and others — each processing the document sequentially. The LLM extract supports JSON schema with recursive-reference detection, auto-repair of malformed output, and fallback models (transformers/llmExtract.ts:49-72 for model selection, lines 774-818 for generateObject with schema).
How are LLMs used, if at all?
answeredPrompting: LLMs are used for structured JSON extraction (performLLMExtract, performQuery, performSummary in transformers/index.ts:632-638), schema-based extraction with generateObject from the Vercel AI SDK (transformers/llmExtract.ts:774-818), text summarization, question answering, highlight generation, and free-form agent extraction (performAgent). The prompt prepends markdown content to system/user prompts with the instruction to ignore data-processing directives (line 489-492).
Chunking of Large Pages: The trimToTokenLimit function (transformers/llmExtract.ts:207-274) uses tiktoken to pre-trim content by characters then precisely slice to the model's max_input_tokens at 80% capacity (line 480-486). It handles both character-based pre-trimming and exact token-aligned trimming with proper UTF-8 boundary handling.
Structured Output / JSON Schema: The system supports jsonSchema from ai, Zod schemas, and raw JSON Schema objects. Recursive schemas (with $ref, $defs) trigger model upgrade to gpt-4.1 (lines 49-72). It includes repair config (experimental_repairText, lines 678-772) that strips markdown code fences and uses LLM-based JSON repair on failure.
Providers: Multiple providers via lib/generic-ai.ts:1-55: OpenAI (default), Anthropic, Ollama, Groq, Google Gemini, OpenRouter, Fireworks, DeepInfra, and Vertex AI (GCP). The default provider switches to Ollama when OLLAMA_BASE_URL is set (line 23). The default model is gpt-4o-mini (line 464), with gpt-4.1-mini as fallback on quota errors (line 449).
Cost Controls: calculateCost tracks per-model pricing from a hardcoded table (modelCosts at lines 278-297) including GPT-4o, GPT-4.1, o3-mini, GPT-5 series, Gemini, Grok, and DeepSeek models. Costs are recorded via the CostTracking system which logs token usage and dollar amounts to billing (lines 527-543, 848-864). Unpriced models are tracked but recorded as $0 with a warning (lines 348-359).
How are anti-bot measures, proxies and fingerprinting handled?
answeredStealth Patches: The system has a multi-layer approach. First, the TLS Client engine (fire-engine;tlsclient) is a custom HTTP client that mimics browser TLS fingerprints — it uses atsv (anti-bot solver) and skipTlsVerification options. The Chrome CDP engine runs real Chrome, so it naturally looks like a browser. The engpicker (lib/engpicker.ts:1-325) is an automated system that scrapes target URLs with multiple engine/proxy combinations, uses GPT-4o-mini to evaluate success vs. antibot pages (lines 70-101), and computes a Rust-based Levenshtein verdict to decide whether tlsclient with or without stealth proxy is adequate for a domain.
Fingerprint Spoofing: Browser profiles (saved authenticated sessions) are supported via lib/browser-profiles.ts:1-15. The mobileProxy flag on Fire Engine requests enables carrier-grade mobile proxies that change the apparent device and network. The forceNonRender option (fire-engine/index.ts:327-344) can route branding scrapes away from the visual render engine to avoid triggering heavy anti-bot challenges.
Proxy Rotation: Three proxy tiers are available: basic (direct, no proxy), stealth (mobile carrier proxy with rotated IP), and auto (escalates from basic to stealth when the engine detects 401/403/429 responses — scrapeURL/index.ts:769-783). Config variables PROXY_SERVER, PROXY_USERNAME, PROXY_PASSWORD (config.ts:408-410) set a forward proxy for the fetch engine. The PARTNER_EGRESS_PROXY_URL (config.ts:197) provides a partner egress proxy for threat-protection scanning.
CAPTCHA Handling: The system detects antibot pages through the scrape result quality check (scrapeURL/index.ts:748-762): short markdown or 401/403/429 status codes trigger AddFeatureError to escalate to stealth proxy engines. The specialty scraper (engines/utils/specialtyHandler.ts:117-130) sniffs browser handoff payloads for bot-blocked content and routes to file-specific parsing.
Rate Limiting: Per-team concurrency limits are enforced via lib/concurrency-limit.ts:1-419 using sorted sets in Redis. Per-minute rate limits are in services/rate-limiter.ts:21-44 covering crawl (15/min), scrape (100/min), search, map, extract, and browser modes with configurable multipliers from the Autumn billing service.
How is crawling at scale implemented?
answeredQueues: Crawls use Redis (BullMQ) via services/queue-service.ts. Jobs are persisted with saveCrawl in lib/crawl-redis.ts:45-79, which stores crawl config with 24-hour TTL. The queue backend supports both PostgreSQL (pg) and FoundationDB (fdb) for crawl job storage (StoredCrawl.queueBackend at line 31). Background extraction jobs use a separate Extract queue (services/extract-queue.ts).
Concurrency: Crawl-scoped concurrency uses sorted sets in Redis (lib/concurrency-limit.ts:174-214). An AB-Test system (services/ab-test.ts) can split or mirror crawl traffic across engines. The getNextConcurrentJob function (line 224-309) pops from a ZSET atomically, checks per-crawl concurrency caps, and re-queues blocked jobs. When a job finishes, concurrentJobDone (line 316-418) promotes backlogged jobs respecting per-team concurrency limits.
URL Dedup: The WebCrawler (crawler.ts:59-165) uses a visited Set for deduplication (line 67). The filterLinks method (line 173-471) applies dedup via both a Rust native function (filterLinks from @mendable/firecrawl-rs, line 206-221) and a JavaScript fallback. URL normalization strips www. prefix and trailing slashes (line 576-581).
Depth/Limits: Configurable maxCrawledDepth (default 10) and a limit cap (default 10000). The getURLDepth utility measures path depth. Links exceeding the depth get descriptive denial reasons (line 345-354). Filtering checks include regex-based exclude/include patterns, backward-crawling restriction, robots.txt, file-type filtering, external-link control, and non-web protocol rejection (lines 310-470).
robots.txt and Politeness: Built-in robots.txt checking via lib/robots-txt.ts:1-80 — fetched on each crawl, cached with 1-day max-age, and parsed with the robots-parser library. The ignoreRobotsTxt option disables it (crawler.ts line 152). Crawl delay from robots.txt is extracted and honored via getRobotsCrawlDelay() (line 550-552) and enforced by concurrentJobDone which pauses between job promotions (line 362-367). Sitemap discovery via tryGetSitemap (line 554-615) fetches sitemap.xml within 120s timeout.
Distributed Workers: Multiple worker processes consume from Redis queues via services/worker/ — the scrapeQueue (NuQ, a custom queue on top of BullMQ) distributes scrapes across worker pods. The concurrency-queue-reconciler.ts handles back-pressure when teams exceed their concurrency limits.
delay option slows a crawl.What is the developer interface?
answeredREST API: The primary interface is a REST API served by Express, with two major versions:
- v1 (
routes/v1.ts): POST endpoints for/scrape,/crawl,/batch/scrape,/search,/map,/extract, plus WebSocket at/crawl/:jobIdfor streaming crawl status. - v2 (
routes/v2.ts): Adds POST/parse,/parse-upload,/agent,/browser,/feedback,/slack,/monitor,/research-proxy,/scrape-browser, and OAuth token introspection. Batch/scrape and crawl via NuQ queues. Both versions use middleware chains for auth, rate-limiting, credit checks, blocklist filtering, idempotency, and country validation.
Language SDKs: Multiple officially maintained SDKs under the monorepo's apps/ directory: python-sdk (FirecrawlApp with sync/async, Watcher, typed errors), js-sdk, go-sdk, rust-sdk, dot-net-sdk, ruby-sdk, php-sdk, java-sdk, elixir-sdk. All cover the core endpoints.
CLI: No standalone CLI is packaged — the SDKs serve as programmatic clients.
MCP Server: An MCP (Model Context Protocol) server is hosted at mcp.firecrawl.dev/v2/mcp with OAuth and API-key auth (services/mcp/action-logs.ts:8-9). Action logs record MCP tool usage with retention, pagination, and team-scoped queries.
UI: An ingestion dashboard UI exists at apps/ui/ingestion-ui/ built with Vue.js/Vite and Tailwind CSS.
Output Formats: The core Document type contains: markdown, rawHtml, rawBase64, html, screenshot, actions, branding, json, extract, summary, answer, highlights, links, images, audio, video, product, menu, changeTracking. All controlled by the formats array parameter.