# firecrawl/firecrawl

> Self-hostable scraping API that races fetch engines in a waterfall, converts pages to Markdown and runs LLM extraction.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/firecrawl/firecrawl (reviewed at commit `8c84d8b6a155495b8088c286d878b2f875de3918`, 2026-10-06)
- Stars: 189121 · Language: TypeScript · License: AGPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/firecrawl/

## Overview

Firecrawl is a web data API: you send it a URL and get back clean Markdown, HTML, links, screenshots or LLM-extracted JSON. You can also crawl a whole site, map its URLs, search the web, or run an "agent" job. The repository is a monorepo. The core is `apps/api`, a TypeScript/Express service with its own queue workers. Around it are SDKs for Python, JavaScript, Go, Rust, .NET, Ruby, PHP, Java and Elixir, a Playwright microservice, a Go HTML-to-Markdown service and a Rust native module.

The code is written for the hosted service first. Much of `apps/api` handles billing, credits, keyless access, zero-data-retention, threat protection and "safe mode". The best scraping path, Fire Engine (a Chrome-over-CDP and TLS-client service with stealth proxies), is not in this repository. The API only calls it over HTTP when `FIRE_ENGINE_BETA_URL` is set ([available.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/fire-engine/available.ts#L1-L5)). A self-hosted Firecrawl therefore runs the bundled Playwright service plus a plain `fetch` fallback, and it loses actions, screenshots and stealth.

The interesting engineering is the scrape pipeline. Engines are ranked by the features each request needs, then raced in a staggered waterfall. A long chain of transformers then turns the winning result into whatever formats were asked for. The AGPL-3.0 licence covers the server. The SDKs are MIT.

## Architecture

```mermaid
flowchart LR
  C["SDK / HTTP client"] --> R["Express routers v1 / v2"]
  R --> MW["auth, credits, blocklist middleware"]
  MW --> SC["scrapeController (inline)"]
  MW --> CC["crawlController"]
  CC --> Q["NuQ queue (Postgres / FDB)"]
  Q --> W["nuq-worker: processJob"]
  SC --> P["scrapeURL pipeline"]
  W --> P
  P --> FL["buildFallbackList"]
  FL --> E["engines: fire-engine, playwright, fetch, pdf, document..."]
  P --> T["executeTransformers"]
  T --> MD["Rust transformHtml + Go/turndown Markdown"]
  T --> LLM["LLM extract (AI SDK)"]
  W --> RD["Redis: crawl state, visited set"]
```

| Component | Path | Role |
|---|---|---|
| HTTP server | `apps/api/src/index.ts`, `routes/v1.ts`, `routes/v2.ts` | Express app; mounts `/v1`, `/v2` and middleware chains per endpoint |
| Controllers | `apps/api/src/controllers/v2/` | `scrape`, `crawl`, `batch-scrape`, `map`, `search`, `extract`, `agent`, `parse` and status endpoints |
| Queue (NuQ) | `apps/api/src/services/worker/nuq.ts`, `apps/nuq-postgres/nuq.sql` | Postgres job table with `FOR UPDATE SKIP LOCKED`; RabbitMQ or `LISTEN` for completion; optional FoundationDB backend |
| Scrape worker | `apps/api/src/services/worker/scrape-worker.ts` | `processJob`: runs a scrape, then for crawls discovers and enqueues links |
| scrapeURL | `apps/api/src/scraper/scrapeURL/index.ts` | Feature flags, engine waterfall, retries on `AddFeatureError` |
| Engines | `apps/api/src/scraper/scrapeURL/engines/` | `fire-engine` client, `playwright`, `fetch`, `pdf`, `document`, `image`, `index`, `wikipedia`, `x-twitter` |
| Transformers | `apps/api/src/scraper/scrapeURL/transformers/` | HTML cleanup, Markdown, links, images, LLM extract, summary, diff |
| Crawler | `apps/api/src/scraper/WebScraper/crawler.ts`, `lib/crawl-redis.ts` | Link filtering (Rust), robots.txt, sitemaps, Redis visited/lock sets |
| Native module | `apps/api/native/` | Rust (napi-rs): HTML transform, link filtering, Markdown post-processing |
| Playwright service | `apps/playwright-service-ts/api.ts` | Self-host browser renderer behind a `/scrape` endpoint |
| SDKs | `apps/*-sdk/` | Thin REST clients for nine languages |

## How a request flows

Take `POST /v2/scrape` with `formats: ["markdown"]`:

1. **Route and gate.** The `/scrape` route runs auth (with keyless access allowed), a country check, a credit check and the scrape blocklist before `scrapeController` ([v2.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/routes/v2.ts#L252-L264)).
2. **Validate and lock.** The controller parses the body with `scrapeRequestSchema` and resolves safe mode ([scrape.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/controllers/v2/scrape.ts#L58-L120)). It then takes a per-team slot in `teamConcurrencySemaphore`, which gets two thirds of the request timeout to acquire ([scrape.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/controllers/v2/scrape.ts#L343-L377)).
3. **Run inline.** A single scrape does not go through the queue. The controller builds a `NuQJob` with `skipNuq: true` and calls `processJobInternal` directly, in the API process ([scrape.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/controllers/v2/scrape.ts#L382-L443)).
4. **Worker wrapper.** `processJob` sets the abort timer, checks for a cancelled crawl, and races `startWebScraperPipeline` against the remaining time ([scrape-worker.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/worker/scrape-worker.ts#L395-L470)). `runWebScraper` calls `scrapeURL` once, or up to three times for crawl pages ([runWebScraper.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/main/runWebScraper.ts#L9-L80)).
5. **Choose engines.** `buildFallbackList` first checks the Exchange data catalogue and the blocklist, then scores every configured engine. Each request flag has a priority. An engine qualifies when it covers at least half the summed priority. The list is then sorted by support score and a fixed `quality` number ([engines/index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L637-L760), [L955-L1070](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L955-L1070)).
6. **Waterfall.** `scrapeURLLoop` starts the first engine. If it has not finished after its "max reasonable time" plus a configured delay, the next engine starts while the first keeps running. The first engine to return an acceptable result wins and the others are aborted ([index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L822-L1000), [L1100-L1125](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L1100-L1125)).
7. **Judge the result.** `scrapeURLLoopIter` converts the HTML to Markdown for a quality check. A non-empty body or a non-2xx status counts as success. A 401, 403 or 429 under `proxy: "auto"` throws `AddFeatureError(["stealthProxy"])` ([index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L663-L800)). The outer loop in `scrapeURL` catches that, adds the flag and rebuilds the waterfall ([index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L1480-L1540)).
8. **Transform.** `executeTransformers` runs a fixed stack in order: HTML cleanup, Markdown, links, images, metadata, then LLM extract, summary, query, agent, diff and format coercion ([transformers/index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/index.ts#L618-L676)). The document goes back in the HTTP response.

## Key components

### Engine waterfall

Engines are declared in one table with feature support and a `quality` score. Index and cache engines rank highest, then Fire Engine Chrome CDP (50), TLS client (10), Playwright (20) and `fetch` (5). Stealth variants have negative quality, so they run only when stealth is required ([engines/index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L52-L140), [L405-L450](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L405-L450)). The list depends on configuration. Without Fire Engine and without Wikipedia or X credentials, it is `playwright`, `fetch`, `pdf`, `document`, and `image` when FirePDF is configured. The self-host Playwright engine declares `actions: false` and `screenshot: false`, so those features fail with `ActionsNotSupportedError` or come back with an "unsupported features" warning ([index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L836-L862)).

Files take a detour. With Fire Engine available, a PDF or DOCX URL goes through a browser engine first. The file comes back via `AddFeatureError` with a prefetch, and the `pdf` or `document` engine then parses those bytes ([engines/index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L795-L860)).

### HTML to Markdown

`htmlTransform` calls `transformHtml` from the Rust module with `includeTags`, `excludeTags` and `onlyMainContent`, and falls back to cheerio on error ([removeUnwantedElements.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/lib/removeUnwantedElements.ts#L69-L110)). `parseMarkdown` tries three converters in order: the Go HTTP service, the Go shared library over `koffi` FFI, then Turndown with GFM. Every path ends in Rust `postProcessMarkdown` ([html-to-markdown.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/html-to-markdown.ts#L1-L140)). If main-content extraction gives empty Markdown, `deriveMarkdownFromHTML` reruns it on the full page ([transformers/index.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/index.ts#L130-L194)).

### LLM extraction

`generateCompletions` defaults to `gpt-4o-mini` with `gpt-4.1-mini` as the retry model. It does not chunk long pages. It trims the Markdown to 80% of the model's input limit and adds a warning. The prompt tells the model to ignore instructions found in the content ([llmExtract.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L440-L495)). Providers come from the Vercel AI SDK: OpenAI, Ollama (the default when `OLLAMA_BASE_URL` is set), Anthropic, Groq, Google, OpenRouter, Fireworks, DeepInfra and Vertex ([generic-ai.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/generic-ai.ts#L1-L55)).

### NuQ queue

Crawls, batch scrapes and async jobs go through NuQ, a custom queue in Postgres. Workers claim jobs with `FOR UPDATE SKIP LOCKED`, ordered by priority and age. Waiters hear about completion through a RabbitMQ reply queue or Postgres `LISTEN` ([nuq.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/worker/nuq.ts#L100-L200)). The schema lives in `apps/nuq-postgres/nuq.sql` ([nuq.sql](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/nuq-postgres/nuq.sql#L1-L60)). `NUQ_BACKEND=fdb` switches to FoundationDB ([queue-jobs.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/queue-jobs.ts#L154-L190)). BullMQ on Redis is still used, but only for side queues such as billing, deep research, llms.txt and precrawl.

### Crawler

`crawlController` fetches robots.txt, creates a NuQ crawl group with `maxConcurrency` and the user's `delay`, saves the crawl in Redis and enqueues one `kickoff` job ([crawl.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/controllers/v2/crawl.ts#L355-L400)). The code that would copy robots.txt `Crawl-delay` into the crawl is commented out, so politeness delay is whatever the caller sets. After each page, the worker extracts links and filters them through Rust `filterLinks` (include/exclude regexes, depth, robots, subdomains, backward crawling) ([crawler.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/WebScraper/crawler.ts#L173-L230)). It runs threat-protection checks and enqueues each link that wins `lockURL` ([scrape-worker.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/worker/scrape-worker.ts#L626-L790)). `lockURL` enforces the page `limit` and dedupes on a normalized URL, or on a permutation key when `deduplicateSimilarURLs` is on ([crawl-redis.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/crawl-redis.ts#L675-L715)).

## Extending it

- **A new engine.** Add a name to the `Engine` union, a handler, a max-reasonable-time function and a feature/quality entry in `engines/index.ts`. The waterfall picks it up from the scores.
- **A new output format.** Add a transformer function and place it in `transformerStack`. Order matters, because later transformers read fields set by earlier ones.
- **Models.** Point `OPENAI_BASE_URL` at any OpenAI-compatible endpoint, or set `OLLAMA_BASE_URL`. Other providers read their standard API-key variables.
- **Clients.** The nine SDKs and the separate CLI and MCP projects all wrap the same REST API, so a new endpoint means a controller plus a route first.

## Running it

- **Docker Compose.** The root `docker-compose.yaml` runs the API (whose harness also starts the workers), the Playwright service, Redis, RabbitMQ and NuQ Postgres. Only port 3002 is published, and the default is `USE_DB_AUTHENTICATION=false`, so the API is unauthenticated ([docker-compose.yaml](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/docker-compose.yaml#L1-L150), [harness.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/harness.ts#L818-L880)).
- **Optional services.** `PROXY_SERVER` for the fetch and Playwright engines, `SEARXNG_ENDPOINT` for search, an LLM key for JSON/summary formats, and Fire Engine or FirePDF URLs if you have access to them.
- **Playwright service.** It runs headless Chromium with a random user agent, routes traffic through an SSRF-checking local proxy, and can block media. It has no stealth patches ([api.ts](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/playwright-service-ts/api.ts#L185-L250)).

## Strengths and caveats

- **Strength: a well-engineered fetch strategy.** The hedged waterfall gives a fast engine a head start without letting a slow one block the request. Flag-based scoring keeps engines that cannot do what was asked out of the list.
- **Strength: wide input coverage.** HTML, PDFs, Office files and images go through one pipeline, with a single document model and many output formats.
- **Strength: built for scale.** The Postgres queue with `SKIP LOCKED`, the Redis visited sets and the per-team semaphores are production designs, not demo code.
- **Caveat: self-host is a different product.** Fire Engine (stealth, actions, screenshots, mobile, location) is a closed external service. Self-hosted, you get plain Playwright and `fetch`.
- **Caveat: SaaS code everywhere.** Billing, credit reservation, Exchange, safe mode and ZDR branches make the core paths long and harder to follow or fork.
- **Caveat: robots.txt is only partly honoured.** Disallow rules are enforced (unless `ignoreRobotsTxt`), but `Crawl-delay` is ignored.
- **Caveat: LLM extraction truncates.** Long pages are trimmed to fit the context window, not chunked and merged, so data near the end of a big page can be lost.

*Sources: code at 8c84d8b, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (41 pages), verified Q&A.*

## How firecrawl/firecrawl answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

**Plain HTTP vs Headless Browser:** Pages are fetched through a waterfall of engines defined in `scrapeURL/engines/index.ts:82-104` — the list includes `fetch` (plain HTTP via `undici`), `playwright` (a Playwright microservice), `fire-engine;chrome-cdp` (proprietary Chrome CDP service), and `fire-engine;tlsclient` (TLS-client engine). The `fetch` engine (`engines/fetch/index.ts:86-229`) uses `undici.fetch` with redirect-following and charset detection from HTTP headers and HTML `<meta>` tags. Headless browser rendering goes through either the Playwright microservice (`engines/playwright/index.ts:8-53`) which POSTs the URL plus wait/timeout/headers, or the Fire Engine's Chrome CDP (`engines/fire-engine/index.ts:346-627`) which sends a rich request with actions, wait, screenshot, mobile, geolocation, and proxy parameters, then polls for completion (`performFireEngineScrape`, lines 70-308).

**JS Rendering:** JS rendering is always-on for the Chrome CDP engine since it runs actual Chrome. The TLS Client engine (`tlsclient`) is a lighter non-JS engine for speed. The `fetch` engine does zero JS rendering.

**Waiting Strategy:** The `waitFor` option controls how long the browser waits before capturing; default is 0ms, but 2000ms for branding scrapes (`fire-engine/index.ts:63`). It's transformed into a `wait` action (lines 389-395). The overall scrape timeout defaults to 300s (`scrapeURL/index.ts:966`) and each engine has its own max-reasonable-time, 30s for Playwright (`engines/playwright/index.ts:65-67`), 15s for fetch (`engines/fetch/index.ts:231-233`).

**Content Types Supported:** The system handles HTML/web pages (all engines), PDFs (`engines/pdf`), Office documents (DOCX, XLSX, PPTX, ODT, CSV, EPUB, RTF via `lib/document-formats.ts:1-32`), images (PNG, JPEG, TIFF, GIF, WebP, AVIF via `lib/image-formats.ts:1-34`), plain text and JSON. PDFs can be processed by FirePDF for OCR. Binary files can be returned as `rawBase64`.


Citations: [apps/api/src/scraper/scrapeURL/engines/index.ts:82-104](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/index.ts#L82-L104) · [apps/api/src/scraper/scrapeURL/engines/fetch/index.ts:86-229](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/fetch/index.ts#L86-L229) · [apps/api/src/scraper/scrapeURL/engines/playwright/index.ts:8-67](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/playwright/index.ts#L8-L67) · [apps/api/src/scraper/scrapeURL/engines/fire-engine/index.ts:346-627](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/fire-engine/index.ts#L346-L627) · [apps/api/src/lib/document-formats.ts:1-32](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/document-formats.ts#L1-L32) · [apps/api/src/lib/image-formats.ts:1-34](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/image-formats.ts#L1-L34)

### How is content extracted or converted? (answered)

**HTML→Markdown Conversion:** The primary conversion is HTML to Markdown via `lib/html-to-markdown.ts`. It supports two backends: an HTTP-based Go HTML-to-Markdown service (`html-to-markdown-client.ts`) or a native Go shared library loaded via FFI (`GoMarkdownConverter`, lines 12-49). Both feed through Rust-based `postProcessMarkdown` from `@mendable/firecrawl-rs` (line 76). The converter is invoked by the `deriveMarkdownFromHTML` transformer (`scrapeURL/transformers/index.ts:86-194`) which applies `htmlTransform` for boilerplate removal and readability processing, then calls `parseMarkdown`.

**Boilerplate Removal:** The `htmlTransform` function in `lib/removeUnwantedElements.ts` strips navigation, ads, and non-content elements. Controlled by the `onlyMainContent` option (defaults to `true`). If the resulting markdown is empty, it falls back to full content extraction (line 166-192).

**Selectors:** Users can pass CSS selectors via `includeTags`/`excludeTags` options. The `actions` system supports `click`, `wait`, `scroll`, `write`, `press`, `scrape`, `executeJavascript`, and `screenshot` actions (`fire-engine/index.ts:313-321`).

**Schema-based Extraction:** JSON/extract schemas for structured data extraction. The `scrapeURL/transformers/index.ts` pipeline (lines 618-643) runs: `performLLMExtract`, `performDeterministicJson`, `performQuery`, `performSummary`, `performAgent`, and others — each processing the document sequentially. The LLM extract supports JSON schema with recursive-reference detection, auto-repair of malformed output, and fallback models (`transformers/llmExtract.ts:49-72` for model selection, lines 774-818 for `generateObject` with schema).


Citations: [apps/api/src/lib/html-to-markdown.ts:1-49](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/html-to-markdown.ts#L1-L49) · [apps/api/src/scraper/scrapeURL/transformers/index.ts:86-194](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/index.ts#L86-L194) · [apps/api/src/scraper/scrapeURL/transformers/index.ts:618-643](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/index.ts#L618-L643) · [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:49-72](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L49-L72) · [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:774-818](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L774-L818)

### How are LLMs used, if at all? (answered)

**Prompting:** LLMs are used for structured JSON extraction (`performLLMExtract`, `performQuery`, `performSummary` in `transformers/index.ts:632-638`), schema-based extraction with `generateObject` from the Vercel AI SDK (`transformers/llmExtract.ts:774-818`), text summarization, question answering, highlight generation, and free-form agent extraction (`performAgent`). The prompt prepends markdown content to system/user prompts with the instruction to ignore data-processing directives (line 489-492).

**Chunking of Large Pages:** The `trimToTokenLimit` function (`transformers/llmExtract.ts:207-274`) uses `tiktoken` to pre-trim content by characters then precisely slice to the model's `max_input_tokens` at 80% capacity (line 480-486). It handles both character-based pre-trimming and exact token-aligned trimming with proper UTF-8 boundary handling.

**Structured Output / JSON Schema:** The system supports `jsonSchema` from `ai`, Zod schemas, and raw JSON Schema objects. Recursive schemas (with `$ref`, `$defs`) trigger model upgrade to `gpt-4.1` (lines 49-72). It includes repair config (`experimental_repairText`, lines 678-772) that strips markdown code fences and uses LLM-based JSON repair on failure.

**Providers:** Multiple providers via `lib/generic-ai.ts:1-55`: OpenAI (default), Anthropic, Ollama, Groq, Google Gemini, OpenRouter, Fireworks, DeepInfra, and Vertex AI (GCP). The default provider switches to Ollama when `OLLAMA_BASE_URL` is set (line 23). The default model is `gpt-4o-mini` (line 464), with `gpt-4.1-mini` as fallback on quota errors (line 449).

**Cost Controls:** `calculateCost` tracks per-model pricing from a hardcoded table (`modelCosts` at lines 278-297) including GPT-4o, GPT-4.1, o3-mini, GPT-5 series, Gemini, Grok, and DeepSeek models. Costs are recorded via the `CostTracking` system which logs token usage and dollar amounts to billing (lines 527-543, 848-864). Unpriced models are tracked but recorded as $0 with a warning (lines 348-359).


Citations: [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:49-72](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L49-L72) · [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:207-274](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L207-L274) · [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:278-297](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L278-L297) · [apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts:440-490](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/llmExtract.ts#L440-L490) · [apps/api/src/lib/generic-ai.ts:1-55](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/generic-ai.ts#L1-L55)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

**Stealth Patches:** The system has a multi-layer approach. First, the TLS Client engine (`fire-engine;tlsclient`) is a custom HTTP client that mimics browser TLS fingerprints — it uses `atsv` (anti-bot solver) and `skipTlsVerification` options. The Chrome CDP engine runs real Chrome, so it naturally looks like a browser. The `engpicker` (`lib/engpicker.ts:1-325`) is an automated system that scrapes target URLs with multiple engine/proxy combinations, uses GPT-4o-mini to evaluate success vs. antibot pages (lines 70-101), and computes a Rust-based Levenshtein verdict to decide whether `tlsclient` with or without stealth proxy is adequate for a domain.

**Fingerprint Spoofing:** Browser profiles (saved authenticated sessions) are supported via `lib/browser-profiles.ts:1-15`. The `mobileProxy` flag on Fire Engine requests enables carrier-grade mobile proxies that change the apparent device and network. The `forceNonRender` option (`fire-engine/index.ts:327-344`) can route branding scrapes away from the visual render engine to avoid triggering heavy anti-bot challenges.

**Proxy Rotation:** Three proxy tiers are available: `basic` (direct, no proxy), `stealth` (mobile carrier proxy with rotated IP), and `auto` (escalates from basic to stealth when the engine detects 401/403/429 responses — `scrapeURL/index.ts:769-783`). Config variables `PROXY_SERVER`, `PROXY_USERNAME`, `PROXY_PASSWORD` (`config.ts:408-410`) set a forward proxy for the `fetch` engine. The `PARTNER_EGRESS_PROXY_URL` (config.ts:197) provides a partner egress proxy for threat-protection scanning.

**CAPTCHA Handling:** The system detects antibot pages through the scrape result quality check (`scrapeURL/index.ts:748-762`): short markdown or 401/403/429 status codes trigger `AddFeatureError` to escalate to stealth proxy engines. The specialty scraper (`engines/utils/specialtyHandler.ts:117-130`) sniffs browser handoff payloads for bot-blocked content and routes to file-specific parsing.

**Rate Limiting:** Per-team concurrency limits are enforced via `lib/concurrency-limit.ts:1-419` using sorted sets in Redis. Per-minute rate limits are in `services/rate-limiter.ts:21-44` covering crawl (15/min), scrape (100/min), search, map, extract, and browser modes with configurable multipliers from the Autumn billing service.

> **Editor's note.** Correction: escalation to stealth proxies happens only when an engine returns 401, 403 or 429 with proxy set to auto; short or empty output just makes the waterfall try the next engine.

Citations: [apps/api/src/lib/engpicker.ts:1-325](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/engpicker.ts#L1-L325) · [apps/api/src/scraper/scrapeURL/index.ts:748-783](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/index.ts#L748-L783) · [apps/api/src/scraper/scrapeURL/engines/fire-engine/index.ts:327-344](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/engines/fire-engine/index.ts#L327-L344) · [apps/api/src/lib/concurrency-limit.ts:1-30](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/concurrency-limit.ts#L1-L30) · [apps/api/src/services/rate-limiter.ts:21-44](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/rate-limiter.ts#L21-L44)

### How is crawling at scale implemented? (answered)

**Queues:** Crawls use Redis (BullMQ) via `services/queue-service.ts`. Jobs are persisted with `saveCrawl` in `lib/crawl-redis.ts:45-79`, which stores crawl config with 24-hour TTL. The queue backend supports both PostgreSQL (`pg`) and FoundationDB (`fdb`) for crawl job storage (`StoredCrawl.queueBackend` at line 31). Background extraction jobs use a separate Extract queue (`services/extract-queue.ts`).

**Concurrency:** Crawl-scoped concurrency uses sorted sets in Redis (`lib/concurrency-limit.ts:174-214`). An `AB-Test` system (`services/ab-test.ts`) can split or mirror crawl traffic across engines. The `getNextConcurrentJob` function (line 224-309) pops from a ZSET atomically, checks per-crawl concurrency caps, and re-queues blocked jobs. When a job finishes, `concurrentJobDone` (line 316-418) promotes backlogged jobs respecting per-team concurrency limits.

**URL Dedup:** The `WebCrawler` (`crawler.ts:59-165`) uses a `visited` Set for deduplication (line 67). The `filterLinks` method (line 173-471) applies dedup via both a Rust native function (`filterLinks` from `@mendable/firecrawl-rs`, line 206-221) and a JavaScript fallback. URL normalization strips `www.` prefix and trailing slashes (line 576-581).

**Depth/Limits:** Configurable `maxCrawledDepth` (default 10) and a `limit` cap (default 10000). The `getURLDepth` utility measures path depth. Links exceeding the depth get descriptive denial reasons (line 345-354). Filtering checks include regex-based exclude/include patterns, backward-crawling restriction, robots.txt, file-type filtering, external-link control, and non-web protocol rejection (lines 310-470).

**robots.txt and Politeness:** Built-in robots.txt checking via `lib/robots-txt.ts:1-80` — fetched on each crawl, cached with 1-day max-age, and parsed with the `robots-parser` library. The `ignoreRobotsTxt` option disables it (crawler.ts line 152). Crawl delay from robots.txt is extracted and honored via `getRobotsCrawlDelay()` (line 550-552) and enforced by `concurrentJobDone` which pauses between job promotions (line 362-367). Sitemap discovery via `tryGetSitemap` (line 554-615) fetches sitemap.xml within 120s timeout.

**Distributed Workers:** Multiple worker processes consume from Redis queues via `services/worker/` — the `scrapeQueue` (NuQ, a custom queue on top of BullMQ) distributes scrapes across worker pods. The `concurrency-queue-reconciler.ts` handles back-pressure when teams exceed their concurrency limits.

> **Editor's note.** Correction: crawl and scrape jobs run on NuQ, a Postgres queue (FOR UPDATE SKIP LOCKED, optional FoundationDB backend, RabbitMQ only for completion signals), not BullMQ. robots.txt Disallow rules are enforced, but Crawl-delay is not: the code that would copy it into the crawl is commented out, so only the caller's `delay` option slows a crawl.

Citations: [apps/api/src/lib/crawl-redis.ts:45-79](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/crawl-redis.ts#L45-L79) · [apps/api/src/lib/concurrency-limit.ts:174-367](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/concurrency-limit.ts#L174-L367) · [apps/api/src/scraper/WebScraper/crawler.ts:59-165](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/WebScraper/crawler.ts#L59-L165) · [apps/api/src/scraper/WebScraper/crawler.ts:173-471](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/WebScraper/crawler.ts#L173-L471) · [apps/api/src/lib/robots-txt.ts:1-80](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/lib/robots-txt.ts#L1-L80)

### What is the developer interface? (answered)

**REST API:** The primary interface is a REST API served by Express, with two major versions:
- **v1** (`routes/v1.ts`): POST endpoints for `/scrape`, `/crawl`, `/batch/scrape`, `/search`, `/map`, `/extract`, plus WebSocket at `/crawl/:jobId` for streaming crawl status.
- **v2** (`routes/v2.ts`): Adds POST `/parse`, `/parse-upload`, `/agent`, `/browser`, `/feedback`, `/slack`, `/monitor`, `/research-proxy`, `/scrape-browser`, and OAuth token introspection. Batch/scrape and crawl via NuQ queues.
Both versions use middleware chains for auth, rate-limiting, credit checks, blocklist filtering, idempotency, and country validation.

**Language SDKs:** Multiple officially maintained SDKs under the monorepo's `apps/` directory: `python-sdk` (FirecrawlApp with sync/async, Watcher, typed errors), `js-sdk`, `go-sdk`, `rust-sdk`, `dot-net-sdk`, `ruby-sdk`, `php-sdk`, `java-sdk`, `elixir-sdk`. All cover the core endpoints.

**CLI:** No standalone CLI is packaged — the SDKs serve as programmatic clients.

**MCP Server:** An MCP (Model Context Protocol) server is hosted at `mcp.firecrawl.dev/v2/mcp` with OAuth and API-key auth (`services/mcp/action-logs.ts:8-9`). Action logs record MCP tool usage with retention, pagination, and team-scoped queries.

**UI:** An ingestion dashboard UI exists at `apps/ui/ingestion-ui/` built with Vue.js/Vite and Tailwind CSS.

**Output Formats:** The core `Document` type contains: `markdown`, `rawHtml`, `rawBase64`, `html`, `screenshot`, `actions`, `branding`, `json`, `extract`, `summary`, `answer`, `highlights`, `links`, `images`, `audio`, `video`, `product`, `menu`, `changeTracking`. All controlled by the formats array parameter.

> **Editor's note.** Correction: the ingestion UI is React + Vite, not Vue. A Firecrawl CLI does exist; it lives in a separate repository that firecrawl-cli/README.md links to.

Citations: [apps/api/src/routes/v1.ts:1-150](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/routes/v1.ts#L1-L150) · [apps/api/src/routes/v2.ts:1-150](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/routes/v2.ts#L1-L150) · [apps/api/src/scraper/scrapeURL/transformers/index.ts:342-616](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/scraper/scrapeURL/transformers/index.ts#L342-L616) · [apps/api/src/services/mcp/action-logs.ts:1-15](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/api/src/services/mcp/action-logs.ts#L1-L15) · [apps/python-sdk/firecrawl/__init__.py:1-44](https://github.com/firecrawl/firecrawl/blob/8c84d8b6a155495b8088c286d878b2f875de3918/apps/python-sdk/firecrawl/__init__.py#L1-L44)
