# unclecode/crawl4ai

> Python library and server that renders pages with Playwright and turns them into Markdown or JSON through pluggable strategies.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/unclecode/crawl4ai (reviewed at commit `8afd0a68064ff7049303c9f9d037ab6228aac43c`, 2026-10-05)
- Stars: 84844 · Language: Python · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/crawl4ai/

## Overview

Crawl4AI is a Python library for turning web pages into Markdown and structured data, built on Playwright. The core is one class, `AsyncWebCrawler`. You call `arun(url, config)` for one page or `arun_many(urls, config)` for a batch, and get back a `CrawlResult` with raw and cleaned HTML, three Markdown variants, links, media, tables, screenshots, PDFs and any extracted JSON. Almost every behaviour is a pluggable strategy: the fetcher, the HTML scraper, the Markdown generator, the content filter, the chunker, the extractor and the deep-crawl traversal.

The same package ships a `crwl` CLI and a FastAPI server under `deploy/docker`. The server wraps the library with a browser pool, a REST API, NDJSON streaming, background jobs backed by Redis, and an MCP bridge. The library is the product, and the server is a thin, security-hardened shell around it.

It suits developers who want a local, self-hosted crawler with no SaaS dependency. Defaults are conservative. Caching is bypassed, robots.txt is not checked, and anti-bot retries are off until you turn them on.

## Architecture

```mermaid
flowchart LR
  U["Your code / crwl / REST"] --> W["AsyncWebCrawler.arun"]
  W --> DC["DeepCrawlDecorator"]
  DC --> BFS["BFS / DFS / BestFirst strategy"]
  BFS --> W
  W --> CA["SQLite cache"]
  W --> CS["AsyncPlaywrightCrawlerStrategy"]
  CS --> BM["BrowserManager + adapter"]
  W --> AB["is_blocked + proxy retry"]
  W --> AP["aprocess_html"]
  AP --> SS["LXMLWebScrapingStrategy"]
  AP --> MG["DefaultMarkdownGenerator + filter"]
  AP --> EX["ExtractionStrategy (CSS / XPath / LLM)"]
  EX --> LL["litellm"]
  M["arun_many"] --> D["MemoryAdaptiveDispatcher"]
  D --> W
```

| Component | Path | Role |
|---|---|---|
| Crawler facade | `crawl4ai/async_webcrawler.py` | `arun`, `arun_many`, cache, robots, anti-bot retries, `aprocess_html` |
| Configs | `crawl4ai/async_configs.py` | `BrowserConfig`, `CrawlerRunConfig`, `LLMConfig`; serialisable for the server |
| Fetch strategies | `crawl4ai/async_crawler_strategy.py` | `AsyncPlaywrightCrawlerStrategy` (default) and `AsyncHTTPCrawlerStrategy` (aiohttp) |
| Browser layer | `crawl4ai/browser_manager.py`, `browser_adapter.py` | Contexts, sessions, CDP/persistent browsers; Playwright, playwright-stealth or patchright adapters |
| Block detection | `crawl4ai/antibot_detector.py` | `is_blocked()`: tiered regexes plus status and structure checks |
| Scraping | `crawl4ai/content_scraping_strategy.py` | `LXMLWebScrapingStrategy`: cleaned HTML, links, media, tables |
| Markdown | `crawl4ai/markdown_generation_strategy.py`, `html2text/` | Raw, cited and "fit" Markdown via a bundled html2text fork |
| Filters | `crawl4ai/content_filter_strategy.py` | `PruningContentFilter`, `BM25ContentFilter`, `LLMContentFilter` |
| Extraction | `crawl4ai/extraction_strategy.py` | `JsonCss`/`JsonXPath`/`JsonLxml`, `Regex`, `Cosine`, `LLMExtractionStrategy` |
| Deep crawl | `crawl4ai/deep_crawling/` | BFS, DFS, best-first; `FilterChain` and URL scorers |
| Dispatch | `crawl4ai/async_dispatcher.py` | `MemoryAdaptiveDispatcher`, `SemaphoreDispatcher`, per-domain `RateLimiter` |
| Server | `deploy/docker/server.py`, `api.py` | FastAPI endpoints, browser pool, job queue, MCP bridge |

## How a request flows

Take `await crawler.arun("https://example.com", CrawlerRunConfig(extraction_strategy=...))`:

1. **Deep-crawl check.** `arun` is wrapped by `DeepCrawlDecorator`. If `deep_crawl_strategy` is set, the call goes to the strategy, which calls `arun` again for each page, guarded by a `ContextVar` ([base_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/deep_crawling/base_strategy.py#L10-L44)).
2. **Cache and proxy.** `arun` reads the SQLite cache unless `cache_mode` says otherwise, and can revalidate with ETag or Last-Modified. It then takes the next proxy from `proxy_rotation_strategy`, with optional sticky sessions ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L216-L380)).
3. **Robots and attempts.** With `check_robots_txt=True`, a disallowed URL returns a synthetic 403 result. Otherwise it loops over `1 + max_retries` attempts and, inside each attempt, over the proxy list ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L380-L480)).
4. **Fetch.** `AsyncPlaywrightCrawlerStrategy.crawl` sends `http(s)` URLs to `_crawl_web`. `file://` and `raw:` inputs skip the browser unless a browser-only feature is requested ([async_crawler_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_crawler_strategy.py#L436-L512)). `_crawl_web` sets a random or fixed user agent with matching `sec-ch-ua`, gets a page from `BrowserManager`, and injects the navigator overrider when `magic`, `simulate_user` or `override_navigator` is on ([async_crawler_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_crawler_strategy.py#L540-L610)). Then it navigates, runs `js_code`, waits, scrolls and captures.
5. **Process.** `aprocess_html` runs the scraping strategy on the HTML ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L780-L800)). It picks the Markdown source (`cleaned_html` by default) and calls `generate_markdown` ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L835-L890)).
6. **Extract.** If an extraction strategy is set, the chosen input format (Markdown, fit Markdown or HTML) is chunked and passed to the strategy's `arun`, or to its sync `run` in a thread ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L905-L960)).
7. **Judge.** Back in `arun`, `is_blocked(status, html)` decides whether the page is a block page. If it is, the next proxy or attempt runs. After the loop, an optional `fallback_fetch_function(url)` gets one last chance, and a still-blocked result is marked failed ([async_webcrawler.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L480-L640)).

## Key components

### Fetching and stealth

The Playwright strategy owns hooks (`before_goto`, `after_goto`, `before_return_html` and others) and builds `BrowserManager` with `use_undetected` set when you pass an `UndetectedAdapter` ([async_crawler_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_crawler_strategy.py#L76-L120)). That adapter runs on patchright and evaluates scripts in an isolated world. `StealthAdapter` applies `playwright_stealth` when `enable_stealth=True` ([browser_adapter.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/browser_adapter.py#L151-L175), [L271-L300](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/browser_adapter.py#L271-L300)). There is also an aiohttp-based `AsyncHTTPCrawlerStrategy` for pages that need no JavaScript.

### Block detection and retries

`is_blocked` is a layered heuristic. Any 429 counts as a block. "Tier 1" vendor markers (Cloudflare, Akamai, PerimeterX, DataDome, Imperva, Kasada and others) are searched in the first 15 KB and in a script-stripped copy of large pages. A 403 or 503 on a non-data page counts as blocked, and so do a near-empty 200 and pages that fail a structural check ([antibot_detector.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/antibot_detector.py#L191-L281)). The retry loop only acts on this when you set `max_retries`, a proxy list or `fallback_fetch_function`. All three default to off ([async_configs.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_configs.py#L1804-L1828)).

### Markdown and filtering

`DefaultMarkdownGenerator` converts HTML with a bundled `CustomHTML2Text`, rewrites links into numbered citations, and, when a content filter is attached, produces `fit_markdown` from the filtered HTML ([markdown_generation_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/markdown_generation_strategy.py#L55-L75)). `PruningContentFilter` scores nodes by text density and link ratio. `BM25ContentFilter` keeps blocks relevant to a query. `LLMContentFilter` asks a model to keep what matters.

### Extraction

Non-LLM extractors take a schema with a `baseSelector` and fields (CSS, XPath or lxml). `generate_schema` can ask an LLM to write that schema once, so later runs need no model. `LLMExtractionStrategy` merges sections into chunks by a word-to-token ratio with overlap. The async path sends every chunk at once with `asyncio.gather` ([extraction_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L972-L1000)). The sync path uses four threads, or runs sequentially for Groq ([extraction_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L774-L830)). Each chunk prompt picks a block, instruction or schema template. Calls go through `litellm` with backoff, token usage is summed, and output is parsed as JSON or `<blocks>` XML with a lenient fallback ([extraction_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L860-L965)). `config.py` lists known provider prefixes and their key variables, and defaults to `openai/gpt-4o` ([config.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/config.py#L1-L40)).

### Deep crawling and dispatch

`BFSDeepCrawlStrategy.link_discovery` normalizes each link, skips visited ones, runs the `FilterChain`, scores with an optional scorer, and trims to the remaining `max_pages` budget, best scores first ([bfs_strategy.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/deep_crawling/bfs_strategy.py#L133-L200)). `arun_many` uses `MemoryAdaptiveDispatcher` by default. It pauses when system memory crosses a threshold and can attach a per-domain `RateLimiter` that backs off on 429 and 503 ([async_dispatcher.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_dispatcher.py#L28-L60)). The state is in-process. There is no distributed queue in the library.

### Docker server

`/crawl` loads browser and run configs with `Provenance.UNTRUSTED`, applies an egress policy and deep-crawl clamps, then takes a browser from the pool ([server.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/server.py#L955-L1006), [api.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/api.py#L660-L760)). The pool keeps one permanent browser plus hot and cold pools keyed by config signature, and recycles them by idle time and memory ([crawler_pool.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/crawler_pool.py#L1-L60)). Background `/crawl/job` and `/llm/job` run on a bounded in-process worker pool with per-caller quotas. Redis stores the job status ([work_queue.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/work_queue.py#L1-L30)). `attach_mcp` exposes the endpoints as MCP tools over WebSocket and SSE, calling back into the API over loopback ([server.py](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/server.py#L1185-L1203)).

## Extending it

- **Strategies.** Subclass `AsyncCrawlerStrategy`, `ContentScrapingStrategy`, `MarkdownGenerationStrategy`, `RelevantContentFilter`, `ChunkingStrategy`, `ExtractionStrategy`, `DeepCrawlStrategy` or the URL filters and scorers, then pass them in the config.
- **Hooks.** `crawler.crawler_strategy.set_hook("before_goto", fn)` and the other hook points give you the Playwright page and context. The server offers declarative hooks instead of arbitrary code by default.
- **Scripting.** C4A-Script (`crawl4ai/script/`) compiles a small DSL (`GO`, `CLICK`, `WAIT`, `IF`, `PROC`) to JavaScript for `js_code`.
- **Models.** Any `litellm` provider string works through `LLMConfig(provider=..., api_token=..., base_url=...)`.

## Running it

- **Library.** `pip install crawl4ai`, then `crawl4ai-setup` to install the Playwright browsers. `crawl4ai-doctor` checks the setup. The cache lives in `~/.crawl4ai/crawl4ai.db` (SQLite).
- **CLI.** `crwl <url>` crawls and prints Markdown. Subcommands manage a persistent browser and CDP connections.
- **Server.** The Dockerfile and `docker-compose.yml` run the FastAPI app with Redis under supervisord. Settings live in `deploy/docker/config.yml` (rate limits, pool sizes, memory threshold, JWT auth).

## Strengths and caveats

- **Strength: composable design.** Every stage is a strategy object, so swapping the Markdown filter or extractor needs no fork.
- **Strength: LLM-optional.** CSS, XPath and regex extractors plus `generate_schema` let you pay for a model once, not on every page.
- **Strength: serious anti-bot plumbing.** Block detection, proxy cascades, a fallback fetcher and three browser adapters cover much more than a plain Playwright wrapper.
- **Caveat: surface area.** `CrawlerRunConfig` has well over a hundred parameters, and legacy kwargs are still accepted. Behaviour depends on flag combinations that are easy to get wrong.
- **Caveat: politeness is opt-in.** robots.txt checks, rate limiting and retries are all off by default.
- **Caveat: LLM chunk fan-out.** The async extractor fires every chunk at once, which can hit provider rate limits on long pages.
- **Caveat: single-node.** Concurrency, dedup and queues are in-process. Scaling out means running several servers behind your own queue.

*Sources: code at 8afd0a6, deepwiki-open wiki (13 pages), OpenDeepWiki wiki (35 pages), verified Q&A.*

## How unclecode/crawl4ai answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

**Plain HTTP vs browser.** `AsyncPlaywrightCrawlerStrategy` (`crawl4ai/async_crawler_strategy.py:46`) is the default, using Playwright to launch Chromium. The `BrowserConfig` (proxy, headless, extra args) is passed to `BrowserManager` which manages the browser lifecycle via `PlaywrightAdapter`. A plain-HTTP path exists via the `HTTPCrawlerConfig` type and the fallback `aiohttp`-based fetch method in the same module. **JS rendering.** Every page is loaded in a real Playwright browser context; `BrowserAdapter` (`crawl4ai/browser_adapter.py`) wraps the Playwright page and runs user-supplied `js_code` before extraction. **Waiting strategy.** `CrawlerRunConfig` supports `wait_for`, `wait_until`, `wait_for_images`, `screenshot_wait_for`; the page is waited on per these directives before extraction begins. The C4A-Script DSL (`crawl4ai/script/c4ai_script.py:139`) adds fine-grained `WAIT` commands (by seconds, CSS selector, or text content) that compile to JavaScript polling loops. **Content types.** PDFs use a separate pipeline: `PDFCrawlerStrategy` (`crawl4ai/processors/pdf/__init__.py:49`) returns a placeholder response, then `PDFContentScrapingStrategy` (line 79) downloads the PDF with `requests`, validates every redirect hop for SSRF, and processes it with `NaivePDFProcessorStrategy`. Screenshots and PDF snapshots of rendered pages are produced via Playwright's native screenshot/PDF methods (`deploy/docker/server.py:742-818`).


Citations: [crawl4ai/async_crawler_strategy.py:36-65](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_crawler_strategy.py#L36-L65) · [crawl4ai/processors/pdf/__init__.py:49-78](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/processors/pdf/__init__.py#L49-L78) · [crawl4ai/processors/pdf/__init__.py:79-130](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/processors/pdf/__init__.py#L79-L130) · [crawl4ai/script/c4ai_script.py:139-147](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/script/c4ai_script.py#L139-L147) · [deploy/docker/server.py:742-818](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/server.py#L742-L818)

### How is content extracted or converted? (answered)

**HTML→Markdown/text.** `LXMLWebScrapingStrategy` (`crawl4ai/content_scraping_strategy.py`) is the default scraper; it lxml-parses the DOM and produces clean HTML. The `DefaultMarkdownGenerator` (`crawl4ai/markdown_generation_strategy.py`) converts that to Markdown in three variants: `raw_markdown` (full), `fit_markdown` (filtered), and `markdown_with_citations`. **Readability-style filtering.** Three content filters (`crawl4ai/content_filter_strategy.py`) sit inside the generator: `PruningContentFilter` (heuristic node pruning based on text density/link ratio), `BM25ContentFilter` (relevance-scoring against a user query), and `LLMContentFilter` (LLM-summarizes the page per an instruction). A faster C-ext variant `PruningContentFilterLXML` lives in `content_filter_strategy_lxml.py`. **Selectors.** `CrawlerRunConfig` accepts `css_selector` for targeted extraction. For schema-based extraction, `JsonCssExtractionStrategy` and `JsonXPathExtractionStrategy` (`crawl4ai/extraction_strategy.py:1989,2449`) extract structured fields from HTML using CSS selectors or XPath expressions. **Schema-based extraction.** `LLMExtractionStrategy` (line 533) accepts a JSON schema dict and an instruction; it prompts the LLM to return JSON matching that schema. `RegexExtractionStrategy` (line 2558) applies regex patterns. `JsonElementExtractionStrategy` (line 1043) is the base for CSS/XPath/LXML structured extraction — you define a `baseSelector` and per-field selectors, and it iterates matches.


Citations: [crawl4ai/extraction_strategy.py:533-600](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L533-L600) · [crawl4ai/extraction_strategy.py:1989-2010](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L1989-L2010) · [crawl4ai/content_filter_strategy.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/content_filter_strategy.py#L1-L30) · [crawl4ai/markdown_generation_strategy.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/markdown_generation_strategy.py#L1-L30)

### How are LLMs used, if at all? (answered)

**LLM extraction.** `LLMExtractionStrategy` (`crawl4ai/extraction_strategy.py:533`) orchestrates all LLM calls. Its `extract()` method (line 641) builds a prompt from templates in `crawl4ai/prompts.py`, sends it to the LLM via `perform_completion_with_backoff` (from `crawl4ai/utils.py`), parses the response (XML `<blocks>` or raw JSON), and tracks token usage. **Chunking.** The strategy chunks pages that exceed `CHUNK_TOKEN_THRESHOLD` (defined in `crawl4ai/config.py`). The `run()` method (line 786) calls `_merge()` → `merge_chunks()` (`crawl4ai/utils.py:170`) which splits documents by token estimate (`word_token_ratio`) with configurable `overlap_rate`. Sections are then dispatched to the LLM via `ThreadPoolExecutor` (max 4 workers) for parallel processing. **Schema support.** When `extraction_type="schema"` and a `schema` dict is provided, the prompt switches to `PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION`, which instructs the LLM to return JSON matching the schema. `force_json_response=True` uses the provider's JSON-mode if available. **Providers.** Supported providers and their API keys are registered in `crawl4ai/config.py:11-32`, including OpenAI (`gpt-4o`, `o1`, `o3-mini`), Anthropic (`claude-3-*`), Gemini (`gemini-pro`, `2.0-flash`), Groq, Ollama, and DeepSeek. The `LLMConfig` type (`crawl4ai/types.py`) holds provider/Api_token/temperature/base_url. **Cost tracking.** Each extraction records a `TokenUsage` object with prompt/completion tokens; `total_usage` accumulates across chunks.

> **Editor's note.** Correction: on the normal async path LLMExtractionStrategy.arun sends all chunks at once with asyncio.gather; the 4-thread ThreadPoolExecutor is only the sync run(). Calls go through litellm, so any litellm provider string works, not only those in config.py.

Citations: [crawl4ai/extraction_strategy.py:533-630](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L533-L630) · [crawl4ai/extraction_strategy.py:760-830](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/extraction_strategy.py#L760-L830) · [crawl4ai/utils.py:170-210](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/utils.py#L170-L210) · [crawl4ai/config.py:1-32](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/config.py#L1-L32) · [crawl4ai/prompts.py:1-50](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/prompts.py#L1-L50)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

**Stealth patches.** `UndetectedAdapter` (`crawl4ai/browser_adapter.py`) wraps Playwright to apply fingerprinting countermeasures — it can be selected via `BrowserConfig(browser_type="undetected")`. Examples in `docs/examples/simple_anti_bot_examples.py` demonstrate bypassing Cloudflare, DataDome, and generic bot detection. **Fingerprint spoofing.** `crawl4ai/user_agent_generator.py` provides `ValidUAGenerator` and `UAGen` for realistic User-Agent rotation. The Docker server also exposes config knobs for browser extra args and language headers. **Proxy rotation.** `crawl4ai/proxy_strategy.py` implements the `ProxyRotationStrategy` base class; the Docker server supports per-request proxy config through `BrowserConfig.proxy_config`. The `auth_proxy_example.py` and `nstproxy_example.py` in `docs/examples/` show authenticated proxy setup. **CAPTCHA handling.** The `docs/examples/capsolver_captcha_solver/` directory has production examples for solving Cloudflare Turnstile/Challenge, reCAPTCHA v2/v3, and AWS WAF via both Capsolver extension integration and API integration. These work by injecting the capsolver extension into Playwright. **Rate limiting.** `RateLimiter` (`crawl4ai/async_dispatcher.py:28`) implements per-domain exponential backoff on 429/503 responses, configurable base delay, max delay, and max retries. The dispatcher (`MemoryAdaptiveDispatcher`) also monitors system memory and pauses crawling when thresholds are exceeded.

> **Editor's note.** Correction: UndetectedAdapter is chosen by passing it as browser_adapter to AsyncPlaywrightCrawlerStrategy (patchright), not via browser_type. The answer also misses the main mechanism: antibot_detector.is_blocked() classifies block pages, and arun() retries over max_retries and a proxy list, then calls an optional fallback_fetch_function; all of it is off by default.

Citations: [crawl4ai/async_dispatcher.py:28-80](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_dispatcher.py#L28-L80) · [crawl4ai/proxy_strategy.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/proxy_strategy.py#L1-L30) · [crawl4ai/user_agent_generator.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/user_agent_generator.py#L1-L30) · [crawl4ai/browser_adapter.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/browser_adapter.py#L1-L30)

### How is crawling at scale implemented? (answered)

**Queues.** `async_dispatcher.py` provides `MemoryAdaptiveDispatcher` which maintains a domain-state map for per-domain rate limiting. The Docker server adds a Redis-backed `WorkQueue` (`deploy/docker/work_queue.py`) with per-principal quotas and bounded-size queues (HTTP 429/503). **Concurrency.** `CrawlerRunConfig.semaphore_count` limits parallel page loads per dispatcher. `MemoryAdaptiveDispatcher` also monitors system RAM via `psutil` and pauses fetching when `memory_threshold_percent` is exceeded. **URL dedup.** `BFSDeepCrawlStrategy.link_discovery()` (`crawl4ai/deep_crawling/bfs_strategy.py:133`) maintains a `visited: Set[str]` and checks `can_process_url()` before adding new URLs. URLs are normalized via `normalize_url_for_deep_crawl()` before dedup. **Depth/limits.** `BFSDeepCrawlStrategy` accepts `max_depth`, `max_pages`, `score_threshold`, and `include_external`. The `FilterChain` (`crawl4ai/deep_crawling/filters.py`) composes URL filters (domain, path, content-type). **Robots.txt/politeness.** The `RateLimiter.wait_if_needed()` method enforces per-domain delays between requests. Crawl strategies accept `should_cancel` callbacks for graceful interruption. **Distributed workers.** The Docker deployment (`deploy/docker/server.py`) manages a pool of browser instances via `BrowserPool` (`crawler_pool.py`). It supports `crawler_configs` — per-URL custom configs with `url_matcher` patterns for arun_many(). Webhook delivery (`deploy/docker/webhook.py`) notifies external services on completion/failure. Crash recovery is supported via `resume_state` callback in the deep crawl strategies.

> **Editor's note.** Correction: the server's WorkQueue is an in-process asyncio queue with per-caller caps (Redis only stores job status). robots.txt is supported: CrawlerRunConfig.check_robots_txt=True makes arun return a 403 result for disallowed URLs, but it defaults to False.

Citations: [crawl4ai/deep_crawling/bfs_strategy.py:16-80](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/deep_crawling/bfs_strategy.py#L16-L80) · [crawl4ai/deep_crawling/bfs_strategy.py:133-160](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/deep_crawling/bfs_strategy.py#L133-L160) · [crawl4ai/async_dispatcher.py:28-80](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_dispatcher.py#L28-L80) · [deploy/docker/api.py:660-750](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/api.py#L660-L750) · [deploy/docker/work_queue.py:1-30](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/work_queue.py#L1-L30)

### What is the developer interface? (answered)

**Library API.** `AsyncWebCrawler` (`crawl4ai/async_webcrawler.py`) is the main class, with `arun()` for single URLs and `arun_many()` for batches. Both accept `CrawlerRunConfig` and return `CrawlResult` objects containing html, markdown (raw/fit/citations), extracted_content, screenshot, pdf, metadata, and more. **CLI.** `crawl4ai/cli.py` provides a command-line interface; `crawl4ai/cloud/cli.py` handles cloud deployments. **REST API.** The Docker server (`deploy/docker/server.py`) exposes POST endpoints: `/crawl` (JSON), `/crawl/stream` (NDJSON), `/md` (Markdown), `/llm` (QA), `/screenshot`, `/pdf`, `/html`, `/execute_js`, `/config/dump`, `/ask` (BM25 context retrieval), and `/schema`. **Streaming protocol.** `/crawl/stream` returns a `StreamingResponse` with `application/x-ndjson` content type. Each result is a JSON line via `stream_results()` (`deploy/docker/api.py:616`), which serializes each `CrawlResult` as it arrives from the async generator, includes a heartbeat mechanism, and finishes with `{"status":"completed"}`. Deep-crawl streaming (`crawl4ai/async_webcrawler.py:1047-1057`) wraps the BFS async generator in the same NDJSON format. **MCP server.** `deploy/docker/mcp_bridge.py` attaches at `/mcp/ws` (WebSocket) and `/mcp/sse` (SSE) transports. The WebSocket handler (`_ws`, line 198) uses `anyio.create_memory_object_stream` to bridge the MCP server's read/write streams, receives JSON-RPC messages from the client, and proxies tool calls to the FastAPI endpoints via `httpx` loopback with a service token. The SSE transport (`_MCPSseApp`, line 247) uses `SseServerTransport.connect_sse` from the MCP SDK. **Monitor WebSocket.** `/monitor/ws` (`deploy/docker/monitor_routes.py:354`) pushes JSON every 2 seconds with health stats, active requests, browser pool, timeline data, and janitor logs. The dashboard JS (`static/monitor/index.html`) implements auto-reconnect with exponential backoff and falls back to HTTP polling. **Output formats.** Results return as Python objects (library) or JSON (API): `CrawlResult.model_dump()` includes all fields. PDF bytes are base64-encoded over the wire. The schema endpoint `GET /schema` exposes `BrowserConfig`/`CrawlerRunConfig` for automated client generation.

## WebSocket/SSE Streaming Details

The server has **three** distinct streaming protocols:

**1. Crawl streaming (NDJSON).** Triggered by setting `stream=true` in `CrawlerRunConfig` or hitting `/crawl/stream`. `handle_stream_crawl_request()` (`deploy/docker/api.py:889`) creates the crawler, configures `stream=True` on the `CrawlerRunConfig`, and returns a `(crawler, async_generator, hooks_info)` tuple. `stream_process()` (`deploy/docker/server.py:1020`) wraps the generator in FastAPI's `StreamingResponse(media_type="application/x-ndjson")`. The `stream_results()` function (api.py:616) iterates the generator, serializes each `CrawlResult` via `model_dump()`, encodes base64 if PDF bytes are present, adds `server_memory_mb`, and yields one JSON line per result followed by `
`. For deep crawling, `arun()` returns an async generator directly (async_webcrawler.py:1047-1057) — the BFS strategy's `_arun_stream()` yields results as pages are crawled level-by-level.

**2. MCP WebSocket (`/mcp/ws`).** Defined in `mcp_bridge.py:197`. The handler creates two `anyio.create_memory_object_stream(100)` pairs (client→server and server→client). Messages are validated against `JSONRPCMessage` schema via pydantic's `TypeAdapter`. Three concurrent tasks run in an `anyio.TaskGroup`: `ws_to_srv()` reads JSON from the WebSocket and sends to the MCP server; `mcp.run()` processes the MCP protocol; `srv_to_ws()` reads MCP responses and sends JSON back to the WebSocket. Authentication is handled by `AuthGateMiddleware` at the ASGI level; WebSocket clients pass `?token=` in the query string.

**3. MCP SSE (`/mcp/sse`).** Uses the MCP SDK's `SseServerTransport` (mcp_bridge.py:241). A raw ASGI callable class `_MCPSseApp` avoids Starlette middleware interference with route wrappers. Messages POST to `/mcp/messages/{session_id}`. The same tool/resource handlers back both transports.

**4. Monitor WebSocket (`/monitor/ws`).** Pushes a flat JSON object every 2 seconds (monitor_routes.py:370-397) aggregating health summary, request counts, browser list, timeline data, janitor log, and error log. The JS dashboard falls back to HTTP polling if WebSocket fails.

## Agent Composition & Multi-Step Extraction Pipelines

The codebase has several composition patterns:

**C4A-Script DSL.** `crawl4ai/script/c4ai_script.py` defines a domain-specific language for browser automation. Scripts compile to JavaScript — the `Compiler` class (line 324) parses via Lark grammar, collects `PROC...ENDPROC` procedure definitions (line 355), inlines `CALL` references (line 362), substitutes `$variable` references (line 372), and emits JavaScript for each command (line 387). Commands include `GO`, `CLICK`, `WAIT`, `TYPE`, `EVAL`, `SCROLL`, `DRAG`, `IF`/`REPEAT` flow control, and `SETVAR` state. This compiles to `js_code` that runs in the Playwright browser — no core crawler modification needed. Example: a "login" procedure defined once, then called after navigation.

**Deep crawl strategy chain.** `DeepCrawlDecorator` (`crawl4ai/deep_crawling/base_strategy.py:10`) wraps `AsyncWebCrawler.arun()`: when `CrawlerRunConfig.deep_crawl_strategy` is set, the decorator intercepts the call, passes it to the strategy's `arun()`, which calls `_arun_stream()` or `_arun_batch()`. The strategies (`BFSDeepCrawlStrategy`, `DFSDeepCrawlStrategy`, `BFFDeepCrawlStrategy`) each implement link discovery with a composable `FilterChain` and optional `URLScorer`. This forms a pipeline: fetch → extract links → filter → score → queue → recurse.

**Extraction strategy pipeline.** Multiple extraction strategies can be composed with `ContentScrapingStrategy` — for example, PDF content uses `PDFCrawlerStrategy` (crawler) → `PDFContentScrapingStrategy` (scraper that downloads + parses + extracts images/links). The LLM extraction internally chains: HTML → `sanitize_html` → `merge_chunks` (splits by token budget) → `ThreadPoolExecutor` with LLM calls → `split_and_parse_json_objects` → merge results. Hooks (`deploy/docker/hook_registry.py`) provide another composition dimension: declarative actions like `block_resources`, `scroll_to_bottom`, or `set_cookie` can be chained before/after crawl phases.

## PDF Support Status (Source-Level)

PDF support is **fully implemented at the source level**, not just README claims. Key classes:

- **`PDFCrawlerStrategy`** (`crawl4ai/processors/pdf/__init__.py:49`): An `AsyncCrawlerStrategy` that returns a placeholder response (`placeholder_html=True`) — it does nothing itself, signaling to the pipeline that `PDFContentScrapingStrategy` will handle actual work.

- **`PDFContentScrapingStrategy`** (line 79): A `ContentScrapingStrategy` that downloads PDFs over HTTP(S) using `requests` with a hand-written redirect follower (`_fetch_with_redirect_checks`, line 259) that validates every hop against an injected URL/peer-IP validator for SSRF protection. It caps download size (`max_pdf_bytes`, default 100 MiB), page count (`max_pdf_pages`, default 2000), and redirect depth (`max_redirects`, default 5). The file is written to a temp directory, then processed. Output is a `ScrapingResult` with `cleaned_html` (wrapping per-page `<div class="pdf-page">` sections), plus `media[images]` and `links[urls]` each tagged with page numbers.

- **`NaivePDFProcessorStrategy`** (`crawl4ai/processors/pdf/processor.py:57`): The processing engine using `pypdf`. It extracts metadata (title, author, producer, creation/modification dates, encryption status, file size — lines 435-457), per-page text (with positional layout info via `visitor_text` callback), images (handling FlateDecode/PNG-predictor, DCTDecode/JPEG, CCITTFaxDecode/TIFF, JPXDecode/JPEG2000 — lines 254-421), and hyperlinks. Images can be saved locally or returned as base64 data URLs. Text is cleaned into both Markdown and HTML per page (via `clean_pdf_text`/`clean_pdf_text_to_html` in `utils.py`). A `process_batch()` method (line 144) uses `ThreadPoolExecutor` for parallel page processing.

- **Integration.** The Docker server imports these in `handle_crawl_request()` (api.py:703) and `handle_stream_crawl_request()` (api.py:931) — when the user's `CrawlerRunConfig.scraping_strategy` is an instance of `PDFContentScrapingStrategy`, it instantiates `PDFCrawlerStrategy` directly (not from the pool, since it has no browser) and wires the URL validator. The egress policy is injected at server boot (server.py:163-190) via `set_url_validator` and `set_peer_ip_validator`, so PDF fetches enforce the same SSRF protection as browser-based fetches. Hooks are **not supported** with PDFs (api.py:707-708: "PDFCrawlerStrategy has no browser page, so hooks can't attach to it").


Citations: [deploy/docker/api.py:616-648](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/api.py#L616-L648) · [deploy/docker/mcp_bridge.py:196-253](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/mcp_bridge.py#L196-L253) · [crawl4ai/script/c4ai_script.py:324-393](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/script/c4ai_script.py#L324-L393) · [crawl4ai/deep_crawling/base_strategy.py:10-44](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/deep_crawling/base_strategy.py#L10-L44) · [crawl4ai/processors/pdf/__init__.py:49-360](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/processors/pdf/__init__.py#L49-L360) · [crawl4ai/processors/pdf/processor.py:57-250](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/processors/pdf/processor.py#L57-L250) · [crawl4ai/async_webcrawler.py:1045-1070](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/crawl4ai/async_webcrawler.py#L1045-L1070) · [deploy/docker/server.py:1020-1055](https://github.com/unclecode/crawl4ai/blob/8afd0a68064ff7049303c9f9d037ab6228aac43c/deploy/docker/server.py#L1020-L1055)
