LLMs Technical Reviews
Home / AI web scraping / crawl4ai

unclecode/crawl4ai

Python library and server that renders pages with Playwright and turns them into Markdown or JSON through pluggable strategies.

GitHub ↗★ 85kPythonApache-2.0commit 8afd0a6 · 2026-10-05homepage ↗

Overview

Crawl4AI is a Python library for turning web pages into Markdown and structured data, built on Playwright. The core is one class, AsyncWebCrawler. You call arun(url, config) for one page or arun_many(urls, config) for a batch, and get back a CrawlResult with raw and cleaned HTML, three Markdown variants, links, media, tables, screenshots, PDFs and any extracted JSON. Almost every behaviour is a pluggable strategy: the fetcher, the HTML scraper, the Markdown generator, the content filter, the chunker, the extractor and the deep-crawl traversal.

The same package ships a crwl CLI and a FastAPI server under deploy/docker. The server wraps the library with a browser pool, a REST API, NDJSON streaming, background jobs backed by Redis, and an MCP bridge. The library is the product, and the server is a thin, security-hardened shell around it.

It suits developers who want a local, self-hosted crawler with no SaaS dependency. Defaults are conservative. Caching is bypassed, robots.txt is not checked, and anti-bot retries are off until you turn them on.

Architecture

flowchart LR
  U["Your code / crwl / REST"] --> W["AsyncWebCrawler.arun"]
  W --> DC["DeepCrawlDecorator"]
  DC --> BFS["BFS / DFS / BestFirst strategy"]
  BFS --> W
  W --> CA["SQLite cache"]
  W --> CS["AsyncPlaywrightCrawlerStrategy"]
  CS --> BM["BrowserManager + adapter"]
  W --> AB["is_blocked + proxy retry"]
  W --> AP["aprocess_html"]
  AP --> SS["LXMLWebScrapingStrategy"]
  AP --> MG["DefaultMarkdownGenerator + filter"]
  AP --> EX["ExtractionStrategy (CSS / XPath / LLM)"]
  EX --> LL["litellm"]
  M["arun_many"] --> D["MemoryAdaptiveDispatcher"]
  D --> W
Component Path Role
Crawler facade crawl4ai/async_webcrawler.py arun, arun_many, cache, robots, anti-bot retries, aprocess_html
Configs crawl4ai/async_configs.py BrowserConfig, CrawlerRunConfig, LLMConfig; serialisable for the server
Fetch strategies crawl4ai/async_crawler_strategy.py AsyncPlaywrightCrawlerStrategy (default) and AsyncHTTPCrawlerStrategy (aiohttp)
Browser layer crawl4ai/browser_manager.py, browser_adapter.py Contexts, sessions, CDP/persistent browsers; Playwright, playwright-stealth or patchright adapters
Block detection crawl4ai/antibot_detector.py is_blocked(): tiered regexes plus status and structure checks
Scraping crawl4ai/content_scraping_strategy.py LXMLWebScrapingStrategy: cleaned HTML, links, media, tables
Markdown crawl4ai/markdown_generation_strategy.py, html2text/ Raw, cited and “fit” Markdown via a bundled html2text fork
Filters crawl4ai/content_filter_strategy.py PruningContentFilter, BM25ContentFilter, LLMContentFilter
Extraction crawl4ai/extraction_strategy.py JsonCss/JsonXPath/JsonLxml, Regex, Cosine, LLMExtractionStrategy
Deep crawl crawl4ai/deep_crawling/ BFS, DFS, best-first; FilterChain and URL scorers
Dispatch crawl4ai/async_dispatcher.py MemoryAdaptiveDispatcher, SemaphoreDispatcher, per-domain RateLimiter
Server deploy/docker/server.py, api.py FastAPI endpoints, browser pool, job queue, MCP bridge

How a request flows

Take await crawler.arun("https://example.com", CrawlerRunConfig(extraction_strategy=...)):

  1. Deep-crawl check. arun is wrapped by DeepCrawlDecorator. If deep_crawl_strategy is set, the call goes to the strategy, which calls arun again for each page, guarded by a ContextVar (base_strategy.py).
  2. Cache and proxy. arun reads the SQLite cache unless cache_mode says otherwise, and can revalidate with ETag or Last-Modified. It then takes the next proxy from proxy_rotation_strategy, with optional sticky sessions (async_webcrawler.py).
  3. Robots and attempts. With check_robots_txt=True, a disallowed URL returns a synthetic 403 result. Otherwise it loops over 1 + max_retries attempts and, inside each attempt, over the proxy list (async_webcrawler.py).
  4. Fetch. AsyncPlaywrightCrawlerStrategy.crawl sends http(s) URLs to _crawl_web. file:// and raw: inputs skip the browser unless a browser-only feature is requested (async_crawler_strategy.py). _crawl_web sets a random or fixed user agent with matching sec-ch-ua, gets a page from BrowserManager, and injects the navigator overrider when magic, simulate_user or override_navigator is on (async_crawler_strategy.py). Then it navigates, runs js_code, waits, scrolls and captures.
  5. Process. aprocess_html runs the scraping strategy on the HTML (async_webcrawler.py). It picks the Markdown source (cleaned_html by default) and calls generate_markdown (async_webcrawler.py).
  6. Extract. If an extraction strategy is set, the chosen input format (Markdown, fit Markdown or HTML) is chunked and passed to the strategy’s arun, or to its sync run in a thread (async_webcrawler.py).
  7. Judge. Back in arun, is_blocked(status, html) decides whether the page is a block page. If it is, the next proxy or attempt runs. After the loop, an optional fallback_fetch_function(url) gets one last chance, and a still-blocked result is marked failed (async_webcrawler.py).

Key components

Fetching and stealth

The Playwright strategy owns hooks (before_goto, after_goto, before_return_html and others) and builds BrowserManager with use_undetected set when you pass an UndetectedAdapter (async_crawler_strategy.py). That adapter runs on patchright and evaluates scripts in an isolated world. StealthAdapter applies playwright_stealth when enable_stealth=True (browser_adapter.py, L271-L300). There is also an aiohttp-based AsyncHTTPCrawlerStrategy for pages that need no JavaScript.

Block detection and retries

is_blocked is a layered heuristic. Any 429 counts as a block. “Tier 1” vendor markers (Cloudflare, Akamai, PerimeterX, DataDome, Imperva, Kasada and others) are searched in the first 15 KB and in a script-stripped copy of large pages. A 403 or 503 on a non-data page counts as blocked, and so do a near-empty 200 and pages that fail a structural check (antibot_detector.py). The retry loop only acts on this when you set max_retries, a proxy list or fallback_fetch_function. All three default to off (async_configs.py).

Markdown and filtering

DefaultMarkdownGenerator converts HTML with a bundled CustomHTML2Text, rewrites links into numbered citations, and, when a content filter is attached, produces fit_markdown from the filtered HTML (markdown_generation_strategy.py). PruningContentFilter scores nodes by text density and link ratio. BM25ContentFilter keeps blocks relevant to a query. LLMContentFilter asks a model to keep what matters.

Extraction

Non-LLM extractors take a schema with a baseSelector and fields (CSS, XPath or lxml). generate_schema can ask an LLM to write that schema once, so later runs need no model. LLMExtractionStrategy merges sections into chunks by a word-to-token ratio with overlap. The async path sends every chunk at once with asyncio.gather (extraction_strategy.py). The sync path uses four threads, or runs sequentially for Groq (extraction_strategy.py). Each chunk prompt picks a block, instruction or schema template. Calls go through litellm with backoff, token usage is summed, and output is parsed as JSON or <blocks> XML with a lenient fallback (extraction_strategy.py). config.py lists known provider prefixes and their key variables, and defaults to openai/gpt-4o (config.py).

Deep crawling and dispatch

BFSDeepCrawlStrategy.link_discovery normalizes each link, skips visited ones, runs the FilterChain, scores with an optional scorer, and trims to the remaining max_pages budget, best scores first (bfs_strategy.py). arun_many uses MemoryAdaptiveDispatcher by default. It pauses when system memory crosses a threshold and can attach a per-domain RateLimiter that backs off on 429 and 503 (async_dispatcher.py). The state is in-process. There is no distributed queue in the library.

Docker server

/crawl loads browser and run configs with Provenance.UNTRUSTED, applies an egress policy and deep-crawl clamps, then takes a browser from the pool (server.py, api.py). The pool keeps one permanent browser plus hot and cold pools keyed by config signature, and recycles them by idle time and memory (crawler_pool.py). Background /crawl/job and /llm/job run on a bounded in-process worker pool with per-caller quotas. Redis stores the job status (work_queue.py). attach_mcp exposes the endpoints as MCP tools over WebSocket and SSE, calling back into the API over loopback (server.py).

Extending it

  • Strategies. Subclass AsyncCrawlerStrategy, ContentScrapingStrategy, MarkdownGenerationStrategy, RelevantContentFilter, ChunkingStrategy, ExtractionStrategy, DeepCrawlStrategy or the URL filters and scorers, then pass them in the config.
  • Hooks. crawler.crawler_strategy.set_hook("before_goto", fn) and the other hook points give you the Playwright page and context. The server offers declarative hooks instead of arbitrary code by default.
  • Scripting. C4A-Script (crawl4ai/script/) compiles a small DSL (GO, CLICK, WAIT, IF, PROC) to JavaScript for js_code.
  • Models. Any litellm provider string works through LLMConfig(provider=..., api_token=..., base_url=...).

Running it

  • Library. pip install crawl4ai, then crawl4ai-setup to install the Playwright browsers. crawl4ai-doctor checks the setup. The cache lives in ~/.crawl4ai/crawl4ai.db (SQLite).
  • CLI. crwl <url> crawls and prints Markdown. Subcommands manage a persistent browser and CDP connections.
  • Server. The Dockerfile and docker-compose.yml run the FastAPI app with Redis under supervisord. Settings live in deploy/docker/config.yml (rate limits, pool sizes, memory threshold, JWT auth).

Strengths and caveats

  • Strength: composable design. Every stage is a strategy object, so swapping the Markdown filter or extractor needs no fork.
  • Strength: LLM-optional. CSS, XPath and regex extractors plus generate_schema let you pay for a model once, not on every page.
  • Strength: serious anti-bot plumbing. Block detection, proxy cascades, a fallback fetcher and three browser adapters cover much more than a plain Playwright wrapper.
  • Caveat: surface area. CrawlerRunConfig has well over a hundred parameters, and legacy kwargs are still accepted. Behaviour depends on flag combinations that are easy to get wrong.
  • Caveat: politeness is opt-in. robots.txt checks, rate limiting and retries are all off by default.
  • Caveat: LLM chunk fan-out. The async extractor fires every chunk at once, which can hit provider rate limits on long pages.
  • Caveat: single-node. Concurrency, dedup and queues are in-process. Scaling out means running several servers behind your own queue.

Sources: code at 8afd0a6, deepwiki-open wiki (13 pages), OpenDeepWiki wiki (35 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 8afd0a6. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Plain HTTP vs browser. AsyncPlaywrightCrawlerStrategy (crawl4ai/async_crawler_strategy.py:46) is the default, using Playwright to launch Chromium. The BrowserConfig (proxy, headless, extra args) is passed to BrowserManager which manages the browser lifecycle via PlaywrightAdapter. A plain-HTTP path exists via the HTTPCrawlerConfig type and the fallback aiohttp-based fetch method in the same module. JS rendering. Every page is loaded in a real Playwright browser context; BrowserAdapter (crawl4ai/browser_adapter.py) wraps the Playwright page and runs user-supplied js_code before extraction. Waiting strategy. CrawlerRunConfig supports wait_for, wait_until, wait_for_images, screenshot_wait_for; the page is waited on per these directives before extraction begins. The C4A-Script DSL (crawl4ai/script/c4ai_script.py:139) adds fine-grained WAIT commands (by seconds, CSS selector, or text content) that compile to JavaScript polling loops. Content types. PDFs use a separate pipeline: PDFCrawlerStrategy (crawl4ai/processors/pdf/__init__.py:49) returns a placeholder response, then PDFContentScrapingStrategy (line 79) downloads the PDF with requests, validates every redirect hop for SSRF, and processes it with NaivePDFProcessorStrategy. Screenshots and PDF snapshots of rendered pages are produced via Playwright's native screenshot/PDF methods (deploy/docker/server.py:742-818).

How is content extracted or converted?

answered

HTML→Markdown/text. LXMLWebScrapingStrategy (crawl4ai/content_scraping_strategy.py) is the default scraper; it lxml-parses the DOM and produces clean HTML. The DefaultMarkdownGenerator (crawl4ai/markdown_generation_strategy.py) converts that to Markdown in three variants: raw_markdown (full), fit_markdown (filtered), and markdown_with_citations. Readability-style filtering. Three content filters (crawl4ai/content_filter_strategy.py) sit inside the generator: PruningContentFilter (heuristic node pruning based on text density/link ratio), BM25ContentFilter (relevance-scoring against a user query), and LLMContentFilter (LLM-summarizes the page per an instruction). A faster C-ext variant PruningContentFilterLXML lives in content_filter_strategy_lxml.py. Selectors. CrawlerRunConfig accepts css_selector for targeted extraction. For schema-based extraction, JsonCssExtractionStrategy and JsonXPathExtractionStrategy (crawl4ai/extraction_strategy.py:1989,2449) extract structured fields from HTML using CSS selectors or XPath expressions. Schema-based extraction. LLMExtractionStrategy (line 533) accepts a JSON schema dict and an instruction; it prompts the LLM to return JSON matching that schema. RegexExtractionStrategy (line 2558) applies regex patterns. JsonElementExtractionStrategy (line 1043) is the base for CSS/XPath/LXML structured extraction — you define a baseSelector and per-field selectors, and it iterates matches.

How are LLMs used, if at all?

answered

LLM extraction. LLMExtractionStrategy (crawl4ai/extraction_strategy.py:533) orchestrates all LLM calls. Its extract() method (line 641) builds a prompt from templates in crawl4ai/prompts.py, sends it to the LLM via perform_completion_with_backoff (from crawl4ai/utils.py), parses the response (XML <blocks> or raw JSON), and tracks token usage. Chunking. The strategy chunks pages that exceed CHUNK_TOKEN_THRESHOLD (defined in crawl4ai/config.py). The run() method (line 786) calls _merge() → merge_chunks() (crawl4ai/utils.py:170) which splits documents by token estimate (word_token_ratio) with configurable overlap_rate. Sections are then dispatched to the LLM via ThreadPoolExecutor (max 4 workers) for parallel processing. Schema support. When extraction_type="schema" and a schema dict is provided, the prompt switches to PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION, which instructs the LLM to return JSON matching the schema. force_json_response=True uses the provider's JSON-mode if available. Providers. Supported providers and their API keys are registered in crawl4ai/config.py:11-32, including OpenAI (gpt-4o, o1, o3-mini), Anthropic (claude-3-*), Gemini (gemini-pro, 2.0-flash), Groq, Ollama, and DeepSeek. The LLMConfig type (crawl4ai/types.py) holds provider/Api_token/temperature/base_url. Cost tracking. Each extraction records a TokenUsage object with prompt/completion tokens; total_usage accumulates across chunks.

Editor's note. Correction: on the normal async path LLMExtractionStrategy.arun sends all chunks at once with asyncio.gather; the 4-thread ThreadPoolExecutor is only the sync run(). Calls go through litellm, so any litellm provider string works, not only those in config.py.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Stealth patches. UndetectedAdapter (crawl4ai/browser_adapter.py) wraps Playwright to apply fingerprinting countermeasures — it can be selected via BrowserConfig(browser_type="undetected"). Examples in docs/examples/simple_anti_bot_examples.py demonstrate bypassing Cloudflare, DataDome, and generic bot detection. Fingerprint spoofing. crawl4ai/user_agent_generator.py provides ValidUAGenerator and UAGen for realistic User-Agent rotation. The Docker server also exposes config knobs for browser extra args and language headers. Proxy rotation. crawl4ai/proxy_strategy.py implements the ProxyRotationStrategy base class; the Docker server supports per-request proxy config through BrowserConfig.proxy_config. The auth_proxy_example.py and nstproxy_example.py in docs/examples/ show authenticated proxy setup. CAPTCHA handling. The docs/examples/capsolver_captcha_solver/ directory has production examples for solving Cloudflare Turnstile/Challenge, reCAPTCHA v2/v3, and AWS WAF via both Capsolver extension integration and API integration. These work by injecting the capsolver extension into Playwright. Rate limiting. RateLimiter (crawl4ai/async_dispatcher.py:28) implements per-domain exponential backoff on 429/503 responses, configurable base delay, max delay, and max retries. The dispatcher (MemoryAdaptiveDispatcher) also monitors system memory and pauses crawling when thresholds are exceeded.

Editor's note. Correction: UndetectedAdapter is chosen by passing it as browser_adapter to AsyncPlaywrightCrawlerStrategy (patchright), not via browser_type. The answer also misses the main mechanism: antibot_detector.is_blocked() classifies block pages, and arun() retries over max_retries and a proxy list, then calls an optional fallback_fetch_function; all of it is off by default.

How is crawling at scale implemented?

answered

Queues. async_dispatcher.py provides MemoryAdaptiveDispatcher which maintains a domain-state map for per-domain rate limiting. The Docker server adds a Redis-backed WorkQueue (deploy/docker/work_queue.py) with per-principal quotas and bounded-size queues (HTTP 429/503). Concurrency. CrawlerRunConfig.semaphore_count limits parallel page loads per dispatcher. MemoryAdaptiveDispatcher also monitors system RAM via psutil and pauses fetching when memory_threshold_percent is exceeded. URL dedup. BFSDeepCrawlStrategy.link_discovery() (crawl4ai/deep_crawling/bfs_strategy.py:133) maintains a visited: Set[str] and checks can_process_url() before adding new URLs. URLs are normalized via normalize_url_for_deep_crawl() before dedup. Depth/limits. BFSDeepCrawlStrategy accepts max_depth, max_pages, score_threshold, and include_external. The FilterChain (crawl4ai/deep_crawling/filters.py) composes URL filters (domain, path, content-type). Robots.txt/politeness. The RateLimiter.wait_if_needed() method enforces per-domain delays between requests. Crawl strategies accept should_cancel callbacks for graceful interruption. Distributed workers. The Docker deployment (deploy/docker/server.py) manages a pool of browser instances via BrowserPool (crawler_pool.py). It supports crawler_configs — per-URL custom configs with url_matcher patterns for arun_many(). Webhook delivery (deploy/docker/webhook.py) notifies external services on completion/failure. Crash recovery is supported via resume_state callback in the deep crawl strategies.

Editor's note. Correction: the server's WorkQueue is an in-process asyncio queue with per-caller caps (Redis only stores job status). robots.txt is supported: CrawlerRunConfig.check_robots_txt=True makes arun return a 403 result for disallowed URLs, but it defaults to False.

What is the developer interface?

answered

Library API. AsyncWebCrawler (crawl4ai/async_webcrawler.py) is the main class, with arun() for single URLs and arun_many() for batches. Both accept CrawlerRunConfig and return CrawlResult objects containing html, markdown (raw/fit/citations), extracted_content, screenshot, pdf, metadata, and more. CLI. crawl4ai/cli.py provides a command-line interface; crawl4ai/cloud/cli.py handles cloud deployments. REST API. The Docker server (deploy/docker/server.py) exposes POST endpoints: /crawl (JSON), /crawl/stream (NDJSON), /md (Markdown), /llm (QA), /screenshot, /pdf, /html, /execute_js, /config/dump, /ask (BM25 context retrieval), and /schema. Streaming protocol. /crawl/stream returns a StreamingResponse with application/x-ndjson content type. Each result is a JSON line via stream_results() (deploy/docker/api.py:616), which serializes each CrawlResult as it arrives from the async generator, includes a heartbeat mechanism, and finishes with {"status":"completed"}. Deep-crawl streaming (crawl4ai/async_webcrawler.py:1047-1057) wraps the BFS async generator in the same NDJSON format. MCP server. deploy/docker/mcp_bridge.py attaches at /mcp/ws (WebSocket) and /mcp/sse (SSE) transports. The WebSocket handler (_ws, line 198) uses anyio.create_memory_object_stream to bridge the MCP server's read/write streams, receives JSON-RPC messages from the client, and proxies tool calls to the FastAPI endpoints via httpx loopback with a service token. The SSE transport (_MCPSseApp, line 247) uses SseServerTransport.connect_sse from the MCP SDK. Monitor WebSocket. /monitor/ws (deploy/docker/monitor_routes.py:354) pushes JSON every 2 seconds with health stats, active requests, browser pool, timeline data, and janitor logs. The dashboard JS (static/monitor/index.html) implements auto-reconnect with exponential backoff and falls back to HTTP polling. Output formats. Results return as Python objects (library) or JSON (API): CrawlResult.model_dump() includes all fields. PDF bytes are base64-encoded over the wire. The schema endpoint GET /schema exposes BrowserConfig/CrawlerRunConfig for automated client generation.

WebSocket/SSE Streaming Details

The server has three distinct streaming protocols:

1. Crawl streaming (NDJSON). Triggered by setting stream=true in CrawlerRunConfig or hitting /crawl/stream. handle_stream_crawl_request() (deploy/docker/api.py:889) creates the crawler, configures stream=True on the CrawlerRunConfig, and returns a (crawler, async_generator, hooks_info) tuple. stream_process() (deploy/docker/server.py:1020) wraps the generator in FastAPI's StreamingResponse(media_type="application/x-ndjson"). The stream_results() function (api.py:616) iterates the generator, serializes each CrawlResult via model_dump(), encodes base64 if PDF bytes are present, adds server_memory_mb, and yields one JSON line per result followed by . For deep crawling, arun() returns an async generator directly (async_webcrawler.py:1047-1057) — the BFS strategy's _arun_stream() yields results as pages are crawled level-by-level.

2. MCP WebSocket (/mcp/ws). Defined in mcp_bridge.py:197. The handler creates two anyio.create_memory_object_stream(100) pairs (client→server and server→client). Messages are validated against JSONRPCMessage schema via pydantic's TypeAdapter. Three concurrent tasks run in an anyio.TaskGroup: ws_to_srv() reads JSON from the WebSocket and sends to the MCP server; mcp.run() processes the MCP protocol; srv_to_ws() reads MCP responses and sends JSON back to the WebSocket. Authentication is handled by AuthGateMiddleware at the ASGI level; WebSocket clients pass ?token= in the query string.

3. MCP SSE (/mcp/sse). Uses the MCP SDK's SseServerTransport (mcp_bridge.py:241). A raw ASGI callable class _MCPSseApp avoids Starlette middleware interference with route wrappers. Messages POST to /mcp/messages/{session_id}. The same tool/resource handlers back both transports.

4. Monitor WebSocket (/monitor/ws). Pushes a flat JSON object every 2 seconds (monitor_routes.py:370-397) aggregating health summary, request counts, browser list, timeline data, janitor log, and error log. The JS dashboard falls back to HTTP polling if WebSocket fails.

Agent Composition & Multi-Step Extraction Pipelines

The codebase has several composition patterns:

C4A-Script DSL. crawl4ai/script/c4ai_script.py defines a domain-specific language for browser automation. Scripts compile to JavaScript — the Compiler class (line 324) parses via Lark grammar, collects PROC...ENDPROC procedure definitions (line 355), inlines CALL references (line 362), substitutes $variable references (line 372), and emits JavaScript for each command (line 387). Commands include GO, CLICK, WAIT, TYPE, EVAL, SCROLL, DRAG, IF/REPEAT flow control, and SETVAR state. This compiles to js_code that runs in the Playwright browser — no core crawler modification needed. Example: a "login" procedure defined once, then called after navigation.

Deep crawl strategy chain. DeepCrawlDecorator (crawl4ai/deep_crawling/base_strategy.py:10) wraps AsyncWebCrawler.arun(): when CrawlerRunConfig.deep_crawl_strategy is set, the decorator intercepts the call, passes it to the strategy's arun(), which calls _arun_stream() or _arun_batch(). The strategies (BFSDeepCrawlStrategy, DFSDeepCrawlStrategy, BFFDeepCrawlStrategy) each implement link discovery with a composable FilterChain and optional URLScorer. This forms a pipeline: fetch → extract links → filter → score → queue → recurse.

Extraction strategy pipeline. Multiple extraction strategies can be composed with ContentScrapingStrategy — for example, PDF content uses PDFCrawlerStrategy (crawler) → PDFContentScrapingStrategy (scraper that downloads + parses + extracts images/links). The LLM extraction internally chains: HTML → sanitize_html → merge_chunks (splits by token budget) → ThreadPoolExecutor with LLM calls → split_and_parse_json_objects → merge results. Hooks (deploy/docker/hook_registry.py) provide another composition dimension: declarative actions like block_resources, scroll_to_bottom, or set_cookie can be chained before/after crawl phases.

PDF Support Status (Source-Level)

PDF support is fully implemented at the source level, not just README claims. Key classes:

  • PDFCrawlerStrategy (crawl4ai/processors/pdf/__init__.py:49): An AsyncCrawlerStrategy that returns a placeholder response (placeholder_html=True) — it does nothing itself, signaling to the pipeline that PDFContentScrapingStrategy will handle actual work.

  • PDFContentScrapingStrategy (line 79): A ContentScrapingStrategy that downloads PDFs over HTTP(S) using requests with a hand-written redirect follower (_fetch_with_redirect_checks, line 259) that validates every hop against an injected URL/peer-IP validator for SSRF protection. It caps download size (max_pdf_bytes, default 100 MiB), page count (max_pdf_pages, default 2000), and redirect depth (max_redirects, default 5). The file is written to a temp directory, then processed. Output is a ScrapingResult with cleaned_html (wrapping per-page <div class="pdf-page"> sections), plus media[images] and links[urls] each tagged with page numbers.

  • NaivePDFProcessorStrategy (crawl4ai/processors/pdf/processor.py:57): The processing engine using pypdf. It extracts metadata (title, author, producer, creation/modification dates, encryption status, file size — lines 435-457), per-page text (with positional layout info via visitor_text callback), images (handling FlateDecode/PNG-predictor, DCTDecode/JPEG, CCITTFaxDecode/TIFF, JPXDecode/JPEG2000 — lines 254-421), and hyperlinks. Images can be saved locally or returned as base64 data URLs. Text is cleaned into both Markdown and HTML per page (via clean_pdf_text/clean_pdf_text_to_html in utils.py). A process_batch() method (line 144) uses ThreadPoolExecutor for parallel page processing.

  • Integration. The Docker server imports these in handle_crawl_request() (api.py:703) and handle_stream_crawl_request() (api.py:931) — when the user's CrawlerRunConfig.scraping_strategy is an instance of PDFContentScrapingStrategy, it instantiates PDFCrawlerStrategy directly (not from the pool, since it has no browser) and wires the URL validator. The egress policy is injected at server boot (server.py:163-190) via set_url_validator and set_peer_ip_validator, so PDF fetches enforce the same SSRF protection as browser-based fetches. Hooks are not supported with PDFs (api.py:707-708: "PDFCrawlerStrategy has no browser page, so hooks can't attach to it").