unclecode/crawl4ai
Python library and server that renders pages with Playwright and turns them into Markdown or JSON through pluggable strategies.
Overview
Crawl4AI is a Python library for turning web pages into Markdown and structured data, built on Playwright. The core is one class, AsyncWebCrawler. You call arun(url, config) for one page or arun_many(urls, config) for a batch, and get back a CrawlResult with raw and cleaned HTML, three Markdown variants, links, media, tables, screenshots, PDFs and any extracted JSON. Almost every behaviour is a pluggable strategy: the fetcher, the HTML scraper, the Markdown generator, the content filter, the chunker, the extractor and the deep-crawl traversal.
The same package ships a crwl CLI and a FastAPI server under deploy/docker. The server wraps the library with a browser pool, a REST API, NDJSON streaming, background jobs backed by Redis, and an MCP bridge. The library is the product, and the server is a thin, security-hardened shell around it.
It suits developers who want a local, self-hosted crawler with no SaaS dependency. Defaults are conservative. Caching is bypassed, robots.txt is not checked, and anti-bot retries are off until you turn them on.
Architecture
flowchart LR
U["Your code / crwl / REST"] --> W["AsyncWebCrawler.arun"]
W --> DC["DeepCrawlDecorator"]
DC --> BFS["BFS / DFS / BestFirst strategy"]
BFS --> W
W --> CA["SQLite cache"]
W --> CS["AsyncPlaywrightCrawlerStrategy"]
CS --> BM["BrowserManager + adapter"]
W --> AB["is_blocked + proxy retry"]
W --> AP["aprocess_html"]
AP --> SS["LXMLWebScrapingStrategy"]
AP --> MG["DefaultMarkdownGenerator + filter"]
AP --> EX["ExtractionStrategy (CSS / XPath / LLM)"]
EX --> LL["litellm"]
M["arun_many"] --> D["MemoryAdaptiveDispatcher"]
D --> W
| Component | Path | Role |
|---|---|---|
| Crawler facade | crawl4ai/async_webcrawler.py |
arun, arun_many, cache, robots, anti-bot retries, aprocess_html |
| Configs | crawl4ai/async_configs.py |
BrowserConfig, CrawlerRunConfig, LLMConfig; serialisable for the server |
| Fetch strategies | crawl4ai/async_crawler_strategy.py |
AsyncPlaywrightCrawlerStrategy (default) and AsyncHTTPCrawlerStrategy (aiohttp) |
| Browser layer | crawl4ai/browser_manager.py, browser_adapter.py |
Contexts, sessions, CDP/persistent browsers; Playwright, playwright-stealth or patchright adapters |
| Block detection | crawl4ai/antibot_detector.py |
is_blocked(): tiered regexes plus status and structure checks |
| Scraping | crawl4ai/content_scraping_strategy.py |
LXMLWebScrapingStrategy: cleaned HTML, links, media, tables |
| Markdown | crawl4ai/markdown_generation_strategy.py, html2text/ |
Raw, cited and “fit” Markdown via a bundled html2text fork |
| Filters | crawl4ai/content_filter_strategy.py |
PruningContentFilter, BM25ContentFilter, LLMContentFilter |
| Extraction | crawl4ai/extraction_strategy.py |
JsonCss/JsonXPath/JsonLxml, Regex, Cosine, LLMExtractionStrategy |
| Deep crawl | crawl4ai/deep_crawling/ |
BFS, DFS, best-first; FilterChain and URL scorers |
| Dispatch | crawl4ai/async_dispatcher.py |
MemoryAdaptiveDispatcher, SemaphoreDispatcher, per-domain RateLimiter |
| Server | deploy/docker/server.py, api.py |
FastAPI endpoints, browser pool, job queue, MCP bridge |
How a request flows
Take await crawler.arun("https://example.com", CrawlerRunConfig(extraction_strategy=...)):
- Deep-crawl check.
arunis wrapped byDeepCrawlDecorator. Ifdeep_crawl_strategyis set, the call goes to the strategy, which callsarunagain for each page, guarded by aContextVar(base_strategy.py). - Cache and proxy.
arunreads the SQLite cache unlesscache_modesays otherwise, and can revalidate with ETag or Last-Modified. It then takes the next proxy fromproxy_rotation_strategy, with optional sticky sessions (async_webcrawler.py). - Robots and attempts. With
check_robots_txt=True, a disallowed URL returns a synthetic 403 result. Otherwise it loops over1 + max_retriesattempts and, inside each attempt, over the proxy list (async_webcrawler.py). - Fetch.
AsyncPlaywrightCrawlerStrategy.crawlsendshttp(s)URLs to_crawl_web.file://andraw:inputs skip the browser unless a browser-only feature is requested (async_crawler_strategy.py)._crawl_websets a random or fixed user agent with matchingsec-ch-ua, gets a page fromBrowserManager, and injects the navigator overrider whenmagic,simulate_useroroverride_navigatoris on (async_crawler_strategy.py). Then it navigates, runsjs_code, waits, scrolls and captures. - Process.
aprocess_htmlruns the scraping strategy on the HTML (async_webcrawler.py). It picks the Markdown source (cleaned_htmlby default) and callsgenerate_markdown(async_webcrawler.py). - Extract. If an extraction strategy is set, the chosen input format (Markdown, fit Markdown or HTML) is chunked and passed to the strategy’s
arun, or to its syncrunin a thread (async_webcrawler.py). - Judge. Back in
arun,is_blocked(status, html)decides whether the page is a block page. If it is, the next proxy or attempt runs. After the loop, an optionalfallback_fetch_function(url)gets one last chance, and a still-blocked result is marked failed (async_webcrawler.py).
Key components
Fetching and stealth
The Playwright strategy owns hooks (before_goto, after_goto, before_return_html and others) and builds BrowserManager with use_undetected set when you pass an UndetectedAdapter (async_crawler_strategy.py). That adapter runs on patchright and evaluates scripts in an isolated world. StealthAdapter applies playwright_stealth when enable_stealth=True (browser_adapter.py, L271-L300). There is also an aiohttp-based AsyncHTTPCrawlerStrategy for pages that need no JavaScript.
Block detection and retries
is_blocked is a layered heuristic. Any 429 counts as a block. “Tier 1” vendor markers (Cloudflare, Akamai, PerimeterX, DataDome, Imperva, Kasada and others) are searched in the first 15 KB and in a script-stripped copy of large pages. A 403 or 503 on a non-data page counts as blocked, and so do a near-empty 200 and pages that fail a structural check (antibot_detector.py). The retry loop only acts on this when you set max_retries, a proxy list or fallback_fetch_function. All three default to off (async_configs.py).
Markdown and filtering
DefaultMarkdownGenerator converts HTML with a bundled CustomHTML2Text, rewrites links into numbered citations, and, when a content filter is attached, produces fit_markdown from the filtered HTML (markdown_generation_strategy.py). PruningContentFilter scores nodes by text density and link ratio. BM25ContentFilter keeps blocks relevant to a query. LLMContentFilter asks a model to keep what matters.
Extraction
Non-LLM extractors take a schema with a baseSelector and fields (CSS, XPath or lxml). generate_schema can ask an LLM to write that schema once, so later runs need no model. LLMExtractionStrategy merges sections into chunks by a word-to-token ratio with overlap. The async path sends every chunk at once with asyncio.gather (extraction_strategy.py). The sync path uses four threads, or runs sequentially for Groq (extraction_strategy.py). Each chunk prompt picks a block, instruction or schema template. Calls go through litellm with backoff, token usage is summed, and output is parsed as JSON or <blocks> XML with a lenient fallback (extraction_strategy.py). config.py lists known provider prefixes and their key variables, and defaults to openai/gpt-4o (config.py).
Deep crawling and dispatch
BFSDeepCrawlStrategy.link_discovery normalizes each link, skips visited ones, runs the FilterChain, scores with an optional scorer, and trims to the remaining max_pages budget, best scores first (bfs_strategy.py). arun_many uses MemoryAdaptiveDispatcher by default. It pauses when system memory crosses a threshold and can attach a per-domain RateLimiter that backs off on 429 and 503 (async_dispatcher.py). The state is in-process. There is no distributed queue in the library.
Docker server
/crawl loads browser and run configs with Provenance.UNTRUSTED, applies an egress policy and deep-crawl clamps, then takes a browser from the pool (server.py, api.py). The pool keeps one permanent browser plus hot and cold pools keyed by config signature, and recycles them by idle time and memory (crawler_pool.py). Background /crawl/job and /llm/job run on a bounded in-process worker pool with per-caller quotas. Redis stores the job status (work_queue.py). attach_mcp exposes the endpoints as MCP tools over WebSocket and SSE, calling back into the API over loopback (server.py).
Extending it
- Strategies. Subclass
AsyncCrawlerStrategy,ContentScrapingStrategy,MarkdownGenerationStrategy,RelevantContentFilter,ChunkingStrategy,ExtractionStrategy,DeepCrawlStrategyor the URL filters and scorers, then pass them in the config. - Hooks.
crawler.crawler_strategy.set_hook("before_goto", fn)and the other hook points give you the Playwright page and context. The server offers declarative hooks instead of arbitrary code by default. - Scripting. C4A-Script (
crawl4ai/script/) compiles a small DSL (GO,CLICK,WAIT,IF,PROC) to JavaScript forjs_code. - Models. Any
litellmprovider string works throughLLMConfig(provider=..., api_token=..., base_url=...).
Running it
- Library.
pip install crawl4ai, thencrawl4ai-setupto install the Playwright browsers.crawl4ai-doctorchecks the setup. The cache lives in~/.crawl4ai/crawl4ai.db(SQLite). - CLI.
crwl <url>crawls and prints Markdown. Subcommands manage a persistent browser and CDP connections. - Server. The Dockerfile and
docker-compose.ymlrun the FastAPI app with Redis under supervisord. Settings live indeploy/docker/config.yml(rate limits, pool sizes, memory threshold, JWT auth).
Strengths and caveats
- Strength: composable design. Every stage is a strategy object, so swapping the Markdown filter or extractor needs no fork.
- Strength: LLM-optional. CSS, XPath and regex extractors plus
generate_schemalet you pay for a model once, not on every page. - Strength: serious anti-bot plumbing. Block detection, proxy cascades, a fallback fetcher and three browser adapters cover much more than a plain Playwright wrapper.
- Caveat: surface area.
CrawlerRunConfighas well over a hundred parameters, and legacy kwargs are still accepted. Behaviour depends on flag combinations that are easy to get wrong. - Caveat: politeness is opt-in. robots.txt checks, rate limiting and retries are all off by default.
- Caveat: LLM chunk fan-out. The async extractor fires every chunk at once, which can hit provider rate limits on long pages.
- Caveat: single-node. Concurrency, dedup and queues are in-process. Scaling out means running several servers behind your own queue.
Sources: code at 8afd0a6, deepwiki-open wiki (13 pages), OpenDeepWiki wiki (35 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 8afd0a6. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPlain HTTP vs browser. AsyncPlaywrightCrawlerStrategy (crawl4ai/async_crawler_strategy.py:46) is the default, using Playwright to launch Chromium. The BrowserConfig (proxy, headless, extra args) is passed to BrowserManager which manages the browser lifecycle via PlaywrightAdapter. A plain-HTTP path exists via the HTTPCrawlerConfig type and the fallback aiohttp-based fetch method in the same module. JS rendering. Every page is loaded in a real Playwright browser context; BrowserAdapter (crawl4ai/browser_adapter.py) wraps the Playwright page and runs user-supplied js_code before extraction. Waiting strategy. CrawlerRunConfig supports wait_for, wait_until, wait_for_images, screenshot_wait_for; the page is waited on per these directives before extraction begins. The C4A-Script DSL (crawl4ai/script/c4ai_script.py:139) adds fine-grained WAIT commands (by seconds, CSS selector, or text content) that compile to JavaScript polling loops. Content types. PDFs use a separate pipeline: PDFCrawlerStrategy (crawl4ai/processors/pdf/__init__.py:49) returns a placeholder response, then PDFContentScrapingStrategy (line 79) downloads the PDF with requests, validates every redirect hop for SSRF, and processes it with NaivePDFProcessorStrategy. Screenshots and PDF snapshots of rendered pages are produced via Playwright's native screenshot/PDF methods (deploy/docker/server.py:742-818).
How is content extracted or converted?
answeredHTML→Markdown/text. LXMLWebScrapingStrategy (crawl4ai/content_scraping_strategy.py) is the default scraper; it lxml-parses the DOM and produces clean HTML. The DefaultMarkdownGenerator (crawl4ai/markdown_generation_strategy.py) converts that to Markdown in three variants: raw_markdown (full), fit_markdown (filtered), and markdown_with_citations. Readability-style filtering. Three content filters (crawl4ai/content_filter_strategy.py) sit inside the generator: PruningContentFilter (heuristic node pruning based on text density/link ratio), BM25ContentFilter (relevance-scoring against a user query), and LLMContentFilter (LLM-summarizes the page per an instruction). A faster C-ext variant PruningContentFilterLXML lives in content_filter_strategy_lxml.py. Selectors. CrawlerRunConfig accepts css_selector for targeted extraction. For schema-based extraction, JsonCssExtractionStrategy and JsonXPathExtractionStrategy (crawl4ai/extraction_strategy.py:1989,2449) extract structured fields from HTML using CSS selectors or XPath expressions. Schema-based extraction. LLMExtractionStrategy (line 533) accepts a JSON schema dict and an instruction; it prompts the LLM to return JSON matching that schema. RegexExtractionStrategy (line 2558) applies regex patterns. JsonElementExtractionStrategy (line 1043) is the base for CSS/XPath/LXML structured extraction — you define a baseSelector and per-field selectors, and it iterates matches.
How are LLMs used, if at all?
answeredLLM extraction. LLMExtractionStrategy (crawl4ai/extraction_strategy.py:533) orchestrates all LLM calls. Its extract() method (line 641) builds a prompt from templates in crawl4ai/prompts.py, sends it to the LLM via perform_completion_with_backoff (from crawl4ai/utils.py), parses the response (XML <blocks> or raw JSON), and tracks token usage. Chunking. The strategy chunks pages that exceed CHUNK_TOKEN_THRESHOLD (defined in crawl4ai/config.py). The run() method (line 786) calls _merge() → merge_chunks() (crawl4ai/utils.py:170) which splits documents by token estimate (word_token_ratio) with configurable overlap_rate. Sections are then dispatched to the LLM via ThreadPoolExecutor (max 4 workers) for parallel processing. Schema support. When extraction_type="schema" and a schema dict is provided, the prompt switches to PROMPT_EXTRACT_SCHEMA_WITH_INSTRUCTION, which instructs the LLM to return JSON matching the schema. force_json_response=True uses the provider's JSON-mode if available. Providers. Supported providers and their API keys are registered in crawl4ai/config.py:11-32, including OpenAI (gpt-4o, o1, o3-mini), Anthropic (claude-3-*), Gemini (gemini-pro, 2.0-flash), Groq, Ollama, and DeepSeek. The LLMConfig type (crawl4ai/types.py) holds provider/Api_token/temperature/base_url. Cost tracking. Each extraction records a TokenUsage object with prompt/completion tokens; total_usage accumulates across chunks.
How are anti-bot measures, proxies and fingerprinting handled?
answeredStealth patches. UndetectedAdapter (crawl4ai/browser_adapter.py) wraps Playwright to apply fingerprinting countermeasures — it can be selected via BrowserConfig(browser_type="undetected"). Examples in docs/examples/simple_anti_bot_examples.py demonstrate bypassing Cloudflare, DataDome, and generic bot detection. Fingerprint spoofing. crawl4ai/user_agent_generator.py provides ValidUAGenerator and UAGen for realistic User-Agent rotation. The Docker server also exposes config knobs for browser extra args and language headers. Proxy rotation. crawl4ai/proxy_strategy.py implements the ProxyRotationStrategy base class; the Docker server supports per-request proxy config through BrowserConfig.proxy_config. The auth_proxy_example.py and nstproxy_example.py in docs/examples/ show authenticated proxy setup. CAPTCHA handling. The docs/examples/capsolver_captcha_solver/ directory has production examples for solving Cloudflare Turnstile/Challenge, reCAPTCHA v2/v3, and AWS WAF via both Capsolver extension integration and API integration. These work by injecting the capsolver extension into Playwright. Rate limiting. RateLimiter (crawl4ai/async_dispatcher.py:28) implements per-domain exponential backoff on 429/503 responses, configurable base delay, max delay, and max retries. The dispatcher (MemoryAdaptiveDispatcher) also monitors system memory and pauses crawling when thresholds are exceeded.
How is crawling at scale implemented?
answeredQueues. async_dispatcher.py provides MemoryAdaptiveDispatcher which maintains a domain-state map for per-domain rate limiting. The Docker server adds a Redis-backed WorkQueue (deploy/docker/work_queue.py) with per-principal quotas and bounded-size queues (HTTP 429/503). Concurrency. CrawlerRunConfig.semaphore_count limits parallel page loads per dispatcher. MemoryAdaptiveDispatcher also monitors system RAM via psutil and pauses fetching when memory_threshold_percent is exceeded. URL dedup. BFSDeepCrawlStrategy.link_discovery() (crawl4ai/deep_crawling/bfs_strategy.py:133) maintains a visited: Set[str] and checks can_process_url() before adding new URLs. URLs are normalized via normalize_url_for_deep_crawl() before dedup. Depth/limits. BFSDeepCrawlStrategy accepts max_depth, max_pages, score_threshold, and include_external. The FilterChain (crawl4ai/deep_crawling/filters.py) composes URL filters (domain, path, content-type). Robots.txt/politeness. The RateLimiter.wait_if_needed() method enforces per-domain delays between requests. Crawl strategies accept should_cancel callbacks for graceful interruption. Distributed workers. The Docker deployment (deploy/docker/server.py) manages a pool of browser instances via BrowserPool (crawler_pool.py). It supports crawler_configs — per-URL custom configs with url_matcher patterns for arun_many(). Webhook delivery (deploy/docker/webhook.py) notifies external services on completion/failure. Crash recovery is supported via resume_state callback in the deep crawl strategies.
What is the developer interface?
answeredLibrary API. AsyncWebCrawler (crawl4ai/async_webcrawler.py) is the main class, with arun() for single URLs and arun_many() for batches. Both accept CrawlerRunConfig and return CrawlResult objects containing html, markdown (raw/fit/citations), extracted_content, screenshot, pdf, metadata, and more. CLI. crawl4ai/cli.py provides a command-line interface; crawl4ai/cloud/cli.py handles cloud deployments. REST API. The Docker server (deploy/docker/server.py) exposes POST endpoints: /crawl (JSON), /crawl/stream (NDJSON), /md (Markdown), /llm (QA), /screenshot, /pdf, /html, /execute_js, /config/dump, /ask (BM25 context retrieval), and /schema. Streaming protocol. /crawl/stream returns a StreamingResponse with application/x-ndjson content type. Each result is a JSON line via stream_results() (deploy/docker/api.py:616), which serializes each CrawlResult as it arrives from the async generator, includes a heartbeat mechanism, and finishes with {"status":"completed"}. Deep-crawl streaming (crawl4ai/async_webcrawler.py:1047-1057) wraps the BFS async generator in the same NDJSON format. MCP server. deploy/docker/mcp_bridge.py attaches at /mcp/ws (WebSocket) and /mcp/sse (SSE) transports. The WebSocket handler (_ws, line 198) uses anyio.create_memory_object_stream to bridge the MCP server's read/write streams, receives JSON-RPC messages from the client, and proxies tool calls to the FastAPI endpoints via httpx loopback with a service token. The SSE transport (_MCPSseApp, line 247) uses SseServerTransport.connect_sse from the MCP SDK. Monitor WebSocket. /monitor/ws (deploy/docker/monitor_routes.py:354) pushes JSON every 2 seconds with health stats, active requests, browser pool, timeline data, and janitor logs. The dashboard JS (static/monitor/index.html) implements auto-reconnect with exponential backoff and falls back to HTTP polling. Output formats. Results return as Python objects (library) or JSON (API): CrawlResult.model_dump() includes all fields. PDF bytes are base64-encoded over the wire. The schema endpoint GET /schema exposes BrowserConfig/CrawlerRunConfig for automated client generation.
WebSocket/SSE Streaming Details
The server has three distinct streaming protocols:
1. Crawl streaming (NDJSON). Triggered by setting stream=true in CrawlerRunConfig or hitting /crawl/stream. handle_stream_crawl_request() (deploy/docker/api.py:889) creates the crawler, configures stream=True on the CrawlerRunConfig, and returns a (crawler, async_generator, hooks_info) tuple. stream_process() (deploy/docker/server.py:1020) wraps the generator in FastAPI's StreamingResponse(media_type="application/x-ndjson"). The stream_results() function (api.py:616) iterates the generator, serializes each CrawlResult via model_dump(), encodes base64 if PDF bytes are present, adds server_memory_mb, and yields one JSON line per result followed by . For deep crawling, arun() returns an async generator directly (async_webcrawler.py:1047-1057) — the BFS strategy's _arun_stream() yields results as pages are crawled level-by-level.
2. MCP WebSocket (/mcp/ws). Defined in mcp_bridge.py:197. The handler creates two anyio.create_memory_object_stream(100) pairs (client→server and server→client). Messages are validated against JSONRPCMessage schema via pydantic's TypeAdapter. Three concurrent tasks run in an anyio.TaskGroup: ws_to_srv() reads JSON from the WebSocket and sends to the MCP server; mcp.run() processes the MCP protocol; srv_to_ws() reads MCP responses and sends JSON back to the WebSocket. Authentication is handled by AuthGateMiddleware at the ASGI level; WebSocket clients pass ?token= in the query string.
3. MCP SSE (/mcp/sse). Uses the MCP SDK's SseServerTransport (mcp_bridge.py:241). A raw ASGI callable class _MCPSseApp avoids Starlette middleware interference with route wrappers. Messages POST to /mcp/messages/{session_id}. The same tool/resource handlers back both transports.
4. Monitor WebSocket (/monitor/ws). Pushes a flat JSON object every 2 seconds (monitor_routes.py:370-397) aggregating health summary, request counts, browser list, timeline data, janitor log, and error log. The JS dashboard falls back to HTTP polling if WebSocket fails.
Agent Composition & Multi-Step Extraction Pipelines
The codebase has several composition patterns:
C4A-Script DSL. crawl4ai/script/c4ai_script.py defines a domain-specific language for browser automation. Scripts compile to JavaScript — the Compiler class (line 324) parses via Lark grammar, collects PROC...ENDPROC procedure definitions (line 355), inlines CALL references (line 362), substitutes $variable references (line 372), and emits JavaScript for each command (line 387). Commands include GO, CLICK, WAIT, TYPE, EVAL, SCROLL, DRAG, IF/REPEAT flow control, and SETVAR state. This compiles to js_code that runs in the Playwright browser — no core crawler modification needed. Example: a "login" procedure defined once, then called after navigation.
Deep crawl strategy chain. DeepCrawlDecorator (crawl4ai/deep_crawling/base_strategy.py:10) wraps AsyncWebCrawler.arun(): when CrawlerRunConfig.deep_crawl_strategy is set, the decorator intercepts the call, passes it to the strategy's arun(), which calls _arun_stream() or _arun_batch(). The strategies (BFSDeepCrawlStrategy, DFSDeepCrawlStrategy, BFFDeepCrawlStrategy) each implement link discovery with a composable FilterChain and optional URLScorer. This forms a pipeline: fetch → extract links → filter → score → queue → recurse.
Extraction strategy pipeline. Multiple extraction strategies can be composed with ContentScrapingStrategy — for example, PDF content uses PDFCrawlerStrategy (crawler) → PDFContentScrapingStrategy (scraper that downloads + parses + extracts images/links). The LLM extraction internally chains: HTML → sanitize_html → merge_chunks (splits by token budget) → ThreadPoolExecutor with LLM calls → split_and_parse_json_objects → merge results. Hooks (deploy/docker/hook_registry.py) provide another composition dimension: declarative actions like block_resources, scroll_to_bottom, or set_cookie can be chained before/after crawl phases.
PDF Support Status (Source-Level)
PDF support is fully implemented at the source level, not just README claims. Key classes:
PDFCrawlerStrategy(crawl4ai/processors/pdf/__init__.py:49): AnAsyncCrawlerStrategythat returns a placeholder response (placeholder_html=True) — it does nothing itself, signaling to the pipeline thatPDFContentScrapingStrategywill handle actual work.PDFContentScrapingStrategy(line 79): AContentScrapingStrategythat downloads PDFs over HTTP(S) usingrequestswith a hand-written redirect follower (_fetch_with_redirect_checks, line 259) that validates every hop against an injected URL/peer-IP validator for SSRF protection. It caps download size (max_pdf_bytes, default 100 MiB), page count (max_pdf_pages, default 2000), and redirect depth (max_redirects, default 5). The file is written to a temp directory, then processed. Output is aScrapingResultwithcleaned_html(wrapping per-page<div class="pdf-page">sections), plusmedia[images]andlinks[urls]each tagged with page numbers.NaivePDFProcessorStrategy(crawl4ai/processors/pdf/processor.py:57): The processing engine usingpypdf. It extracts metadata (title, author, producer, creation/modification dates, encryption status, file size — lines 435-457), per-page text (with positional layout info viavisitor_textcallback), images (handling FlateDecode/PNG-predictor, DCTDecode/JPEG, CCITTFaxDecode/TIFF, JPXDecode/JPEG2000 — lines 254-421), and hyperlinks. Images can be saved locally or returned as base64 data URLs. Text is cleaned into both Markdown and HTML per page (viaclean_pdf_text/clean_pdf_text_to_htmlinutils.py). Aprocess_batch()method (line 144) usesThreadPoolExecutorfor parallel page processing.Integration. The Docker server imports these in
handle_crawl_request()(api.py:703) andhandle_stream_crawl_request()(api.py:931) — when the user'sCrawlerRunConfig.scraping_strategyis an instance ofPDFContentScrapingStrategy, it instantiatesPDFCrawlerStrategydirectly (not from the pool, since it has no browser) and wires the URL validator. The egress policy is injected at server boot (server.py:163-190) viaset_url_validatorandset_peer_ip_validator, so PDF fetches enforce the same SSRF protection as browser-based fetches. Hooks are not supported with PDFs (api.py:707-708: "PDFCrawlerStrategy has no browser page, so hooks can't attach to it").