D4Vinci/Scrapling
Python scraping library pairing an lxml parser that re-finds moved elements with curl_cffi, Playwright and stealth-browser fetchers.
Overview
Scrapling is a Python scraping library with three parts that work together. The first is a fast lxml-based parser, Selector, with a Scrapy/Parsel-style API and an “adaptive” mode that can find an element again after the page layout changes. The second is a family of fetchers: a curl_cffi HTTP client that copies browser TLS fingerprints, a Playwright browser, and a Patchright “stealthy” browser that can click through Cloudflare Turnstile. The third is a Scrapy-like async spider framework with a scheduler, per-domain throttling, robots.txt support and checkpoints.
It is not an LLM extraction tool. Nothing in the library calls a model. Its “AI” surface is an MCP server that lets an agent fetch pages and get back Markdown, HTML or text. Before that, the server removes hidden elements that could carry prompt injections. Structured extraction stays in your code, written with CSS or XPath selectors.
The core install needs only lxml, cssselect, orjson, tld and w3lib. The fetchers, MCP server and shell are optional extras (pyproject.toml). This suits developers who want Scrapy-style control with built-in anti-bot fetching, in one package.
Architecture
flowchart LR
U["Your code / CLI / MCP"] --> F["Fetcher / AsyncFetcher"]
U --> DF["DynamicFetcher"]
U --> SF["StealthyFetcher"]
U --> SP["Spider"]
F --> CC["curl_cffi session"]
DF --> PW["Playwright Chromium"]
SF --> PR["Patchright Chromium + CF solver"]
SP --> EN["CrawlerEngine"]
EN --> SCH["Scheduler (priority + dedup)"]
EN --> SM["SessionManager"]
SM --> CC
SM --> PW
SM --> PR
CC --> R["Response = Selector"]
PW --> R
PR --> R
R --> ST["SQLite adaptive storage"]
| Component | Path | Role |
|---|---|---|
| Parser | scrapling/parser.py |
Selector/Selectors: CSS, XPath, text/regex search, find_similar, adaptive relocation |
| Adaptive storage | scrapling/core/storage.py |
SQLite table of element fingerprints keyed by domain and identifier |
| HTTP engine | scrapling/engines/static.py |
FetcherSession on curl_cffi: impersonation, stealth headers, retries, proxy rotation |
| Browser engines | scrapling/engines/_browsers/ |
DynamicSession (Playwright), StealthySession (Patchright), page pool, validators |
| Toolbelt | scrapling/engines/toolbelt/ |
ResponseFactory, browserforge fingerprints, ProxyRotator, ad-domain list |
| Fetcher facades | scrapling/fetchers/ |
Fetcher, AsyncFetcher, DynamicFetcher, StealthyFetcher one-shot classes |
| Spiders | scrapling/spiders/ |
Spider, CrawlerEngine, Scheduler, AutoThrottle, robots.txt, checkpoints, cache |
| Templates | scrapling/spiders/templates/ |
Crawl, sitemap, feed, Shopify and site-to-Markdown spiders |
| MCP server | scrapling/core/ai.py |
ScraplingMCPServer: fetch tools, sessions, screenshots |
| Shell and CLI | scrapling/core/shell.py, scrapling/cli.py |
IPython shell, curl-to-Scrapling converter, extract command, Markdown conversion |
How a request flows
Take a spider whose parse does response.css(".price::text", adaptive=True):
- Schedule.
Scheduler.enqueuefingerprints eachRequest(SHA-1 of session id, method, canonical URL and body, plus kwargs or headers if configured) (request.py). It drops ones it has seen unlessdont_filteris set, and pushes the rest onto anasyncio.PriorityQueue(scheduler.py). - Politeness.
CrawlerEngine._process_requestchecks robots.txt whenrobots_txt_obeyis on. It takes the larger ofdownload_delay,Crawl-delayandRequest-rateas the floor, and may answer from the development cache (engine.py). - Fetch. Inside a global or per-domain
CapacityLimiter, the engine sleeps for the AutoThrottle delay. It then callsSessionManager.fetch, which picks the session named by the request’ssid: an HTTP session or a browser session (engine.py, session.py). - HTTP path.
_make_requestpicks a proxy from the rotator, merges headers (a Google referer plus browserforge headers when impersonation is off), sends the request throughcurl_cffi, and retries onCurlError, switching proxy when the error looks like a proxy failure (static.py). - Browser path.
StealthySession.fetchopens a pooled page, navigates with a Google referer, waits for load and network idle, runs the Cloudflare solver if asked, thenpage_actionandwait_selector, and builds theResponse(_stealth.py). - Blocked?
spider.is_blockedtreats 401, 403, 407, 429, 444 and 5xx as blocked by default. A blocked request is copied with lower priority, stripped of its proxy, and re-queued up tomax_blocked_retriestimes (spider.py, L204-L212). - Parse. The callback gets a
Response, which is aSelector.css()compiles to XPath. If nothing matches andadaptive=True, the stored fingerprint for that selector is loaded and the whole tree is scored to relocate the element (parser.py). Yielded dicts go to the item list or stream, and yieldedRequests go back to step 1.
Key components
Adaptive selection
With auto_save=True, the first match of a selector is stored as a dictionary: tag, text, attributes, DOM path, parent name, attributes and text, and sibling tags. It goes into a SQLite table keyed by the site’s domain and the selector or a custom identifier (parser.py, storage.py). When the selector later finds nothing, relocate visits every element in the page and scores it against that record with difflib.SequenceMatcher, averaging tag, text, attribute, class/id/href/src, path, parent and sibling similarity. It returns all elements tied at the top score if that score is at least percentage (default 40) (parser.py, L822-L895). This is a brute-force O(n) pass per lookup. It runs only on a miss, which is the right trade-off.
Fetchers and stealth
Fetcher.get and its siblings are class methods on a shared client instance (requests.py). Browser sessions start Playwright or Patchright Chromium. They connect over CDP if you pass cdp_url, launch a plain browser when a proxy rotator needs per-proxy contexts, and otherwise use a persistent context with a temporary profile (_stealth.py). The stealth tier adds Chromium flags, not JavaScript patches: WebRTC limited to the proxy, optional WebGL disabling, and Chromium’s own canvas noise flag (_base.py). Firefox and WebKit are not used.
Cloudflare solver
_cloudflare_solver classifies the page as non-interactive, managed or embedded Turnstile. It waits out the non-interactive kind. For the others it finds the challenges.cloudflare.com iframe or a fallback box and clicks about 26 px into it at a jittered point with a random press delay. It re-checks and recurses, giving up after three attempts (_stealth.py, L108-L193). There is no general CAPTCHA solving.
AutoThrottle
When enabled, AutoThrottle.record moves each domain’s delay toward latency / target_concurrency. On a block it doubles the delay or honours Retry-After, and a block never lowers the delay. The result is clamped between the spider’s floor and max_delay (throttle.py).
MCP server and AI-safe output
MCP tools such as make_request, fetch and stealthy_fetch wrap a session and return a ResponseModel. They default to Markdown and main_content_only=True (ai.py). In Convertor._extract_content, main-content mode keeps <body>, drops script/style/noscript/svg, and runs _sanitize_for_ai. That removes CSS-hidden and aria-hidden elements, <template>, comments, zero-width and control characters before the optional css_selector is applied. Markdown comes from markdownify (shell.py).
Extending it
- Spider hooks. Override
parse,is_blocked,retry_blocked_request,on_scraped_itemandon_error. Usesidon aRequestto send some pages through a browser session and others through HTTP. - Storage. Implement
StorageSystemMixinto keep adaptive fingerprints somewhere other than SQLite. - Browser automation.
page_setup(before navigation) andpage_action(after) receive the raw Playwright page.init_script,extra_flagsandadditional_argsreach the browser and context. - Scrapy.
scrapling/integrations/scrapy.pylets existing Scrapy callbacks parse withSelector.
Running it
pip install scraplinggives the parser only.pip install "scrapling[fetchers]"thenscrapling installadds the fetchers and browsers.[ai]adds the MCP server (scrapling mcporscrapling-mcp, stdio or streamable HTTP with optional bearer auth).[shell]adds the IPython shell.- The CLI
scrapling extract get <url> out.mdfetches a page and writes Markdown, HTML or text based on the file extension. - Docker images ship with the browsers installed. No external service is required. Adaptive data lives in a local SQLite file.
Strengths and caveats
- Strength: adaptive selectors. Relocating elements by similarity after a redesign is unusual and practical, and it costs nothing until a selector misses.
- Strength: one API across fetch tiers. HTTP, browser and stealth browser all return the same
Response/Selector, so moving a site to a stronger tier is a one-line change. - Strength: injection-aware output. Removing hidden text before handing pages to an agent is a sensible default that few scrapers have.
- Caveat: no LLM extraction. No schema-to-JSON, chunking or model calls. You write selectors or let the calling agent do the reading.
- Caveat: Chromium only. The stealth tier is Patchright plus launch flags and a Turnstile clicker. Akamai, DataDome or Kasada pages need outside help.
- Caveat: single process. The scheduler, dedup set and throttle are in memory. Checkpoints give pause and resume, not distribution.
- Caveat: politeness is opt-in.
robots_txt_obeyand AutoThrottle default to off, andconcurrent_requestsis 4 with no per-domain cap.
Sources: code at 43dee00, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (17 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 43dee00. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredScrapling has three fetching tiers in scrapling/fetchers/:
Plain HTTP (Fetcher/AsyncFetcher/FetcherSession) — Built on curl_cffi in scrapling/engines/static.py. Impersonates browser TLS fingerprints via impersonate (defaults to latest Chrome). Supports GET/POST/PUT/DELETE, HTTP/3, session persistence, 3 retries by default, and SSRF-safe redirects. No JS rendering.
Dynamic browser (DynamicFetcher/DynamicSession) — Playwright Chromium via scrapling/engines/_browsers/_controllers.py. Runs a real browser (headless default), loads JS, supports load_dom, network_idle (500ms idle wait), wait_selector, and custom page_action. Connects to remote browsers via CDP (cdp_url). Ad blocking across ~3,500 domains from scrapling/engines/toolbelt/ad_domains.py. Page pooling in scrapling/engines/_browsers/_page.py.
Stealthy browser (StealthyFetcher/StealthySession) — Playwright/Patchright via scrapling/engines/_browsers/_stealth.py. Adds solve_cloudflare (detects and clicks turnstile/interstitial challenges at randomized coordinates, retries up to 3 times), canvas noise (hide_canvas), WebRTC proxy locking (block_webrtc), and WebGL preservation. Uses Patchright (Playwright fork) for undetectable automation. Realistic user-agent generation via browserforge in scrapling/engines/toolbelt/fingerprints.py.
Content types: ResponseFactory in scrapling/engines/toolbelt/convertor.py handles HTML by extracting DOM content and non-HTML (PDFs, images) by returning raw response body. All return a unified Response object (subclass of Selector).
How is content extracted or converted?
answeredExtraction uses the Selector parser in scrapling/parser.py and the Convertor class in scrapling/core/shell.py.
Parsing — Selector wraps lxml's HtmlElement and provides CSS selectors (via cssselect), XPath, BS4-style find_all() by tag/class/attrs, find_by_text(), and regex-based find_by_regex(). Pseudo-elements ::text, ::attr(name) (Scrapy/Parsel-compatible) are supported. Element data persists in SQLite (scrapling/core/storage.py) for adaptive selection.
Adaptive selection — auto_save=True saves element fingerprints; adaptive=True re-locates changed elements using difflib.SequenceMatcher. find_similar() finds structurally similar elements.
HTML to Markdown — Convertor._convert_to_markdown() uses markdownify (optional [rag] dependency). The pipeline in _extract_content() (shell.py:622-660): (1) optional CSS selector narrowing, (2) main_content_only scopes to <body>, (3) _strip_noise_tags() removes <script>/<style>/<svg>, (4) _sanitize_for_ai() strips CSS-hidden elements (display:none, visibility:hidden, opacity:0), aria-hidden, <template>, HTML comments, zero-width Unicode, and control characters (prompt-injection defense), (5) markdownify converts cleaned HTML, or returns raw HTML or plain text per extraction_type.
Schema-based extraction — Not built in. Users build structured extraction with .get(), .getall(), .attrib. The MCP server returns ResponseModel with status, url, and content lists in the chosen format.
How are LLMs used, if at all?
answeredScrapling does not call any LLM in its core pipeline — no API calls to OpenAI, Anthropic, or others. It provides LLM-adjacent infrastructure:
MCP Server (scrapling/core/ai.py, class ScraplingMCPServer) — Implements MCP to let AI agents (Claude, Cursor) use Scrapling as a tool. ~12 tools: make_request/bulk_get (HTTP), fetch/bulk_fetch (Playwright), stealthy_fetch/bulk_stealthy_fetch (stealth browser), session management, and screenshot. All fetch tools return a ResponseModel with status, URL, and content (Markdown by default). Server instructions tell agents to use css_selector to narrow content and save tokens. Supports bearer-auth-protected HTTP transport or stdio.
Prompt-injection sanitization — Convertor._sanitize_for_ai() (shell.py:604-619) strips hidden elements, <template> tags, HTML comments, zero-width Unicode, and control characters before content reaches the LLM, preventing hidden injection text.
RAG-ready Markdown — SiteToMarkdownSpider template (scrapling/spiders/templates/site_to_markdown.py) crawls entire sites to Markdown for RAG ingestion, with css_selector, main_content_only, and output_dir controls.
Agent Skill — A skill file at agent-skill/Scrapling-Skill/ teaching coding agents the current API so generated code doesn't guess outdated interfaces.
No chunking, cost controls, or structured output enforcement — those are the calling agent's responsibility.
How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-bot bypass is a first-class feature with several layers:
TLS fingerprint impersonation (scrapling/engines/static.py) — The impersonate parameter lets curl_cffi mimic Chrome/Firefox TLS fingerprints at the wire level. Accepts a single browser string or a list for random selection. HTTP/3 available. stealthy_headers=True (default) sets real browser headers and a Google referer via _headers_job().
Stealth browser patches (scrapling/engines/_browsers/_stealth.py and _base.py) — hide_canvas adds random noise to canvas fingerprinting; block_webrtc forces WebRTC through the proxy to prevent WebRTC-based IP leaks; allow_webgl keeps WebGL active (some WAFs check for it). Uses Patchright (a Playwright fork) for undetectable automation. Realistic user-agent generation via browserforge (scrapling/engines/toolbelt/fingerprints.py) keyed to the detected OS and Chromium version.
Cloudflare Turnstile solver (scrapling/engines/_browsers/_stealth.py:108-193) — solve_cloudflare=True detects the challenge type (non-interactive, standard turnstile, embedded turnstile), locates the Cloudflare iframe, and clicks the checkbox at randomized coordinates with human-like mouse delays. Retries up to 3 times.
Proxy rotation (scrapling/engines/toolbelt/proxy_rotation.py) — ProxyRotator is thread-safe with pluggable rotation strategies (default cyclic). On connection errors matching _PROXY_ERROR_INDICATORS, the HTTP engine retries with the next proxy.
DNS leak prevention — Available via dns_over_https (Cloudflare DoH) in browser sessions.
Rate limiting — AutoThrottle (scrapling/spiders/throttle.py) doubles delay on blocked responses (status codes 401/403/407/429/444/500/502/503/504 from spider.py:16), respects Retry-After headers, and reduces latency on healthy responses.
No native CAPTCHA-solving beyond Cloudflare Turnstile. Docs direct to a partner API for Akamai/DataDome/Kasada/Incapsula.
How is crawling at scale implemented?
answeredCrawling is implemented in scrapling/spiders/ with a Scrapy-inspired architecture:
Spider base class (scrapling/spiders/spider.py) — Users subclass Spider (or CrawlSpider/SitemapSpider/ShopifySpider templates) defining name, start_urls, allowed_domains, and an async parse(response) yielding items or Request objects.
Scheduler (scrapling/spiders/scheduler.py) — asyncio.PriorityQueue-based with URL deduplication via SHA-1 fingerprints. Duplicates dropped unless dont_filter=True. Supports snapshot/restore for checkpoint persistence (scheduler.py:31-80).
Engine (scrapling/spiders/engine.py) — CrawlerEngine orchestrates. Concurrency: anyio.CapacityLimiter — global (concurrent_requests, default 4) and per-domain (concurrent_requests_per_domain). Items stream via anyio.create_memory_object_stream for real-time stream() iteration.
AutoThrottle (scrapling/spiders/throttle.py) — Tunes per-domain delays from observed response latency. record() doubles delay on blocked responses or respects Retry-After headers; speeds back up on healthy responses. Delays clamped between start_delay (5s default) and max_delay (60s default).
Robots.txt (scrapling/spiders/robotstxt.py) — Optional robots_txt_obey flag. Fetches and caches per-domain robots.txt via Protego, checking can_fetch(), Crawl-delay, and Request-rate directives.
Pause/Resume (scrapling/spiders/checkpoint.py) — CheckpointManager saves scheduler state to disk periodically (default 5min) and on graceful shutdown. Restarting with the same crawldir resumes.
Multi-session routing (scrapling/spiders/session.py) — SessionManager handles different session types by ID. Requests carry sid to route through HTTP or browser sessions.
Templates — CrawlSpider with rule-based link following via LinkExtractor (allow/deny, domain/extension filters), SitemapSpider, XMLFeedSpider/CSVFeedSpider, ShopifySpider, and SiteToMarkdownSpider.
No distributed workers — single-process with async concurrency.
What is the developer interface?
answeredScrapling provides multiple interfaces:
Python library API — Primary interface. Import fetchers (Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher) for stateless one-off use, or session classes (FetcherSession, StealthySession, DynamicSession) with with/async with. All return a Response object (subclass of Selector parser) for CSS/XPath/text extraction. Standalone parser: from scrapling.parser import Selector.
Spider framework — Scrapy-like class-based API with async parse() methods yielding items/requests. Built-in export: result.items.to_json(), to_jsonl(), to_csv(), to_xml(). Streaming via async for item in spider.stream().
CLI (scrapling/cli.py) — scrapling shell (IPython-based interactive shell with Scrapling integration and curl-to-Scrapling conversion), scrapling extract <method> <url> <output_file> (format auto-detected from extension: .html/.md/.txt), scrapling install (browser deps).
MCP Server (scrapling/core/ai.py) — Invoked via scrapling-mcp. ~12 tools over stdio or streamable-http with optional bearer auth. Session management with create/fetch/close/list lifecycle. Screenshot tool returns images.
Agent Skill — agent-skill/Scrapling-Skill/ file teaching coding agents the current API so generated code is accurate.
Docker — pyd4vinci/scrapling (DockerHub) and ghcr.io/d4vinci/scrapling:latest (GHCR) with all browsers pre-installed.
Scrapy integration (scrapling/integrations/scrapy.py) — scrapling_response decorator lets existing Scrapy callbacks parse with Scrapling's parser.
Output formats — Items export to JSON/JSONL/CSV/XML. Page content exports to HTML/Markdown/plain text via file extension or extraction_type.