LLMs Technical Reviews
Home / AI web scraping / Scrapling

D4Vinci/Scrapling

Python scraping library pairing an lxml parser that re-finds moved elements with curl_cffi, Playwright and stealth-browser fetchers.

GitHub ↗★ 86kPythonBSD-3-Clausecommit 43dee00 · 2026-10-06homepage ↗

Overview

Scrapling is a Python scraping library with three parts that work together. The first is a fast lxml-based parser, Selector, with a Scrapy/Parsel-style API and an “adaptive” mode that can find an element again after the page layout changes. The second is a family of fetchers: a curl_cffi HTTP client that copies browser TLS fingerprints, a Playwright browser, and a Patchright “stealthy” browser that can click through Cloudflare Turnstile. The third is a Scrapy-like async spider framework with a scheduler, per-domain throttling, robots.txt support and checkpoints.

It is not an LLM extraction tool. Nothing in the library calls a model. Its “AI” surface is an MCP server that lets an agent fetch pages and get back Markdown, HTML or text. Before that, the server removes hidden elements that could carry prompt injections. Structured extraction stays in your code, written with CSS or XPath selectors.

The core install needs only lxml, cssselect, orjson, tld and w3lib. The fetchers, MCP server and shell are optional extras (pyproject.toml). This suits developers who want Scrapy-style control with built-in anti-bot fetching, in one package.

Architecture

flowchart LR
  U["Your code / CLI / MCP"] --> F["Fetcher / AsyncFetcher"]
  U --> DF["DynamicFetcher"]
  U --> SF["StealthyFetcher"]
  U --> SP["Spider"]
  F --> CC["curl_cffi session"]
  DF --> PW["Playwright Chromium"]
  SF --> PR["Patchright Chromium + CF solver"]
  SP --> EN["CrawlerEngine"]
  EN --> SCH["Scheduler (priority + dedup)"]
  EN --> SM["SessionManager"]
  SM --> CC
  SM --> PW
  SM --> PR
  CC --> R["Response = Selector"]
  PW --> R
  PR --> R
  R --> ST["SQLite adaptive storage"]
Component Path Role
Parser scrapling/parser.py Selector/Selectors: CSS, XPath, text/regex search, find_similar, adaptive relocation
Adaptive storage scrapling/core/storage.py SQLite table of element fingerprints keyed by domain and identifier
HTTP engine scrapling/engines/static.py FetcherSession on curl_cffi: impersonation, stealth headers, retries, proxy rotation
Browser engines scrapling/engines/_browsers/ DynamicSession (Playwright), StealthySession (Patchright), page pool, validators
Toolbelt scrapling/engines/toolbelt/ ResponseFactory, browserforge fingerprints, ProxyRotator, ad-domain list
Fetcher facades scrapling/fetchers/ Fetcher, AsyncFetcher, DynamicFetcher, StealthyFetcher one-shot classes
Spiders scrapling/spiders/ Spider, CrawlerEngine, Scheduler, AutoThrottle, robots.txt, checkpoints, cache
Templates scrapling/spiders/templates/ Crawl, sitemap, feed, Shopify and site-to-Markdown spiders
MCP server scrapling/core/ai.py ScraplingMCPServer: fetch tools, sessions, screenshots
Shell and CLI scrapling/core/shell.py, scrapling/cli.py IPython shell, curl-to-Scrapling converter, extract command, Markdown conversion

How a request flows

Take a spider whose parse does response.css(".price::text", adaptive=True):

  1. Schedule. Scheduler.enqueue fingerprints each Request (SHA-1 of session id, method, canonical URL and body, plus kwargs or headers if configured) (request.py). It drops ones it has seen unless dont_filter is set, and pushes the rest onto an asyncio.PriorityQueue (scheduler.py).
  2. Politeness. CrawlerEngine._process_request checks robots.txt when robots_txt_obey is on. It takes the larger of download_delay, Crawl-delay and Request-rate as the floor, and may answer from the development cache (engine.py).
  3. Fetch. Inside a global or per-domain CapacityLimiter, the engine sleeps for the AutoThrottle delay. It then calls SessionManager.fetch, which picks the session named by the request’s sid: an HTTP session or a browser session (engine.py, session.py).
  4. HTTP path. _make_request picks a proxy from the rotator, merges headers (a Google referer plus browserforge headers when impersonation is off), sends the request through curl_cffi, and retries on CurlError, switching proxy when the error looks like a proxy failure (static.py).
  5. Browser path. StealthySession.fetch opens a pooled page, navigates with a Google referer, waits for load and network idle, runs the Cloudflare solver if asked, then page_action and wait_selector, and builds the Response (_stealth.py).
  6. Blocked? spider.is_blocked treats 401, 403, 407, 429, 444 and 5xx as blocked by default. A blocked request is copied with lower priority, stripped of its proxy, and re-queued up to max_blocked_retries times (spider.py, L204-L212).
  7. Parse. The callback gets a Response, which is a Selector. css() compiles to XPath. If nothing matches and adaptive=True, the stored fingerprint for that selector is loaded and the whole tree is scored to relocate the element (parser.py). Yielded dicts go to the item list or stream, and yielded Requests go back to step 1.

Key components

Adaptive selection

With auto_save=True, the first match of a selector is stored as a dictionary: tag, text, attributes, DOM path, parent name, attributes and text, and sibling tags. It goes into a SQLite table keyed by the site’s domain and the selector or a custom identifier (parser.py, storage.py). When the selector later finds nothing, relocate visits every element in the page and scores it against that record with difflib.SequenceMatcher, averaging tag, text, attribute, class/id/href/src, path, parent and sibling similarity. It returns all elements tied at the top score if that score is at least percentage (default 40) (parser.py, L822-L895). This is a brute-force O(n) pass per lookup. It runs only on a miss, which is the right trade-off.

Fetchers and stealth

Fetcher.get and its siblings are class methods on a shared client instance (requests.py). Browser sessions start Playwright or Patchright Chromium. They connect over CDP if you pass cdp_url, launch a plain browser when a proxy rotator needs per-proxy contexts, and otherwise use a persistent context with a temporary profile (_stealth.py). The stealth tier adds Chromium flags, not JavaScript patches: WebRTC limited to the proxy, optional WebGL disabling, and Chromium’s own canvas noise flag (_base.py). Firefox and WebKit are not used.

Cloudflare solver

_cloudflare_solver classifies the page as non-interactive, managed or embedded Turnstile. It waits out the non-interactive kind. For the others it finds the challenges.cloudflare.com iframe or a fallback box and clicks about 26 px into it at a jittered point with a random press delay. It re-checks and recurses, giving up after three attempts (_stealth.py, L108-L193). There is no general CAPTCHA solving.

AutoThrottle

When enabled, AutoThrottle.record moves each domain’s delay toward latency / target_concurrency. On a block it doubles the delay or honours Retry-After, and a block never lowers the delay. The result is clamped between the spider’s floor and max_delay (throttle.py).

MCP server and AI-safe output

MCP tools such as make_request, fetch and stealthy_fetch wrap a session and return a ResponseModel. They default to Markdown and main_content_only=True (ai.py). In Convertor._extract_content, main-content mode keeps <body>, drops script/style/noscript/svg, and runs _sanitize_for_ai. That removes CSS-hidden and aria-hidden elements, <template>, comments, zero-width and control characters before the optional css_selector is applied. Markdown comes from markdownify (shell.py).

Extending it

  • Spider hooks. Override parse, is_blocked, retry_blocked_request, on_scraped_item and on_error. Use sid on a Request to send some pages through a browser session and others through HTTP.
  • Storage. Implement StorageSystemMixin to keep adaptive fingerprints somewhere other than SQLite.
  • Browser automation. page_setup (before navigation) and page_action (after) receive the raw Playwright page. init_script, extra_flags and additional_args reach the browser and context.
  • Scrapy. scrapling/integrations/scrapy.py lets existing Scrapy callbacks parse with Selector.

Running it

  • pip install scrapling gives the parser only. pip install "scrapling[fetchers]" then scrapling install adds the fetchers and browsers. [ai] adds the MCP server (scrapling mcp or scrapling-mcp, stdio or streamable HTTP with optional bearer auth). [shell] adds the IPython shell.
  • The CLI scrapling extract get <url> out.md fetches a page and writes Markdown, HTML or text based on the file extension.
  • Docker images ship with the browsers installed. No external service is required. Adaptive data lives in a local SQLite file.

Strengths and caveats

  • Strength: adaptive selectors. Relocating elements by similarity after a redesign is unusual and practical, and it costs nothing until a selector misses.
  • Strength: one API across fetch tiers. HTTP, browser and stealth browser all return the same Response/Selector, so moving a site to a stronger tier is a one-line change.
  • Strength: injection-aware output. Removing hidden text before handing pages to an agent is a sensible default that few scrapers have.
  • Caveat: no LLM extraction. No schema-to-JSON, chunking or model calls. You write selectors or let the calling agent do the reading.
  • Caveat: Chromium only. The stealth tier is Patchright plus launch flags and a Turnstile clicker. Akamai, DataDome or Kasada pages need outside help.
  • Caveat: single process. The scheduler, dedup set and throttle are in memory. Checkpoints give pause and resume, not distribution.
  • Caveat: politeness is opt-in. robots_txt_obey and AutoThrottle default to off, and concurrent_requests is 4 with no per-domain cap.

Sources: code at 43dee00, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (17 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 43dee00. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Scrapling has three fetching tiers in scrapling/fetchers/:

Plain HTTP (Fetcher/AsyncFetcher/FetcherSession) — Built on curl_cffi in scrapling/engines/static.py. Impersonates browser TLS fingerprints via impersonate (defaults to latest Chrome). Supports GET/POST/PUT/DELETE, HTTP/3, session persistence, 3 retries by default, and SSRF-safe redirects. No JS rendering.

Dynamic browser (DynamicFetcher/DynamicSession) — Playwright Chromium via scrapling/engines/_browsers/_controllers.py. Runs a real browser (headless default), loads JS, supports load_dom, network_idle (500ms idle wait), wait_selector, and custom page_action. Connects to remote browsers via CDP (cdp_url). Ad blocking across ~3,500 domains from scrapling/engines/toolbelt/ad_domains.py. Page pooling in scrapling/engines/_browsers/_page.py.

Stealthy browser (StealthyFetcher/StealthySession) — Playwright/Patchright via scrapling/engines/_browsers/_stealth.py. Adds solve_cloudflare (detects and clicks turnstile/interstitial challenges at randomized coordinates, retries up to 3 times), canvas noise (hide_canvas), WebRTC proxy locking (block_webrtc), and WebGL preservation. Uses Patchright (Playwright fork) for undetectable automation. Realistic user-agent generation via browserforge in scrapling/engines/toolbelt/fingerprints.py.

Content types: ResponseFactory in scrapling/engines/toolbelt/convertor.py handles HTML by extracting DOM content and non-HTML (PDFs, images) by returning raw response body. All return a unified Response object (subclass of Selector).

How is content extracted or converted?

answered

Extraction uses the Selector parser in scrapling/parser.py and the Convertor class in scrapling/core/shell.py.

Parsing — Selector wraps lxml's HtmlElement and provides CSS selectors (via cssselect), XPath, BS4-style find_all() by tag/class/attrs, find_by_text(), and regex-based find_by_regex(). Pseudo-elements ::text, ::attr(name) (Scrapy/Parsel-compatible) are supported. Element data persists in SQLite (scrapling/core/storage.py) for adaptive selection.

Adaptive selection — auto_save=True saves element fingerprints; adaptive=True re-locates changed elements using difflib.SequenceMatcher. find_similar() finds structurally similar elements.

HTML to Markdown — Convertor._convert_to_markdown() uses markdownify (optional [rag] dependency). The pipeline in _extract_content() (shell.py:622-660): (1) optional CSS selector narrowing, (2) main_content_only scopes to <body>, (3) _strip_noise_tags() removes <script>/<style>/<svg>, (4) _sanitize_for_ai() strips CSS-hidden elements (display:none, visibility:hidden, opacity:0), aria-hidden, <template>, HTML comments, zero-width Unicode, and control characters (prompt-injection defense), (5) markdownify converts cleaned HTML, or returns raw HTML or plain text per extraction_type.

Schema-based extraction — Not built in. Users build structured extraction with .get(), .getall(), .attrib. The MCP server returns ResponseModel with status, url, and content lists in the chosen format.

How are LLMs used, if at all?

answered

Scrapling does not call any LLM in its core pipeline — no API calls to OpenAI, Anthropic, or others. It provides LLM-adjacent infrastructure:

MCP Server (scrapling/core/ai.py, class ScraplingMCPServer) — Implements MCP to let AI agents (Claude, Cursor) use Scrapling as a tool. ~12 tools: make_request/bulk_get (HTTP), fetch/bulk_fetch (Playwright), stealthy_fetch/bulk_stealthy_fetch (stealth browser), session management, and screenshot. All fetch tools return a ResponseModel with status, URL, and content (Markdown by default). Server instructions tell agents to use css_selector to narrow content and save tokens. Supports bearer-auth-protected HTTP transport or stdio.

Prompt-injection sanitization — Convertor._sanitize_for_ai() (shell.py:604-619) strips hidden elements, <template> tags, HTML comments, zero-width Unicode, and control characters before content reaches the LLM, preventing hidden injection text.

RAG-ready Markdown — SiteToMarkdownSpider template (scrapling/spiders/templates/site_to_markdown.py) crawls entire sites to Markdown for RAG ingestion, with css_selector, main_content_only, and output_dir controls.

Agent Skill — A skill file at agent-skill/Scrapling-Skill/ teaching coding agents the current API so generated code doesn't guess outdated interfaces.

No chunking, cost controls, or structured output enforcement — those are the calling agent's responsibility.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot bypass is a first-class feature with several layers:

TLS fingerprint impersonation (scrapling/engines/static.py) — The impersonate parameter lets curl_cffi mimic Chrome/Firefox TLS fingerprints at the wire level. Accepts a single browser string or a list for random selection. HTTP/3 available. stealthy_headers=True (default) sets real browser headers and a Google referer via _headers_job().

Stealth browser patches (scrapling/engines/_browsers/_stealth.py and _base.py) — hide_canvas adds random noise to canvas fingerprinting; block_webrtc forces WebRTC through the proxy to prevent WebRTC-based IP leaks; allow_webgl keeps WebGL active (some WAFs check for it). Uses Patchright (a Playwright fork) for undetectable automation. Realistic user-agent generation via browserforge (scrapling/engines/toolbelt/fingerprints.py) keyed to the detected OS and Chromium version.

Cloudflare Turnstile solver (scrapling/engines/_browsers/_stealth.py:108-193) — solve_cloudflare=True detects the challenge type (non-interactive, standard turnstile, embedded turnstile), locates the Cloudflare iframe, and clicks the checkbox at randomized coordinates with human-like mouse delays. Retries up to 3 times.

Proxy rotation (scrapling/engines/toolbelt/proxy_rotation.py) — ProxyRotator is thread-safe with pluggable rotation strategies (default cyclic). On connection errors matching _PROXY_ERROR_INDICATORS, the HTTP engine retries with the next proxy.

DNS leak prevention — Available via dns_over_https (Cloudflare DoH) in browser sessions.

Rate limiting — AutoThrottle (scrapling/spiders/throttle.py) doubles delay on blocked responses (status codes 401/403/407/429/444/500/502/503/504 from spider.py:16), respects Retry-After headers, and reduces latency on healthy responses.

No native CAPTCHA-solving beyond Cloudflare Turnstile. Docs direct to a partner API for Akamai/DataDome/Kasada/Incapsula.

How is crawling at scale implemented?

answered

Crawling is implemented in scrapling/spiders/ with a Scrapy-inspired architecture:

Spider base class (scrapling/spiders/spider.py) — Users subclass Spider (or CrawlSpider/SitemapSpider/ShopifySpider templates) defining name, start_urls, allowed_domains, and an async parse(response) yielding items or Request objects.

Scheduler (scrapling/spiders/scheduler.py) — asyncio.PriorityQueue-based with URL deduplication via SHA-1 fingerprints. Duplicates dropped unless dont_filter=True. Supports snapshot/restore for checkpoint persistence (scheduler.py:31-80).

Engine (scrapling/spiders/engine.py) — CrawlerEngine orchestrates. Concurrency: anyio.CapacityLimiter — global (concurrent_requests, default 4) and per-domain (concurrent_requests_per_domain). Items stream via anyio.create_memory_object_stream for real-time stream() iteration.

AutoThrottle (scrapling/spiders/throttle.py) — Tunes per-domain delays from observed response latency. record() doubles delay on blocked responses or respects Retry-After headers; speeds back up on healthy responses. Delays clamped between start_delay (5s default) and max_delay (60s default).

Robots.txt (scrapling/spiders/robotstxt.py) — Optional robots_txt_obey flag. Fetches and caches per-domain robots.txt via Protego, checking can_fetch(), Crawl-delay, and Request-rate directives.

Pause/Resume (scrapling/spiders/checkpoint.py) — CheckpointManager saves scheduler state to disk periodically (default 5min) and on graceful shutdown. Restarting with the same crawldir resumes.

Multi-session routing (scrapling/spiders/session.py) — SessionManager handles different session types by ID. Requests carry sid to route through HTTP or browser sessions.

Templates — CrawlSpider with rule-based link following via LinkExtractor (allow/deny, domain/extension filters), SitemapSpider, XMLFeedSpider/CSVFeedSpider, ShopifySpider, and SiteToMarkdownSpider.

No distributed workers — single-process with async concurrency.

What is the developer interface?

answered

Scrapling provides multiple interfaces:

Python library API — Primary interface. Import fetchers (Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher) for stateless one-off use, or session classes (FetcherSession, StealthySession, DynamicSession) with with/async with. All return a Response object (subclass of Selector parser) for CSS/XPath/text extraction. Standalone parser: from scrapling.parser import Selector.

Spider framework — Scrapy-like class-based API with async parse() methods yielding items/requests. Built-in export: result.items.to_json(), to_jsonl(), to_csv(), to_xml(). Streaming via async for item in spider.stream().

CLI (scrapling/cli.py) — scrapling shell (IPython-based interactive shell with Scrapling integration and curl-to-Scrapling conversion), scrapling extract <method> <url> <output_file> (format auto-detected from extension: .html/.md/.txt), scrapling install (browser deps).

MCP Server (scrapling/core/ai.py) — Invoked via scrapling-mcp. ~12 tools over stdio or streamable-http with optional bearer auth. Session management with create/fetch/close/list lifecycle. Screenshot tool returns images.

Agent Skill — agent-skill/Scrapling-Skill/ file teaching coding agents the current API so generated code is accurate.

Docker — pyd4vinci/scrapling (DockerHub) and ghcr.io/d4vinci/scrapling:latest (GHCR) with all browsers pre-installed.

Scrapy integration (scrapling/integrations/scrapy.py) — scrapling_response decorator lets existing Scrapy callbacks parse with Scrapling's parser.

Output formats — Items export to JSON/JSONL/CSV/XML. Page content exports to HTML/Markdown/plain text via file extension or extraction_type.