# D4Vinci/Scrapling

> Python scraping library pairing an lxml parser that re-finds moved elements with curl_cffi, Playwright and stealth-browser fetchers.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/D4Vinci/Scrapling (reviewed at commit `43dee004866e1c46843a9ac293cc1494aa7915d6`, 2026-10-06)
- Stars: 85964 · Language: Python · License: BSD-3-Clause
- Canonical page: https://llms-technical-reviews.com/p/scrapling/

## Overview

Scrapling is a Python scraping library with three parts that work together. The first is a fast lxml-based parser, `Selector`, with a Scrapy/Parsel-style API and an "adaptive" mode that can find an element again after the page layout changes. The second is a family of fetchers: a `curl_cffi` HTTP client that copies browser TLS fingerprints, a Playwright browser, and a Patchright "stealthy" browser that can click through Cloudflare Turnstile. The third is a Scrapy-like async spider framework with a scheduler, per-domain throttling, robots.txt support and checkpoints.

It is not an LLM extraction tool. Nothing in the library calls a model. Its "AI" surface is an MCP server that lets an agent fetch pages and get back Markdown, HTML or text. Before that, the server removes hidden elements that could carry prompt injections. Structured extraction stays in your code, written with CSS or XPath selectors.

The core install needs only `lxml`, `cssselect`, `orjson`, `tld` and `w3lib`. The fetchers, MCP server and shell are optional extras ([pyproject.toml](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/pyproject.toml#L60-L100)). This suits developers who want Scrapy-style control with built-in anti-bot fetching, in one package.

## Architecture

```mermaid
flowchart LR
  U["Your code / CLI / MCP"] --> F["Fetcher / AsyncFetcher"]
  U --> DF["DynamicFetcher"]
  U --> SF["StealthyFetcher"]
  U --> SP["Spider"]
  F --> CC["curl_cffi session"]
  DF --> PW["Playwright Chromium"]
  SF --> PR["Patchright Chromium + CF solver"]
  SP --> EN["CrawlerEngine"]
  EN --> SCH["Scheduler (priority + dedup)"]
  EN --> SM["SessionManager"]
  SM --> CC
  SM --> PW
  SM --> PR
  CC --> R["Response = Selector"]
  PW --> R
  PR --> R
  R --> ST["SQLite adaptive storage"]
```

| Component | Path | Role |
|---|---|---|
| Parser | `scrapling/parser.py` | `Selector`/`Selectors`: CSS, XPath, text/regex search, `find_similar`, adaptive relocation |
| Adaptive storage | `scrapling/core/storage.py` | SQLite table of element fingerprints keyed by domain and identifier |
| HTTP engine | `scrapling/engines/static.py` | `FetcherSession` on `curl_cffi`: impersonation, stealth headers, retries, proxy rotation |
| Browser engines | `scrapling/engines/_browsers/` | `DynamicSession` (Playwright), `StealthySession` (Patchright), page pool, validators |
| Toolbelt | `scrapling/engines/toolbelt/` | `ResponseFactory`, browserforge fingerprints, `ProxyRotator`, ad-domain list |
| Fetcher facades | `scrapling/fetchers/` | `Fetcher`, `AsyncFetcher`, `DynamicFetcher`, `StealthyFetcher` one-shot classes |
| Spiders | `scrapling/spiders/` | `Spider`, `CrawlerEngine`, `Scheduler`, `AutoThrottle`, robots.txt, checkpoints, cache |
| Templates | `scrapling/spiders/templates/` | Crawl, sitemap, feed, Shopify and site-to-Markdown spiders |
| MCP server | `scrapling/core/ai.py` | `ScraplingMCPServer`: fetch tools, sessions, screenshots |
| Shell and CLI | `scrapling/core/shell.py`, `scrapling/cli.py` | IPython shell, curl-to-Scrapling converter, `extract` command, Markdown conversion |

## How a request flows

Take a spider whose `parse` does `response.css(".price::text", adaptive=True)`:

1. **Schedule.** `Scheduler.enqueue` fingerprints each `Request` (SHA-1 of session id, method, canonical URL and body, plus kwargs or headers if configured) ([request.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/request.py#L95-L125)). It drops ones it has seen unless `dont_filter` is set, and pushes the rest onto an `asyncio.PriorityQueue` ([scheduler.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/scheduler.py#L12-L60)).
2. **Politeness.** `CrawlerEngine._process_request` checks robots.txt when `robots_txt_obey` is on. It takes the larger of `download_delay`, `Crawl-delay` and `Request-rate` as the floor, and may answer from the development cache ([engine.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/engine.py#L106-L210)).
3. **Fetch.** Inside a global or per-domain `CapacityLimiter`, the engine sleeps for the AutoThrottle delay. It then calls `SessionManager.fetch`, which picks the session named by the request's `sid`: an HTTP session or a browser session ([engine.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/engine.py#L210-L268), [session.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/session.py#L103-L128)).
4. **HTTP path.** `_make_request` picks a proxy from the rotator, merges headers (a Google referer plus browserforge headers when impersonation is off), sends the request through `curl_cffi`, and retries on `CurlError`, switching proxy when the error looks like a proxy failure ([static.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/static.py#L168-L280)).
5. **Browser path.** `StealthySession.fetch` opens a pooled page, navigates with a Google referer, waits for load and network idle, runs the Cloudflare solver if asked, then `page_action` and `wait_selector`, and builds the `Response` ([_stealth.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L195-L312)).
6. **Blocked?** `spider.is_blocked` treats 401, 403, 407, 429, 444 and 5xx as blocked by default. A blocked request is copied with lower priority, stripped of its proxy, and re-queued up to `max_blocked_retries` times ([spider.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/spider.py#L65-L100), [L204-L212](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/spider.py#L204-L212)).
7. **Parse.** The callback gets a `Response`, which is a `Selector`. `css()` compiles to XPath. If nothing matches and `adaptive=True`, the stored fingerprint for that selector is loaded and the whole tree is scored to relocate the element ([parser.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/parser.py#L579-L708)). Yielded dicts go to the item list or stream, and yielded `Request`s go back to step 1.

## Key components

### Adaptive selection

With `auto_save=True`, the first match of a selector is stored as a dictionary: tag, text, attributes, DOM path, parent name, attributes and text, and sibling tags. It goes into a SQLite table keyed by the site's domain and the selector or a custom `identifier` ([parser.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/parser.py#L896-L930), [storage.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/storage.py#L74-L155)). When the selector later finds nothing, `relocate` visits every element in the page and scores it against that record with `difflib.SequenceMatcher`, averaging tag, text, attribute, class/id/href/src, path, parent and sibling similarity. It returns all elements tied at the top score if that score is at least `percentage` (default 40) ([parser.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/parser.py#L530-L578), [L822-L895](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/parser.py#L822-L895)). This is a brute-force O(n) pass per lookup. It runs only on a miss, which is the right trade-off.

### Fetchers and stealth

`Fetcher.get` and its siblings are class methods on a shared client instance ([requests.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/fetchers/requests.py#L1-L60)). Browser sessions start Playwright or Patchright Chromium. They connect over CDP if you pass `cdp_url`, launch a plain browser when a proxy rotator needs per-proxy contexts, and otherwise use a persistent context with a temporary profile ([_stealth.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L39-L106)). The stealth tier adds Chromium flags, not JavaScript patches: WebRTC limited to the proxy, optional WebGL disabling, and Chromium's own canvas noise flag ([_base.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_base.py#L540-L600)). Firefox and WebKit are not used.

### Cloudflare solver

`_cloudflare_solver` classifies the page as non-interactive, managed or embedded Turnstile. It waits out the non-interactive kind. For the others it finds the `challenges.cloudflare.com` iframe or a fallback box and clicks about 26 px into it at a jittered point with a random press delay. It re-checks and recurses, giving up after three attempts ([_stealth.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L1-L22), [L108-L193](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L108-L193)). There is no general CAPTCHA solving.

### AutoThrottle

When enabled, `AutoThrottle.record` moves each domain's delay toward `latency / target_concurrency`. On a block it doubles the delay or honours `Retry-After`, and a block never lowers the delay. The result is clamped between the spider's floor and `max_delay` ([throttle.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/throttle.py#L31-L99)).

### MCP server and AI-safe output

MCP tools such as `make_request`, `fetch` and `stealthy_fetch` wrap a session and return a `ResponseModel`. They default to Markdown and `main_content_only=True` ([ai.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/ai.py#L427-L506)). In `Convertor._extract_content`, main-content mode keeps `<body>`, drops `script`/`style`/`noscript`/`svg`, and runs `_sanitize_for_ai`. That removes CSS-hidden and `aria-hidden` elements, `<template>`, comments, zero-width and control characters before the optional `css_selector` is applied. Markdown comes from `markdownify` ([shell.py](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/shell.py#L574-L665)).

## Extending it

- **Spider hooks.** Override `parse`, `is_blocked`, `retry_blocked_request`, `on_scraped_item` and `on_error`. Use `sid` on a `Request` to send some pages through a browser session and others through HTTP.
- **Storage.** Implement `StorageSystemMixin` to keep adaptive fingerprints somewhere other than SQLite.
- **Browser automation.** `page_setup` (before navigation) and `page_action` (after) receive the raw Playwright page. `init_script`, `extra_flags` and `additional_args` reach the browser and context.
- **Scrapy.** `scrapling/integrations/scrapy.py` lets existing Scrapy callbacks parse with `Selector`.

## Running it

- `pip install scrapling` gives the parser only. `pip install "scrapling[fetchers]"` then `scrapling install` adds the fetchers and browsers. `[ai]` adds the MCP server (`scrapling mcp` or `scrapling-mcp`, stdio or streamable HTTP with optional bearer auth). `[shell]` adds the IPython shell.
- The CLI `scrapling extract get <url> out.md` fetches a page and writes Markdown, HTML or text based on the file extension.
- Docker images ship with the browsers installed. No external service is required. Adaptive data lives in a local SQLite file.

## Strengths and caveats

- **Strength: adaptive selectors.** Relocating elements by similarity after a redesign is unusual and practical, and it costs nothing until a selector misses.
- **Strength: one API across fetch tiers.** HTTP, browser and stealth browser all return the same `Response`/`Selector`, so moving a site to a stronger tier is a one-line change.
- **Strength: injection-aware output.** Removing hidden text before handing pages to an agent is a sensible default that few scrapers have.
- **Caveat: no LLM extraction.** No schema-to-JSON, chunking or model calls. You write selectors or let the calling agent do the reading.
- **Caveat: Chromium only.** The stealth tier is Patchright plus launch flags and a Turnstile clicker. Akamai, DataDome or Kasada pages need outside help.
- **Caveat: single process.** The scheduler, dedup set and throttle are in memory. Checkpoints give pause and resume, not distribution.
- **Caveat: politeness is opt-in.** `robots_txt_obey` and AutoThrottle default to off, and `concurrent_requests` is 4 with no per-domain cap.

*Sources: code at 43dee00, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (17 pages), verified Q&A.*

## How D4Vinci/Scrapling answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

Scrapling has three fetching tiers in `scrapling/fetchers/`:

**Plain HTTP (Fetcher/AsyncFetcher/FetcherSession)** — Built on `curl_cffi` in `scrapling/engines/static.py`. Impersonates browser TLS fingerprints via `impersonate` (defaults to latest Chrome). Supports GET/POST/PUT/DELETE, HTTP/3, session persistence, 3 retries by default, and SSRF-safe redirects. No JS rendering.

**Dynamic browser (DynamicFetcher/DynamicSession)** — Playwright Chromium via `scrapling/engines/_browsers/_controllers.py`. Runs a real browser (headless default), loads JS, supports `load_dom`, `network_idle` (500ms idle wait), `wait_selector`, and custom `page_action`. Connects to remote browsers via CDP (`cdp_url`). Ad blocking across ~3,500 domains from `scrapling/engines/toolbelt/ad_domains.py`. Page pooling in `scrapling/engines/_browsers/_page.py`.

**Stealthy browser (StealthyFetcher/StealthySession)** — Playwright/Patchright via `scrapling/engines/_browsers/_stealth.py`. Adds `solve_cloudflare` (detects and clicks turnstile/interstitial challenges at randomized coordinates, retries up to 3 times), canvas noise (`hide_canvas`), WebRTC proxy locking (`block_webrtc`), and WebGL preservation. Uses Patchright (Playwright fork) for undetectable automation. Realistic user-agent generation via `browserforge` in `scrapling/engines/toolbelt/fingerprints.py`.

**Content types**: `ResponseFactory` in `scrapling/engines/toolbelt/convertor.py` handles HTML by extracting DOM content and non-HTML (PDFs, images) by returning raw response body. All return a unified `Response` object (subclass of `Selector`).


Citations: [scrapling/engines/static.py:224-279](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/static.py#L224-L279) · [scrapling/engines/_browsers/_controllers.py:72-98](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_controllers.py#L72-L98) · [scrapling/engines/_browsers/_page.py:46-100](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_page.py#L46-L100) · [scrapling/engines/_browsers/_stealth.py:108-193](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L108-L193) · [scrapling/engines/toolbelt/convertor.py:27-326](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/toolbelt/convertor.py#L27-L326) · [scrapling/engines/toolbelt/fingerprints.py:1-60](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/toolbelt/fingerprints.py#L1-L60)

### How is content extracted or converted? (answered)

Extraction uses the `Selector` parser in `scrapling/parser.py` and the `Convertor` class in `scrapling/core/shell.py`.

**Parsing** — `Selector` wraps lxml's `HtmlElement` and provides CSS selectors (via cssselect), XPath, BS4-style `find_all()` by tag/class/attrs, `find_by_text()`, and regex-based `find_by_regex()`. Pseudo-elements `::text`, `::attr(name)` (Scrapy/Parsel-compatible) are supported. Element data persists in SQLite (`scrapling/core/storage.py`) for adaptive selection.

**Adaptive selection** — `auto_save=True` saves element fingerprints; `adaptive=True` re-locates changed elements using `difflib.SequenceMatcher`. `find_similar()` finds structurally similar elements.

**HTML to Markdown** — `Convertor._convert_to_markdown()` uses markdownify (optional `[rag]` dependency). The pipeline in `_extract_content()` (shell.py:622-660): (1) optional CSS selector narrowing, (2) `main_content_only` scopes to `<body>`, (3) `_strip_noise_tags()` removes `<script>/<style>/<svg>`, (4) `_sanitize_for_ai()` strips CSS-hidden elements (`display:none`, `visibility:hidden`, `opacity:0`), `aria-hidden`, `<template>`, HTML comments, zero-width Unicode, and control characters (prompt-injection defense), (5) markdownify converts cleaned HTML, or returns raw HTML or plain text per `extraction_type`.

**Schema-based extraction** — Not built in. Users build structured extraction with `.get()`, `.getall()`, `.attrib`. The MCP server returns `ResponseModel` with `status`, `url`, and `content` lists in the chosen format.


Citations: [scrapling/parser.py:77-120](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/parser.py#L77-L120) · [scrapling/core/shell.py:604-619](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/shell.py#L604-L619) · [scrapling/core/storage.py:1-30](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/storage.py#L1-L30) · [scrapling/core/ai.py:96-153](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/ai.py#L96-L153)

### How are LLMs used, if at all? (answered)

Scrapling does **not** call any LLM in its core pipeline — no API calls to OpenAI, Anthropic, or others. It provides LLM-adjacent infrastructure:

**MCP Server** (`scrapling/core/ai.py`, class `ScraplingMCPServer`) — Implements MCP to let AI agents (Claude, Cursor) use Scrapling as a tool. ~12 tools: `make_request`/`bulk_get` (HTTP), `fetch`/`bulk_fetch` (Playwright), `stealthy_fetch`/`bulk_stealthy_fetch` (stealth browser), session management, and `screenshot`. All fetch tools return a `ResponseModel` with status, URL, and content (Markdown by default). Server instructions tell agents to use `css_selector` to narrow content and save tokens. Supports bearer-auth-protected HTTP transport or stdio.

**Prompt-injection sanitization** — `Convertor._sanitize_for_ai()` (shell.py:604-619) strips hidden elements, `<template>` tags, HTML comments, zero-width Unicode, and control characters before content reaches the LLM, preventing hidden injection text.

**RAG-ready Markdown** — `SiteToMarkdownSpider` template (`scrapling/spiders/templates/site_to_markdown.py`) crawls entire sites to Markdown for RAG ingestion, with `css_selector`, `main_content_only`, and `output_dir` controls.

**Agent Skill** — A skill file at `agent-skill/Scrapling-Skill/` teaching coding agents the current API so generated code doesn't guess outdated interfaces.

**No chunking, cost controls, or structured output enforcement** — those are the calling agent's responsibility.


Citations: [scrapling/core/shell.py:604-619](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/shell.py#L604-L619) · [scrapling/core/ai.py:1082-1096](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/ai.py#L1082-L1096) · [scrapling/spiders/templates/site_to_markdown.py:30-58](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/templates/site_to_markdown.py#L30-L58)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot bypass is a first-class feature with several layers:

**TLS fingerprint impersonation** (`scrapling/engines/static.py`) — The `impersonate` parameter lets `curl_cffi` mimic Chrome/Firefox TLS fingerprints at the wire level. Accepts a single browser string or a list for random selection. HTTP/3 available. `stealthy_headers=True` (default) sets real browser headers and a Google referer via `_headers_job()`.

**Stealth browser patches** (`scrapling/engines/_browsers/_stealth.py` and `_base.py`) — `hide_canvas` adds random noise to canvas fingerprinting; `block_webrtc` forces WebRTC through the proxy to prevent WebRTC-based IP leaks; `allow_webgl` keeps WebGL active (some WAFs check for it). Uses Patchright (a Playwright fork) for undetectable automation. Realistic user-agent generation via `browserforge` (`scrapling/engines/toolbelt/fingerprints.py`) keyed to the detected OS and Chromium version.

**Cloudflare Turnstile solver** (`scrapling/engines/_browsers/_stealth.py:108-193`) — `solve_cloudflare=True` detects the challenge type (non-interactive, standard turnstile, embedded turnstile), locates the Cloudflare iframe, and clicks the checkbox at randomized coordinates with human-like mouse delays. Retries up to 3 times.

**Proxy rotation** (`scrapling/engines/toolbelt/proxy_rotation.py`) — `ProxyRotator` is thread-safe with pluggable rotation strategies (default cyclic). On connection errors matching `_PROXY_ERROR_INDICATORS`, the HTTP engine retries with the next proxy.

**DNS leak prevention** — Available via `dns_over_https` (Cloudflare DoH) in browser sessions.

**Rate limiting** — `AutoThrottle` (`scrapling/spiders/throttle.py`) doubles delay on blocked responses (status codes 401/403/407/429/444/500/502/503/504 from `spider.py:16`), respects `Retry-After` headers, and reduces latency on healthy responses.

**No native CAPTCHA-solving** beyond Cloudflare Turnstile. Docs direct to a partner API for Akamai/DataDome/Kasada/Incapsula.


Citations: [scrapling/engines/static.py:102-191](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/static.py#L102-L191) · [scrapling/engines/_browsers/_stealth.py:108-193](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/_browsers/_stealth.py#L108-L193) · [scrapling/engines/toolbelt/fingerprints.py:1-60](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/toolbelt/fingerprints.py#L1-L60) · [scrapling/spiders/throttle.py:31-85](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/throttle.py#L31-L85) · [scrapling/spiders/spider.py:16-16](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/spider.py#L16-L16)

### How is crawling at scale implemented? (answered)

Crawling is implemented in `scrapling/spiders/` with a Scrapy-inspired architecture:

**Spider base class** (`scrapling/spiders/spider.py`) — Users subclass `Spider` (or `CrawlSpider`/`SitemapSpider`/`ShopifySpider` templates) defining `name`, `start_urls`, `allowed_domains`, and an async `parse(response)` yielding items or `Request` objects.

**Scheduler** (`scrapling/spiders/scheduler.py`) — `asyncio.PriorityQueue`-based with URL deduplication via SHA-1 fingerprints. Duplicates dropped unless `dont_filter=True`. Supports snapshot/restore for checkpoint persistence (scheduler.py:31-80).

**Engine** (`scrapling/spiders/engine.py`) — `CrawlerEngine` orchestrates. Concurrency: `anyio.CapacityLimiter` — global (`concurrent_requests`, default 4) and per-domain (`concurrent_requests_per_domain`). Items stream via `anyio.create_memory_object_stream` for real-time `stream()` iteration.

**AutoThrottle** (`scrapling/spiders/throttle.py`) — Tunes per-domain delays from observed response latency. `record()` doubles delay on blocked responses or respects `Retry-After` headers; speeds back up on healthy responses. Delays clamped between `start_delay` (5s default) and `max_delay` (60s default).

**Robots.txt** (`scrapling/spiders/robotstxt.py`) — Optional `robots_txt_obey` flag. Fetches and caches per-domain robots.txt via `Protego`, checking `can_fetch()`, `Crawl-delay`, and `Request-rate` directives.

**Pause/Resume** (`scrapling/spiders/checkpoint.py`) — `CheckpointManager` saves scheduler state to disk periodically (default 5min) and on graceful shutdown. Restarting with the same `crawldir` resumes.

**Multi-session routing** (`scrapling/spiders/session.py`) — `SessionManager` handles different session types by ID. Requests carry `sid` to route through HTTP or browser sessions.

**Templates** — `CrawlSpider` with rule-based link following via `LinkExtractor` (allow/deny, domain/extension filters), `SitemapSpider`, `XMLFeedSpider`/`CSVFeedSpider`, `ShopifySpider`, and `SiteToMarkdownSpider`.

**No distributed workers** — single-process with async concurrency.


Citations: [scrapling/spiders/scheduler.py:12-80](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/scheduler.py#L12-L80) · [scrapling/spiders/engine.py:29-145](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/engine.py#L29-L145) · [scrapling/spiders/throttle.py:31-85](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/throttle.py#L31-L85) · [scrapling/spiders/robotstxt.py:10-60](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/robotstxt.py#L10-L60) · [scrapling/spiders/spider.py:65-95](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/spider.py#L65-L95) · [scrapling/spiders/templates/site_to_markdown.py:30-58](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/spiders/templates/site_to_markdown.py#L30-L58)

### What is the developer interface? (answered)

Scrapling provides multiple interfaces:

**Python library API** — Primary interface. Import fetchers (`Fetcher`, `AsyncFetcher`, `StealthyFetcher`, `DynamicFetcher`) for stateless one-off use, or session classes (`FetcherSession`, `StealthySession`, `DynamicSession`) with `with`/`async with`. All return a `Response` object (subclass of `Selector` parser) for CSS/XPath/text extraction. Standalone parser: `from scrapling.parser import Selector`.

**Spider framework** — Scrapy-like class-based API with async `parse()` methods yielding items/requests. Built-in export: `result.items.to_json()`, `to_jsonl()`, `to_csv()`, `to_xml()`. Streaming via `async for item in spider.stream()`.

**CLI** (`scrapling/cli.py`) — `scrapling shell` (IPython-based interactive shell with Scrapling integration and curl-to-Scrapling conversion), `scrapling extract <method> <url> <output_file>` (format auto-detected from extension: `.html`/`.md`/`.txt`), `scrapling install` (browser deps).

**MCP Server** (`scrapling/core/ai.py`) — Invoked via `scrapling-mcp`. ~12 tools over stdio or streamable-http with optional bearer auth. Session management with create/fetch/close/list lifecycle. Screenshot tool returns images.

**Agent Skill** — `agent-skill/Scrapling-Skill/` file teaching coding agents the current API so generated code is accurate.

**Docker** — `pyd4vinci/scrapling` (DockerHub) and `ghcr.io/d4vinci/scrapling:latest` (GHCR) with all browsers pre-installed.

**Scrapy integration** (`scrapling/integrations/scrapy.py`) — `scrapling_response` decorator lets existing Scrapy callbacks parse with Scrapling's parser.

**Output formats** — Items export to JSON/JSONL/CSV/XML. Page content exports to HTML/Markdown/plain text via file extension or `extraction_type`.


Citations: [scrapling/cli.py:1-62](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/cli.py#L1-L62) · [scrapling/core/ai.py:182-240](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/ai.py#L182-L240) · [scrapling/core/ai.py:1069-1188](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/core/ai.py#L1069-L1188) · [scrapling/engines/toolbelt/custom.py:28-50](https://github.com/D4Vinci/Scrapling/blob/43dee004866e1c46843a9ac293cc1494aa7915d6/scrapling/engines/toolbelt/custom.py#L28-L50)
