# scraperai/scraperai

> Python library and CLI that uses GPT-4o once to generate an XPath scraping config, then scrapes catalogs with lxml and no LLM calls.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/scraperai/scraperai (reviewed at commit `5d914ef0f00d322d86acd159e3a0e100745c8da0`, 2025-09-18)
- Stars: 485 · Language: HTML · License: GPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/scraperai/

## Overview

ScraperAI uses an LLM to *write* a scraper, then runs that scraper without the LLM. You point it at a catalog page (a product list, a job board, a directory) or a single detail page. GPT-4o classifies the page, finds the "next page" button, finds the XPath of one repeated card, and proposes XPaths for the fields inside that card. All of this lands in a `ScraperConfig`, a Pydantic model you can save as JSON. From then on, `Scraper` walks the pages with plain `lxml` XPath queries. No model calls happen during the scrape.

That split is the project's main idea, and it matters for cost. LLM spend is paid once per site layout, not once per page. Thousands of rows cost the same in tokens as ten. The trade-off is the usual one for selector-based scraping: when the site's markup changes, the XPaths break and you re-run detection.

The code base is about 2,800 lines of Python. It has a Python API and an interactive Click CLI (`scraperai --url ...`) that walks you through each detection step, highlights what the model found in a live Chrome window, and lets you correct it. It is a desktop tool for an analyst, not a crawling service. Fetching is synchronous Selenium or `requests`, one page at a time. The last commit is from September 2025, and `setup.py` declares version `0.0.3`.

## Architecture

```mermaid
flowchart LR
  U["CLI Controller or your code"] --> C["Crawler: Selenium / requests"]
  U --> PA["ParserAI"]
  C --> HTML["page_source + screenshot"]
  HTML --> PA
  PA --> CL["Page type classifier"]
  PA --> PG["PaginationDetector"]
  PA --> CI["CatalogItemDetector"]
  PA --> DF["DataFieldsExtractor"]
  CL --> LLM["OpenAI gpt-4o (JSON / vision)"]
  PG --> LLM
  CI --> LLM
  DF --> LLM
  PA --> CFG["ScraperConfig (XPaths)"]
  CFG --> S["Scraper"]
  C --> S
  S --> OUT["rows -> JSON / CSV / XLSX"]
```

| Component | Path | Role |
|---|---|---|
| Config models | `scraperai/models.py` | `ScraperConfig`, `Pagination`, `CatalogItem`, `StaticField`, `DynamicField`, `WebpageType` |
| Detection facade | `scraperai/parsers/parserai.py` | `ParserAI`: wires default OpenAI models into the detectors and tracks cost |
| Retry loop | `scraperai/parsers/agent.py` | `ChatModelAgent.query_with_validation`: re-prompts with the validation error |
| Detectors | `scraperai/parsers/` | Page type, pagination, catalog card, field XPaths, relevant-HTML pruning |
| LLM wrappers | `scraperai/llm/` | `BaseJsonLM` / `BaseVision` interfaces, `JsonOpenAI` / `VisionOpenAI` implementations |
| HTML utilities | `scraperai/utils/html.py` | `minify_html`, `split_html`, XPath field extraction |
| Runtime | `scraperai/scraper.py` | `Scraper.scrape()`: LLM-free generator over pages and items |
| Crawlers | `scraperai/crawlers/` | `SeleniumCrawler`, `RequestsCrawler`, Chrome and Selenoid webdrivers |
| CLI | `scraperai/cli/` | `Controller` wizard, `View` prompts, config cache in the user data dir |

## How a request flows

The CLI path for a catalog page, which is also the order a library user follows:

1. **Start.** `Controller.run` starts a `SeleniumCrawler` (a visible Chrome window), reads `OPENAI_API_KEY` or prompts for it, and builds `ParserAI`. It then checks the user data directory for a saved `*.scraperai.json` with the same start URL and offers to reuse it, skipping all LLM steps ([controller.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L266-L326)).
2. **Classify.** `detect_page_type` loads the URL and takes a screenshot at 60 % zoom. With a screenshot, `WebpageVisionClassifier` asks the vision model for `catalog` or `detailed_page`. Without one, `WebpageTextClassifier` sends the first 12k tokens of minified HTML and accepts four labels, including `captcha` and `other` ([parserai.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/parserai.py#L54-L61), [webpage_classifier.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/webpage_classifier.py#L13-L91)).
3. **Pagination.** `PaginationDetector.find_pagination` asks for the XPath of a "Next" or "Load more" button. If none is returned, it assumes infinite scroll ([pagination_detector.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/pagination_detector.py#L23-L73)).
4. **Find the card.** `CatalogItemDetector.detect_catalog_item` sends minified HTML (only `class`, `href`, `id` kept) and asks for a `card` XPath and a `url` XPath. The validator runs both against the tree. It rejects a card XPath that matches nothing, or a pair whose match counts differ, and feeds that message back for up to five retries ([catalog_item_detector.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/catalog_item_detector.py#L30-L91)).
5. **Find fields.** The first card's HTML goes to `DataFieldsExtractor.extract_fields`. One call returns "static" fields (name plus XPath). A second call returns "dynamic" sections, which are label/value XPath pairs such as a spec table. Each XPath is evaluated to show a sample value ([data_fields_extractor.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/data_fields_extractor.py#L32-L147)). If the user chooses to open nested pages, the first detail URL is loaded and pruned first (see below).
6. **Human review.** At each step the CLI highlights the XPaths in Chrome with coloured borders and lets the user accept, type new XPaths, or add a hint that is sent back to the model ([controller.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L105-L198)).
7. **Scrape.** `Scraper.scrape_catalog_items` parses `page_source` with `lxml`, re-parses each card node as a fragment, applies the field XPaths relative to it, yields a dict, then calls `crawler.switch_page`. It stops at `max_pages`, `max_rows` or when pagination fails ([scraper.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/scraper.py#L19-L85)).
8. **Export.** Rows are written to `results_<timestamp>.json`, `.csv` or `.xlsx` with pandas.

## Key components

### Validation-driven retries

Every detector inherits `ChatModelAgent.query_with_validation`. It calls the model, runs a validator, and if that returns an error string, appends the model's answer plus "Your previous response was wrong because …" and recurses ([agent.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/agent.py#L13-L39)). The validators check real behaviour (does the XPath select nodes, do the counts line up), not just JSON shape. That is the best idea in the code base. One caveat: if the validator itself raises, the response is accepted as valid.

### HTML minification

`minify_html` removes `script`, `style`, `meta` and `noscript`, strips every attribute outside an allow-list, collapses whitespace with `htmlmin`, and can replace long `<p>` texts with placeholders ([html.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/utils/html.py#L24-L66)). Detectors then truncate to the first token chunk: 12k for classification and pagination, 16k for fields, 32k for cards. Pagination is the only detector that looks at more than the first chunk. It keeps the first and last chunks, but because the loop returns on its first pass, only the last chunk is actually sent.

### Detail-page pruning

For detail pages, `summarize_details_page_as_valid_html` calls `split_html`, which recursively splits the DOM into parts of at most 4,000 text tokens, each keyed by its XPath. One LLM call describes every part, a second picks the relevant XPaths, and the rest are removed before field detection ([webpage_descriptor.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/webpage_descriptor.py#L54-L144)). A screenshot description is also generated with the vision model and passed in as `context`, but `split_and_describe` never uses that argument. That vision call is paid for and then thrown away.

### Crawlers

`BaseCrawler` needs only `get`, `page_source` and `switch_page` ([base.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/base.py#L9-L21)). `SeleniumCrawler` sleeps 1 s after each load and 3 s after a pagination click, and scrolls 500 px for infinite scroll ([selenium.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/selenium.py#L16-L81)). `DefaultChromeWebdriver` is headful. It disables `AutomationControlled`, hides `navigator.webdriver`, loads a bundled `cookies.crx` extension, and overrides the user agent with a random one before each `get` ([local.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/local.py#L15-L55)). `RequestsCrawler` is a bare `requests.get` and only supports URL-list pagination. `WebdriversManager` picks a Selenoid grid with free slots for remote Chrome or Firefox.

## Extending it

- **Other models.** `ParserAI` accepts any `BaseJsonLM` (returns a dict) and `BaseVision` (returns a string). Each is a one-method interface, so wrapping Anthropic, a local model or a LangChain chat model is a few lines ([base.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/llm/base.py#L12-L27)). Cost tracking only works for the OpenAI classes.
- **Other fetchers.** Subclass `BaseCrawler` for Playwright, httpx or a proxy-backed client. Screenshots, highlighting and click/scroll pagination are Selenium-only features of the CLI.
- **Skip the LLM entirely.** A `ScraperConfig` is plain JSON. You can write or edit XPaths by hand and run `Scraper(config, crawler).scrape()`.
- **Pagination by URL list.** `Pagination(type='urls', urls=[...])` works with both crawlers and suits `?page=N` sites.

## Running it

- `pip install scraperai`, then `scraperai --url https://example.com/catalog`. It needs Chrome locally; `webdriver-manager` downloads the matching driver. A desktop session is needed, because the default driver opens a visible window.
- Set `OPENAI_API_KEY` in the environment or `.env`. If it is missing, the CLI asks and can write it to `.env` in the current directory.
- Saved configs go to the platform user-data directory (via `appdirs`) as `<host>_<timestamp>.scraperai.json`.
- Library use follows the notebooks in `examples/`: build a crawler and `ParserAI`, call the `detect_*` and `extract_fields` methods, assemble a `ScraperConfig`, iterate `Scraper.scrape()`.

## Strengths and caveats

- **Strength: LLM cost is per layout, not per page.** The output is a reusable XPath recipe, and the scrape itself is deterministic and free of model calls.
- **Strength: self-checking prompts.** Validators execute the proposed XPaths and push concrete errors back to the model. This catches many hallucinated selectors before a human sees them.
- **Strength: human-in-the-loop CLI.** Visual highlighting and plain-language corrections make it usable by non-programmers.
- **Caveat: brittle by design.** XPaths (often class-based) break on redesigns or hashed CSS class names. Nothing detects drift during a run; fields just come back `None`.
- **Caveat: truncation.** Most detectors see only the first 12k to 32k tokens of minified HTML. A catalog that starts deep in a heavy page can be missed.
- **Caveat: rough edges.** If the user gives a prompt on the pagination screen, the CLI calls `ParserAI.generate_pagination_urls`, which does not exist ([controller.py](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L92-L103)). `StaticField.multiple` is always set to `False`, a bundled `PythonCodeOpenAI` model is created but never used, and the vision description is discarded.
- **Caveat: no scale features.** It is synchronous and single-browser. It has no proxy support, retries, robots.txt handling or CAPTCHA solving (a `captcha` page just prompts the user to solve it by hand). Dependencies are pinned to `langchain==0.1.16` and `selenium==4.9.1`, both old.

*Sources: code at 5d914ef, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (15 pages), verified Q&A.*

## How scraperai/scraperai answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

ScraperAI offers two crawler implementations. **RequestsCrawler** (`scraperai/crawlers/requests.py:7-40`) uses plain `requests.get(url).text` — no JavaScript rendering, no browser, just raw HTTP. **SeleniumCrawler** (`scraperai/crawlers/selenium.py:16-82`) wraps Selenium WebDriver (Chrome by default) and provides full browser execution including JS rendering, clicking, scrolling, and screenshot capture. The default driver is `DefaultChromeWebdriver` (`scraperai/crawlers/webdriver/local.py:15-49`), which creates a non-headless Chrome window (headless is commented out). After each `driver.get(url)` call, the `SeleniumCrawler` applies a blanket `time.sleep(1)` — no explicit waiting strategy beyond that single-second sleep. For pagination, it supports XPath-based button clicking (`switch_page` with type `xpath`), infinite scroll (`type='scroll'`), and URL-list iteration. Screenshots are taken at 60% zoom via `driver.execute_script("document.body.style.zoom='60%'")`. The project has **no PDF or image content-type handling** — it fetches HTML pages only.


Citations: [scraperai/crawlers/requests.py:7-40](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/requests.py#L7-L40) · [scraperai/crawlers/webdriver/local.py:15-49](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/local.py#L15-L49) · [scraperai/models.py:46-60](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/models.py#L46-L60)

### How is content extracted or converted? (answered)

Content is extracted from HTML via XPath-based field extraction, not by converting HTML to Markdown. The core extraction pipeline is in `scraperai/parsers/utils.py:32-47`: given a set of `StaticField` and `DynamicField` objects (each containing an XPath string), it calls `extract_field_by_xpath` and `extract_dynamic_fields_by_xpath` from `scraperai/utils/html.py:168-198`. Static fields extract single or multiple text values from a node; dynamic fields pair two XPaths (one for labels, one for values) to produce key–value dictionaries. The `DataFieldsExtractor` (`scraperai/parsers/data_fields_extractor.py:26-154`) uses an LLM (gpt-4o via `JsonOpenAI`) to *discover* these XPaths from a raw HTML snippet. It prompts the model to return JSON with XPaths, then validates the XPaths actually resolve in the HTML. There is **no readability-style boilerplate removal** (no `trafilatura`, `readability`, or `goose`). Instead, a separate LLM-based mechanism (`WebpagePartsDescriptor` in `scraperai/parsers/webpage_descriptor.py:54-144`) splits the page into chunks via `split_html` (`scraperai/utils/html.py:90-130`), has GPT-4 describe each chunk, then mark chunks irrelevant and remove them by XPath — a form of AI-powered content pruning. HTML is pre-minified using `htmlmin` and BeautifulSoup to strip script/style/meta/noscript tags and remove non-essential attributes before being sent to the LLM (`scraperai/utils/html.py:24-66`).


Citations: [scraperai/parsers/utils.py:32-47](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/utils.py#L32-L47) · [scraperai/utils/html.py:168-198](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/utils/html.py#L168-L198) · [scraperai/parsers/data_fields_extractor.py:26-154](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/data_fields_extractor.py#L26-L154) · [scraperai/parsers/webpage_descriptor.py:54-144](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/webpage_descriptor.py#L54-L144) · [scraperai/utils/html.py:24-66](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/utils/html.py#L24-L66)

### How are LLMs used, if at all? (answered)

LLMs are the central intelligence of ScraperAI — every detection and extraction step calls OpenAI GPT-4. The only supported provider is **OpenAI** (`scraperai/llm/openai.py`). Three model abstractions exist: `JsonOpenAI` (`scraperai/llm/openai.py:43-82`) uses `gpt-4o` with `response_format: {"type": "json_object"}` and optionally `with_structured_output(schema, method='json_mode')` for Pydantic-validated JSON responses. `VisionOpenAI` (`scraperai/llm/openai.py:84-106`) uses `gpt-4o` for vision tasks (page classification, webpage description from screenshots). `PythonCodeOpenAI` (`scraperai/llm/openai.py:16-40`) extracts Python code from responses for code generation tasks. All models track token cost via `langchain_community.callbacks.get_openai_callback`. **Page chunking** is done with `TokenTextSplitter` from `langchain_text_splitters` at sizes of 12,000–16,000 tokens (e.g., `scraperai/parsers/data_fields_extractor.py:30` sets `max_chunk_size = 16000`). When the HTML exceeds the chunk size, only the first chunk is sent. For pagination detection, the first and last chunks are both sent (`scraperai/parsers/pagination_detector.py:64-65`). **Structured output** is enforced by `ChatModelAgent.query_with_validation` (`scraperai/parsers/agent.py:13-39`), which validates the model's JSON response against Pydantic schemas and retries up to 3–5 times with the validation error as feedback. **No cost controls** exist beyond the retry limit — there is no token budget, no model fallback, and no caching. The `ParserAI.total_cost` property (`scraperai/parsers/parserai.py:46-52`) sums costs from both `JsonOpenAI` and `VisionOpenAI` instances. The `.env.example` shows that only `OPENAI_API_KEY` is required.

> **Editor's note.** Correction: OpenAI is the bundled default, not the only option. `ParserAI` takes any `BaseJsonLM` and `BaseVision` implementation (one `invoke` method each, scraperai/llm/base.py L12-L27), so other providers plug in; only cost tracking is OpenAI-specific. `PythonCodeOpenAI` is constructed but never called.

Citations: [scraperai/llm/openai.py:43-82](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/llm/openai.py#L43-L82) · [scraperai/llm/openai.py:84-106](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/llm/openai.py#L84-L106) · [scraperai/parsers/agent.py:13-39](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/agent.py#L13-L39) · [scraperai/parsers/data_fields_extractor.py:26-54](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/data_fields_extractor.py#L26-L54) · [scraperai/parsers/parserai.py:46-52](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/parsers/parserai.py#L46-L52)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot measures are **minimal and basic**. The `DefaultChromeWebdriver` (`scraperai/crawlers/webdriver/local.py:15-49`) applies the following stealth patches on Chrome options and runtime:

- `--disable-blink-features=AutomationControlled` — removes the "Chrome is being controlled" banner
- `excludeSwitches: ["enable-automation"]` — prevents Selenium from advertising automation
- `useAutomationExtension: False` — disables the automation extension
- JavaScript injection after driver init: `Object.defineProperty(navigator, 'webdriver', {get: () => undefined})` — masks the `navigator.webdriver` flag
- A random user-agent is set **per request** in the `get()` method via CDP: `execute_cdp_cmd("Network.setUserAgentOverride", {"userAgent": get_random_useragent()})` — agents are drawn from a text file (`scraperai/crawlers/webdriver/useragents.py:1-12`)

The same patterns are duplicated in `RemoteWebdriver` (`scraperai/crawlers/webdriver/remote.py:46-61`) for Selenoid-based remote drivers.

**CAPTCHA handling:** The `WebpageType.CAPTCHA` enum value exists (`scraperai/models.py:66`) and the CLI flow detects CAPTCHA pages. However, `controller.py:290-292` simply shows a warning message and asks the user to solve it manually (`view.py:45-47`: "Captcha detected! Please solve it before we can continue"). There is **no automatic CAPTCHA solving** (roadmap mentions "anti-captcha integration" as a future item).

**No proxy rotation, no rate limiting, no request throttling, and no fingerprint-level stealth** (no TLS fingerprint spoofing, no TCP-level masking). The `requests.get()`-based `RequestsCrawler` sets no custom headers at all — just the library defaults.


Citations: [scraperai/crawlers/webdriver/local.py:15-49](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/local.py#L15-L49) · [scraperai/crawlers/webdriver/remote.py:46-88](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/remote.py#L46-L88) · [scraperai/crawlers/webdriver/useragents.py:1-12](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/useragents.py#L1-L12) · [scraperai/models.py:63-71](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/models.py#L63-L71) · [scraperai/cli/controller.py:286-293](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L286-L293)

### How is crawling at scale implemented? (answered)

ScraperAI is **not designed for crawling at scale** — all implementations are single-threaded and synchronous. The `Scraper` class (`scraperai/scraper.py:14-86`) coordinates the crawl via a pull-based generator pattern, yielding one row at a time. For **catalog pages**, it loops: fetch page, parse HTML with lxml, extract items from each card (matching `catalog_item.card_xpath`), then call `crawler.switch_page(pagination)` to advance. The three pagination types (`xpath`, `scroll`, `urls`) are documented in `scraperai/models.py:46-60`. For **nested detail pages**, the scraper first collects all URLs from the catalog via `scrape_nested_items_urls` (extracting hrefs by `url_xpath`), then iterates over them calling `crawler.get(url)` per item. **Concurrency: zero.** The `SeleniumCrawler` holds one browser window. No async, no threading, no parallel workers. There is a `WebdriversManager` (`scraperai/crawlers/webdriver/manager.py:27-61`) that can distribute Selenium sessions across a pool of Selenoid servers (checking `/status` for active session counts) — but this is only used to *create* a single driver, not to parallelize work. **URL deduplication** happens only at the CLI level via a `set()` (`controller.py:215`). **No robots.txt** parsing exists anywhere in the codebase. **No politeness delays** — the only waits are the hard-coded `time.sleep(1)` in `SeleniumCrawler.get()` and `time.sleep(3)` after pagination clicks. **No distributed workers, no message queues, no crawl frontier.** The README roadmap lists "Add httpx and aiohttp crawlers" as a future improvement, confirming async/distributed is not yet implemented.


Citations: [scraperai/models.py:46-60](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/models.py#L46-L60) · [scraperai/crawlers/webdriver/manager.py:27-61](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/webdriver/manager.py#L27-L61) · [scraperai/cli/controller.py:214-225](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L214-L225) · [scraperai/crawlers/selenium.py:60-81](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/crawlers/selenium.py#L60-L81)

### What is the developer interface? (answered)

The project provides two main developer interfaces and a serializable config format.

**Python Library API** — the primary interface. Users import `Scraper`, `ParserAI`, `SeleniumCrawler`, and model classes (`scraperai/__init__.py:1-4`). The typical flow: create a `SeleniumCrawler`, instantiate `ParserAI(openai_api_key=...)`, call `detect_page_type()`, `detect_pagination()`, `detect_catalog_item()`, `extract_fields()`, assemble a `ScraperConfig`, and pass it with the crawler to `Scraper(config, crawler)`. The scraper exposes a generator-based `scrape()` method. This is shown in detail in `examples/ycombinator_full.ipynb`.

**CLI** — built with Click (`scraperai/cli/app.py`), entry point `scraperai` (via `console_scripts` in `setup.py:28-30`). Run `scraperai --url <url>`. It drives the same library classes through a step-by-step interactive wizard (`Controller` in `scraperai/cli/controller.py:21-334`): checks for cached configs, auto-detects page type/pagination/catalog items/fields, lets the user edit each detection, then scrapes and exports. The `View` class (`scraperai/cli/view.py:11-204`) handles all terminal I/O with Click prompts and `pandas`-based markdown tables for field display.

**ScraperConfig serialization** (`scraperai/models.py:74-83`) — the full scraping configuration (URL, page type, pagination settings, catalog item XPaths, field definitions, limits) is a Pydantic model that dumps to JSON. Configs are saved as `.scraperai.json` files in the user's app data directory (`scraperai/cli/utils.py:9`) and are loaded on subsequent runs to skip re-detection (`controller.py:42-51`).

**No REST service, no MCP server, no web UI.** The README roadmap lists "Release SaaS web app" as a future goal. **No REST API** — the library is consumed in-process only. **Language bindings:** Python only (package on PyPI). Output formats are JSON, XLSX, and CSV.


Citations: [scraperai/__init__.py:1-5](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/__init__.py#L1-L5) · [scraperai/cli/app.py:28-51](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/app.py#L28-L51) · [scraperai/cli/controller.py:21-86](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/scraperai/cli/controller.py#L21-L86) · [setup.py:13-32](https://github.com/scraperai/scraperai/blob/5d914ef0f00d322d86acd159e3a0e100745c8da0/setup.py#L13-L32)
