scraperai/scraperai
Python library and CLI that uses GPT-4o once to generate an XPath scraping config, then scrapes catalogs with lxml and no LLM calls.
Overview
ScraperAI uses an LLM to write a scraper, then runs that scraper without the LLM. You point it at a catalog page (a product list, a job board, a directory) or a single detail page. GPT-4o classifies the page, finds the “next page” button, finds the XPath of one repeated card, and proposes XPaths for the fields inside that card. All of this lands in a ScraperConfig, a Pydantic model you can save as JSON. From then on, Scraper walks the pages with plain lxml XPath queries. No model calls happen during the scrape.
That split is the project’s main idea, and it matters for cost. LLM spend is paid once per site layout, not once per page. Thousands of rows cost the same in tokens as ten. The trade-off is the usual one for selector-based scraping: when the site’s markup changes, the XPaths break and you re-run detection.
The code base is about 2,800 lines of Python. It has a Python API and an interactive Click CLI (scraperai --url ...) that walks you through each detection step, highlights what the model found in a live Chrome window, and lets you correct it. It is a desktop tool for an analyst, not a crawling service. Fetching is synchronous Selenium or requests, one page at a time. The last commit is from September 2025, and setup.py declares version 0.0.3.
Architecture
flowchart LR
U["CLI Controller or your code"] --> C["Crawler: Selenium / requests"]
U --> PA["ParserAI"]
C --> HTML["page_source + screenshot"]
HTML --> PA
PA --> CL["Page type classifier"]
PA --> PG["PaginationDetector"]
PA --> CI["CatalogItemDetector"]
PA --> DF["DataFieldsExtractor"]
CL --> LLM["OpenAI gpt-4o (JSON / vision)"]
PG --> LLM
CI --> LLM
DF --> LLM
PA --> CFG["ScraperConfig (XPaths)"]
CFG --> S["Scraper"]
C --> S
S --> OUT["rows -> JSON / CSV / XLSX"]
| Component | Path | Role |
|---|---|---|
| Config models | scraperai/models.py |
ScraperConfig, Pagination, CatalogItem, StaticField, DynamicField, WebpageType |
| Detection facade | scraperai/parsers/parserai.py |
ParserAI: wires default OpenAI models into the detectors and tracks cost |
| Retry loop | scraperai/parsers/agent.py |
ChatModelAgent.query_with_validation: re-prompts with the validation error |
| Detectors | scraperai/parsers/ |
Page type, pagination, catalog card, field XPaths, relevant-HTML pruning |
| LLM wrappers | scraperai/llm/ |
BaseJsonLM / BaseVision interfaces, JsonOpenAI / VisionOpenAI implementations |
| HTML utilities | scraperai/utils/html.py |
minify_html, split_html, XPath field extraction |
| Runtime | scraperai/scraper.py |
Scraper.scrape(): LLM-free generator over pages and items |
| Crawlers | scraperai/crawlers/ |
SeleniumCrawler, RequestsCrawler, Chrome and Selenoid webdrivers |
| CLI | scraperai/cli/ |
Controller wizard, View prompts, config cache in the user data dir |
How a request flows
The CLI path for a catalog page, which is also the order a library user follows:
- Start.
Controller.runstarts aSeleniumCrawler(a visible Chrome window), readsOPENAI_API_KEYor prompts for it, and buildsParserAI. It then checks the user data directory for a saved*.scraperai.jsonwith the same start URL and offers to reuse it, skipping all LLM steps (controller.py). - Classify.
detect_page_typeloads the URL and takes a screenshot at 60 % zoom. With a screenshot,WebpageVisionClassifierasks the vision model forcatalogordetailed_page. Without one,WebpageTextClassifiersends the first 12k tokens of minified HTML and accepts four labels, includingcaptchaandother(parserai.py, webpage_classifier.py). - Pagination.
PaginationDetector.find_paginationasks for the XPath of a “Next” or “Load more” button. If none is returned, it assumes infinite scroll (pagination_detector.py). - Find the card.
CatalogItemDetector.detect_catalog_itemsends minified HTML (onlyclass,href,idkept) and asks for acardXPath and aurlXPath. The validator runs both against the tree. It rejects a card XPath that matches nothing, or a pair whose match counts differ, and feeds that message back for up to five retries (catalog_item_detector.py). - Find fields. The first card’s HTML goes to
DataFieldsExtractor.extract_fields. One call returns “static” fields (name plus XPath). A second call returns “dynamic” sections, which are label/value XPath pairs such as a spec table. Each XPath is evaluated to show a sample value (data_fields_extractor.py). If the user chooses to open nested pages, the first detail URL is loaded and pruned first (see below). - Human review. At each step the CLI highlights the XPaths in Chrome with coloured borders and lets the user accept, type new XPaths, or add a hint that is sent back to the model (controller.py).
- Scrape.
Scraper.scrape_catalog_itemsparsespage_sourcewithlxml, re-parses each card node as a fragment, applies the field XPaths relative to it, yields a dict, then callscrawler.switch_page. It stops atmax_pages,max_rowsor when pagination fails (scraper.py). - Export. Rows are written to
results_<timestamp>.json,.csvor.xlsxwith pandas.
Key components
Validation-driven retries
Every detector inherits ChatModelAgent.query_with_validation. It calls the model, runs a validator, and if that returns an error string, appends the model’s answer plus “Your previous response was wrong because …” and recurses (agent.py). The validators check real behaviour (does the XPath select nodes, do the counts line up), not just JSON shape. That is the best idea in the code base. One caveat: if the validator itself raises, the response is accepted as valid.
HTML minification
minify_html removes script, style, meta and noscript, strips every attribute outside an allow-list, collapses whitespace with htmlmin, and can replace long <p> texts with placeholders (html.py). Detectors then truncate to the first token chunk: 12k for classification and pagination, 16k for fields, 32k for cards. Pagination is the only detector that looks at more than the first chunk. It keeps the first and last chunks, but because the loop returns on its first pass, only the last chunk is actually sent.
Detail-page pruning
For detail pages, summarize_details_page_as_valid_html calls split_html, which recursively splits the DOM into parts of at most 4,000 text tokens, each keyed by its XPath. One LLM call describes every part, a second picks the relevant XPaths, and the rest are removed before field detection (webpage_descriptor.py). A screenshot description is also generated with the vision model and passed in as context, but split_and_describe never uses that argument. That vision call is paid for and then thrown away.
Crawlers
BaseCrawler needs only get, page_source and switch_page (base.py). SeleniumCrawler sleeps 1 s after each load and 3 s after a pagination click, and scrolls 500 px for infinite scroll (selenium.py). DefaultChromeWebdriver is headful. It disables AutomationControlled, hides navigator.webdriver, loads a bundled cookies.crx extension, and overrides the user agent with a random one before each get (local.py). RequestsCrawler is a bare requests.get and only supports URL-list pagination. WebdriversManager picks a Selenoid grid with free slots for remote Chrome or Firefox.
Extending it
- Other models.
ParserAIaccepts anyBaseJsonLM(returns a dict) andBaseVision(returns a string). Each is a one-method interface, so wrapping Anthropic, a local model or a LangChain chat model is a few lines (base.py). Cost tracking only works for the OpenAI classes. - Other fetchers. Subclass
BaseCrawlerfor Playwright, httpx or a proxy-backed client. Screenshots, highlighting and click/scroll pagination are Selenium-only features of the CLI. - Skip the LLM entirely. A
ScraperConfigis plain JSON. You can write or edit XPaths by hand and runScraper(config, crawler).scrape(). - Pagination by URL list.
Pagination(type='urls', urls=[...])works with both crawlers and suits?page=Nsites.
Running it
pip install scraperai, thenscraperai --url https://example.com/catalog. It needs Chrome locally;webdriver-managerdownloads the matching driver. A desktop session is needed, because the default driver opens a visible window.- Set
OPENAI_API_KEYin the environment or.env. If it is missing, the CLI asks and can write it to.envin the current directory. - Saved configs go to the platform user-data directory (via
appdirs) as<host>_<timestamp>.scraperai.json. - Library use follows the notebooks in
examples/: build a crawler andParserAI, call thedetect_*andextract_fieldsmethods, assemble aScraperConfig, iterateScraper.scrape().
Strengths and caveats
- Strength: LLM cost is per layout, not per page. The output is a reusable XPath recipe, and the scrape itself is deterministic and free of model calls.
- Strength: self-checking prompts. Validators execute the proposed XPaths and push concrete errors back to the model. This catches many hallucinated selectors before a human sees them.
- Strength: human-in-the-loop CLI. Visual highlighting and plain-language corrections make it usable by non-programmers.
- Caveat: brittle by design. XPaths (often class-based) break on redesigns or hashed CSS class names. Nothing detects drift during a run; fields just come back
None. - Caveat: truncation. Most detectors see only the first 12k to 32k tokens of minified HTML. A catalog that starts deep in a heavy page can be missed.
- Caveat: rough edges. If the user gives a prompt on the pagination screen, the CLI calls
ParserAI.generate_pagination_urls, which does not exist (controller.py).StaticField.multipleis always set toFalse, a bundledPythonCodeOpenAImodel is created but never used, and the vision description is discarded. - Caveat: no scale features. It is synchronous and single-browser. It has no proxy support, retries, robots.txt handling or CAPTCHA solving (a
captchapage just prompts the user to solve it by hand). Dependencies are pinned tolangchain==0.1.16andselenium==4.9.1, both old.
Sources: code at 5d914ef, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (15 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 5d914ef. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredScraperAI offers two crawler implementations. RequestsCrawler (scraperai/crawlers/requests.py:7-40) uses plain requests.get(url).text — no JavaScript rendering, no browser, just raw HTTP. SeleniumCrawler (scraperai/crawlers/selenium.py:16-82) wraps Selenium WebDriver (Chrome by default) and provides full browser execution including JS rendering, clicking, scrolling, and screenshot capture. The default driver is DefaultChromeWebdriver (scraperai/crawlers/webdriver/local.py:15-49), which creates a non-headless Chrome window (headless is commented out). After each driver.get(url) call, the SeleniumCrawler applies a blanket time.sleep(1) — no explicit waiting strategy beyond that single-second sleep. For pagination, it supports XPath-based button clicking (switch_page with type xpath), infinite scroll (type='scroll'), and URL-list iteration. Screenshots are taken at 60% zoom via driver.execute_script("document.body.style.zoom='60%'"). The project has no PDF or image content-type handling — it fetches HTML pages only.
How is content extracted or converted?
answeredContent is extracted from HTML via XPath-based field extraction, not by converting HTML to Markdown. The core extraction pipeline is in scraperai/parsers/utils.py:32-47: given a set of StaticField and DynamicField objects (each containing an XPath string), it calls extract_field_by_xpath and extract_dynamic_fields_by_xpath from scraperai/utils/html.py:168-198. Static fields extract single or multiple text values from a node; dynamic fields pair two XPaths (one for labels, one for values) to produce key–value dictionaries. The DataFieldsExtractor (scraperai/parsers/data_fields_extractor.py:26-154) uses an LLM (gpt-4o via JsonOpenAI) to discover these XPaths from a raw HTML snippet. It prompts the model to return JSON with XPaths, then validates the XPaths actually resolve in the HTML. There is no readability-style boilerplate removal (no trafilatura, readability, or goose). Instead, a separate LLM-based mechanism (WebpagePartsDescriptor in scraperai/parsers/webpage_descriptor.py:54-144) splits the page into chunks via split_html (scraperai/utils/html.py:90-130), has GPT-4 describe each chunk, then mark chunks irrelevant and remove them by XPath — a form of AI-powered content pruning. HTML is pre-minified using htmlmin and BeautifulSoup to strip script/style/meta/noscript tags and remove non-essential attributes before being sent to the LLM (scraperai/utils/html.py:24-66).
How are LLMs used, if at all?
answeredLLMs are the central intelligence of ScraperAI — every detection and extraction step calls OpenAI GPT-4. The only supported provider is OpenAI (scraperai/llm/openai.py). Three model abstractions exist: JsonOpenAI (scraperai/llm/openai.py:43-82) uses gpt-4o with response_format: {"type": "json_object"} and optionally with_structured_output(schema, method='json_mode') for Pydantic-validated JSON responses. VisionOpenAI (scraperai/llm/openai.py:84-106) uses gpt-4o for vision tasks (page classification, webpage description from screenshots). PythonCodeOpenAI (scraperai/llm/openai.py:16-40) extracts Python code from responses for code generation tasks. All models track token cost via langchain_community.callbacks.get_openai_callback. Page chunking is done with TokenTextSplitter from langchain_text_splitters at sizes of 12,000–16,000 tokens (e.g., scraperai/parsers/data_fields_extractor.py:30 sets max_chunk_size = 16000). When the HTML exceeds the chunk size, only the first chunk is sent. For pagination detection, the first and last chunks are both sent (scraperai/parsers/pagination_detector.py:64-65). Structured output is enforced by ChatModelAgent.query_with_validation (scraperai/parsers/agent.py:13-39), which validates the model's JSON response against Pydantic schemas and retries up to 3–5 times with the validation error as feedback. No cost controls exist beyond the retry limit — there is no token budget, no model fallback, and no caching. The ParserAI.total_cost property (scraperai/parsers/parserai.py:46-52) sums costs from both JsonOpenAI and VisionOpenAI instances. The .env.example shows that only OPENAI_API_KEY is required.
ParserAI takes any BaseJsonLM and BaseVision implementation (one invoke method each, scraperai/llm/base.py L12-L27), so other providers plug in; only cost tracking is OpenAI-specific. PythonCodeOpenAI is constructed but never called.How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-bot measures are minimal and basic. The DefaultChromeWebdriver (scraperai/crawlers/webdriver/local.py:15-49) applies the following stealth patches on Chrome options and runtime:
--disable-blink-features=AutomationControlled— removes the "Chrome is being controlled" bannerexcludeSwitches: ["enable-automation"]— prevents Selenium from advertising automationuseAutomationExtension: False— disables the automation extension- JavaScript injection after driver init:
Object.defineProperty(navigator, 'webdriver', {get: () => undefined})— masks thenavigator.webdriverflag - A random user-agent is set per request in the
get()method via CDP:execute_cdp_cmd("Network.setUserAgentOverride", {"userAgent": get_random_useragent()})— agents are drawn from a text file (scraperai/crawlers/webdriver/useragents.py:1-12)
The same patterns are duplicated in RemoteWebdriver (scraperai/crawlers/webdriver/remote.py:46-61) for Selenoid-based remote drivers.
CAPTCHA handling: The WebpageType.CAPTCHA enum value exists (scraperai/models.py:66) and the CLI flow detects CAPTCHA pages. However, controller.py:290-292 simply shows a warning message and asks the user to solve it manually (view.py:45-47: "Captcha detected! Please solve it before we can continue"). There is no automatic CAPTCHA solving (roadmap mentions "anti-captcha integration" as a future item).
No proxy rotation, no rate limiting, no request throttling, and no fingerprint-level stealth (no TLS fingerprint spoofing, no TCP-level masking). The requests.get()-based RequestsCrawler sets no custom headers at all — just the library defaults.
How is crawling at scale implemented?
answeredScraperAI is not designed for crawling at scale — all implementations are single-threaded and synchronous. The Scraper class (scraperai/scraper.py:14-86) coordinates the crawl via a pull-based generator pattern, yielding one row at a time. For catalog pages, it loops: fetch page, parse HTML with lxml, extract items from each card (matching catalog_item.card_xpath), then call crawler.switch_page(pagination) to advance. The three pagination types (xpath, scroll, urls) are documented in scraperai/models.py:46-60. For nested detail pages, the scraper first collects all URLs from the catalog via scrape_nested_items_urls (extracting hrefs by url_xpath), then iterates over them calling crawler.get(url) per item. Concurrency: zero. The SeleniumCrawler holds one browser window. No async, no threading, no parallel workers. There is a WebdriversManager (scraperai/crawlers/webdriver/manager.py:27-61) that can distribute Selenium sessions across a pool of Selenoid servers (checking /status for active session counts) — but this is only used to create a single driver, not to parallelize work. URL deduplication happens only at the CLI level via a set() (controller.py:215). No robots.txt parsing exists anywhere in the codebase. No politeness delays — the only waits are the hard-coded time.sleep(1) in SeleniumCrawler.get() and time.sleep(3) after pagination clicks. No distributed workers, no message queues, no crawl frontier. The README roadmap lists "Add httpx and aiohttp crawlers" as a future improvement, confirming async/distributed is not yet implemented.
What is the developer interface?
answeredThe project provides two main developer interfaces and a serializable config format.
Python Library API — the primary interface. Users import Scraper, ParserAI, SeleniumCrawler, and model classes (scraperai/__init__.py:1-4). The typical flow: create a SeleniumCrawler, instantiate ParserAI(openai_api_key=...), call detect_page_type(), detect_pagination(), detect_catalog_item(), extract_fields(), assemble a ScraperConfig, and pass it with the crawler to Scraper(config, crawler). The scraper exposes a generator-based scrape() method. This is shown in detail in examples/ycombinator_full.ipynb.
CLI — built with Click (scraperai/cli/app.py), entry point scraperai (via console_scripts in setup.py:28-30). Run scraperai --url <url>. It drives the same library classes through a step-by-step interactive wizard (Controller in scraperai/cli/controller.py:21-334): checks for cached configs, auto-detects page type/pagination/catalog items/fields, lets the user edit each detection, then scrapes and exports. The View class (scraperai/cli/view.py:11-204) handles all terminal I/O with Click prompts and pandas-based markdown tables for field display.
ScraperConfig serialization (scraperai/models.py:74-83) — the full scraping configuration (URL, page type, pagination settings, catalog item XPaths, field definitions, limits) is a Pydantic model that dumps to JSON. Configs are saved as .scraperai.json files in the user's app data directory (scraperai/cli/utils.py:9) and are loaded on subsequent runs to skip re-detection (controller.py:42-51).
No REST service, no MCP server, no web UI. The README roadmap lists "Release SaaS web app" as a future goal. No REST API — the library is consumed in-process only. Language bindings: Python only (package on PyPI). Output formats are JSON, XLSX, and CSV.