# AI web scraping: comparison

> Crawlers, extractors and stealth browsers that turn web pages into LLM-ready data.

Canonical page: https://llms-technical-reviews.com/compare/ai-scraping/

## How are pages fetched and rendered?

The two published projects sit at opposite ends of the fetching spectrum, and neither one manages a browser for you.

[AutoScraper](/p/autoscraper/) makes one synchronous `requests.get` per call (`_fetch_html`), with a hard-coded Chrome 84 User-Agent and whatever you put in `request_args`. It never runs JavaScript, has no waiting strategy, and only handles HTML text: PDFs, images and other binary responses are not supported. Every entry point also accepts an `html=` string, so you can render elsewhere and pass the result in.

[llm-scraper](/p/llm-scraper/) does no fetching at all. You launch Playwright, navigate, and wait for the page to be ready, then pass in a live `Page`. The library only reads from that page, as raw HTML, cleaned HTML, Markdown, Readability text or a screenshot. It inherits full JS rendering from Playwright but adds no readiness logic of its own. Note that its default `html` format strips elements out of the live DOM before reading it.

Choose AutoScraper's built-in fetch for static, server-rendered pages where speed and zero infrastructure matter. Choose llm-scraper (or feed AutoScraper pre-rendered HTML) when content appears only after JavaScript runs, or when you need logged-in sessions, since in both cases the browser is your code's job.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/fetching/index.md

## How is content extracted or converted?

The two projects extract content in completely different ways: learned DOM paths versus an LLM filling a schema.

[AutoScraper](/p/autoscraper/) learns by example. You give it values visible on a page. `build()` finds the elements that hold them, by exact, regex or `SequenceMatcher`-fuzzy text or attribute match, and records each element's path from the root as `(tag, {class, style}, sibling_index)` steps. Replay is pure BeautifulSoup traversal. "Similar" mode fans out across siblings to collect whole lists, and "exact" mode follows the recorded positions. There is no Markdown conversion and no boilerplate removal.

[llm-scraper](/p/llm-scraper/) has no selectors. It converts the page to cleaned HTML (an in-page `cleanup` that removes 30 tag types and noisy attributes), Turndown Markdown, Mozilla Readability text, a screenshot, or a custom string. It then asks the model to fill an AI SDK `Output` schema (Zod or JSON Schema) in one call. Its `generate()` mode instead asks the model to write a JavaScript extractor you can reuse.

AutoScraper suits repetitive, stable layouts where you can show it an example and want deterministic, free extraction. It breaks when the page structure changes. llm-scraper suits varied or unfamiliar pages and nested typed output, at the cost of model calls and non-determinism.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/extraction/index.md

## How are LLMs used, if at all?

[AutoScraper](/p/autoscraper/) is not_applicable here: it uses no LLM. Its dependencies are only `requests`, `bs4` and `lxml`, and its "learning" is DOM-path recording plus `difflib` string similarity.

[llm-scraper](/p/llm-scraper/) is built around one AI SDK call. `run()` sends the whole preprocessed page as a single user message (text or an image part) under a fixed system prompt, "You are a sophisticated web scraper...". It passes your `Output` to `generateText`, so the SDK handles structured output and parsing. `stream()` does the same through `streamText` and yields partial objects. `generate()` sends the URL, the JSON Schema and the content, and asks for an IIFE scraper. Any AI SDK `LanguageModel` works: the examples use OpenAI and Ollama, and the README adds Anthropic, Google and Groq. You can override `system`, append `messages`, and pass `CallSettings` through. There is no chunking, token counting or truncation. Cost control means choosing a compact format (`markdown`, `text`, cleaned `html`) or using `generate()` once and running the resulting code with `page.evaluate` on later pages.

If you need zero model cost and deterministic output, the non-LLM approach wins. If pages vary or you want typed nested data from a schema, llm-scraper's thin wrapper is easy to read, but you must handle long pages and budgets yourself.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/llm-usage/index.md

## How are anti-bot measures, proxies and fingerprinting handled?

Neither published project does real anti-bot work. Both leave it to the caller.

[AutoScraper](/p/autoscraper/) sends a fixed Chrome 84 User-Agent and sets `Host` from the URL. That is the full extent of its disguise. `request_args` is passed straight to `requests.get`, so you can supply a proxy dict, cookies, headers or timeouts. But there is no proxy rotation, retry, backoff, throttling, CAPTCHA handling or TLS fingerprinting, and the dated UA string is easy to flag.

[llm-scraper](/p/llm-scraper/) is not_applicable: it has no anti-bot features at all. Its `cleanup` step removes scripts and attributes only to reduce tokens. Because you create the Playwright browser and context yourself, you can add a stealth plugin, a proxy at launch, persistent auth state or your own pacing before you hand the `Page` over. The library neither helps nor gets in the way.

For protected targets, pair either library with a separate access layer. A proxy-routed or stealth-patched Playwright context fits llm-scraper naturally. With AutoScraper, the practical route is fetching through your own client and passing `html=`, since its built-in fetch has no hooks beyond `request_args`.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/anti-bot/index.md

## How is crawling at scale implemented?

Neither published project crawls. Both answers are not_applicable.

[AutoScraper](/p/autoscraper/) processes exactly one page per `build()` or `get_result_*()` call. It does not extract or follow links, and it has no queue, concurrency, deduplication of URLs, depth limits or robots.txt handling. The parts that do carry across pages are its learned rules. These are JSON, so you can apply one saved rule set to every URL in your own loop.

[llm-scraper](/p/llm-scraper/) is likewise one `Page` per call. The examples and tests all follow a single launch, `goto`, `run` sequence. Crawling means writing your own loop around `page.goto()` and `scraper.run()`, and adding concurrency through multiple Playwright pages. Every page then costs one model call, unless you use `generate()` once and reuse the code.

If you need a crawler, these are extraction components to plug into one, not crawlers themselves. AutoScraper is the cheaper per-page step for large runs over same-template pages. llm-scraper fits low-volume crawls over varied pages, where per-page LLM cost is acceptable.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/crawling/index.md

## What is the developer interface?

Both published projects are libraries only, with no CLI, REST service, MCP server or UI.

[AutoScraper](/p/autoscraper/) is a single Python class. `build()` learns rules. `get_result_similar()`, `get_result_exact()` and `get_result()` (both together) apply them. `save()`/`load()` persist rules as JSON, and `remove_rules`/`keep_rules`/`set_rule_aliases` edit them by `stack_id`. Output is a plain list of strings, or a dict when `grouped=True` (keyed by rule) or `group_by_alias=True` (keyed by your aliases from `wanted_dict`). Values are always strings, so you build any typing or nesting yourself.

[llm-scraper](/p/llm-scraper/) is an ESM TypeScript class with three async methods. `run()` returns `{data, url}`, `stream()` returns `{stream, url}` with partial objects, and `generate()` returns `{code, url}`. Output is shaped by an AI SDK `Output`, from Zod or JSON Schema, so nested typed objects and arrays come naturally. Options select the input format and pass AI SDK call settings through. The repo also shows wrapping it as an AI SDK `tool()` for agents.

Choose AutoScraper for Python scripts that need flat fields and reusable rule files. Choose llm-scraper for Node/TypeScript stacks that want typed objects or an agent tool. Neither offers a hosted API, so for that you would wrap either one yourself.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/ai-scraping/q/interface/index.md

## Projects

- [alirezamika/autoscraper](https://llms-technical-reviews.com/p/autoscraper/index.md) — Python library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages.
- [mishushakov/llm-scraper](https://llms-technical-reviews.com/p/llm-scraper/index.md) — TypeScript library that turns an already-open Playwright page into schema-typed data, or generated scraper code, via the Vercel AI SDK.