LLMs Technical Reviews
Home / AI web scraping / Comparison

AI web scraping: autoscraper vs llm-scraper

Crawlers, extractors and stealth browsers that turn web pages into LLM-ready data. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.

At a glance

● answered from code · — not applicable (the project does not do this) · ? insufficient evidence

How are pages fetched and rendered?

The two published projects sit at opposite ends of the fetching spectrum, and neither one manages a browser for you.

AutoScraper makes one synchronous requests.get per call (_fetch_html), with a hard-coded Chrome 84 User-Agent and whatever you put in request_args. It never runs JavaScript, has no waiting strategy, and only handles HTML text: PDFs, images and other binary responses are not supported. Every entry point also accepts an html= string, so you can render elsewhere and pass the result in.

llm-scraper does no fetching at all. You launch Playwright, navigate, and wait for the page to be ready, then pass in a live Page. The library only reads from that page, as raw HTML, cleaned HTML, Markdown, Readability text or a screenshot. It inherits full JS rendering from Playwright but adds no readiness logic of its own. Note that its default html format strips elements out of the live DOM before reading it.

Choose AutoScraper's built-in fetch for static, server-rendered pages where speed and zero infrastructure matter. Choose llm-scraper (or feed AutoScraper pre-rendered HTML) when content appears only after JavaScript runs, or when you need logged-in sessions, since in both cases the browser is your code's job.

More projects in this category are being researched.

How is content extracted or converted?

The two projects extract content in completely different ways: learned DOM paths versus an LLM filling a schema.

AutoScraper learns by example. You give it values visible on a page. build() finds the elements that hold them, by exact, regex or SequenceMatcher-fuzzy text or attribute match, and records each element's path from the root as (tag, {class, style}, sibling_index) steps. Replay is pure BeautifulSoup traversal. "Similar" mode fans out across siblings to collect whole lists, and "exact" mode follows the recorded positions. There is no Markdown conversion and no boilerplate removal.

llm-scraper has no selectors. It converts the page to cleaned HTML (an in-page cleanup that removes 30 tag types and noisy attributes), Turndown Markdown, Mozilla Readability text, a screenshot, or a custom string. It then asks the model to fill an AI SDK Output schema (Zod or JSON Schema) in one call. Its generate() mode instead asks the model to write a JavaScript extractor you can reuse.

AutoScraper suits repetitive, stable layouts where you can show it an example and want deterministic, free extraction. It breaks when the page structure changes. llm-scraper suits varied or unfamiliar pages and nested typed output, at the cost of model calls and non-determinism.

More projects in this category are being researched.

How are LLMs used, if at all?

AutoScraper is not_applicable here: it uses no LLM. Its dependencies are only requests, bs4 and lxml, and its "learning" is DOM-path recording plus difflib string similarity.

llm-scraper is built around one AI SDK call. run() sends the whole preprocessed page as a single user message (text or an image part) under a fixed system prompt, "You are a sophisticated web scraper...". It passes your Output to generateText, so the SDK handles structured output and parsing. stream() does the same through streamText and yields partial objects. generate() sends the URL, the JSON Schema and the content, and asks for an IIFE scraper. Any AI SDK LanguageModel works: the examples use OpenAI and Ollama, and the README adds Anthropic, Google and Groq. You can override system, append messages, and pass CallSettings through. There is no chunking, token counting or truncation. Cost control means choosing a compact format (markdown, text, cleaned html) or using generate() once and running the resulting code with page.evaluate on later pages.

If you need zero model cost and deterministic output, the non-LLM approach wins. If pages vary or you want typed nested data from a schema, llm-scraper's thin wrapper is easy to read, but you must handle long pages and budgets yourself.

More projects in this category are being researched.

How are anti-bot measures, proxies and fingerprinting handled?

Neither published project does real anti-bot work. Both leave it to the caller.

AutoScraper sends a fixed Chrome 84 User-Agent and sets Host from the URL. That is the full extent of its disguise. request_args is passed straight to requests.get, so you can supply a proxy dict, cookies, headers or timeouts. But there is no proxy rotation, retry, backoff, throttling, CAPTCHA handling or TLS fingerprinting, and the dated UA string is easy to flag.

llm-scraper is not_applicable: it has no anti-bot features at all. Its cleanup step removes scripts and attributes only to reduce tokens. Because you create the Playwright browser and context yourself, you can add a stealth plugin, a proxy at launch, persistent auth state or your own pacing before you hand the Page over. The library neither helps nor gets in the way.

For protected targets, pair either library with a separate access layer. A proxy-routed or stealth-patched Playwright context fits llm-scraper naturally. With AutoScraper, the practical route is fetching through your own client and passing html=, since its built-in fetch has no hooks beyond request_args.

More projects in this category are being researched.

How is crawling at scale implemented?

Neither published project crawls. Both answers are not_applicable.

AutoScraper processes exactly one page per build() or get_result_*() call. It does not extract or follow links, and it has no queue, concurrency, deduplication of URLs, depth limits or robots.txt handling. The parts that do carry across pages are its learned rules. These are JSON, so you can apply one saved rule set to every URL in your own loop.

llm-scraper is likewise one Page per call. The examples and tests all follow a single launch, goto, run sequence. Crawling means writing your own loop around page.goto() and scraper.run(), and adding concurrency through multiple Playwright pages. Every page then costs one model call, unless you use generate() once and reuse the code.

If you need a crawler, these are extraction components to plug into one, not crawlers themselves. AutoScraper is the cheaper per-page step for large runs over same-template pages. llm-scraper fits low-volume crawls over varied pages, where per-page LLM cost is acceptable.

More projects in this category are being researched.

What is the developer interface?

Both published projects are libraries only, with no CLI, REST service, MCP server or UI.

AutoScraper is a single Python class. build() learns rules. get_result_similar(), get_result_exact() and get_result() (both together) apply them. save()/load() persist rules as JSON, and remove_rules/keep_rules/set_rule_aliases edit them by stack_id. Output is a plain list of strings, or a dict when grouped=True (keyed by rule) or group_by_alias=True (keyed by your aliases from wanted_dict). Values are always strings, so you build any typing or nesting yourself.

llm-scraper is an ESM TypeScript class with three async methods. run() returns {data, url}, stream() returns {stream, url} with partial objects, and generate() returns {code, url}. Output is shaped by an AI SDK Output, from Zod or JSON Schema, so nested typed objects and arrays come naturally. Options select the input format and pass AI SDK call settings through. The repo also shows wrapping it as an AI SDK tool() for agents.

Choose AutoScraper for Python scripts that need flat fields and reusable rule files. Choose llm-scraper for Node/TypeScript stacks that want typed objects or an agent tool. Neither offers a hosted API, so for that you would wrap either one yourself.

More projects in this category are being researched.

The projects