How is crawling at scale implemented?
Queues; concurrency; URL dedup; depth/limits; robots.txt and politeness; distributed workers.
Verdict
Neither published project crawls. Both answers are not_applicable.
AutoScraper processes exactly one page per build() or get_result_*() call. It does not extract or follow links, and it has no queue, concurrency, deduplication of URLs, depth limits or robots.txt handling. The parts that do carry across pages are its learned rules. These are JSON, so you can apply one saved rule set to every URL in your own loop.
llm-scraper is likewise one Page per call. The examples and tests all follow a single launch, goto, run sequence. Crawling means writing your own loop around page.goto() and scraper.run(), and adding concurrency through multiple Playwright pages. Every page then costs one model call, unless you use generate() once and reuse the code.
If you need a crawler, these are extraction components to plug into one, not crawlers themselves. AutoScraper is the cheaper per-page step for large runs over same-template pages. llm-scraper fits low-volume crawls over varied pages, where per-page LLM cost is acceptable.
More projects in this category are being researched.
Per-project answers
alirezamika/autoscraper
not applicableAutoScraper has no crawling capability. It is a single-page scraper: each call to build(), get_result_similar(), or get_result_exact() fetches exactly one URL and extracts data from that page. There are no URL queues, no concurrency/threading, no URL-deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker architecture. The library does not even iterate over links on a page — it has no link-extraction logic and no "follow" mechanism. The entire source is two small files focusing purely on learning extraction rules from one page and applying them to another single page.
mishushakov/llm-scraper
not applicableThe library does not implement crawling. It is a single-page extraction tool with no URL queues, no concurrent-page management, no deduplication, no depth/limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. A caller who needs crawling must build their own loop: iterate over URLs, call page.goto() for each, and call scraper.run() per page. The test infrastructure (tests/index.ts:6-12) and all examples show this pattern — one browser, one page, one scraper.run(). No crawl() or spider() method exists.
← How are anti-bot measures, proxies and fingerprinting handled? · What is the developer interface? →