# How is crawling at scale implemented?

> AI web scraping — a good answer covers: Queues; concurrency; URL dedup; depth/limits; robots.txt and politeness; distributed workers.

Canonical page: https://llms-technical-reviews.com/ai-scraping/q/crawling/

## Verdict

Neither published project crawls. Both answers are not_applicable.

[AutoScraper](/p/autoscraper/) processes exactly one page per `build()` or `get_result_*()` call. It does not extract or follow links, and it has no queue, concurrency, deduplication of URLs, depth limits or robots.txt handling. The parts that do carry across pages are its learned rules. These are JSON, so you can apply one saved rule set to every URL in your own loop.

[llm-scraper](/p/llm-scraper/) is likewise one `Page` per call. The examples and tests all follow a single launch, `goto`, `run` sequence. Crawling means writing your own loop around `page.goto()` and `scraper.run()`, and adding concurrency through multiple Playwright pages. Every page then costs one model call, unless you use `generate()` once and reuse the code.

If you need a crawler, these are extraction components to plug into one, not crawlers themselves. AutoScraper is the cheaper per-page step for large runs over same-template pages. llm-scraper fits low-volume crawls over varied pages, where per-page LLM cost is acceptable.

More projects in this category are being researched.

## Per-project answers

### alirezamika/autoscraper (not applicable)

AutoScraper has no crawling capability. It is a single-page scraper: each call to `build()`, `get_result_similar()`, or `get_result_exact()` fetches exactly one URL and extracts data from that page. There are no URL queues, no concurrency/threading, no URL-deduplication, no crawl-depth limits, no `robots.txt` parsing, no politeness delays, and no distributed-worker architecture. The library does not even iterate over links on a page — it has no link-extraction logic and no "follow" mechanism. The entire source is two small files focusing purely on learning extraction rules from one page and applying them to another single page.



### mishushakov/llm-scraper (not applicable)

The library does **not** implement crawling. It is a single-page extraction tool with no URL queues, no concurrent-page management, no deduplication, no depth/limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. A caller who needs crawling must build their own loop: iterate over URLs, call `page.goto()` for each, and call `scraper.run()` per page. The test infrastructure (`tests/index.ts:6-12`) and all examples show this pattern — one browser, one page, one `scraper.run()`. No `crawl()` or `spider()` method exists.


Citations: [src/index.ts:24-56](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L24-L56) · [tests/index.ts:6-13](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L6-L13) · [examples/hn.ts:8-44](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/hn.ts#L8-L44)
