AI web scraping
Crawlers, extractors and stealth browsers that turn web pages into LLM-ready data.
These tools get data out of web pages so it can be stored, searched or fed to an LLM. The category covers several layers. Some tools fetch and render pages (plain HTTP or a headless browser). Some extract content, by learned selectors, readability-style text, Markdown conversion, or an LLM filling a schema. Some crawl many pages, and some handle access with stealth browsers, proxies and fingerprinting. Few projects cover every layer, so first decide which ones you actually need.
When choosing, check four things. Does it render JavaScript, or does it expect you to supply the browser? Is extraction deterministic and free per page, or does every page cost a model call? How are long pages chunked or trimmed? And does it include crawling and anti-bot support, or only extraction you plug into your own pipeline?
Projects (2)
| Project | Stars |
|---|---|
| alirezamika/autoscraperPython library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages. | ★ 8.0k |
| mishushakov/llm-scraperTypeScript library that turns an already-open Playwright page into schema-typed data, or generated scraper code, via the Vercel AI SDK. | ★ 6.9k |
In the research queue: firecrawl/firecrawl, unclecode/crawl4ai, D4Vinci/Scrapling, ScrapeGraphAI/Scrapegraph-ai, raznem/parsera, itsOwen/CyberScraper-2077, adbar/trafilatura, adithya-s-k/omniparse, scraperai/scraperai, browserless/browserless, BrowserBox/BrowserBox, vinyzu-archive/Botright, ttlns/Selenium-Driverless, crawlab-team/crawlab, ssssssss-team/spider-flow, mixmark-io/turndown.
Comparison questions
Each question is answered separately for every project in this category, from that project's source code.
- How are pages fetched and rendered?Plain HTTP vs headless browser; JS rendering; waiting strategy; supported content types (PDF, images).
- How is content extracted or converted?HTML→Markdown/text heuristics; readability-style boilerplate removal; selectors; schema-based extraction.
- How are LLMs used, if at all?Prompting; chunking of large pages; structured output / JSON schema; which providers; cost controls.
- How are anti-bot measures, proxies and fingerprinting handled?Stealth patches; fingerprint spoofing; proxy rotation; CAPTCHA handling; rate limiting.
- How is crawling at scale implemented?Queues; concurrency; URL dedup; depth/limits; robots.txt and politeness; distributed workers.
- What is the developer interface?Library API, CLI, REST service, MCP server, UI; output formats; language bindings.