LLMs Technical Reviews
Home / AI web scraping

AI web scraping

Crawlers, extractors and stealth browsers that turn web pages into LLM-ready data.

These tools get data out of web pages so it can be stored, searched or fed to an LLM. The category covers several layers. Some tools fetch and render pages (plain HTTP or a headless browser). Some extract content, by learned selectors, readability-style text, Markdown conversion, or an LLM filling a schema. Some crawl many pages, and some handle access with stealth browsers, proxies and fingerprinting. Few projects cover every layer, so first decide which ones you actually need.

When choosing, check four things. Does it render JavaScript, or does it expect you to supply the browser? Is extraction deterministic and free per page, or does every page cost a model call? How are long pages chunked or trimmed? And does it include crawling and anti-bot support, or only extraction you plug into your own pipeline?

Projects (2)

ProjectStarsLanguageLicense
alirezamika/autoscraperPython library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages.★ 8.0kPythonMIT
mishushakov/llm-scraperTypeScript library that turns an already-open Playwright page into schema-typed data, or generated scraper code, via the Vercel AI SDK.★ 6.9kTypeScriptMIT

In the research queue: firecrawl/firecrawl, unclecode/crawl4ai, D4Vinci/Scrapling, ScrapeGraphAI/Scrapegraph-ai, raznem/parsera, itsOwen/CyberScraper-2077, adbar/trafilatura, adithya-s-k/omniparse, scraperai/scraperai, browserless/browserless, BrowserBox/BrowserBox, vinyzu-archive/Botright, ttlns/Selenium-Driverless, crawlab-team/crawlab, ssssssss-team/spider-flow, mixmark-io/turndown.

Comparison questions

Each question is answered separately for every project in this category, from that project's source code.

All verdicts on one page →

  1. How are pages fetched and rendered?Plain HTTP vs headless browser; JS rendering; waiting strategy; supported content types (PDF, images).
  2. How is content extracted or converted?HTML→Markdown/text heuristics; readability-style boilerplate removal; selectors; schema-based extraction.
  3. How are LLMs used, if at all?Prompting; chunking of large pages; structured output / JSON schema; which providers; cost controls.
  4. How are anti-bot measures, proxies and fingerprinting handled?Stealth patches; fingerprint spoofing; proxy rotation; CAPTCHA handling; rate limiting.
  5. How is crawling at scale implemented?Queues; concurrency; URL dedup; depth/limits; robots.txt and politeness; distributed workers.
  6. What is the developer interface?Library API, CLI, REST service, MCP server, UI; output formats; language bindings.