LLMs Technical Reviews

How are anti-bot measures, proxies and fingerprinting handled?

Stealth patches; fingerprint spoofing; proxy rotation; CAPTCHA handling; rate limiting.

Verdict

Neither published project does real anti-bot work. Both leave it to the caller.

AutoScraper sends a fixed Chrome 84 User-Agent and sets Host from the URL. That is the full extent of its disguise. request_args is passed straight to requests.get, so you can supply a proxy dict, cookies, headers or timeouts. But there is no proxy rotation, retry, backoff, throttling, CAPTCHA handling or TLS fingerprinting, and the dated UA string is easy to flag.

llm-scraper is not_applicable: it has no anti-bot features at all. Its cleanup step removes scripts and attributes only to reduce tokens. Because you create the Playwright browser and context yourself, you can add a stealth plugin, a proxy at launch, persistent auth state or your own pacing before you hand the Page over. The library neither helps nor gets in the way.

For protected targets, pair either library with a separate access layer. A proxy-routed or stealth-patched Playwright context fits llm-scraper naturally. With AutoScraper, the practical route is fetching through your own client and passing html=, since its built-in fetch has no hooks beyond request_args.

More projects in this category are being researched.

Per-project answers

alirezamika/autoscraper

answered

Anti-bot measures are minimal and largely absent. The only built-in fingerprinting is a default Chrome 84 User-Agent header set at auto_scraper.py:45-48. The _fetch_html() method at auto_scraper.py:96-110 also sets the Host header from the URL's netloc. Beyond that, there are no stealth patches (no TLS fingerprint spoofing, no browser-emulation of headers beyond UA), no proxy rotation system, no CAPTCHA handling, and no rate-limiting logic — no retry-on-403, no exponential backoff, no request throttling. The code does accept a request_args parameter (auto_scraper.py:97) that is unpacked as **kwargs to requests.get(), which lets the calling user manually supply proxy dictionaries, custom headers, cookies, or timeouts. The README (README.md:86-94) demonstrates passing proxies via this mechanism. This is a pass-through, not a built-in rotation or management strategy.

mishushakov/llm-scraper

not applicable

The library does not implement any anti-bot evasion. There are no stealth patches (no user-agent spoofing, no TLS fingerprint manipulation, no browser-fingerprint alteration), no proxy rotation, no CAPTCHA handling, no request rate limiting, and no retry logic. The cleanup() function removes script, iframe, svg, aria-* attributes and event handlers, but this is aimed at reducing token count in LLM prompts, not evading detection. The caller manually configures Playwright (chromium.launch()) and could pass their own stealth settings there — but the library wraps none of that.

← How are LLMs used, if at all? · How is crawling at scale implemented? →