How are pages fetched and rendered?
Plain HTTP vs headless browser; JS rendering; waiting strategy; supported content types (PDF, images).
Verdict
The two published projects sit at opposite ends of the fetching spectrum, and neither one manages a browser for you.
AutoScraper makes one synchronous requests.get per call (_fetch_html), with a hard-coded Chrome 84 User-Agent and whatever you put in request_args. It never runs JavaScript, has no waiting strategy, and only handles HTML text: PDFs, images and other binary responses are not supported. Every entry point also accepts an html= string, so you can render elsewhere and pass the result in.
llm-scraper does no fetching at all. You launch Playwright, navigate, and wait for the page to be ready, then pass in a live Page. The library only reads from that page, as raw HTML, cleaned HTML, Markdown, Readability text or a screenshot. It inherits full JS rendering from Playwright but adds no readiness logic of its own. Note that its default html format strips elements out of the live DOM before reading it.
Choose AutoScraper’s built-in fetch for static, server-rendered pages where speed and zero infrastructure matter. Choose llm-scraper (or feed AutoScraper pre-rendered HTML) when content appears only after JavaScript runs, or when you need logged-in sessions, since in both cases the browser is your code’s job.
More projects in this category are being researched.
Per-project answers
alirezamika/autoscraper
answeredPages are fetched via plain synchronous HTTP using the requests library. The classmethod _fetch_html() at auto_scraper.py:96-110 performs a requests.get(url, headers=headers, **request_args) call. A default Chrome 84 User-Agent header is set at auto_scraper.py:45-48. No headless browser, no JavaScript rendering engine, and no waiting strategy is implemented — if the target page relies on JS to populate content, AutoScraper will only see the raw HTML. request_args is forwarded as **kwargs to requests.get(), allowing users to manually pass proxies, custom headers, cookies, or timeouts. Content-type detection for encoding is minimal: if the server declares ISO-8859-1 but the actual Content-Type header doesn't contain it, the apparent encoding is used instead (auto_scraper.py:105-108). The only content type consumed is text/HTML; PDFs, images, or binary responses are not supported at all — res.text is called unconditionally and passed through str.strip() + unicodedata.normalize('NFKD', ...) (utils.py:29-32). The fallback path _get_soup() at auto_scraper.py:113-121 also accepts a raw html string parameter, so the caller can pre-fetch with any tool and feed the HTML directly.
mishushakov/llm-scraper
answeredThe library does not fetch or render pages itself. It delegates entirely to a Playwright Page object that the caller creates and navigates. The caller launches Playwright (chromium.launch()) and navigates (page.goto(url)) before passing the page to the scraper — all JS rendering, redirects, and dynamic loading are handled by Playwright's full browser engine. There is no built-in wait strategy; the caller must ensure the page is ready (e.g. await page.waitForSelector(...)) before calling scraper.run(). Once given a Page, the preprocess() function (src/preprocess.ts:25-83) reads content in six modes: raw_html (plain page.content(), line 34), html (runs a cleanup pass then page.content(), lines 55-58), markdown (extracts <body> innerHTML and converts via Turndown, lines 38-40), text (evaluates Mozilla Readability.js loaded from CDN inside the browser to extract main-article text, lines 43-53), image (takes a Playwright page.screenshot() and returns base64, lines 61-65), and custom (user-provided formatFunction(page), lines 67-76). There is no support for PDFs or inline image/media extraction — images inside pages are removed by the cleanup step.