# How are pages fetched and rendered?

> AI web scraping — a good answer covers: Plain HTTP vs headless browser; JS rendering; waiting strategy; supported content types (PDF, images).

Canonical page: https://llms-technical-reviews.com/ai-scraping/q/fetching/

## Verdict

The two published projects sit at opposite ends of the fetching spectrum, and neither one manages a browser for you.

[AutoScraper](/p/autoscraper/) makes one synchronous `requests.get` per call (`_fetch_html`), with a hard-coded Chrome 84 User-Agent and whatever you put in `request_args`. It never runs JavaScript, has no waiting strategy, and only handles HTML text: PDFs, images and other binary responses are not supported. Every entry point also accepts an `html=` string, so you can render elsewhere and pass the result in.

[llm-scraper](/p/llm-scraper/) does no fetching at all. You launch Playwright, navigate, and wait for the page to be ready, then pass in a live `Page`. The library only reads from that page, as raw HTML, cleaned HTML, Markdown, Readability text or a screenshot. It inherits full JS rendering from Playwright but adds no readiness logic of its own. Note that its default `html` format strips elements out of the live DOM before reading it.

Choose AutoScraper's built-in fetch for static, server-rendered pages where speed and zero infrastructure matter. Choose llm-scraper (or feed AutoScraper pre-rendered HTML) when content appears only after JavaScript runs, or when you need logged-in sessions, since in both cases the browser is your code's job.

More projects in this category are being researched.

## Per-project answers

### alirezamika/autoscraper (answered)

Pages are fetched via plain synchronous HTTP using the `requests` library. The classmethod `_fetch_html()` at `auto_scraper.py:96-110` performs a `requests.get(url, headers=headers, **request_args)` call. A default Chrome 84 User-Agent header is set at `auto_scraper.py:45-48`. No headless browser, no JavaScript rendering engine, and no waiting strategy is implemented — if the target page relies on JS to populate content, AutoScraper will only see the raw HTML. `request_args` is forwarded as `**kwargs` to `requests.get()`, allowing users to manually pass proxies, custom headers, cookies, or timeouts. Content-type detection for encoding is minimal: if the server declares `ISO-8859-1` but the actual `Content-Type` header doesn't contain it, the apparent encoding is used instead (`auto_scraper.py:105-108`). The only content type consumed is text/HTML; PDFs, images, or binary responses are not supported at all — `res.text` is called unconditionally and passed through `str.strip()` + `unicodedata.normalize('NFKD', ...)` (`utils.py:29-32`). The fallback path `_get_soup()` at `auto_scraper.py:113-121` also accepts a raw `html` string parameter, so the caller can pre-fetch with any tool and feed the HTML directly.


Citations: [autoscraper/auto_scraper.py:96-110](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L96-L110) · [autoscraper/auto_scraper.py:113-121](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L113-L121) · [autoscraper/auto_scraper.py:45-48](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L45-L48) · [autoscraper/utils.py:29-32](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L29-L32)

### mishushakov/llm-scraper (answered)

The library does **not** fetch or render pages itself. It delegates entirely to a **Playwright `Page` object** that the caller creates and navigates. The caller launches Playwright (`chromium.launch()`) and navigates (`page.goto(url)`) before passing the page to the scraper — all JS rendering, redirects, and dynamic loading are handled by Playwright's full browser engine. There is **no built-in wait strategy**; the caller must ensure the page is ready (e.g. `await page.waitForSelector(...)`) before calling `scraper.run()`. Once given a `Page`, the `preprocess()` function (`src/preprocess.ts:25-83`) reads content in six modes: **raw_html** (plain `page.content()`, line 34), **html** (runs a cleanup pass then `page.content()`, lines 55-58), **markdown** (extracts `<body>` innerHTML and converts via Turndown, lines 38-40), **text** (evaluates Mozilla Readability.js loaded from CDN inside the browser to extract main-article text, lines 43-53), **image** (takes a Playwright `page.screenshot()` and returns base64, lines 61-65), and **custom** (user-provided `formatFunction(page)`, lines 67-76). There is **no support for PDFs or inline image/media extraction** — images inside pages are removed by the cleanup step.


Citations: [src/preprocess.ts:25-83](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L25-L83) · [tests/index.ts:8-13](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L8-L13) · [src/cleanup.ts:1-60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60) · [src/index.ts:29-36](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L29-L36)
