# mishushakov/llm-scraper

> TypeScript library that turns an already-open Playwright page into schema-typed data, or generated scraper code, via the Vercel AI SDK.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/mishushakov/llm-scraper (reviewed at commit `2b43d999a17eac040f7cc315fc21dd526870564f`, 2026-06-15)
- Stars: 6936 · Language: TypeScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/llm-scraper/

## Overview

llm-scraper is a thin TypeScript layer between a Playwright `Page` and the Vercel AI SDK. You open and navigate the page yourself, describe the data you want as an AI SDK `Output` (usually `Output.object({ schema })` with a Zod schema), and call `run`, `stream` or `generate`. The library turns the page into one model input (cleaned HTML, raw HTML, Markdown, Readability text, a screenshot, or your own string) and makes one `generateText` or `streamText` call.

The whole library is four source files and about 330 lines. It has no fetching, waiting, crawling, retry, chunking or anti-bot logic. Those are left to Playwright and to the caller. Its runtime dependencies are only `ai`, `@ai-sdk/provider` and `turndown`. Playwright and Zod are dev dependencies, because the library only uses Playwright's `Page` type and never launches a browser ([package.json](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/package.json)).

It suits developers who already drive a browser from Node and want typed extraction from arbitrary pages without writing selectors. The `generate` mode is the cost-saving option: the model writes a JavaScript extractor once, and you run it with `page.evaluate` on later pages.

## Architecture

```mermaid
flowchart LR
  C["Caller: Playwright page + Output schema"] --> L["LLMScraper (index.ts)"]
  L --> P["preprocess (preprocess.ts)"]
  P --> CL["cleanup (in-page DOM strip)"]
  P --> TD["Turndown markdown"]
  P --> RD["Readability via CDN import"]
  P --> SS["page.screenshot base64"]
  L --> M["models.ts"]
  M --> GT["generateText (run)"]
  M --> ST["streamText (stream)"]
  M --> GC["generateText code prompt (generate)"]
  GT --> D["data, url"]
  ST --> S["partialOutputStream, url"]
  GC --> K["code, url"]
```

| Component | Path | Role |
|---|---|---|
| `LLMScraper` | `src/index.ts` | Public class: `run`, `stream`, `generate`, all wrapping `preprocess` and one model call |
| Preprocessor | `src/preprocess.ts` | Turns the live page into `{url, content, format}` in one of six formats |
| DOM cleanup | `src/cleanup.ts` | Runs inside the page. Removes noisy elements and attributes before HTML capture |
| Model calls | `src/models.ts` | Builds messages and system prompts, calls the AI SDK, strips code fences |
| Examples | `examples/*.ts` | HN extraction, streaming, codegen, Ollama, tool use |
| Tests | `tests/` | Live Vitest runs against Hacker News and example.com with `gpt-4o-mini` |

## How a request flows

For `scraper.run(page, Output.object({ schema }), { format: 'markdown' })`:

1. `LLMScraper.run` calls `preprocess(page, options)`, then `generateAISDKCompletions` with the model you passed to the constructor ([index.ts#L24-L56](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L24-L56)).
2. `preprocess` reads `page.url()` and picks the format, `html` by default ([preprocess.ts#L25-L83](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L25-L83)):
   - `raw_html`: `page.content()`.
   - `markdown`: `page.innerHTML('body')` converted with `new Turndown().turndown(...)` in Node.
   - `text`: inside the page, dynamically `import('https://cdn.skypack.dev/@mozilla/readability')`, run `Readability(document).parse()`, and return `Page Title: ...` plus `textContent`.
   - `html`: `page.evaluate(cleanup)`, then `page.content()`.
   - `image`: `page.screenshot({ fullPage })` as base64.
   - `custom`: `await options.formatFunction(page)`. It throws if the function is missing.
3. `prepareAISDKPage` wraps the content as a single text part, or an image part for `image` ([models.ts#L25-L35](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L25-L35)).
4. `generateAISDKCompletions` builds `messages = [user: page content, ...options.messages]`, uses the default system prompt "You are a sophisticated web scraper. Extract the contents of the webpage" unless you override `system`, and spreads the remaining options into `generateText({ model, output, system, messages, ...rest })` ([models.ts#L12-L63](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L12-L63)).
5. It returns `{ data: result.output, url }`. Schema enforcement and parsing come from the AI SDK's `Output` handling. The library itself never calls `schema.parse`.

`stream` is the same, except that it calls `streamText` and returns `{ stream: partialOutputStream, url }` for `for await` consumption ([models.ts#L65-L91](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L65-L91)).

## Key components

### Preprocessing and cleanup

`cleanup` is a function serialised into the page by `page.evaluate`. It walks `document.querySelectorAll('*')`, removes 30 tag types (`script`, `style`, `svg`, `img`, `form`, `input`, `button`, `header`, `footer`, `nav`, `aside`, `head` and others), and strips attributes whose names start with `style`, `src`, `alt`, `title`, `role`, `aria-`, `tabindex`, `on` or `data-` ([cleanup.ts#L1-L60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60)). The goal is fewer tokens, not boilerplate detection. `class`, `id` and `href` survive. The cleanup **mutates the live page**: after an `html`-mode call, the same `Page` has no scripts, forms, nav or images. This matters if you keep using the page, or if you `generate` code in `html` mode and then evaluate it on the stripped DOM. The bundled codegen example uses `raw_html`.

`text` mode loads Readability from a public CDN at run time inside the page. It needs outbound network access from the browser, and a page Content Security Policy can block it.

### Model integration

The library accepts any AI SDK `LanguageModel`, so providers are whatever AI SDK adapters you install. The examples use `@ai-sdk/openai` and `ollama-ai-provider-v2`, and the README also shows Anthropic, Google and Groq. `ScraperLLMOptions` is `CallSettings` plus `system` and extra `messages`, so temperature, max tokens and retries go to the SDK unchanged ([index.ts#L11-L22](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L11-L22)). The whole preprocessed page goes into one message. Nothing counts tokens, chunks or truncates, so a long page in `raw_html` can exceed the context window.

### Code generation

`generateAISDKCode` awaits `output.responseFormat` to get the JSON Schema that the AI SDK derived from your Zod schema. It then sends `Website / Schema / Content` as one user message under a system prompt that asks for a comment-free IIFE, and strips Markdown fences from the reply with `stripMarkdownBackticks` ([models.ts#L93-L129](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L93-L129)). It returns `{ code, url }` and does not run the code. The caller runs `page.evaluate(code)` and should validate the result, as [examples/codegen.ts](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/codegen.ts#L35-L42) does with `schema.parse`. `ScraperGenerateOptions` limits `format` to `html | raw_html` at the type level.

## Extending it

- **Custom input**: `format: 'custom'` with a `formatFunction(page)` lets you send just one region, an accessibility snapshot, or anything else you build.
- **Prompting**: override `system` or append `messages` (for example few-shot hints) through options.
- **Providers and settings**: any AI SDK model and any `CallSettings`.
- **Agent tool**: [examples/toolUse.ts](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/toolUse.ts#L15-L51) wraps `scraper.run` in an AI SDK `tool()`. Note that the example ignores the `url` argument the model passes in and always opens Hacker News.
- **Stealth, auth, waiting**: configure Playwright before you hand over the `Page` (contexts, cookies, `waitForSelector`, a stealth plugin). The library touches none of this.

## Running it

`npm i zod playwright llm-scraper` plus an AI SDK provider package such as `@ai-sdk/openai`, and that provider's API key in the environment. The package is ESM (`"type": "module"`) and builds with `tsc`. `npm test` runs Vitest tests that launch Chromium, hit `news.ycombinator.com` and `example.com`, and call `gpt-4o-mini`, so they need network access and an OpenAI key ([tests/index.ts](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L1-L38)). No server or database is involved.

## Strengths and caveats

- **Strength: very small surface.** Three methods, one model call each, about 330 lines. Easy to audit or vendor.
- **Strength: typed output and multimodal input.** AI SDK `Output` gives schema-shaped `data`. `image` mode lets vision models read layouts that HTML hides.
- **Strength: codegen escape hatch.** `generate` trades per-page LLM cost for a reusable extractor you can review.
- **Caveat: no fetching or rendering logic.** Readiness, navigation, retries and blocking are all the caller's job.
- **Caveat: one-shot context.** No chunking or token budgeting. Cost and context limits scale with page size and format choice.
- **Caveat: side effects and external fetches.** `html` mode rewrites the live DOM, and `text` mode pulls code from `cdn.skypack.dev` at run time.
- **Caveat: generated code is untrusted.** The `generate` output is run in the page as-is. Nothing sandboxes or checks it beyond what you add.

*Sources: code at 2b43d99, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (9 pages), verified Q&A.*

## How mishushakov/llm-scraper answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

The library does **not** fetch or render pages itself. It delegates entirely to a **Playwright `Page` object** that the caller creates and navigates. The caller launches Playwright (`chromium.launch()`) and navigates (`page.goto(url)`) before passing the page to the scraper — all JS rendering, redirects, and dynamic loading are handled by Playwright's full browser engine. There is **no built-in wait strategy**; the caller must ensure the page is ready (e.g. `await page.waitForSelector(...)`) before calling `scraper.run()`. Once given a `Page`, the `preprocess()` function (`src/preprocess.ts:25-83`) reads content in six modes: **raw_html** (plain `page.content()`, line 34), **html** (runs a cleanup pass then `page.content()`, lines 55-58), **markdown** (extracts `<body>` innerHTML and converts via Turndown, lines 38-40), **text** (evaluates Mozilla Readability.js loaded from CDN inside the browser to extract main-article text, lines 43-53), **image** (takes a Playwright `page.screenshot()` and returns base64, lines 61-65), and **custom** (user-provided `formatFunction(page)`, lines 67-76). There is **no support for PDFs or inline image/media extraction** — images inside pages are removed by the cleanup step.


Citations: [src/preprocess.ts:25-83](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L25-L83) · [tests/index.ts:8-13](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L8-L13) · [src/cleanup.ts:1-60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60) · [src/index.ts:29-36](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L29-L36)

### How is content extracted or converted? (answered)

Extraction is **LLM-driven**, not selector-based. There are no CSS/XPath selectors for data extraction. The flow is: preprocess → send to LLM with a schema → parse the LLM's structured response. **HTML→Markdown** uses the Turndown library (`src/preprocess.ts:2,38-40`). **HTML→Text** uses Mozilla Readability.js dynamically imported from SkyPack CDN inside the browser context, returning the parsed article title + text content (`src/preprocess.ts:43-53`). **Boilerplate removal** is done server-side via `cleanup()` (`src/cleanup.ts:1-60`), which strips 33 element types including `script`, `style`, `nav`, `header`, `footer`, `aside`, `form`, `iframe`, `svg`, `img`, and `canvas`, plus removes attributes like `style`, `src`, `aria-*`, `data-*`, and `on*` event handlers. This cleanup runs by default in `html` mode. **Schema-based extraction** uses Vercel AI SDK's `Output.object({schema})` / `Output.array({element: schema})` — schemas can be **Zod objects** or raw **JSON Schema** (via `jsonSchema()` helper from the AI SDK, per `tests/scraper.test.ts:75-113`). The LLM receives the preprocessed content with the system prompt "You are a sophisticated web scraper" and must return data matching the schema shape. There is no chunking — the entire page content goes in one LLM call.

> **Editor's note.** Correction: `cleanup` removes 30 element types (not 33). It runs inside the browser via `page.evaluate`, not server-side, so in the default `html` mode it mutates the live Playwright page.

Citations: [src/cleanup.ts:1-60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60) · [src/preprocess.ts:37-53](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L37-L53) · [src/models.ts:12-16](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L12-L16) · [tests/scraper.test.ts:75-113](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/scraper.test.ts#L75-L113)

### How are LLMs used, if at all? (answered)

LLMs are the core extraction engine. The **system prompt** is hardcoded to `"You are a sophisticated web scraper. Extract the contents of the webpage"` (`src/models.ts:12-13`). The **user message** contains the preprocessed page content (as text, markdown, or a multimodal image via `prepareAISDKPage()` at `src/models.ts:25-35`). The LLM call goes through Vercel AI SDK's `generateText()`, which supports **structured output / JSON Schema** through the `output` parameter — the SDK constrains the model to produce valid JSON matching the provided Zod or JSON Schema (`src/models.ts:51-57`). **No chunking** exists; large pages are sent whole. **No cost controls, token counting, or trimming** are implemented in the library. **Providers** are any Vercel AI SDK-compatible `LanguageModel`: examples show OpenAI (`@ai-sdk/openai`), Anthropic (`@ai-sdk/anthropic`), Google (`@ai-sdk/google`), Groq (custom `createOpenAI` with baseURL), Ollama (`ollama-ai-provider-v2`). **Streaming** uses `streamText()` from the AI SDK, returning a `partialOutputStream` the caller iterates with `for await` (`src/models.ts:65-91`). **Code generation mode** (`generate()`) instead prompts the LLM to produce a JavaScript IIFE that extracts data — the code is then `page.evaluate()`'d by the caller, bypassing the LLM for repeated runs over similar pages.


Citations: [src/models.ts:12-63](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L12-L63) · [src/models.ts:65-91](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L65-L91) · [src/models.ts:93-129](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L93-L129) · [src/index.ts:48-55](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L48-L55)

### How are anti-bot measures, proxies and fingerprinting handled? (not applicable)

The library does **not** implement any anti-bot evasion. There are no stealth patches (no user-agent spoofing, no TLS fingerprint manipulation, no browser-fingerprint alteration), no proxy rotation, no CAPTCHA handling, no request rate limiting, and no retry logic. The `cleanup()` function removes `script`, `iframe`, `svg`, `aria-*` attributes and event handlers, but this is aimed at **reducing token count** in LLM prompts, not evading detection. The caller manually configures Playwright (`chromium.launch()`) and could pass their own stealth settings there — but the library wraps none of that.


Citations: [src/cleanup.ts:1-60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60) · [tests/index.ts:8-13](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L8-L13)

### How is crawling at scale implemented? (not applicable)

The library does **not** implement crawling. It is a single-page extraction tool with no URL queues, no concurrent-page management, no deduplication, no depth/limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. A caller who needs crawling must build their own loop: iterate over URLs, call `page.goto()` for each, and call `scraper.run()` per page. The test infrastructure (`tests/index.ts:6-12`) and all examples show this pattern — one browser, one page, one `scraper.run()`. No `crawl()` or `spider()` method exists.


Citations: [src/index.ts:24-56](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L24-L56) · [tests/index.ts:6-13](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/index.ts#L6-L13) · [examples/hn.ts:8-44](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/hn.ts#L8-L44)

### What is the developer interface? (answered)

The developer interface is a **single TypeScript library** published as the `llm-scraper` npm package (`package.json:2`). There is **no CLI, no REST service, no MCP server, and no UI** — only the programmatic API. The export is a default class `LLMScraper` (`src/index.ts:24-56`) with three methods:

- **`run(page, output, options?)`** → `{ data, url }` — synchronous extraction. `options.format` selects the preprocessing mode: `'html'` (default, cleaned), `'raw_html'`, `'markdown'`, `'text'` (Readability.js), `'image'` (base64 screenshot), or `'custom'` (user function). `options` also supports **any Vercel AI SDK `CallSettings`** — temperature, maxTokens, etc.
- **`stream(page, output, options?)`** → `{ stream, url }` — same as `run` but returns a `partialOutputStream` the caller iterates with `for await` for incremental results.
- **`generate(page, output, options?)`** → `{ code, url }` — asks the LLM to write a JavaScript IIFE that extracts data according to the schema; the user then runs it with `page.evaluate(code)`.

**Output formats** are structured JSON via LLM (typed by the schema). **Schema definitions** support both Zod (`z.object(...)`) and raw JSON Schema (Vercel AI SDK's `jsonSchema()` helper). The library is **TypeScript-only** — no other language bindings. All examples use a `run()` → inspect result pattern, and the tool-use example (`toolUse.ts:15-48`) shows embedding the scraper inside a Vercel AI SDK `tool()` for agentic workflows — it calls `scraper.run()` inside the tool's `execute` handler.


Citations: [src/index.ts:24-56](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L24-L56) · [src/preprocess.ts:6-17](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L6-L17) · [tests/scraper.test.ts:75-113](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/scraper.test.ts#L75-L113) · [examples/toolUse.ts:15-51](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/toolUse.ts#L15-L51)
