mishushakov/llm-scraper
TypeScript library that turns an already-open Playwright page into schema-typed data, or generated scraper code, via the Vercel AI SDK.
Overview
llm-scraper is a thin TypeScript layer between a Playwright Page and the Vercel AI SDK. You open and navigate the page yourself, describe the data you want as an AI SDK Output (usually Output.object({ schema }) with a Zod schema), and call run, stream or generate. The library turns the page into one model input (cleaned HTML, raw HTML, Markdown, Readability text, a screenshot, or your own string) and makes one generateText or streamText call.
The whole library is four source files and about 330 lines. It has no fetching, waiting, crawling, retry, chunking or anti-bot logic. Those are left to Playwright and to the caller. Its runtime dependencies are only ai, @ai-sdk/provider and turndown. Playwright and Zod are dev dependencies, because the library only uses Playwright’s Page type and never launches a browser (package.json).
It suits developers who already drive a browser from Node and want typed extraction from arbitrary pages without writing selectors. The generate mode is the cost-saving option: the model writes a JavaScript extractor once, and you run it with page.evaluate on later pages.
Architecture
flowchart LR
C["Caller: Playwright page + Output schema"] --> L["LLMScraper (index.ts)"]
L --> P["preprocess (preprocess.ts)"]
P --> CL["cleanup (in-page DOM strip)"]
P --> TD["Turndown markdown"]
P --> RD["Readability via CDN import"]
P --> SS["page.screenshot base64"]
L --> M["models.ts"]
M --> GT["generateText (run)"]
M --> ST["streamText (stream)"]
M --> GC["generateText code prompt (generate)"]
GT --> D["data, url"]
ST --> S["partialOutputStream, url"]
GC --> K["code, url"]
| Component | Path | Role |
|---|---|---|
LLMScraper |
src/index.ts |
Public class: run, stream, generate, all wrapping preprocess and one model call |
| Preprocessor | src/preprocess.ts |
Turns the live page into {url, content, format} in one of six formats |
| DOM cleanup | src/cleanup.ts |
Runs inside the page. Removes noisy elements and attributes before HTML capture |
| Model calls | src/models.ts |
Builds messages and system prompts, calls the AI SDK, strips code fences |
| Examples | examples/*.ts |
HN extraction, streaming, codegen, Ollama, tool use |
| Tests | tests/ |
Live Vitest runs against Hacker News and example.com with gpt-4o-mini |
How a request flows
For scraper.run(page, Output.object({ schema }), { format: 'markdown' }):
LLMScraper.runcallspreprocess(page, options), thengenerateAISDKCompletionswith the model you passed to the constructor (index.ts#L24-L56).preprocessreadspage.url()and picks the format,htmlby default (preprocess.ts#L25-L83):raw_html:page.content().markdown:page.innerHTML('body')converted withnew Turndown().turndown(...)in Node.text: inside the page, dynamicallyimport('https://cdn.skypack.dev/@mozilla/readability'), runReadability(document).parse(), and returnPage Title: ...plustextContent.html:page.evaluate(cleanup), thenpage.content().image:page.screenshot({ fullPage })as base64.custom:await options.formatFunction(page). It throws if the function is missing.
prepareAISDKPagewraps the content as a single text part, or an image part forimage(models.ts#L25-L35).generateAISDKCompletionsbuildsmessages = [user: page content, ...options.messages], uses the default system prompt “You are a sophisticated web scraper. Extract the contents of the webpage” unless you overridesystem, and spreads the remaining options intogenerateText({ model, output, system, messages, ...rest })(models.ts#L12-L63).- It returns
{ data: result.output, url }. Schema enforcement and parsing come from the AI SDK’sOutputhandling. The library itself never callsschema.parse.
stream is the same, except that it calls streamText and returns { stream: partialOutputStream, url } for for await consumption (models.ts#L65-L91).
Key components
Preprocessing and cleanup
cleanup is a function serialised into the page by page.evaluate. It walks document.querySelectorAll('*'), removes 30 tag types (script, style, svg, img, form, input, button, header, footer, nav, aside, head and others), and strips attributes whose names start with style, src, alt, title, role, aria-, tabindex, on or data- (cleanup.ts#L1-L60). The goal is fewer tokens, not boilerplate detection. class, id and href survive. The cleanup mutates the live page: after an html-mode call, the same Page has no scripts, forms, nav or images. This matters if you keep using the page, or if you generate code in html mode and then evaluate it on the stripped DOM. The bundled codegen example uses raw_html.
text mode loads Readability from a public CDN at run time inside the page. It needs outbound network access from the browser, and a page Content Security Policy can block it.
Model integration
The library accepts any AI SDK LanguageModel, so providers are whatever AI SDK adapters you install. The examples use @ai-sdk/openai and ollama-ai-provider-v2, and the README also shows Anthropic, Google and Groq. ScraperLLMOptions is CallSettings plus system and extra messages, so temperature, max tokens and retries go to the SDK unchanged (index.ts#L11-L22). The whole preprocessed page goes into one message. Nothing counts tokens, chunks or truncates, so a long page in raw_html can exceed the context window.
Code generation
generateAISDKCode awaits output.responseFormat to get the JSON Schema that the AI SDK derived from your Zod schema. It then sends Website / Schema / Content as one user message under a system prompt that asks for a comment-free IIFE, and strips Markdown fences from the reply with stripMarkdownBackticks (models.ts#L93-L129). It returns { code, url } and does not run the code. The caller runs page.evaluate(code) and should validate the result, as examples/codegen.ts does with schema.parse. ScraperGenerateOptions limits format to html | raw_html at the type level.
Extending it
- Custom input:
format: 'custom'with aformatFunction(page)lets you send just one region, an accessibility snapshot, or anything else you build. - Prompting: override
systemor appendmessages(for example few-shot hints) through options. - Providers and settings: any AI SDK model and any
CallSettings. - Agent tool: examples/toolUse.ts wraps
scraper.runin an AI SDKtool(). Note that the example ignores theurlargument the model passes in and always opens Hacker News. - Stealth, auth, waiting: configure Playwright before you hand over the
Page(contexts, cookies,waitForSelector, a stealth plugin). The library touches none of this.
Running it
npm i zod playwright llm-scraper plus an AI SDK provider package such as @ai-sdk/openai, and that provider’s API key in the environment. The package is ESM ("type": "module") and builds with tsc. npm test runs Vitest tests that launch Chromium, hit news.ycombinator.com and example.com, and call gpt-4o-mini, so they need network access and an OpenAI key (tests/index.ts). No server or database is involved.
Strengths and caveats
- Strength: very small surface. Three methods, one model call each, about 330 lines. Easy to audit or vendor.
- Strength: typed output and multimodal input. AI SDK
Outputgives schema-shapeddata.imagemode lets vision models read layouts that HTML hides. - Strength: codegen escape hatch.
generatetrades per-page LLM cost for a reusable extractor you can review. - Caveat: no fetching or rendering logic. Readiness, navigation, retries and blocking are all the caller’s job.
- Caveat: one-shot context. No chunking or token budgeting. Cost and context limits scale with page size and format choice.
- Caveat: side effects and external fetches.
htmlmode rewrites the live DOM, andtextmode pulls code fromcdn.skypack.devat run time. - Caveat: generated code is untrusted. The
generateoutput is run in the page as-is. Nothing sandboxes or checks it beyond what you add.
Sources: code at 2b43d99, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (9 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 2b43d99. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredThe library does not fetch or render pages itself. It delegates entirely to a Playwright Page object that the caller creates and navigates. The caller launches Playwright (chromium.launch()) and navigates (page.goto(url)) before passing the page to the scraper — all JS rendering, redirects, and dynamic loading are handled by Playwright's full browser engine. There is no built-in wait strategy; the caller must ensure the page is ready (e.g. await page.waitForSelector(...)) before calling scraper.run(). Once given a Page, the preprocess() function (src/preprocess.ts:25-83) reads content in six modes: raw_html (plain page.content(), line 34), html (runs a cleanup pass then page.content(), lines 55-58), markdown (extracts <body> innerHTML and converts via Turndown, lines 38-40), text (evaluates Mozilla Readability.js loaded from CDN inside the browser to extract main-article text, lines 43-53), image (takes a Playwright page.screenshot() and returns base64, lines 61-65), and custom (user-provided formatFunction(page), lines 67-76). There is no support for PDFs or inline image/media extraction — images inside pages are removed by the cleanup step.
How is content extracted or converted?
answeredExtraction is LLM-driven, not selector-based. There are no CSS/XPath selectors for data extraction. The flow is: preprocess → send to LLM with a schema → parse the LLM's structured response. HTML→Markdown uses the Turndown library (src/preprocess.ts:2,38-40). HTML→Text uses Mozilla Readability.js dynamically imported from SkyPack CDN inside the browser context, returning the parsed article title + text content (src/preprocess.ts:43-53). Boilerplate removal is done server-side via cleanup() (src/cleanup.ts:1-60), which strips 33 element types including script, style, nav, header, footer, aside, form, iframe, svg, img, and canvas, plus removes attributes like style, src, aria-*, data-*, and on* event handlers. This cleanup runs by default in html mode. Schema-based extraction uses Vercel AI SDK's Output.object({schema}) / Output.array({element: schema}) — schemas can be Zod objects or raw JSON Schema (via jsonSchema() helper from the AI SDK, per tests/scraper.test.ts:75-113). The LLM receives the preprocessed content with the system prompt "You are a sophisticated web scraper" and must return data matching the schema shape. There is no chunking — the entire page content goes in one LLM call.
cleanup removes 30 element types (not 33). It runs inside the browser via page.evaluate, not server-side, so in the default html mode it mutates the live Playwright page.How are LLMs used, if at all?
answeredLLMs are the core extraction engine. The system prompt is hardcoded to "You are a sophisticated web scraper. Extract the contents of the webpage" (src/models.ts:12-13). The user message contains the preprocessed page content (as text, markdown, or a multimodal image via prepareAISDKPage() at src/models.ts:25-35). The LLM call goes through Vercel AI SDK's generateText(), which supports structured output / JSON Schema through the output parameter — the SDK constrains the model to produce valid JSON matching the provided Zod or JSON Schema (src/models.ts:51-57). No chunking exists; large pages are sent whole. No cost controls, token counting, or trimming are implemented in the library. Providers are any Vercel AI SDK-compatible LanguageModel: examples show OpenAI (@ai-sdk/openai), Anthropic (@ai-sdk/anthropic), Google (@ai-sdk/google), Groq (custom createOpenAI with baseURL), Ollama (ollama-ai-provider-v2). Streaming uses streamText() from the AI SDK, returning a partialOutputStream the caller iterates with for await (src/models.ts:65-91). Code generation mode (generate()) instead prompts the LLM to produce a JavaScript IIFE that extracts data — the code is then page.evaluate()'d by the caller, bypassing the LLM for repeated runs over similar pages.
How are anti-bot measures, proxies and fingerprinting handled?
not applicableThe library does not implement any anti-bot evasion. There are no stealth patches (no user-agent spoofing, no TLS fingerprint manipulation, no browser-fingerprint alteration), no proxy rotation, no CAPTCHA handling, no request rate limiting, and no retry logic. The cleanup() function removes script, iframe, svg, aria-* attributes and event handlers, but this is aimed at reducing token count in LLM prompts, not evading detection. The caller manually configures Playwright (chromium.launch()) and could pass their own stealth settings there — but the library wraps none of that.
How is crawling at scale implemented?
not applicableThe library does not implement crawling. It is a single-page extraction tool with no URL queues, no concurrent-page management, no deduplication, no depth/limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. A caller who needs crawling must build their own loop: iterate over URLs, call page.goto() for each, and call scraper.run() per page. The test infrastructure (tests/index.ts:6-12) and all examples show this pattern — one browser, one page, one scraper.run(). No crawl() or spider() method exists.
What is the developer interface?
answeredThe developer interface is a single TypeScript library published as the llm-scraper npm package (package.json:2). There is no CLI, no REST service, no MCP server, and no UI — only the programmatic API. The export is a default class LLMScraper (src/index.ts:24-56) with three methods:
run(page, output, options?)→{ data, url }— synchronous extraction.options.formatselects the preprocessing mode:'html'(default, cleaned),'raw_html','markdown','text'(Readability.js),'image'(base64 screenshot), or'custom'(user function).optionsalso supports any Vercel AI SDKCallSettings— temperature, maxTokens, etc.stream(page, output, options?)→{ stream, url }— same asrunbut returns apartialOutputStreamthe caller iterates withfor awaitfor incremental results.generate(page, output, options?)→{ code, url }— asks the LLM to write a JavaScript IIFE that extracts data according to the schema; the user then runs it withpage.evaluate(code).
Output formats are structured JSON via LLM (typed by the schema). Schema definitions support both Zod (z.object(...)) and raw JSON Schema (Vercel AI SDK's jsonSchema() helper). The library is TypeScript-only — no other language bindings. All examples use a run() → inspect result pattern, and the tool-use example (toolUse.ts:15-48) shows embedding the scraper inside a Vercel AI SDK tool() for agentic workflows — it calls scraper.run() inside the tool's execute handler.