# How are LLMs used, if at all?

> AI web scraping — a good answer covers: Prompting; chunking of large pages; structured output / JSON schema; which providers; cost controls.

Canonical page: https://llms-technical-reviews.com/ai-scraping/q/llm-usage/

## Verdict

[AutoScraper](/p/autoscraper/) is not_applicable here: it uses no LLM. Its dependencies are only `requests`, `bs4` and `lxml`, and its "learning" is DOM-path recording plus `difflib` string similarity.

[llm-scraper](/p/llm-scraper/) is built around one AI SDK call. `run()` sends the whole preprocessed page as a single user message (text or an image part) under a fixed system prompt, "You are a sophisticated web scraper...". It passes your `Output` to `generateText`, so the SDK handles structured output and parsing. `stream()` does the same through `streamText` and yields partial objects. `generate()` sends the URL, the JSON Schema and the content, and asks for an IIFE scraper. Any AI SDK `LanguageModel` works: the examples use OpenAI and Ollama, and the README adds Anthropic, Google and Groq. You can override `system`, append `messages`, and pass `CallSettings` through. There is no chunking, token counting or truncation. Cost control means choosing a compact format (`markdown`, `text`, cleaned `html`) or using `generate()` once and running the resulting code with `page.evaluate` on later pages.

If you need zero model cost and deterministic output, the non-LLM approach wins. If pages vary or you want typed nested data from a schema, llm-scraper's thin wrapper is easy to read, but you must handle long pages and budgets yourself.

More projects in this category are being researched.

## Per-project answers

### alirezamika/autoscraper (not applicable)

AutoScraper does not use any LLM. There are no prompts, no chunking logic, no structured-output JSON schemas, no provider API keys, and no cost controls anywhere in the codebase. The project's dependencies (`setup.py:29`) are limited to `requests`, `bs4`, and `lxml` — none of which relate to language models. All extraction is purely algorithmic: pattern-matching and DOM traversal with BeautifulSoup. The fuzzy matching in `utils.py:35-40` uses Python's `difflib.SequenceMatcher`, not any learned or generative model.


Citations: [setup.py:29-29](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/setup.py#L29-L29) · [autoscraper/utils.py:35-40](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L35-L40)

### mishushakov/llm-scraper (answered)

LLMs are the core extraction engine. The **system prompt** is hardcoded to `"You are a sophisticated web scraper. Extract the contents of the webpage"` (`src/models.ts:12-13`). The **user message** contains the preprocessed page content (as text, markdown, or a multimodal image via `prepareAISDKPage()` at `src/models.ts:25-35`). The LLM call goes through Vercel AI SDK's `generateText()`, which supports **structured output / JSON Schema** through the `output` parameter — the SDK constrains the model to produce valid JSON matching the provided Zod or JSON Schema (`src/models.ts:51-57`). **No chunking** exists; large pages are sent whole. **No cost controls, token counting, or trimming** are implemented in the library. **Providers** are any Vercel AI SDK-compatible `LanguageModel`: examples show OpenAI (`@ai-sdk/openai`), Anthropic (`@ai-sdk/anthropic`), Google (`@ai-sdk/google`), Groq (custom `createOpenAI` with baseURL), Ollama (`ollama-ai-provider-v2`). **Streaming** uses `streamText()` from the AI SDK, returning a `partialOutputStream` the caller iterates with `for await` (`src/models.ts:65-91`). **Code generation mode** (`generate()`) instead prompts the LLM to produce a JavaScript IIFE that extracts data — the code is then `page.evaluate()`'d by the caller, bypassing the LLM for repeated runs over similar pages.


Citations: [src/models.ts:12-63](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L12-L63) · [src/models.ts:65-91](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L65-L91) · [src/models.ts:93-129](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L93-L129) · [src/index.ts:48-55](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L48-L55)
