LLMs Technical Reviews

How are LLMs used, if at all?

Prompting; chunking of large pages; structured output / JSON schema; which providers; cost controls.

Verdict

AutoScraper is not_applicable here: it uses no LLM. Its dependencies are only requests, bs4 and lxml, and its “learning” is DOM-path recording plus difflib string similarity.

llm-scraper is built around one AI SDK call. run() sends the whole preprocessed page as a single user message (text or an image part) under a fixed system prompt, “You are a sophisticated web scraper…”. It passes your Output to generateText, so the SDK handles structured output and parsing. stream() does the same through streamText and yields partial objects. generate() sends the URL, the JSON Schema and the content, and asks for an IIFE scraper. Any AI SDK LanguageModel works: the examples use OpenAI and Ollama, and the README adds Anthropic, Google and Groq. You can override system, append messages, and pass CallSettings through. There is no chunking, token counting or truncation. Cost control means choosing a compact format (markdown, text, cleaned html) or using generate() once and running the resulting code with page.evaluate on later pages.

If you need zero model cost and deterministic output, the non-LLM approach wins. If pages vary or you want typed nested data from a schema, llm-scraper’s thin wrapper is easy to read, but you must handle long pages and budgets yourself.

More projects in this category are being researched.

Per-project answers

mishushakov/llm-scraper

answered

LLMs are the core extraction engine. The system prompt is hardcoded to "You are a sophisticated web scraper. Extract the contents of the webpage" (src/models.ts:12-13). The user message contains the preprocessed page content (as text, markdown, or a multimodal image via prepareAISDKPage() at src/models.ts:25-35). The LLM call goes through Vercel AI SDK's generateText(), which supports structured output / JSON Schema through the output parameter — the SDK constrains the model to produce valid JSON matching the provided Zod or JSON Schema (src/models.ts:51-57). No chunking exists; large pages are sent whole. No cost controls, token counting, or trimming are implemented in the library. Providers are any Vercel AI SDK-compatible LanguageModel: examples show OpenAI (@ai-sdk/openai), Anthropic (@ai-sdk/anthropic), Google (@ai-sdk/google), Groq (custom createOpenAI with baseURL), Ollama (ollama-ai-provider-v2). Streaming uses streamText() from the AI SDK, returning a partialOutputStream the caller iterates with for await (src/models.ts:65-91). Code generation mode (generate()) instead prompts the LLM to produce a JavaScript IIFE that extracts data — the code is then page.evaluate()'d by the caller, bypassing the LLM for repeated runs over similar pages.

alirezamika/autoscraper

not applicable

AutoScraper does not use any LLM. There are no prompts, no chunking logic, no structured-output JSON schemas, no provider API keys, and no cost controls anywhere in the codebase. The project's dependencies (setup.py:29) are limited to requests, bs4, and lxml — none of which relate to language models. All extraction is purely algorithmic: pattern-matching and DOM traversal with BeautifulSoup. The fuzzy matching in utils.py:35-40 uses Python's difflib.SequenceMatcher, not any learned or generative model.

← How is content extracted or converted? · How are anti-bot measures, proxies and fingerprinting handled? →