What is the developer interface?
Library API, CLI, REST service, MCP server, UI; output formats; language bindings.
Verdict
Both published projects are libraries only, with no CLI, REST service, MCP server or UI.
AutoScraper is a single Python class. build() learns rules. get_result_similar(), get_result_exact() and get_result() (both together) apply them. save()/load() persist rules as JSON, and remove_rules/keep_rules/set_rule_aliases edit them by stack_id. Output is a plain list of strings, or a dict when grouped=True (keyed by rule) or group_by_alias=True (keyed by your aliases from wanted_dict). Values are always strings, so you build any typing or nesting yourself.
llm-scraper is an ESM TypeScript class with three async methods. run() returns {data, url}, stream() returns {stream, url} with partial objects, and generate() returns {code, url}. Output is shaped by an AI SDK Output, from Zod or JSON Schema, so nested typed objects and arrays come naturally. Options select the input format and pass AI SDK call settings through. The repo also shows wrapping it as an AI SDK tool() for agents.
Choose AutoScraper for Python scripts that need flat fields and reusable rule files. Choose llm-scraper for Node/TypeScript stacks that want typed objects or an agent tool. Neither offers a hosted API, so for that you would wrap either one yourself.
More projects in this category are being researched.
Per-project answers
alirezamika/autoscraper
answeredAutoScraper exposes a pure Python library API only. The single entry point is the AutoScraper class imported from autoscraper.auto_scraper (auto_scraper.py:21). Key methods: build() learns rules from a URL/HTML and wanted_list; get_result_similar() applies rules broadly to extract siblings; get_result_exact() applies index-position-based rules for precise ordering; get_result() returns both; save()/load() persist/restore rules as JSON (auto_scraper.py:53-92); remove_rules()/keep_rules()/set_rule_aliases() manage the rule set (auto_scraper.py:673-721). There is no CLI tool, no REST service, no MCP server, and no graphical UI. Output formats are plain Python lists (or dicts when grouped=True or group_by_alias=True). There is only one language binding (Python). The generate_python_code() method at auto_scraper.py:723-725 is deprecated and prints a message directing users to save() and load() instead.
mishushakov/llm-scraper
answeredThe developer interface is a single TypeScript library published as the llm-scraper npm package (package.json:2). There is no CLI, no REST service, no MCP server, and no UI — only the programmatic API. The export is a default class LLMScraper (src/index.ts:24-56) with three methods:
run(page, output, options?)→{ data, url }— synchronous extraction.options.formatselects the preprocessing mode:'html'(default, cleaned),'raw_html','markdown','text'(Readability.js),'image'(base64 screenshot), or'custom'(user function).optionsalso supports any Vercel AI SDKCallSettings— temperature, maxTokens, etc.stream(page, output, options?)→{ stream, url }— same asrunbut returns apartialOutputStreamthe caller iterates withfor awaitfor incremental results.generate(page, output, options?)→{ code, url }— asks the LLM to write a JavaScript IIFE that extracts data according to the schema; the user then runs it withpage.evaluate(code).
Output formats are structured JSON via LLM (typed by the schema). Schema definitions support both Zod (z.object(...)) and raw JSON Schema (Vercel AI SDK's jsonSchema() helper). The library is TypeScript-only — no other language bindings. All examples use a run() → inspect result pattern, and the tool-use example (toolUse.ts:15-48) shows embedding the scraper inside a Vercel AI SDK tool() for agentic workflows — it calls scraper.run() inside the tool's execute handler.