# What is the developer interface?

> AI web scraping — a good answer covers: Library API, CLI, REST service, MCP server, UI; output formats; language bindings.

Canonical page: https://llms-technical-reviews.com/ai-scraping/q/interface/

## Verdict

Both published projects are libraries only, with no CLI, REST service, MCP server or UI.

[AutoScraper](/p/autoscraper/) is a single Python class. `build()` learns rules. `get_result_similar()`, `get_result_exact()` and `get_result()` (both together) apply them. `save()`/`load()` persist rules as JSON, and `remove_rules`/`keep_rules`/`set_rule_aliases` edit them by `stack_id`. Output is a plain list of strings, or a dict when `grouped=True` (keyed by rule) or `group_by_alias=True` (keyed by your aliases from `wanted_dict`). Values are always strings, so you build any typing or nesting yourself.

[llm-scraper](/p/llm-scraper/) is an ESM TypeScript class with three async methods. `run()` returns `{data, url}`, `stream()` returns `{stream, url}` with partial objects, and `generate()` returns `{code, url}`. Output is shaped by an AI SDK `Output`, from Zod or JSON Schema, so nested typed objects and arrays come naturally. Options select the input format and pass AI SDK call settings through. The repo also shows wrapping it as an AI SDK `tool()` for agents.

Choose AutoScraper for Python scripts that need flat fields and reusable rule files. Choose llm-scraper for Node/TypeScript stacks that want typed objects or an agent tool. Neither offers a hosted API, so for that you would wrap either one yourself.

More projects in this category are being researched.

## Per-project answers

### alirezamika/autoscraper (answered)

AutoScraper exposes a pure Python library API only. The single entry point is the `AutoScraper` class imported from `autoscraper.auto_scraper` (`auto_scraper.py:21`). Key methods: `build()` learns rules from a URL/HTML and `wanted_list`; `get_result_similar()` applies rules broadly to extract siblings; `get_result_exact()` applies index-position-based rules for precise ordering; `get_result()` returns both; `save()`/`load()` persist/restore rules as JSON (`auto_scraper.py:53-92`); `remove_rules()`/`keep_rules()`/`set_rule_aliases()` manage the rule set (`auto_scraper.py:673-721`). There is no CLI tool, no REST service, no MCP server, and no graphical UI. Output formats are plain Python lists (or dicts when `grouped=True` or `group_by_alias=True`). There is only one language binding (Python). The `generate_python_code()` method at `auto_scraper.py:723-725` is deprecated and prints a message directing users to `save()` and `load()` instead.


Citations: [autoscraper/auto_scraper.py:21-48](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L21-L48) · [autoscraper/auto_scraper.py:53-92](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L53-L92) · [autoscraper/auto_scraper.py:673-725](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L673-L725)

### mishushakov/llm-scraper (answered)

The developer interface is a **single TypeScript library** published as the `llm-scraper` npm package (`package.json:2`). There is **no CLI, no REST service, no MCP server, and no UI** — only the programmatic API. The export is a default class `LLMScraper` (`src/index.ts:24-56`) with three methods:

- **`run(page, output, options?)`** → `{ data, url }` — synchronous extraction. `options.format` selects the preprocessing mode: `'html'` (default, cleaned), `'raw_html'`, `'markdown'`, `'text'` (Readability.js), `'image'` (base64 screenshot), or `'custom'` (user function). `options` also supports **any Vercel AI SDK `CallSettings`** — temperature, maxTokens, etc.
- **`stream(page, output, options?)`** → `{ stream, url }` — same as `run` but returns a `partialOutputStream` the caller iterates with `for await` for incremental results.
- **`generate(page, output, options?)`** → `{ code, url }` — asks the LLM to write a JavaScript IIFE that extracts data according to the schema; the user then runs it with `page.evaluate(code)`.

**Output formats** are structured JSON via LLM (typed by the schema). **Schema definitions** support both Zod (`z.object(...)`) and raw JSON Schema (Vercel AI SDK's `jsonSchema()` helper). The library is **TypeScript-only** — no other language bindings. All examples use a `run()` → inspect result pattern, and the tool-use example (`toolUse.ts:15-48`) shows embedding the scraper inside a Vercel AI SDK `tool()` for agentic workflows — it calls `scraper.run()` inside the tool's `execute` handler.


Citations: [src/index.ts:24-56](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/index.ts#L24-L56) · [src/preprocess.ts:6-17](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L6-L17) · [tests/scraper.test.ts:75-113](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/scraper.test.ts#L75-L113) · [examples/toolUse.ts:15-51](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/examples/toolUse.ts#L15-L51)
