# How is content extracted or converted?

> AI web scraping — a good answer covers: HTML→Markdown/text heuristics; readability-style boilerplate removal; selectors; schema-based extraction.

Canonical page: https://llms-technical-reviews.com/ai-scraping/q/extraction/

## Verdict

The two projects extract content in completely different ways: learned DOM paths versus an LLM filling a schema.

[AutoScraper](/p/autoscraper/) learns by example. You give it values visible on a page. `build()` finds the elements that hold them, by exact, regex or `SequenceMatcher`-fuzzy text or attribute match, and records each element's path from the root as `(tag, {class, style}, sibling_index)` steps. Replay is pure BeautifulSoup traversal. "Similar" mode fans out across siblings to collect whole lists, and "exact" mode follows the recorded positions. There is no Markdown conversion and no boilerplate removal.

[llm-scraper](/p/llm-scraper/) has no selectors. It converts the page to cleaned HTML (an in-page `cleanup` that removes 30 tag types and noisy attributes), Turndown Markdown, Mozilla Readability text, a screenshot, or a custom string. It then asks the model to fill an AI SDK `Output` schema (Zod or JSON Schema) in one call. Its `generate()` mode instead asks the model to write a JavaScript extractor you can reuse.

AutoScraper suits repetitive, stable layouts where you can show it an example and want deterministic, free extraction. It breaks when the page structure changes. llm-scraper suits varied or unfamiliar pages and nested typed output, at the cost of model calls and non-determinism.

More projects in this category are being researched.

## Per-project answers

### alirezamika/autoscraper (answered)

Content extraction is done entirely via BeautifulSoup's tree traversal. The library does **not** convert HTML to Markdown, does **not** apply readability-style boilerplate removal, and has no schema-based extraction DSL. Its core mechanism is **rule learning by example**: given a `wanted_list` of target strings (or compiled regexps) and a URL, `build()` at `auto_scraper.py:177-258` finds every child element whose text or attribute value matches the sample (via `_child_has_text()` at `auto_scraper.py:136-168` — using exact equality by default, or fuzzy `difflib.SequenceMatcher` when `text_fuzz_ratio < 1.0`), then walks upward through the DOM to build a traversal stack of `(tag_name, attrs, child_index)` tuples (`_build_stack()` at `auto_scraper.py:261-297`). Each rule is stored as a dict with `content` (the stack), `wanted_attr` (e.g. `"href"` for links), `is_full_url`, and a SHA-256-based `stack_id`. Later, `_get_result_with_stack()` at `auto_scraper.py:330-370` replays the stack on a new page by iteratively calling `BeautifulSoup.findAll()` with the rule's tag names and attributes, descending level by level. Two retrieval modes exist: **similar** (sibling-aware, uses `findAll` broadly) via `get_result_similar()` at `auto_scraper.py:471-545`, and **exact** (index-position-based via `_get_result_with_stack_index_based()` at `auto_scraper.py:372-404`) via `get_result_exact()` at `auto_scraper.py:547-611`. Attribute matching can also be fuzzied via `FuzzyText` (`utils.py:52-59`). Duplicates are removed while preserving order via `unique_hashable()` (`utils.py:20-22`). Results can be grouped by rule ID (`grouped=True`) or by alias (`group_by_alias=True`).


Citations: [autoscraper/auto_scraper.py:177-258](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L177-L258) · [autoscraper/auto_scraper.py:136-168](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L136-L168) · [autoscraper/auto_scraper.py:261-297](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L261-L297) · [autoscraper/auto_scraper.py:330-370](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L330-L370) · [autoscraper/utils.py:20-22](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L20-L22) · [autoscraper/utils.py:52-59](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L52-L59)

### mishushakov/llm-scraper (answered)

Extraction is **LLM-driven**, not selector-based. There are no CSS/XPath selectors for data extraction. The flow is: preprocess → send to LLM with a schema → parse the LLM's structured response. **HTML→Markdown** uses the Turndown library (`src/preprocess.ts:2,38-40`). **HTML→Text** uses Mozilla Readability.js dynamically imported from SkyPack CDN inside the browser context, returning the parsed article title + text content (`src/preprocess.ts:43-53`). **Boilerplate removal** is done server-side via `cleanup()` (`src/cleanup.ts:1-60`), which strips 33 element types including `script`, `style`, `nav`, `header`, `footer`, `aside`, `form`, `iframe`, `svg`, `img`, and `canvas`, plus removes attributes like `style`, `src`, `aria-*`, `data-*`, and `on*` event handlers. This cleanup runs by default in `html` mode. **Schema-based extraction** uses Vercel AI SDK's `Output.object({schema})` / `Output.array({element: schema})` — schemas can be **Zod objects** or raw **JSON Schema** (via `jsonSchema()` helper from the AI SDK, per `tests/scraper.test.ts:75-113`). The LLM receives the preprocessed content with the system prompt "You are a sophisticated web scraper" and must return data matching the schema shape. There is no chunking — the entire page content goes in one LLM call.

> **Editor's note.** Correction: `cleanup` removes 30 element types (not 33). It runs inside the browser via `page.evaluate`, not server-side, so in the default `html` mode it mutates the live Playwright page.

Citations: [src/cleanup.ts:1-60](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/cleanup.ts#L1-L60) · [src/preprocess.ts:37-53](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/preprocess.ts#L37-L53) · [src/models.ts:12-16](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/src/models.ts#L12-L16) · [tests/scraper.test.ts:75-113](https://github.com/mishushakov/llm-scraper/blob/2b43d999a17eac040f7cc315fc21dd526870564f/tests/scraper.test.ts#L75-L113)
