How is content extracted or converted?
HTML→Markdown/text heuristics; readability-style boilerplate removal; selectors; schema-based extraction.
Verdict
The two projects extract content in completely different ways: learned DOM paths versus an LLM filling a schema.
AutoScraper learns by example. You give it values visible on a page. build() finds the elements that hold them, by exact, regex or SequenceMatcher-fuzzy text or attribute match, and records each element’s path from the root as (tag, {class, style}, sibling_index) steps. Replay is pure BeautifulSoup traversal. “Similar” mode fans out across siblings to collect whole lists, and “exact” mode follows the recorded positions. There is no Markdown conversion and no boilerplate removal.
llm-scraper has no selectors. It converts the page to cleaned HTML (an in-page cleanup that removes 30 tag types and noisy attributes), Turndown Markdown, Mozilla Readability text, a screenshot, or a custom string. It then asks the model to fill an AI SDK Output schema (Zod or JSON Schema) in one call. Its generate() mode instead asks the model to write a JavaScript extractor you can reuse.
AutoScraper suits repetitive, stable layouts where you can show it an example and want deterministic, free extraction. It breaks when the page structure changes. llm-scraper suits varied or unfamiliar pages and nested typed output, at the cost of model calls and non-determinism.
More projects in this category are being researched.
Per-project answers
alirezamika/autoscraper
answeredContent extraction is done entirely via BeautifulSoup's tree traversal. The library does not convert HTML to Markdown, does not apply readability-style boilerplate removal, and has no schema-based extraction DSL. Its core mechanism is rule learning by example: given a wanted_list of target strings (or compiled regexps) and a URL, build() at auto_scraper.py:177-258 finds every child element whose text or attribute value matches the sample (via _child_has_text() at auto_scraper.py:136-168 — using exact equality by default, or fuzzy difflib.SequenceMatcher when text_fuzz_ratio < 1.0), then walks upward through the DOM to build a traversal stack of (tag_name, attrs, child_index) tuples (_build_stack() at auto_scraper.py:261-297). Each rule is stored as a dict with content (the stack), wanted_attr (e.g. "href" for links), is_full_url, and a SHA-256-based stack_id. Later, _get_result_with_stack() at auto_scraper.py:330-370 replays the stack on a new page by iteratively calling BeautifulSoup.findAll() with the rule's tag names and attributes, descending level by level. Two retrieval modes exist: similar (sibling-aware, uses findAll broadly) via get_result_similar() at auto_scraper.py:471-545, and exact (index-position-based via _get_result_with_stack_index_based() at auto_scraper.py:372-404) via get_result_exact() at auto_scraper.py:547-611. Attribute matching can also be fuzzied via FuzzyText (utils.py:52-59). Duplicates are removed while preserving order via unique_hashable() (utils.py:20-22). Results can be grouped by rule ID (grouped=True) or by alias (group_by_alias=True).
mishushakov/llm-scraper
answeredExtraction is LLM-driven, not selector-based. There are no CSS/XPath selectors for data extraction. The flow is: preprocess → send to LLM with a schema → parse the LLM's structured response. HTML→Markdown uses the Turndown library (src/preprocess.ts:2,38-40). HTML→Text uses Mozilla Readability.js dynamically imported from SkyPack CDN inside the browser context, returning the parsed article title + text content (src/preprocess.ts:43-53). Boilerplate removal is done server-side via cleanup() (src/cleanup.ts:1-60), which strips 33 element types including script, style, nav, header, footer, aside, form, iframe, svg, img, and canvas, plus removes attributes like style, src, aria-*, data-*, and on* event handlers. This cleanup runs by default in html mode. Schema-based extraction uses Vercel AI SDK's Output.object({schema}) / Output.array({element: schema}) — schemas can be Zod objects or raw JSON Schema (via jsonSchema() helper from the AI SDK, per tests/scraper.test.ts:75-113). The LLM receives the preprocessed content with the system prompt "You are a sophisticated web scraper" and must return data matching the schema shape. There is no chunking — the entire page content goes in one LLM call.
cleanup removes 30 element types (not 33). It runs inside the browser via page.evaluate, not server-side, so in the default html mode it mutates the live Playwright page.← How are pages fetched and rendered? · How are LLMs used, if at all? →