LLMs Technical Reviews
Home / AI web scraping / autoscraper

alirezamika/autoscraper

Python library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages.

GitHub ↗★ 8.0kPythonMITcommit 68a8181 · 2026-07-29

Overview

AutoScraper is a small Python library (one class, about 725 lines, plus a 59-line utils.py) that writes scraping rules for you. You give it a page and a few values you can see on that page, for example a product title and a price. It finds the elements that hold those values, records the path from the document root down to each one, and stores those paths as rules. Later you apply the rules to another page with the same layout and get back the matching values.

Despite the category it sits in, there is no machine learning and no LLM here. “Learning” means a DOM walk plus string comparison, with optional fuzzy matching from difflib.SequenceMatcher. Pages are fetched with a plain requests.get, so JavaScript-rendered content is invisible unless you pass in rendered HTML yourself. There is no crawler, CLI or service: it is a library for people who want a quick extractor for a list or detail page without writing CSS selectors.

The interesting part is the rule format. A rule is a JSON-serialisable dict: a list of (tag, {class, style}, sibling_index) steps plus a few flags. This makes rules easy to save, diff, prune and share.

Architecture

flowchart LR
  U["Caller: url or html + wanted_list"] --> S["_get_soup"]
  S --> F["_fetch_html (requests.get)"]
  S --> BS["BeautifulSoup lxml tree"]
  BS --> GC["_get_children / _child_has_text"]
  GC --> BST["_build_stack"]
  BST --> SL["stack_list (rules)"]
  SL --> J["save / load JSON"]
  SL --> SIM["_get_result_with_stack (similar)"]
  SL --> EX["_get_result_with_stack_index_based (exact)"]
  SIM --> CR["_clean_result"]
  EX --> CR
  CR --> R["list or dict of strings"]
Component Path Role
AutoScraper class autoscraper/auto_scraper.py Public API: build, get_result_similar, get_result_exact, get_result, rule management, save/load
Fetch and parse auto_scraper.py _fetch_html, _get_soup One requests.get, encoding fix-up, NFKD normalisation, BeautifulSoup(html, "lxml")
Matcher auto_scraper.py _child_has_text, _get_children Finds elements whose text, direct text or attribute equals a wanted value
Rule builder auto_scraper.py _build_stack Walks from a matched element up to the root and records the path
Rule replay auto_scraper.py _get_result_with_stack, _get_result_with_stack_index_based Re-walks a path on a new tree, broadly or by exact position
Utilities autoscraper/utils.py text_match, FuzzyText, normalize, ResultItem, order-preserving dedupe

How a request flows

Learning rules with build():

  1. build(url=..., wanted_list=[...]) checks that some target was given, then calls _get_soup, which either parses the html you passed or calls _fetch_html. That method does one requests.get with a hard-coded Chrome 84 User-Agent, a Host header taken from the URL, and any request_args you supplied; it fixes a wrongly declared ISO-8859-1 encoding (auto_scraper.py#L95-L121).
  2. A plain wanted_list becomes wanted_dict = {"": wanted_list}, so aliased and unaliased targets share one loop. Unless update=True, existing rules are dropped (auto_scraper.py#L226-L258).
  3. For each wanted value, _get_children scans every element in reverse document order. _child_has_text accepts an element if its full text matches (but not when its parent has the same text, so the outermost element with exactly that text wins), if its own non-recursive text matches, or if an attribute value matches. For href/src it also tries the URL resolved against the page URL (auto_scraper.py#L135-L175).
  4. _build_stack climbs from the match to the root. At each level it stores the parent tag, its class and style only, and the index of the child among siblings with the same tag and attributes. The rule gets a SHA-256 hash and a stack_id of rule_ plus 8 hex chars (auto_scraper.py#L260-L297).
  5. Each new rule is replayed at once on the same tree with the “similar” strategy, and those values are what build() returns. Rules are deduplicated by hash with unique_stack_list.

Applying rules with get_result_similar() / get_result_exact():

  1. Both go through _get_result_by_func, which parses the page once, optionally tags every element with a child_index (needed for keep_order and group_by_alias), and runs the chosen replay function for each rule (auto_scraper.py#L406-L469).
  2. Similar replay keeps every matching child at every level (findAll(tag, attrs, recursive=False)), which is how one example item generalises to all rows of a list. Only at the leaf does it pick by stored index, unless contain_sibling_leaves=True (auto_scraper.py#L330-L370).
  3. Exact replay follows the stored sibling index at every level and returns at most one value per rule (auto_scraper.py#L372-L404).
  4. _clean_result flattens, sorts by DOM position if asked, and removes duplicates. Or it returns a dict keyed by stack_id (grouped=True) or by alias (group_by_alias=True).

Key components

Matching

text_match in utils.py#L35-L40 is the only comparison primitive. A compiled regex is tested with fullmatch. A ratio of 1.0 (the default) means exact string equality. Anything lower uses SequenceMatcher.ratio(). That is text_fuzz_ratio at build time. At replay time, attr_fuzz_ratio < 1.0 wraps the stored class/style values in FuzzyText, whose search() method BeautifulSoup calls as a matcher (utils.py#L52-L59). Fuzzy matching is O(elements x targets) SequenceMatcher calls, so it gets slow on large pages.

Rule format

_get_valid_attrs keeps only class and style, and fills in empty strings when they are missing (auto_scraper.py#L123-L133). id, data-* and other attributes never appear in a rule. Besides the path, a rule stores wanted_attr (which attribute to read, or None for text), is_full_url, is_non_rec_text, alias, and the page url when the value was a resolved link.

Rule management and persistence

save() writes {"stack_list": [...]} as JSON. load() also accepts the older bare-list format (auto_scraper.py#L53-L93). remove_rules, keep_rules and set_rule_aliases edit the list by stack_id. generate_python_code() still exists but only prints a deprecation notice (auto_scraper.py#L673-L725). The usual workflow: build, inspect get_result_similar(grouped=True) to see which rule produced which values, keep_rules the good ones, then save.

Combined call

get_result() parses once and returns a (similar, exact) tuple. It does not take keep_order or contain_sibling_leaves (auto_scraper.py#L613-L671).

Extending it

There are no plugin hooks. The practical extension points are:

  • Bring your own fetcher. Every entry point accepts html=, and get_result_similar/get_result_exact also accept a pre-built soup=. Render with Playwright or fetch through your own proxy layer, then pass the HTML in.
  • request_args goes straight to requests.get, so proxies, cookies, timeouts and extra headers work. Your headers are merged over the default UA.
  • Subclassing. _get_valid_attrs and request_headers are class-level. Overriding them changes which attributes rules depend on, or the default UA.
  • Rules as data. The JSON rule file can be edited or generated by other tools.

Running it

pip install autoscraper (or from git). Its dependencies are requests, bs4 and lxml, Python 3.6+ (setup.py#L1-L30). No services, keys or browsers are needed. Tests use pytest under tests/unit and tests/integration.

Strengths and caveats

  • Strength: tiny and dependency-light. The whole algorithm fits in one file you can read in an hour, and rules are plain JSON.
  • Strength: no per-page cost. After build, extraction is pure tree traversal. It is fast and deterministic, with no API calls.
  • Strength: list generalisation. One example row often yields every row, because similar-mode replay fans out at every non-leaf level.
  • Caveat: structural brittleness. Rules depend on exact tag paths, class/style values and sibling positions from the root down. A wrapper div added to the layout breaks every rule. attr_fuzz_ratio only helps with class-name drift.
  • Caveat: noisy rule sets. Every element that matches a wanted value produces a rule. Common strings, such as a price that appears twice, create extra rules you must prune by hand.
  • Caveat: static HTML only. No JS rendering, retries, rate limiting or rotation. The built-in UA is a 2020 Chrome string.
  • Caveat: single page per call. No link following or crawling; you loop over URLs yourself.

Sources: code at 68a8181, deepwiki-open wiki (16 pages), OpenDeepWiki wiki (6 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 68a8181. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched via plain synchronous HTTP using the requests library. The classmethod _fetch_html() at auto_scraper.py:96-110 performs a requests.get(url, headers=headers, **request_args) call. A default Chrome 84 User-Agent header is set at auto_scraper.py:45-48. No headless browser, no JavaScript rendering engine, and no waiting strategy is implemented — if the target page relies on JS to populate content, AutoScraper will only see the raw HTML. request_args is forwarded as **kwargs to requests.get(), allowing users to manually pass proxies, custom headers, cookies, or timeouts. Content-type detection for encoding is minimal: if the server declares ISO-8859-1 but the actual Content-Type header doesn't contain it, the apparent encoding is used instead (auto_scraper.py:105-108). The only content type consumed is text/HTML; PDFs, images, or binary responses are not supported at all — res.text is called unconditionally and passed through str.strip() + unicodedata.normalize('NFKD', ...) (utils.py:29-32). The fallback path _get_soup() at auto_scraper.py:113-121 also accepts a raw html string parameter, so the caller can pre-fetch with any tool and feed the HTML directly.

How is content extracted or converted?

answered

Content extraction is done entirely via BeautifulSoup's tree traversal. The library does not convert HTML to Markdown, does not apply readability-style boilerplate removal, and has no schema-based extraction DSL. Its core mechanism is rule learning by example: given a wanted_list of target strings (or compiled regexps) and a URL, build() at auto_scraper.py:177-258 finds every child element whose text or attribute value matches the sample (via _child_has_text() at auto_scraper.py:136-168 — using exact equality by default, or fuzzy difflib.SequenceMatcher when text_fuzz_ratio < 1.0), then walks upward through the DOM to build a traversal stack of (tag_name, attrs, child_index) tuples (_build_stack() at auto_scraper.py:261-297). Each rule is stored as a dict with content (the stack), wanted_attr (e.g. "href" for links), is_full_url, and a SHA-256-based stack_id. Later, _get_result_with_stack() at auto_scraper.py:330-370 replays the stack on a new page by iteratively calling BeautifulSoup.findAll() with the rule's tag names and attributes, descending level by level. Two retrieval modes exist: similar (sibling-aware, uses findAll broadly) via get_result_similar() at auto_scraper.py:471-545, and exact (index-position-based via _get_result_with_stack_index_based() at auto_scraper.py:372-404) via get_result_exact() at auto_scraper.py:547-611. Attribute matching can also be fuzzied via FuzzyText (utils.py:52-59). Duplicates are removed while preserving order via unique_hashable() (utils.py:20-22). Results can be grouped by rule ID (grouped=True) or by alias (group_by_alias=True).

How are LLMs used, if at all?

not applicable

AutoScraper does not use any LLM. There are no prompts, no chunking logic, no structured-output JSON schemas, no provider API keys, and no cost controls anywhere in the codebase. The project's dependencies (setup.py:29) are limited to requests, bs4, and lxml — none of which relate to language models. All extraction is purely algorithmic: pattern-matching and DOM traversal with BeautifulSoup. The fuzzy matching in utils.py:35-40 uses Python's difflib.SequenceMatcher, not any learned or generative model.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot measures are minimal and largely absent. The only built-in fingerprinting is a default Chrome 84 User-Agent header set at auto_scraper.py:45-48. The _fetch_html() method at auto_scraper.py:96-110 also sets the Host header from the URL's netloc. Beyond that, there are no stealth patches (no TLS fingerprint spoofing, no browser-emulation of headers beyond UA), no proxy rotation system, no CAPTCHA handling, and no rate-limiting logic — no retry-on-403, no exponential backoff, no request throttling. The code does accept a request_args parameter (auto_scraper.py:97) that is unpacked as **kwargs to requests.get(), which lets the calling user manually supply proxy dictionaries, custom headers, cookies, or timeouts. The README (README.md:86-94) demonstrates passing proxies via this mechanism. This is a pass-through, not a built-in rotation or management strategy.

How is crawling at scale implemented?

not applicable

AutoScraper has no crawling capability. It is a single-page scraper: each call to build(), get_result_similar(), or get_result_exact() fetches exactly one URL and extracts data from that page. There are no URL queues, no concurrency/threading, no URL-deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker architecture. The library does not even iterate over links on a page — it has no link-extraction logic and no "follow" mechanism. The entire source is two small files focusing purely on learning extraction rules from one page and applying them to another single page.

What is the developer interface?

answered

AutoScraper exposes a pure Python library API only. The single entry point is the AutoScraper class imported from autoscraper.auto_scraper (auto_scraper.py:21). Key methods: build() learns rules from a URL/HTML and wanted_list; get_result_similar() applies rules broadly to extract siblings; get_result_exact() applies index-position-based rules for precise ordering; get_result() returns both; save()/load() persist/restore rules as JSON (auto_scraper.py:53-92); remove_rules()/keep_rules()/set_rule_aliases() manage the rule set (auto_scraper.py:673-721). There is no CLI tool, no REST service, no MCP server, and no graphical UI. Output formats are plain Python lists (or dicts when grouped=True or group_by_alias=True). There is only one language binding (Python). The generate_python_code() method at auto_scraper.py:723-725 is deprecated and prints a message directing users to save() and load() instead.