# alirezamika/autoscraper

> Python library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/alirezamika/autoscraper (reviewed at commit `68a818158c673bf320a8569da10a8b979c1d23fe`, 2026-07-29)
- Stars: 8004 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/autoscraper/

## Overview

AutoScraper is a small Python library (one class, about 725 lines, plus a 59-line `utils.py`) that writes scraping rules for you. You give it a page and a few values you can see on that page, for example a product title and a price. It finds the elements that hold those values, records the path from the document root down to each one, and stores those paths as rules. Later you apply the rules to another page with the same layout and get back the matching values.

Despite the category it sits in, there is no machine learning and no LLM here. "Learning" means a DOM walk plus string comparison, with optional fuzzy matching from `difflib.SequenceMatcher`. Pages are fetched with a plain `requests.get`, so JavaScript-rendered content is invisible unless you pass in rendered HTML yourself. There is no crawler, CLI or service: it is a library for people who want a quick extractor for a list or detail page without writing CSS selectors.

The interesting part is the rule format. A rule is a JSON-serialisable dict: a list of `(tag, {class, style}, sibling_index)` steps plus a few flags. This makes rules easy to save, diff, prune and share.

## Architecture

```mermaid
flowchart LR
  U["Caller: url or html + wanted_list"] --> S["_get_soup"]
  S --> F["_fetch_html (requests.get)"]
  S --> BS["BeautifulSoup lxml tree"]
  BS --> GC["_get_children / _child_has_text"]
  GC --> BST["_build_stack"]
  BST --> SL["stack_list (rules)"]
  SL --> J["save / load JSON"]
  SL --> SIM["_get_result_with_stack (similar)"]
  SL --> EX["_get_result_with_stack_index_based (exact)"]
  SIM --> CR["_clean_result"]
  EX --> CR
  CR --> R["list or dict of strings"]
```

| Component | Path | Role |
|---|---|---|
| `AutoScraper` class | `autoscraper/auto_scraper.py` | Public API: `build`, `get_result_similar`, `get_result_exact`, `get_result`, rule management, `save`/`load` |
| Fetch and parse | `auto_scraper.py` `_fetch_html`, `_get_soup` | One `requests.get`, encoding fix-up, NFKD normalisation, `BeautifulSoup(html, "lxml")` |
| Matcher | `auto_scraper.py` `_child_has_text`, `_get_children` | Finds elements whose text, direct text or attribute equals a wanted value |
| Rule builder | `auto_scraper.py` `_build_stack` | Walks from a matched element up to the root and records the path |
| Rule replay | `auto_scraper.py` `_get_result_with_stack`, `_get_result_with_stack_index_based` | Re-walks a path on a new tree, broadly or by exact position |
| Utilities | `autoscraper/utils.py` | `text_match`, `FuzzyText`, `normalize`, `ResultItem`, order-preserving dedupe |

## How a request flows

Learning rules with `build()`:

1. `build(url=..., wanted_list=[...])` checks that some target was given, then calls `_get_soup`, which either parses the `html` you passed or calls `_fetch_html`. That method does one `requests.get` with a hard-coded Chrome 84 User-Agent, a `Host` header taken from the URL, and any `request_args` you supplied; it fixes a wrongly declared ISO-8859-1 encoding ([auto_scraper.py#L95-L121](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L95-L121)).
2. A plain `wanted_list` becomes `wanted_dict = {"": wanted_list}`, so aliased and unaliased targets share one loop. Unless `update=True`, existing rules are dropped ([auto_scraper.py#L226-L258](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L226-L258)).
3. For each wanted value, `_get_children` scans every element in reverse document order. `_child_has_text` accepts an element if its full text matches (but not when its parent has the same text, so the outermost element with exactly that text wins), if its own non-recursive text matches, or if an attribute value matches. For `href`/`src` it also tries the URL resolved against the page URL ([auto_scraper.py#L135-L175](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L135-L175)).
4. `_build_stack` climbs from the match to the root. At each level it stores the parent tag, its `class` and `style` only, and the index of the child among siblings with the same tag and attributes. The rule gets a SHA-256 `hash` and a `stack_id` of `rule_` plus 8 hex chars ([auto_scraper.py#L260-L297](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L260-L297)).
5. Each new rule is replayed at once on the same tree with the "similar" strategy, and those values are what `build()` returns. Rules are deduplicated by hash with `unique_stack_list`.

Applying rules with `get_result_similar()` / `get_result_exact()`:

1. Both go through `_get_result_by_func`, which parses the page once, optionally tags every element with a `child_index` (needed for `keep_order` and `group_by_alias`), and runs the chosen replay function for each rule ([auto_scraper.py#L406-L469](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L406-L469)).
2. **Similar** replay keeps every matching child at every level (`findAll(tag, attrs, recursive=False)`), which is how one example item generalises to all rows of a list. Only at the leaf does it pick by stored index, unless `contain_sibling_leaves=True` ([auto_scraper.py#L330-L370](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L330-L370)).
3. **Exact** replay follows the stored sibling index at every level and returns at most one value per rule ([auto_scraper.py#L372-L404](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L372-L404)).
4. `_clean_result` flattens, sorts by DOM position if asked, and removes duplicates. Or it returns a dict keyed by `stack_id` (`grouped=True`) or by alias (`group_by_alias=True`).

## Key components

### Matching

`text_match` in [utils.py#L35-L40](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L35-L40) is the only comparison primitive. A compiled regex is tested with `fullmatch`. A ratio of 1.0 (the default) means exact string equality. Anything lower uses `SequenceMatcher.ratio()`. That is `text_fuzz_ratio` at build time. At replay time, `attr_fuzz_ratio < 1.0` wraps the stored `class`/`style` values in `FuzzyText`, whose `search()` method BeautifulSoup calls as a matcher ([utils.py#L52-L59](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L52-L59)). Fuzzy matching is O(elements x targets) `SequenceMatcher` calls, so it gets slow on large pages.

### Rule format

`_get_valid_attrs` keeps only `class` and `style`, and fills in empty strings when they are missing ([auto_scraper.py#L123-L133](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L123-L133)). `id`, `data-*` and other attributes never appear in a rule. Besides the path, a rule stores `wanted_attr` (which attribute to read, or `None` for text), `is_full_url`, `is_non_rec_text`, `alias`, and the page `url` when the value was a resolved link.

### Rule management and persistence

`save()` writes `{"stack_list": [...]}` as JSON. `load()` also accepts the older bare-list format ([auto_scraper.py#L53-L93](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L53-L93)). `remove_rules`, `keep_rules` and `set_rule_aliases` edit the list by `stack_id`. `generate_python_code()` still exists but only prints a deprecation notice ([auto_scraper.py#L673-L725](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L673-L725)). The usual workflow: `build`, inspect `get_result_similar(grouped=True)` to see which rule produced which values, `keep_rules` the good ones, then `save`.

### Combined call

`get_result()` parses once and returns a `(similar, exact)` tuple. It does not take `keep_order` or `contain_sibling_leaves` ([auto_scraper.py#L613-L671](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L613-L671)).

## Extending it

There are no plugin hooks. The practical extension points are:

- **Bring your own fetcher.** Every entry point accepts `html=`, and `get_result_similar`/`get_result_exact` also accept a pre-built `soup=`. Render with Playwright or fetch through your own proxy layer, then pass the HTML in.
- **`request_args`** goes straight to `requests.get`, so proxies, cookies, timeouts and extra headers work. Your `headers` are merged over the default UA.
- **Subclassing.** `_get_valid_attrs` and `request_headers` are class-level. Overriding them changes which attributes rules depend on, or the default UA.
- **Rules as data.** The JSON rule file can be edited or generated by other tools.

## Running it

`pip install autoscraper` (or from git). Its dependencies are `requests`, `bs4` and `lxml`, Python 3.6+ ([setup.py#L1-L30](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/setup.py#L1-L30)). No services, keys or browsers are needed. Tests use pytest under `tests/unit` and `tests/integration`.

## Strengths and caveats

- **Strength: tiny and dependency-light.** The whole algorithm fits in one file you can read in an hour, and rules are plain JSON.
- **Strength: no per-page cost.** After `build`, extraction is pure tree traversal. It is fast and deterministic, with no API calls.
- **Strength: list generalisation.** One example row often yields every row, because similar-mode replay fans out at every non-leaf level.
- **Caveat: structural brittleness.** Rules depend on exact tag paths, `class`/`style` values and sibling positions from the root down. A wrapper `div` added to the layout breaks every rule. `attr_fuzz_ratio` only helps with class-name drift.
- **Caveat: noisy rule sets.** Every element that matches a wanted value produces a rule. Common strings, such as a price that appears twice, create extra rules you must prune by hand.
- **Caveat: static HTML only.** No JS rendering, retries, rate limiting or rotation. The built-in UA is a 2020 Chrome string.
- **Caveat: single page per call.** No link following or crawling; you loop over URLs yourself.

*Sources: code at 68a8181, deepwiki-open wiki (16 pages), OpenDeepWiki wiki (6 pages), verified Q&A.*

## How alirezamika/autoscraper answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

Pages are fetched via plain synchronous HTTP using the `requests` library. The classmethod `_fetch_html()` at `auto_scraper.py:96-110` performs a `requests.get(url, headers=headers, **request_args)` call. A default Chrome 84 User-Agent header is set at `auto_scraper.py:45-48`. No headless browser, no JavaScript rendering engine, and no waiting strategy is implemented — if the target page relies on JS to populate content, AutoScraper will only see the raw HTML. `request_args` is forwarded as `**kwargs` to `requests.get()`, allowing users to manually pass proxies, custom headers, cookies, or timeouts. Content-type detection for encoding is minimal: if the server declares `ISO-8859-1` but the actual `Content-Type` header doesn't contain it, the apparent encoding is used instead (`auto_scraper.py:105-108`). The only content type consumed is text/HTML; PDFs, images, or binary responses are not supported at all — `res.text` is called unconditionally and passed through `str.strip()` + `unicodedata.normalize('NFKD', ...)` (`utils.py:29-32`). The fallback path `_get_soup()` at `auto_scraper.py:113-121` also accepts a raw `html` string parameter, so the caller can pre-fetch with any tool and feed the HTML directly.


Citations: [autoscraper/auto_scraper.py:96-110](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L96-L110) · [autoscraper/auto_scraper.py:113-121](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L113-L121) · [autoscraper/auto_scraper.py:45-48](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L45-L48) · [autoscraper/utils.py:29-32](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L29-L32)

### How is content extracted or converted? (answered)

Content extraction is done entirely via BeautifulSoup's tree traversal. The library does **not** convert HTML to Markdown, does **not** apply readability-style boilerplate removal, and has no schema-based extraction DSL. Its core mechanism is **rule learning by example**: given a `wanted_list` of target strings (or compiled regexps) and a URL, `build()` at `auto_scraper.py:177-258` finds every child element whose text or attribute value matches the sample (via `_child_has_text()` at `auto_scraper.py:136-168` — using exact equality by default, or fuzzy `difflib.SequenceMatcher` when `text_fuzz_ratio < 1.0`), then walks upward through the DOM to build a traversal stack of `(tag_name, attrs, child_index)` tuples (`_build_stack()` at `auto_scraper.py:261-297`). Each rule is stored as a dict with `content` (the stack), `wanted_attr` (e.g. `"href"` for links), `is_full_url`, and a SHA-256-based `stack_id`. Later, `_get_result_with_stack()` at `auto_scraper.py:330-370` replays the stack on a new page by iteratively calling `BeautifulSoup.findAll()` with the rule's tag names and attributes, descending level by level. Two retrieval modes exist: **similar** (sibling-aware, uses `findAll` broadly) via `get_result_similar()` at `auto_scraper.py:471-545`, and **exact** (index-position-based via `_get_result_with_stack_index_based()` at `auto_scraper.py:372-404`) via `get_result_exact()` at `auto_scraper.py:547-611`. Attribute matching can also be fuzzied via `FuzzyText` (`utils.py:52-59`). Duplicates are removed while preserving order via `unique_hashable()` (`utils.py:20-22`). Results can be grouped by rule ID (`grouped=True`) or by alias (`group_by_alias=True`).


Citations: [autoscraper/auto_scraper.py:177-258](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L177-L258) · [autoscraper/auto_scraper.py:136-168](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L136-L168) · [autoscraper/auto_scraper.py:261-297](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L261-L297) · [autoscraper/auto_scraper.py:330-370](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L330-L370) · [autoscraper/utils.py:20-22](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L20-L22) · [autoscraper/utils.py:52-59](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L52-L59)

### How are LLMs used, if at all? (not applicable)

AutoScraper does not use any LLM. There are no prompts, no chunking logic, no structured-output JSON schemas, no provider API keys, and no cost controls anywhere in the codebase. The project's dependencies (`setup.py:29`) are limited to `requests`, `bs4`, and `lxml` — none of which relate to language models. All extraction is purely algorithmic: pattern-matching and DOM traversal with BeautifulSoup. The fuzzy matching in `utils.py:35-40` uses Python's `difflib.SequenceMatcher`, not any learned or generative model.


Citations: [setup.py:29-29](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/setup.py#L29-L29) · [autoscraper/utils.py:35-40](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/utils.py#L35-L40)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot measures are minimal and largely absent. The only built-in fingerprinting is a default Chrome 84 User-Agent header set at `auto_scraper.py:45-48`. The `_fetch_html()` method at `auto_scraper.py:96-110` also sets the `Host` header from the URL's netloc. Beyond that, there are no stealth patches (no `TLS` fingerprint spoofing, no browser-emulation of headers beyond UA), no proxy rotation system, no CAPTCHA handling, and no rate-limiting logic — no retry-on-403, no exponential backoff, no request throttling. The code does accept a `request_args` parameter (`auto_scraper.py:97`) that is unpacked as `**kwargs` to `requests.get()`, which lets the calling user manually supply proxy dictionaries, custom headers, cookies, or timeouts. The README (`README.md:86-94`) demonstrates passing proxies via this mechanism. This is a pass-through, not a built-in rotation or management strategy.


Citations: [autoscraper/auto_scraper.py:45-48](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L45-L48) · [autoscraper/auto_scraper.py:96-110](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L96-L110) · [README.md:86-94](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/README.md#L86-L94)

### How is crawling at scale implemented? (not applicable)

AutoScraper has no crawling capability. It is a single-page scraper: each call to `build()`, `get_result_similar()`, or `get_result_exact()` fetches exactly one URL and extracts data from that page. There are no URL queues, no concurrency/threading, no URL-deduplication, no crawl-depth limits, no `robots.txt` parsing, no politeness delays, and no distributed-worker architecture. The library does not even iterate over links on a page — it has no link-extraction logic and no "follow" mechanism. The entire source is two small files focusing purely on learning extraction rules from one page and applying them to another single page.



### What is the developer interface? (answered)

AutoScraper exposes a pure Python library API only. The single entry point is the `AutoScraper` class imported from `autoscraper.auto_scraper` (`auto_scraper.py:21`). Key methods: `build()` learns rules from a URL/HTML and `wanted_list`; `get_result_similar()` applies rules broadly to extract siblings; `get_result_exact()` applies index-position-based rules for precise ordering; `get_result()` returns both; `save()`/`load()` persist/restore rules as JSON (`auto_scraper.py:53-92`); `remove_rules()`/`keep_rules()`/`set_rule_aliases()` manage the rule set (`auto_scraper.py:673-721`). There is no CLI tool, no REST service, no MCP server, and no graphical UI. Output formats are plain Python lists (or dicts when `grouped=True` or `group_by_alias=True`). There is only one language binding (Python). The `generate_python_code()` method at `auto_scraper.py:723-725` is deprecated and prints a message directing users to `save()` and `load()` instead.


Citations: [autoscraper/auto_scraper.py:21-48](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L21-L48) · [autoscraper/auto_scraper.py:53-92](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L53-L92) · [autoscraper/auto_scraper.py:673-725](https://github.com/alirezamika/autoscraper/blob/68a818158c673bf320a8569da10a8b979c1d23fe/autoscraper/auto_scraper.py#L673-L725)
