alirezamika/autoscraper
Python library that learns BeautifulSoup traversal rules from example values on one page and replays them on similar pages.
Overview
AutoScraper is a small Python library (one class, about 725 lines, plus a 59-line utils.py) that writes scraping rules for you. You give it a page and a few values you can see on that page, for example a product title and a price. It finds the elements that hold those values, records the path from the document root down to each one, and stores those paths as rules. Later you apply the rules to another page with the same layout and get back the matching values.
Despite the category it sits in, there is no machine learning and no LLM here. “Learning” means a DOM walk plus string comparison, with optional fuzzy matching from difflib.SequenceMatcher. Pages are fetched with a plain requests.get, so JavaScript-rendered content is invisible unless you pass in rendered HTML yourself. There is no crawler, CLI or service: it is a library for people who want a quick extractor for a list or detail page without writing CSS selectors.
The interesting part is the rule format. A rule is a JSON-serialisable dict: a list of (tag, {class, style}, sibling_index) steps plus a few flags. This makes rules easy to save, diff, prune and share.
Architecture
flowchart LR
U["Caller: url or html + wanted_list"] --> S["_get_soup"]
S --> F["_fetch_html (requests.get)"]
S --> BS["BeautifulSoup lxml tree"]
BS --> GC["_get_children / _child_has_text"]
GC --> BST["_build_stack"]
BST --> SL["stack_list (rules)"]
SL --> J["save / load JSON"]
SL --> SIM["_get_result_with_stack (similar)"]
SL --> EX["_get_result_with_stack_index_based (exact)"]
SIM --> CR["_clean_result"]
EX --> CR
CR --> R["list or dict of strings"]
| Component | Path | Role |
|---|---|---|
AutoScraper class |
autoscraper/auto_scraper.py |
Public API: build, get_result_similar, get_result_exact, get_result, rule management, save/load |
| Fetch and parse | auto_scraper.py _fetch_html, _get_soup |
One requests.get, encoding fix-up, NFKD normalisation, BeautifulSoup(html, "lxml") |
| Matcher | auto_scraper.py _child_has_text, _get_children |
Finds elements whose text, direct text or attribute equals a wanted value |
| Rule builder | auto_scraper.py _build_stack |
Walks from a matched element up to the root and records the path |
| Rule replay | auto_scraper.py _get_result_with_stack, _get_result_with_stack_index_based |
Re-walks a path on a new tree, broadly or by exact position |
| Utilities | autoscraper/utils.py |
text_match, FuzzyText, normalize, ResultItem, order-preserving dedupe |
How a request flows
Learning rules with build():
build(url=..., wanted_list=[...])checks that some target was given, then calls_get_soup, which either parses thehtmlyou passed or calls_fetch_html. That method does onerequests.getwith a hard-coded Chrome 84 User-Agent, aHostheader taken from the URL, and anyrequest_argsyou supplied; it fixes a wrongly declared ISO-8859-1 encoding (auto_scraper.py#L95-L121).- A plain
wanted_listbecomeswanted_dict = {"": wanted_list}, so aliased and unaliased targets share one loop. Unlessupdate=True, existing rules are dropped (auto_scraper.py#L226-L258). - For each wanted value,
_get_childrenscans every element in reverse document order._child_has_textaccepts an element if its full text matches (but not when its parent has the same text, so the outermost element with exactly that text wins), if its own non-recursive text matches, or if an attribute value matches. Forhref/srcit also tries the URL resolved against the page URL (auto_scraper.py#L135-L175). _build_stackclimbs from the match to the root. At each level it stores the parent tag, itsclassandstyleonly, and the index of the child among siblings with the same tag and attributes. The rule gets a SHA-256hashand astack_idofrule_plus 8 hex chars (auto_scraper.py#L260-L297).- Each new rule is replayed at once on the same tree with the “similar” strategy, and those values are what
build()returns. Rules are deduplicated by hash withunique_stack_list.
Applying rules with get_result_similar() / get_result_exact():
- Both go through
_get_result_by_func, which parses the page once, optionally tags every element with achild_index(needed forkeep_orderandgroup_by_alias), and runs the chosen replay function for each rule (auto_scraper.py#L406-L469). - Similar replay keeps every matching child at every level (
findAll(tag, attrs, recursive=False)), which is how one example item generalises to all rows of a list. Only at the leaf does it pick by stored index, unlesscontain_sibling_leaves=True(auto_scraper.py#L330-L370). - Exact replay follows the stored sibling index at every level and returns at most one value per rule (auto_scraper.py#L372-L404).
_clean_resultflattens, sorts by DOM position if asked, and removes duplicates. Or it returns a dict keyed bystack_id(grouped=True) or by alias (group_by_alias=True).
Key components
Matching
text_match in utils.py#L35-L40 is the only comparison primitive. A compiled regex is tested with fullmatch. A ratio of 1.0 (the default) means exact string equality. Anything lower uses SequenceMatcher.ratio(). That is text_fuzz_ratio at build time. At replay time, attr_fuzz_ratio < 1.0 wraps the stored class/style values in FuzzyText, whose search() method BeautifulSoup calls as a matcher (utils.py#L52-L59). Fuzzy matching is O(elements x targets) SequenceMatcher calls, so it gets slow on large pages.
Rule format
_get_valid_attrs keeps only class and style, and fills in empty strings when they are missing (auto_scraper.py#L123-L133). id, data-* and other attributes never appear in a rule. Besides the path, a rule stores wanted_attr (which attribute to read, or None for text), is_full_url, is_non_rec_text, alias, and the page url when the value was a resolved link.
Rule management and persistence
save() writes {"stack_list": [...]} as JSON. load() also accepts the older bare-list format (auto_scraper.py#L53-L93). remove_rules, keep_rules and set_rule_aliases edit the list by stack_id. generate_python_code() still exists but only prints a deprecation notice (auto_scraper.py#L673-L725). The usual workflow: build, inspect get_result_similar(grouped=True) to see which rule produced which values, keep_rules the good ones, then save.
Combined call
get_result() parses once and returns a (similar, exact) tuple. It does not take keep_order or contain_sibling_leaves (auto_scraper.py#L613-L671).
Extending it
There are no plugin hooks. The practical extension points are:
- Bring your own fetcher. Every entry point accepts
html=, andget_result_similar/get_result_exactalso accept a pre-builtsoup=. Render with Playwright or fetch through your own proxy layer, then pass the HTML in. request_argsgoes straight torequests.get, so proxies, cookies, timeouts and extra headers work. Yourheadersare merged over the default UA.- Subclassing.
_get_valid_attrsandrequest_headersare class-level. Overriding them changes which attributes rules depend on, or the default UA. - Rules as data. The JSON rule file can be edited or generated by other tools.
Running it
pip install autoscraper (or from git). Its dependencies are requests, bs4 and lxml, Python 3.6+ (setup.py#L1-L30). No services, keys or browsers are needed. Tests use pytest under tests/unit and tests/integration.
Strengths and caveats
- Strength: tiny and dependency-light. The whole algorithm fits in one file you can read in an hour, and rules are plain JSON.
- Strength: no per-page cost. After
build, extraction is pure tree traversal. It is fast and deterministic, with no API calls. - Strength: list generalisation. One example row often yields every row, because similar-mode replay fans out at every non-leaf level.
- Caveat: structural brittleness. Rules depend on exact tag paths,
class/stylevalues and sibling positions from the root down. A wrapperdivadded to the layout breaks every rule.attr_fuzz_ratioonly helps with class-name drift. - Caveat: noisy rule sets. Every element that matches a wanted value produces a rule. Common strings, such as a price that appears twice, create extra rules you must prune by hand.
- Caveat: static HTML only. No JS rendering, retries, rate limiting or rotation. The built-in UA is a 2020 Chrome string.
- Caveat: single page per call. No link following or crawling; you loop over URLs yourself.
Sources: code at 68a8181, deepwiki-open wiki (16 pages), OpenDeepWiki wiki (6 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 68a8181. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPages are fetched via plain synchronous HTTP using the requests library. The classmethod _fetch_html() at auto_scraper.py:96-110 performs a requests.get(url, headers=headers, **request_args) call. A default Chrome 84 User-Agent header is set at auto_scraper.py:45-48. No headless browser, no JavaScript rendering engine, and no waiting strategy is implemented — if the target page relies on JS to populate content, AutoScraper will only see the raw HTML. request_args is forwarded as **kwargs to requests.get(), allowing users to manually pass proxies, custom headers, cookies, or timeouts. Content-type detection for encoding is minimal: if the server declares ISO-8859-1 but the actual Content-Type header doesn't contain it, the apparent encoding is used instead (auto_scraper.py:105-108). The only content type consumed is text/HTML; PDFs, images, or binary responses are not supported at all — res.text is called unconditionally and passed through str.strip() + unicodedata.normalize('NFKD', ...) (utils.py:29-32). The fallback path _get_soup() at auto_scraper.py:113-121 also accepts a raw html string parameter, so the caller can pre-fetch with any tool and feed the HTML directly.
How is content extracted or converted?
answeredContent extraction is done entirely via BeautifulSoup's tree traversal. The library does not convert HTML to Markdown, does not apply readability-style boilerplate removal, and has no schema-based extraction DSL. Its core mechanism is rule learning by example: given a wanted_list of target strings (or compiled regexps) and a URL, build() at auto_scraper.py:177-258 finds every child element whose text or attribute value matches the sample (via _child_has_text() at auto_scraper.py:136-168 — using exact equality by default, or fuzzy difflib.SequenceMatcher when text_fuzz_ratio < 1.0), then walks upward through the DOM to build a traversal stack of (tag_name, attrs, child_index) tuples (_build_stack() at auto_scraper.py:261-297). Each rule is stored as a dict with content (the stack), wanted_attr (e.g. "href" for links), is_full_url, and a SHA-256-based stack_id. Later, _get_result_with_stack() at auto_scraper.py:330-370 replays the stack on a new page by iteratively calling BeautifulSoup.findAll() with the rule's tag names and attributes, descending level by level. Two retrieval modes exist: similar (sibling-aware, uses findAll broadly) via get_result_similar() at auto_scraper.py:471-545, and exact (index-position-based via _get_result_with_stack_index_based() at auto_scraper.py:372-404) via get_result_exact() at auto_scraper.py:547-611. Attribute matching can also be fuzzied via FuzzyText (utils.py:52-59). Duplicates are removed while preserving order via unique_hashable() (utils.py:20-22). Results can be grouped by rule ID (grouped=True) or by alias (group_by_alias=True).
How are LLMs used, if at all?
not applicableAutoScraper does not use any LLM. There are no prompts, no chunking logic, no structured-output JSON schemas, no provider API keys, and no cost controls anywhere in the codebase. The project's dependencies (setup.py:29) are limited to requests, bs4, and lxml — none of which relate to language models. All extraction is purely algorithmic: pattern-matching and DOM traversal with BeautifulSoup. The fuzzy matching in utils.py:35-40 uses Python's difflib.SequenceMatcher, not any learned or generative model.
How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-bot measures are minimal and largely absent. The only built-in fingerprinting is a default Chrome 84 User-Agent header set at auto_scraper.py:45-48. The _fetch_html() method at auto_scraper.py:96-110 also sets the Host header from the URL's netloc. Beyond that, there are no stealth patches (no TLS fingerprint spoofing, no browser-emulation of headers beyond UA), no proxy rotation system, no CAPTCHA handling, and no rate-limiting logic — no retry-on-403, no exponential backoff, no request throttling. The code does accept a request_args parameter (auto_scraper.py:97) that is unpacked as **kwargs to requests.get(), which lets the calling user manually supply proxy dictionaries, custom headers, cookies, or timeouts. The README (README.md:86-94) demonstrates passing proxies via this mechanism. This is a pass-through, not a built-in rotation or management strategy.
How is crawling at scale implemented?
not applicableAutoScraper has no crawling capability. It is a single-page scraper: each call to build(), get_result_similar(), or get_result_exact() fetches exactly one URL and extracts data from that page. There are no URL queues, no concurrency/threading, no URL-deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker architecture. The library does not even iterate over links on a page — it has no link-extraction logic and no "follow" mechanism. The entire source is two small files focusing purely on learning extraction rules from one page and applying them to another single page.
What is the developer interface?
answeredAutoScraper exposes a pure Python library API only. The single entry point is the AutoScraper class imported from autoscraper.auto_scraper (auto_scraper.py:21). Key methods: build() learns rules from a URL/HTML and wanted_list; get_result_similar() applies rules broadly to extract siblings; get_result_exact() applies index-position-based rules for precise ordering; get_result() returns both; save()/load() persist/restore rules as JSON (auto_scraper.py:53-92); remove_rules()/keep_rules()/set_rule_aliases() manage the rule set (auto_scraper.py:673-721). There is no CLI tool, no REST service, no MCP server, and no graphical UI. Output formats are plain Python lists (or dicts when grouped=True or group_by_alias=True). There is only one language binding (Python). The generate_python_code() method at auto_scraper.py:723-725 is deprecated and prints a message directing users to save() and load() instead.