LLMs Technical Reviews
Home / AI web scraping / trafilatura

adbar/trafilatura

Rule-based Python library and CLI that fetches pages over HTTP and extracts main text, comments and metadata without a browser or ML.

GitHub ↗★ 6.9kPythonApache-2.0commit 6c1977a · 2026-10-02homepage ↗

Overview

Trafilatura is a Python library and CLI that downloads web pages over plain HTTP and extracts the main text, comments and metadata with deterministic rules. It does not use a browser, a model, or an LLM. Inside it parses HTML with lxml, converts it to a small XML vocabulary (p, head, list, quote, code, table, graphic, ref), runs a cascade of extractors, and serialises the result as TXT, Markdown, CSV, JSON, HTML, XML or XML-TEI.

It is the quiet workhorse of this category. Corpus builders, RAG ingestion jobs and research crawlers use it because it is fast, needs no GPU or API key, and copes with a long tail of templates through accumulated heuristics (the pinned version is 2.3.0). The cost of the rule-based design is that it only sees what the server sends: JavaScript-rendered pages, PDFs and paywalled content are out of scope.

Beyond extraction, it ships its own discovery layer: sitemap and feed parsers, a polite focused crawler built on courlan’s UrlStore, and a CLI that downloads many URLs in parallel with per-domain back-off. In this category it is a pre-processor for LLM pipelines, not an AI scraper itself.

Architecture

flowchart TD
  IN["URL, file or HTML string"] --> DL["downloads.fetch_url (urllib3 or pycurl)"]
  DL --> LOAD["utils.load_html (lxml)"]
  IN --> LOAD
  LOAD --> META["metadata.extract_metadata"]
  LOAD --> SEQ["core.trafilatura_sequence"]
  SEQ --> CLEAN["htmlprocessing: tree_cleaning + convert_tags"]
  CLEAN --> MAIN["main_extractor.extract_content"]
  MAIN --> CMP["external.compare_extraction (readability fork, jusText)"]
  CMP --> BASE["baseline rescue"]
  BASE --> ESC["recall escalation"]
  ESC --> OUT["determine_returnstring: txt, md, json, csv, html, xml, tei"]
  META --> OUT
  DISC["sitemaps, feeds, spider"] --> DL
Component Path Role
Public API trafilatura/__init__.py, core.py fetch_url, extract, extract_with_metadata, bare_extraction, the extraction cascade and output dispatch
Options trafilatura/settings.py, settings.cfg Extractor options object, Document result, config-file defaults
Cleaning and conversion trafilatura/htmlprocessing.py Tag removal, link-density pruning, HTML to internal XML
Main extractor trafilatura/main_extractor.py Content-area XPaths, per-element handlers, wild-text recovery, comments
XPath rules trafilatura/xpaths.py BODY_XPATH, discard lists for sidebars, teasers, comments and ads
Fallbacks trafilatura/external.py, readability_lxml.py, baseline.py Vendored readability fork, jusText, and a last-resort baseline
Metadata trafilatura/metadata.py, json_metadata.py Meta tags, JSON-LD, title, author, date via htmldate
Downloads trafilatura/downloads.py Pooled HTTP with SSRF guard, retries, size cap, optional SOCKS proxy
Discovery trafilatura/spider.py, sitemaps.py, feeds.py Focused crawler, sitemap and feed URL discovery
CLI trafilatura/cli.py, cli_utils.py Batch processing, parallel downloads, crawl and explore modes

How a request flows

Take extract(fetch_url("https://example.org/post"), with_metadata=True, output_format="markdown"):

  1. Download. fetch_response picks pycurl when it is installed and urllib3 otherwise. On an SSL error it retries without certificate checks unless INSECURE_SSL_FALLBACK is off (downloads.py). The urllib3 path streams the body in 128 KiB chunks capped at MAX_FILE_SIZE, with retry and redirect limits from the config (downloads.py). fetch_url returns decoded HTML only for a 200 of acceptable length (downloads.py).
  2. Options and parse. bare_extraction packs the keyword arguments into an Extractor, parses the HTML, optionally checks the <html lang> attribute, and runs extract_metadata (core.py). Metadata comes from meta tags, then JSON-LD, then title and author heuristics, with dates from htmldate.find_date (metadata.py).
  3. Prepare. trafilatura_sequence prunes appended articles, share widgets and (unless comments are wanted) comment sections, then tree_cleaning deletes unwanted tags and convert_tags maps HTML to the internal vocabulary (core.py, htmlprocessing.py).
  4. Main extractor. _extract tries each BODY_XPATH expression in order, from itemprop='articleBody'-style matches through <article> and story/content classes to main. For each match it prunes discard XPaths and link-heavy blocks, converts children through handle_textelem, and stops at the first subtree with real content (main_extractor.py, xpaths.py). If the result is short, recover_wild_text scavenges loose paragraphs.
  5. Compare. Unless fast=True, compare_extraction runs the vendored readability on the raw tree and switches to it when the own result is empty, much shorter, or structurally poor; jusText is then tried for unclean or short outputs (external.py).
  6. Rescue and escalate. If text is still under MIN_EXTRACTED_SIZE (250 chars), baseline tries JSON-LD bodies, <article> blocks, paragraphs, then the whole body. In balanced mode, a result under 3,000 characters that also covers under 20% of the page text triggers a recall-mode retry plus a jusText candidate (core.py).
  7. Filter and serialise. Size, duplicate and language checks can discard the document (the function returns None). determine_returnstring then writes Markdown with a YAML front-matter header built from the metadata (core.py).

Key components

The extraction cascade

The design idea is “each stage engages only if the previous one under-delivered” (core.py). The own extractor is precise but template-bound; readability is good on classic articles; jusText classifies paragraphs by density and stopwords and reaches content buried in generic divs; baseline is a blunt safety net. favor_precision and favor_recall shift each stage’s thresholds rather than choosing a different algorithm. There is special handling for forum threads (detected by schema.org DiscussionForumPosting), where comment containers are kept as content.

Cleaning and the internal XML

tree_cleaning removes a fixed tag list, keeps or drops tables and images depending on options, and in recall mode undoes the deletion if it would remove every paragraph (htmlprocessing.py). The XML intermediate is what makes the output formats cheap: Markdown, HTML and TEI are just different walks over the same tree.

Downloads and politeness

The default user agent is trafilatura/<version> with a project URL; USER_AGENTS and COOKIE in the config override it, picking one agent at random per request (downloads.py). By default a _SafePoolManager rejects private, loopback and link-local addresses on every hop. If http_proxy is set, a urllib3 SOCKSProxyManager is used instead, and that path skips the SSRF pool (downloads.py). Defaults: 30 s timeout, 20 MB cap, 2 redirects, 5 s between requests to one domain (settings.cfg).

Focused crawler

focused_crawler(homepage) reads robots.txt, keeps links inside the start URL’s path prefix, puts navigation pages at the front of the queue, sleeps for the robots crawl delay or SLEEP_TIME between fetches, and stops after max_seen_urls (default 10) pages (spider.py, L308-L351). It returns two URL lists, to visit and known; it does not extract text. The frontier is a module-level URL_STORE, so state persists across calls in one process.

CLI

trafilatura accepts a URL, an input file, a directory or stdin, plus --feed, --sitemap, --crawl, --explore and --probe. URL lists go through buffered_downloads, a thread pool fed by UrlStore.get_download_urls with per-domain back-off (downloads.py). --archived retries failed URLs through the Wayback Machine (cli_utils.py).

Extending it

  • Tune without code. prune_xpath removes site-specific noise before extraction; --config-file or a ConfigParser changes size thresholds, timeouts, user agents and SSRF behaviour.
  • Get structure, not text. bare_extraction returns a Document with the lxml body tree, so you can walk headings, lists and tables yourself before chunking for a RAG index.
  • Custom rules. MANUALLY_CLEANED and the XPath lists in xpaths.py are module-level and can be mutated at runtime; the code explicitly honours a user who removes "form" from the cleaning list.
  • JavaScript pages. Render with a browser of your choice, then pass the HTML string to extract. The library accepts strings, bytes or lxml trees.

Running it

  • Install. pip install trafilatura pulls only lxml, urllib3, courlan, htmldate, justext, charset_normalizer and certifi. trafilatura[all] adds pycurl, SOCKS support, py3langid language detection, Brotli/zstd and faster charset detection.
  • Python. extract(fetch_url(url)) is the one-liner; extract_with_metadata returns a Document with text and metadata fields (init.py).
  • CLI. trafilatura -u URL --output-format markdown --with-metadata, or trafilatura -i urls.txt -o out/ --parallel 8 for batches.

Strengths and caveats

  • Strength: fast, local and cheap. No browser, no model and no network calls beyond the fetch, so it runs anywhere Python and lxml run.
  • Strength: layered fallbacks. Four extraction strategies, gated by measured text length, give good recall on odd templates without hurting typical articles.
  • Strength: rich output. Metadata, comments, tables, formatting, links and images are optional, and TEI and YAML-headed Markdown are first-class.
  • Caveat: static HTML only. There is no JavaScript execution and no wait strategy; client-rendered sites return little or nothing.
  • Caveat: heuristics can misfire. Content-area detection relies on class and id patterns. A page that matches the wrong container yields confident but partial text, and the result silently becomes None when filters reject it.
  • Caveat: transparent by default. The bot-identifying user agent, robots-aware crawler and lack of stealth are good citizenship, but mean protected sites will block it. The default insecure-SSL retry is convenient but worth turning off for sensitive pipelines.
  • Caveat: crawler scope. The spider is a small single-site link discoverer with in-memory state, not a distributed crawler.

Sources: code at 6c1977a, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 6c1977a. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Plain HTTP(S) only, no headless browser. Pages are fetched via one of two backends:

  • urllib3 (default): _send_urllib_request in downloads.py:217 uses a urllib3.PoolManager with configurable retry strategy and decompression. The response body is streamed in 128KB chunks and capped at MAX_FILE_SIZE (20MB default) via _capped().
  • pycurl (optional, imported at downloads.py:42): _send_pycurl_request at downloads.py:477 provides the same functionality via libcurl with shared DNS/SSL session caches across requests. Both backends support SOCKS proxy via the http_proxy env var and enforce an SSRF protection layer (_SafePoolManager) that rejects non-global IP addresses.

No JavaScript rendering. The tool does not use a headless browser, Selenium, or Playwright. It receives raw HTML only. There is no waiting strategy for dynamic content beyond HTTP-level retries on transient status codes (429, 5xx). Both backends follow redirects (configurable via MAX_REDIRECTS). Headers are stored from the response; the User-Agent string defaults to trafilatura/<version> or rotates through user-supplied agents from the config file (downloads.py:173).

Content types: The system accepts HTML and any text-based response. There is no built-in support for downloading or extracting PDFs, images, or binary files; IMAGE_EXTENSION at utils.py:123 is used only to detect whether a URL points to an image, which is then filtered rather than rendered.

How is content extracted or converted?

answered

Rule-based cascading extractor with no ML component. The core pipeline lives in core.py:trafilatura_sequence() (line 165) and runs five stages:

  1. Tag conversion (htmlprocessing.py:408): HTML elements are converted to a simplified XML vocabulary — headings become <head>, lists become <list>/<item>, code blocks become <code>, quotes become <quote>, inline formatting becomes <hi>, <br> → <lb>, <img> → <graphic>, etc. XPath-based content-area detection via BODY_XPATH (xpaths.py:62) locates the main content frame by matching common article ID/class patterns (entry-content, post-body, articleBody).

  2. Boilerplate removal (htmlprocessing.py:81 tree_cleaning): Unwanted elements (nav, footer, iframe, script, aside, form, etc.) are deleted. XPath expressions (OVERALL_DISCARD_XPATH in xpaths.py:274) prune sidebar, breadcrumb, share-button, paywall, ad, and cookie-consent sections by ID/class heuristics. Link-density tests (htmlprocessing.py:177) remove sections rich in links and low in text.

  3. Main extractor (main_extractor.py:686 extract_content): Iterates BODY_XPATH expressions, extracts matching subtrees, and processes child elements via handle_textelem() which dispatches to paragraph/list/table/code/quote handlers.

  4. Cascade: If short extraction → recover_wild_text() scavenges orphan <p>/<code>/<div> elements. Then compare_extraction() (external.py:84) evaluates readability-lxml and justext as fallbacks. A baseline() rescue (baseline.py:167) tries JSON-LD embedded content, <article> tags, paragraphs, and finally the full body. A recall escalation retries in high-recall mode if output covers <20% of page length.

  5. Output conversion (core.py:80 determine_returnstring): The XML body is serialized to TXT/Markdown (with YAML metadata header), CSV, JSON, HTML, XML, or XML-TEI. Metadata (title, author, date, categories, tags) is extracted via metadata.py:1 using XPath selectors and htmldate.

How are LLMs used, if at all?

not applicable

This repository does not use LLMs at all. A full case-insensitive grep across the entire source for openai, anthropic, claude, gpt, gemini, langchain, llm, and huggingface returned zero matches. The project relies entirely on deterministic rule-based extraction: XPath heuristics, tag conversion, link-density analysis, and three classical algorithmic extractors (readability-lxml, jusText, and its own cascading rule engine). No prompting, chunking, structured output generation via LLM, provider integration, or cost controls exist.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Minimal anti-bot evasion. The project does not implement stealth patches, CAPTCHA handling, or browser fingerprint spoofing. Its anti-detection measures are:

  • User-agent rotation (downloads.py:169-174): The _determine_headers() function accepts a multi-line USER_AGENTS config value, picks one at random per request. The default agent identifies itself as trafilatura/<version> (transparent, non-stealthy).
  • SOCKS proxy (downloads.py:37): If http_proxy env var is set, all requests route through a SOCKSProxyManager (urllib3) or PRE_PROXY (pycurl). No built-in proxy rotation or proxy list management.
  • Rate limiting: The crawler sleeps SLEEP_TIME seconds (default 5.0, configurable in settings.cfg:11) between requests, derived from the domain's crawl delay (spider.py:339).
  • robots.txt (spider.py:152-170): The crawler fetches and parses robots.txt via urllib.robotparser.RobotFileParser and honours can_fetch() rules before visiting links.
  • SSRF protection (downloads.py:118-133): By default, connections to non-public IP addresses (loopback, private, link-local) are blocked via a custom _SafePoolManager with per-connection peer vetting.

There is no proxy rotation, no CAPTCHA-solving integration, no TLS fingerprint spoofing, and no referer/header randomization apart from the user-agent draw.

How is crawling at scale implemented?

answered

A focused, single-domain crawler with URL management via the courlan library. The primary entry point is focused_crawler() in spider.py:308.

  • URL queue & dedup: The UrlStore class (from courlan) manages a priority queue of discovered URLs with domain-aware deduplication. URLs are tracked in three states: known, visited, and unvisited. Navigation pages (indices, archives) get priority placement in the queue (spider.py:224-229).
  • Limits: max_seen_urls defaults to 10; max_known_urls defaults to 100,000 (spider.py:39). The primary loop at spider.py:342 stops when either limit is reached or the domain URL store is exhausted.
  • robots.txt & politeness: Each domain's robots.txt is fetched and parsed on init (spider.py:60). is_valid_link() (spider.py:88-90) checks rules.can_fetch("*", link), that the link stays within the reference domain, and that the URL is crawlable. A configurable SLEEP_TIME (5s default) separates requests (spider.py:339).
  • Discovery modes: URLs can be seeded from sitemaps (sitemaps.py supports XML sitemaps, sitemap indexes, and robots.txt sitemap directives) and feeds (feeds.py supports ATOM, JSON, RSS). The --explore CLI option combines both.
  • Concurrency: Crawling itself is single-threaded per domain. For batch URL processing, buffered_downloads() in downloads.py:438 uses a ThreadPoolExecutor (default up to 16 threads) to parallelize fetches. No distributed worker architecture exists.

There is no support for multiple concurrent crawl workers, distributed crawling, or queue persistence across restarts.

What is the developer interface?

answered

Library API, CLI — no REST service, MCP server, or web UI.

Python API (exposed in __init__.py:15-32):

  • fetch_url(url) → downloads a page and returns a decoded string.
  • extract(html, ...) → the main extraction function, returns a string in the chosen format.
  • extract_with_metadata(...) → returns a Document object containing both text and metadata.
  • bare_extraction(...) → returns a Document with the lxml body tree (unsafe for serialization).
  • fetch_response(url) → returns the raw Response object.
  • All functions accept 20+ keyword arguments for fine-grained control (output format, language filtering, deduplication, precision/recall, formatting, images, links, tables, comments, pruning XPath, custom config).

CLI (cli.py:49-168): Invoked as trafilatura (registered in pyproject.toml:83). Input options: --input-file, --input-dir, --URL, stdin. Navigation options: --feed, --sitemap, --crawl, --explore, --probe. Extraction options: --fast, --precision, --recall, --with-metadata, --target-language, --deduplicate. Parallel processing via --parallel.

Output formats (settings.py:29): txt (plain text with optional YAML metadata header), markdown, csv, json, html, xml, xmltei (TEI-conformant XML).

Configuration: A settings.cfg file (settings.py:41-49) controls defaults for download timeout, file sizes, sleep time, user agents, SSRF protection, date search, and extraction thresholds. Users can supply overrides via --config-file or a ConfigParser object in the API.

There is no REST web service, no MCP server integration, no graphical user interface. Language bindings are Python-only.