adbar/trafilatura
Rule-based Python library and CLI that fetches pages over HTTP and extracts main text, comments and metadata without a browser or ML.
Overview
Trafilatura is a Python library and CLI that downloads web pages over plain HTTP and extracts the main text, comments and metadata with deterministic rules. It does not use a browser, a model, or an LLM. Inside it parses HTML with lxml, converts it to a small XML vocabulary (p, head, list, quote, code, table, graphic, ref), runs a cascade of extractors, and serialises the result as TXT, Markdown, CSV, JSON, HTML, XML or XML-TEI.
It is the quiet workhorse of this category. Corpus builders, RAG ingestion jobs and research crawlers use it because it is fast, needs no GPU or API key, and copes with a long tail of templates through accumulated heuristics (the pinned version is 2.3.0). The cost of the rule-based design is that it only sees what the server sends: JavaScript-rendered pages, PDFs and paywalled content are out of scope.
Beyond extraction, it ships its own discovery layer: sitemap and feed parsers, a polite focused crawler built on courlan’s UrlStore, and a CLI that downloads many URLs in parallel with per-domain back-off. In this category it is a pre-processor for LLM pipelines, not an AI scraper itself.
Architecture
flowchart TD
IN["URL, file or HTML string"] --> DL["downloads.fetch_url (urllib3 or pycurl)"]
DL --> LOAD["utils.load_html (lxml)"]
IN --> LOAD
LOAD --> META["metadata.extract_metadata"]
LOAD --> SEQ["core.trafilatura_sequence"]
SEQ --> CLEAN["htmlprocessing: tree_cleaning + convert_tags"]
CLEAN --> MAIN["main_extractor.extract_content"]
MAIN --> CMP["external.compare_extraction (readability fork, jusText)"]
CMP --> BASE["baseline rescue"]
BASE --> ESC["recall escalation"]
ESC --> OUT["determine_returnstring: txt, md, json, csv, html, xml, tei"]
META --> OUT
DISC["sitemaps, feeds, spider"] --> DL
| Component | Path | Role |
|---|---|---|
| Public API | trafilatura/__init__.py, core.py |
fetch_url, extract, extract_with_metadata, bare_extraction, the extraction cascade and output dispatch |
| Options | trafilatura/settings.py, settings.cfg |
Extractor options object, Document result, config-file defaults |
| Cleaning and conversion | trafilatura/htmlprocessing.py |
Tag removal, link-density pruning, HTML to internal XML |
| Main extractor | trafilatura/main_extractor.py |
Content-area XPaths, per-element handlers, wild-text recovery, comments |
| XPath rules | trafilatura/xpaths.py |
BODY_XPATH, discard lists for sidebars, teasers, comments and ads |
| Fallbacks | trafilatura/external.py, readability_lxml.py, baseline.py |
Vendored readability fork, jusText, and a last-resort baseline |
| Metadata | trafilatura/metadata.py, json_metadata.py |
Meta tags, JSON-LD, title, author, date via htmldate |
| Downloads | trafilatura/downloads.py |
Pooled HTTP with SSRF guard, retries, size cap, optional SOCKS proxy |
| Discovery | trafilatura/spider.py, sitemaps.py, feeds.py |
Focused crawler, sitemap and feed URL discovery |
| CLI | trafilatura/cli.py, cli_utils.py |
Batch processing, parallel downloads, crawl and explore modes |
How a request flows
Take extract(fetch_url("https://example.org/post"), with_metadata=True, output_format="markdown"):
- Download.
fetch_responsepicks pycurl when it is installed and urllib3 otherwise. On an SSL error it retries without certificate checks unlessINSECURE_SSL_FALLBACKis off (downloads.py). The urllib3 path streams the body in 128 KiB chunks capped atMAX_FILE_SIZE, with retry and redirect limits from the config (downloads.py).fetch_urlreturns decoded HTML only for a 200 of acceptable length (downloads.py). - Options and parse.
bare_extractionpacks the keyword arguments into anExtractor, parses the HTML, optionally checks the<html lang>attribute, and runsextract_metadata(core.py). Metadata comes from meta tags, then JSON-LD, then title and author heuristics, with dates fromhtmldate.find_date(metadata.py). - Prepare.
trafilatura_sequenceprunes appended articles, share widgets and (unless comments are wanted) comment sections, thentree_cleaningdeletes unwanted tags andconvert_tagsmaps HTML to the internal vocabulary (core.py, htmlprocessing.py). - Main extractor.
_extracttries eachBODY_XPATHexpression in order, fromitemprop='articleBody'-style matches through<article>and story/content classes tomain. For each match it prunes discard XPaths and link-heavy blocks, converts children throughhandle_textelem, and stops at the first subtree with real content (main_extractor.py, xpaths.py). If the result is short,recover_wild_textscavenges loose paragraphs. - Compare. Unless
fast=True,compare_extractionruns the vendored readability on the raw tree and switches to it when the own result is empty, much shorter, or structurally poor; jusText is then tried for unclean or short outputs (external.py). - Rescue and escalate. If text is still under
MIN_EXTRACTED_SIZE(250 chars),baselinetries JSON-LD bodies,<article>blocks, paragraphs, then the whole body. In balanced mode, a result under 3,000 characters that also covers under 20% of the page text triggers a recall-mode retry plus a jusText candidate (core.py). - Filter and serialise. Size, duplicate and language checks can discard the document (the function returns
None).determine_returnstringthen writes Markdown with a YAML front-matter header built from the metadata (core.py).
Key components
The extraction cascade
The design idea is “each stage engages only if the previous one under-delivered” (core.py). The own extractor is precise but template-bound; readability is good on classic articles; jusText classifies paragraphs by density and stopwords and reaches content buried in generic divs; baseline is a blunt safety net. favor_precision and favor_recall shift each stage’s thresholds rather than choosing a different algorithm. There is special handling for forum threads (detected by schema.org DiscussionForumPosting), where comment containers are kept as content.
Cleaning and the internal XML
tree_cleaning removes a fixed tag list, keeps or drops tables and images depending on options, and in recall mode undoes the deletion if it would remove every paragraph (htmlprocessing.py). The XML intermediate is what makes the output formats cheap: Markdown, HTML and TEI are just different walks over the same tree.
Downloads and politeness
The default user agent is trafilatura/<version> with a project URL; USER_AGENTS and COOKIE in the config override it, picking one agent at random per request (downloads.py). By default a _SafePoolManager rejects private, loopback and link-local addresses on every hop. If http_proxy is set, a urllib3 SOCKSProxyManager is used instead, and that path skips the SSRF pool (downloads.py). Defaults: 30 s timeout, 20 MB cap, 2 redirects, 5 s between requests to one domain (settings.cfg).
Focused crawler
focused_crawler(homepage) reads robots.txt, keeps links inside the start URL’s path prefix, puts navigation pages at the front of the queue, sleeps for the robots crawl delay or SLEEP_TIME between fetches, and stops after max_seen_urls (default 10) pages (spider.py, L308-L351). It returns two URL lists, to visit and known; it does not extract text. The frontier is a module-level URL_STORE, so state persists across calls in one process.
CLI
trafilatura accepts a URL, an input file, a directory or stdin, plus --feed, --sitemap, --crawl, --explore and --probe. URL lists go through buffered_downloads, a thread pool fed by UrlStore.get_download_urls with per-domain back-off (downloads.py). --archived retries failed URLs through the Wayback Machine (cli_utils.py).
Extending it
- Tune without code.
prune_xpathremoves site-specific noise before extraction;--config-fileor aConfigParserchanges size thresholds, timeouts, user agents and SSRF behaviour. - Get structure, not text.
bare_extractionreturns aDocumentwith the lxmlbodytree, so you can walk headings, lists and tables yourself before chunking for a RAG index. - Custom rules.
MANUALLY_CLEANEDand the XPath lists inxpaths.pyare module-level and can be mutated at runtime; the code explicitly honours a user who removes"form"from the cleaning list. - JavaScript pages. Render with a browser of your choice, then pass the HTML string to
extract. The library accepts strings, bytes or lxml trees.
Running it
- Install.
pip install trafilaturapulls onlylxml,urllib3,courlan,htmldate,justext,charset_normalizerandcertifi.trafilatura[all]adds pycurl, SOCKS support,py3langidlanguage detection, Brotli/zstd and faster charset detection. - Python.
extract(fetch_url(url))is the one-liner;extract_with_metadatareturns aDocumentwith text and metadata fields (init.py). - CLI.
trafilatura -u URL --output-format markdown --with-metadata, ortrafilatura -i urls.txt -o out/ --parallel 8for batches.
Strengths and caveats
- Strength: fast, local and cheap. No browser, no model and no network calls beyond the fetch, so it runs anywhere Python and lxml run.
- Strength: layered fallbacks. Four extraction strategies, gated by measured text length, give good recall on odd templates without hurting typical articles.
- Strength: rich output. Metadata, comments, tables, formatting, links and images are optional, and TEI and YAML-headed Markdown are first-class.
- Caveat: static HTML only. There is no JavaScript execution and no wait strategy; client-rendered sites return little or nothing.
- Caveat: heuristics can misfire. Content-area detection relies on class and id patterns. A page that matches the wrong container yields confident but partial text, and the result silently becomes
Nonewhen filters reject it. - Caveat: transparent by default. The bot-identifying user agent, robots-aware crawler and lack of stealth are good citizenship, but mean protected sites will block it. The default insecure-SSL retry is convenient but worth turning off for sensitive pipelines.
- Caveat: crawler scope. The spider is a small single-site link discoverer with in-memory state, not a distributed crawler.
Sources: code at 6c1977a, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 6c1977a. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPlain HTTP(S) only, no headless browser. Pages are fetched via one of two backends:
- urllib3 (default):
_send_urllib_requestindownloads.py:217uses aurllib3.PoolManagerwith configurable retry strategy and decompression. The response body is streamed in 128KB chunks and capped atMAX_FILE_SIZE(20MB default) via_capped(). - pycurl (optional, imported at
downloads.py:42):_send_pycurl_requestatdownloads.py:477provides the same functionality via libcurl with shared DNS/SSL session caches across requests. Both backends support SOCKS proxy via thehttp_proxyenv var and enforce an SSRF protection layer (_SafePoolManager) that rejects non-global IP addresses.
No JavaScript rendering. The tool does not use a headless browser, Selenium, or Playwright. It receives raw HTML only. There is no waiting strategy for dynamic content beyond HTTP-level retries on transient status codes (429, 5xx). Both backends follow redirects (configurable via MAX_REDIRECTS). Headers are stored from the response; the User-Agent string defaults to trafilatura/<version> or rotates through user-supplied agents from the config file (downloads.py:173).
Content types: The system accepts HTML and any text-based response. There is no built-in support for downloading or extracting PDFs, images, or binary files; IMAGE_EXTENSION at utils.py:123 is used only to detect whether a URL points to an image, which is then filtered rather than rendered.
How is content extracted or converted?
answeredRule-based cascading extractor with no ML component. The core pipeline lives in core.py:trafilatura_sequence() (line 165) and runs five stages:
Tag conversion (
htmlprocessing.py:408): HTML elements are converted to a simplified XML vocabulary — headings become<head>, lists become<list>/<item>, code blocks become<code>, quotes become<quote>, inline formatting becomes<hi>,<br>→<lb>,<img>→<graphic>, etc. XPath-based content-area detection viaBODY_XPATH(xpaths.py:62) locates the main content frame by matching common article ID/class patterns (entry-content,post-body,articleBody).Boilerplate removal (
htmlprocessing.py:81tree_cleaning): Unwanted elements (nav,footer,iframe,script,aside,form, etc.) are deleted. XPath expressions (OVERALL_DISCARD_XPATHinxpaths.py:274) prune sidebar, breadcrumb, share-button, paywall, ad, and cookie-consent sections by ID/class heuristics. Link-density tests (htmlprocessing.py:177) remove sections rich in links and low in text.Main extractor (
main_extractor.py:686extract_content): IteratesBODY_XPATHexpressions, extracts matching subtrees, and processes child elements viahandle_textelem()which dispatches to paragraph/list/table/code/quote handlers.Cascade: If short extraction →
recover_wild_text()scavenges orphan<p>/<code>/<div>elements. Thencompare_extraction()(external.py:84) evaluates readability-lxml and justext as fallbacks. Abaseline()rescue (baseline.py:167) tries JSON-LD embedded content,<article>tags, paragraphs, and finally the full body. A recall escalation retries in high-recall mode if output covers <20% of page length.Output conversion (
core.py:80determine_returnstring): The XML body is serialized to TXT/Markdown (with YAML metadata header), CSV, JSON, HTML, XML, or XML-TEI. Metadata (title, author, date, categories, tags) is extracted viametadata.py:1using XPath selectors and htmldate.
How are LLMs used, if at all?
not applicableThis repository does not use LLMs at all. A full case-insensitive grep across the entire source for openai, anthropic, claude, gpt, gemini, langchain, llm, and huggingface returned zero matches. The project relies entirely on deterministic rule-based extraction: XPath heuristics, tag conversion, link-density analysis, and three classical algorithmic extractors (readability-lxml, jusText, and its own cascading rule engine). No prompting, chunking, structured output generation via LLM, provider integration, or cost controls exist.
How are anti-bot measures, proxies and fingerprinting handled?
answeredMinimal anti-bot evasion. The project does not implement stealth patches, CAPTCHA handling, or browser fingerprint spoofing. Its anti-detection measures are:
- User-agent rotation (
downloads.py:169-174): The_determine_headers()function accepts a multi-lineUSER_AGENTSconfig value, picks one at random per request. The default agent identifies itself astrafilatura/<version>(transparent, non-stealthy). - SOCKS proxy (
downloads.py:37): Ifhttp_proxyenv var is set, all requests route through aSOCKSProxyManager(urllib3) orPRE_PROXY(pycurl). No built-in proxy rotation or proxy list management. - Rate limiting: The crawler sleeps
SLEEP_TIMEseconds (default 5.0, configurable insettings.cfg:11) between requests, derived from the domain's crawl delay (spider.py:339). - robots.txt (
spider.py:152-170): The crawler fetches and parsesrobots.txtviaurllib.robotparser.RobotFileParserand honourscan_fetch()rules before visiting links. - SSRF protection (
downloads.py:118-133): By default, connections to non-public IP addresses (loopback, private, link-local) are blocked via a custom_SafePoolManagerwith per-connection peer vetting.
There is no proxy rotation, no CAPTCHA-solving integration, no TLS fingerprint spoofing, and no referer/header randomization apart from the user-agent draw.
How is crawling at scale implemented?
answeredA focused, single-domain crawler with URL management via the courlan library. The primary entry point is focused_crawler() in spider.py:308.
- URL queue & dedup: The
UrlStoreclass (fromcourlan) manages a priority queue of discovered URLs with domain-aware deduplication. URLs are tracked in three states: known, visited, and unvisited. Navigation pages (indices, archives) get priority placement in the queue (spider.py:224-229). - Limits:
max_seen_urlsdefaults to 10;max_known_urlsdefaults to 100,000 (spider.py:39). The primary loop atspider.py:342stops when either limit is reached or the domain URL store is exhausted. - robots.txt & politeness: Each domain's
robots.txtis fetched and parsed on init (spider.py:60).is_valid_link()(spider.py:88-90) checksrules.can_fetch("*", link), that the link stays within the reference domain, and that the URL is crawlable. A configurableSLEEP_TIME(5s default) separates requests (spider.py:339). - Discovery modes: URLs can be seeded from sitemaps (
sitemaps.pysupports XML sitemaps, sitemap indexes, and robots.txt sitemap directives) and feeds (feeds.pysupports ATOM, JSON, RSS). The--exploreCLI option combines both. - Concurrency: Crawling itself is single-threaded per domain. For batch URL processing,
buffered_downloads()indownloads.py:438uses aThreadPoolExecutor(default up to 16 threads) to parallelize fetches. No distributed worker architecture exists.
There is no support for multiple concurrent crawl workers, distributed crawling, or queue persistence across restarts.
What is the developer interface?
answeredLibrary API, CLI — no REST service, MCP server, or web UI.
Python API (exposed in __init__.py:15-32):
fetch_url(url)→ downloads a page and returns a decoded string.extract(html, ...)→ the main extraction function, returns a string in the chosen format.extract_with_metadata(...)→ returns aDocumentobject containing both text and metadata.bare_extraction(...)→ returns aDocumentwith the lxml body tree (unsafe for serialization).fetch_response(url)→ returns the rawResponseobject.- All functions accept 20+ keyword arguments for fine-grained control (output format, language filtering, deduplication, precision/recall, formatting, images, links, tables, comments, pruning XPath, custom config).
CLI (cli.py:49-168): Invoked as trafilatura (registered in pyproject.toml:83). Input options: --input-file, --input-dir, --URL, stdin. Navigation options: --feed, --sitemap, --crawl, --explore, --probe. Extraction options: --fast, --precision, --recall, --with-metadata, --target-language, --deduplicate. Parallel processing via --parallel.
Output formats (settings.py:29): txt (plain text with optional YAML metadata header), markdown, csv, json, html, xml, xmltei (TEI-conformant XML).
Configuration: A settings.cfg file (settings.py:41-49) controls defaults for download timeout, file sizes, sleep time, user agents, SSRF protection, date search, and extraction thresholds. Users can supply overrides via --config-file or a ConfigParser object in the API.
There is no REST web service, no MCP server integration, no graphical user interface. Language bindings are Python-only.