# adbar/trafilatura

> Rule-based Python library and CLI that fetches pages over HTTP and extracts main text, comments and metadata without a browser or ML.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/adbar/trafilatura (reviewed at commit `6c1977a00b4e82ebb7e6fad1192467adaf24d430`, 2026-10-02)
- Stars: 6920 · Language: Python · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/trafilatura/

## Overview

Trafilatura is a Python library and CLI that downloads web pages over plain HTTP and extracts the main text, comments and metadata with deterministic rules. It does not use a browser, a model, or an LLM. Inside it parses HTML with lxml, converts it to a small XML vocabulary (`p`, `head`, `list`, `quote`, `code`, `table`, `graphic`, `ref`), runs a cascade of extractors, and serialises the result as TXT, Markdown, CSV, JSON, HTML, XML or XML-TEI.

It is the quiet workhorse of this category. Corpus builders, RAG ingestion jobs and research crawlers use it because it is fast, needs no GPU or API key, and copes with a long tail of templates through accumulated heuristics (the pinned version is 2.3.0). The cost of the rule-based design is that it only sees what the server sends: JavaScript-rendered pages, PDFs and paywalled content are out of scope.

Beyond extraction, it ships its own discovery layer: sitemap and feed parsers, a polite focused crawler built on `courlan`'s `UrlStore`, and a CLI that downloads many URLs in parallel with per-domain back-off. In this category it is a pre-processor for LLM pipelines, not an AI scraper itself.

## Architecture

```mermaid
flowchart TD
  IN["URL, file or HTML string"] --> DL["downloads.fetch_url (urllib3 or pycurl)"]
  DL --> LOAD["utils.load_html (lxml)"]
  IN --> LOAD
  LOAD --> META["metadata.extract_metadata"]
  LOAD --> SEQ["core.trafilatura_sequence"]
  SEQ --> CLEAN["htmlprocessing: tree_cleaning + convert_tags"]
  CLEAN --> MAIN["main_extractor.extract_content"]
  MAIN --> CMP["external.compare_extraction (readability fork, jusText)"]
  CMP --> BASE["baseline rescue"]
  BASE --> ESC["recall escalation"]
  ESC --> OUT["determine_returnstring: txt, md, json, csv, html, xml, tei"]
  META --> OUT
  DISC["sitemaps, feeds, spider"] --> DL
```

| Component | Path | Role |
|---|---|---|
| Public API | `trafilatura/__init__.py`, `core.py` | `fetch_url`, `extract`, `extract_with_metadata`, `bare_extraction`, the extraction cascade and output dispatch |
| Options | `trafilatura/settings.py`, `settings.cfg` | `Extractor` options object, `Document` result, config-file defaults |
| Cleaning and conversion | `trafilatura/htmlprocessing.py` | Tag removal, link-density pruning, HTML to internal XML |
| Main extractor | `trafilatura/main_extractor.py` | Content-area XPaths, per-element handlers, wild-text recovery, comments |
| XPath rules | `trafilatura/xpaths.py` | `BODY_XPATH`, discard lists for sidebars, teasers, comments and ads |
| Fallbacks | `trafilatura/external.py`, `readability_lxml.py`, `baseline.py` | Vendored readability fork, jusText, and a last-resort baseline |
| Metadata | `trafilatura/metadata.py`, `json_metadata.py` | Meta tags, JSON-LD, title, author, date via `htmldate` |
| Downloads | `trafilatura/downloads.py` | Pooled HTTP with SSRF guard, retries, size cap, optional SOCKS proxy |
| Discovery | `trafilatura/spider.py`, `sitemaps.py`, `feeds.py` | Focused crawler, sitemap and feed URL discovery |
| CLI | `trafilatura/cli.py`, `cli_utils.py` | Batch processing, parallel downloads, crawl and explore modes |

## How a request flows

Take `extract(fetch_url("https://example.org/post"), with_metadata=True, output_format="markdown")`:

1. **Download.** `fetch_response` picks pycurl when it is installed and urllib3 otherwise. On an SSL error it retries without certificate checks unless `INSECURE_SSL_FALLBACK` is off ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L290-L326)). The urllib3 path streams the body in 128 KiB chunks capped at `MAX_FILE_SIZE`, with retry and redirect limits from the config ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L182-L253)). `fetch_url` returns decoded HTML only for a 200 of acceptable length ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L256-L287)).
2. **Options and parse.** `bare_extraction` packs the keyword arguments into an `Extractor`, parses the HTML, optionally checks the `<html lang>` attribute, and runs `extract_metadata` ([core.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L348-L407)). Metadata comes from meta tags, then JSON-LD, then title and author heuristics, with dates from `htmldate.find_date` ([metadata.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/metadata.py#L447-L517)).
3. **Prepare.** `trafilatura_sequence` prunes appended articles, share widgets and (unless comments are wanted) comment sections, then `tree_cleaning` deletes unwanted tags and `convert_tags` maps HTML to the internal vocabulary ([core.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L139-L208), [htmlprocessing.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/htmlprocessing.py#L408-L467)).
4. **Main extractor.** `_extract` tries each `BODY_XPATH` expression in order, from `itemprop='articleBody'`-style matches through `<article>` and story/content classes to `main`. For each match it prunes discard XPaths and link-heavy blocks, converts children through `handle_textelem`, and stops at the first subtree with real content ([main_extractor.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/main_extractor.py#L579-L683), [xpaths.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/xpaths.py#L62-L107)). If the result is short, `recover_wild_text` scavenges loose paragraphs.
5. **Compare.** Unless `fast=True`, `compare_extraction` runs the vendored readability on the raw tree and switches to it when the own result is empty, much shorter, or structurally poor; jusText is then tried for unclean or short outputs ([external.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/external.py#L47-L138)).
6. **Rescue and escalate.** If text is still under `MIN_EXTRACTED_SIZE` (250 chars), `baseline` tries JSON-LD bodies, `<article>` blocks, paragraphs, then the whole body. In balanced mode, a result under 3,000 characters that also covers under 20% of the page text triggers a recall-mode retry plus a jusText candidate ([core.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L210-L266)).
7. **Filter and serialise.** Size, duplicate and language checks can discard the document (the function returns `None`). `determine_returnstring` then writes Markdown with a YAML front-matter header built from the metadata ([core.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L80-L114)).

## Key components

### The extraction cascade

The design idea is "each stage engages only if the previous one under-delivered" ([core.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L165-L183)). The own extractor is precise but template-bound; readability is good on classic articles; jusText classifies paragraphs by density and stopwords and reaches content buried in generic `div`s; baseline is a blunt safety net. `favor_precision` and `favor_recall` shift each stage's thresholds rather than choosing a different algorithm. There is special handling for forum threads (detected by schema.org `DiscussionForumPosting`), where comment containers are kept as content.

### Cleaning and the internal XML

`tree_cleaning` removes a fixed tag list, keeps or drops tables and images depending on options, and in recall mode undoes the deletion if it would remove every paragraph ([htmlprocessing.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/htmlprocessing.py#L81-L124)). The XML intermediate is what makes the output formats cheap: Markdown, HTML and TEI are just different walks over the same tree.

### Downloads and politeness

The default user agent is `trafilatura/<version>` with a project URL; `USER_AGENTS` and `COOKIE` in the config override it, picking one agent at random per request ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L169-L179)). By default a `_SafePoolManager` rejects private, loopback and link-local addresses on every hop. If `http_proxy` is set, a urllib3 `SOCKSProxyManager` is used instead, and that path skips the SSRF pool ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L34-L68)). Defaults: 30 s timeout, 20 MB cap, 2 redirects, 5 s between requests to one domain ([settings.cfg](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/settings.cfg#L5-L30)).

### Focused crawler

`focused_crawler(homepage)` reads `robots.txt`, keeps links inside the start URL's path prefix, puts navigation pages at the front of the queue, sleeps for the robots crawl delay or `SLEEP_TIME` between fetches, and stops after `max_seen_urls` (default 10) pages ([spider.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L189-L229), [L308-L351](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L308-L351)). It returns two URL lists, to visit and known; it does not extract text. The frontier is a module-level `URL_STORE`, so state persists across calls in one process.

### CLI

`trafilatura` accepts a URL, an input file, a directory or stdin, plus `--feed`, `--sitemap`, `--crawl`, `--explore` and `--probe`. URL lists go through `buffered_downloads`, a thread pool fed by `UrlStore.get_download_urls` with per-domain back-off ([downloads.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L414-L446)). `--archived` retries failed URLs through the Wayback Machine ([cli_utils.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/cli_utils.py#L402-L429)).

## Extending it

- **Tune without code.** `prune_xpath` removes site-specific noise before extraction; `--config-file` or a `ConfigParser` changes size thresholds, timeouts, user agents and SSRF behaviour.
- **Get structure, not text.** `bare_extraction` returns a `Document` with the lxml `body` tree, so you can walk headings, lists and tables yourself before chunking for a RAG index.
- **Custom rules.** `MANUALLY_CLEANED` and the XPath lists in `xpaths.py` are module-level and can be mutated at runtime; the code explicitly honours a user who removes `"form"` from the cleaning list.
- **JavaScript pages.** Render with a browser of your choice, then pass the HTML string to `extract`. The library accepts strings, bytes or lxml trees.

## Running it

- **Install.** `pip install trafilatura` pulls only `lxml`, `urllib3`, `courlan`, `htmldate`, `justext`, `charset_normalizer` and `certifi`. `trafilatura[all]` adds pycurl, SOCKS support, `py3langid` language detection, Brotli/zstd and faster charset detection.
- **Python.** `extract(fetch_url(url))` is the one-liner; `extract_with_metadata` returns a `Document` with text and metadata fields ([__init__.py](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/__init__.py#L13-L33)).
- **CLI.** `trafilatura -u URL --output-format markdown --with-metadata`, or `trafilatura -i urls.txt -o out/ --parallel 8` for batches.

## Strengths and caveats

- **Strength: fast, local and cheap.** No browser, no model and no network calls beyond the fetch, so it runs anywhere Python and lxml run.
- **Strength: layered fallbacks.** Four extraction strategies, gated by measured text length, give good recall on odd templates without hurting typical articles.
- **Strength: rich output.** Metadata, comments, tables, formatting, links and images are optional, and TEI and YAML-headed Markdown are first-class.
- **Caveat: static HTML only.** There is no JavaScript execution and no wait strategy; client-rendered sites return little or nothing.
- **Caveat: heuristics can misfire.** Content-area detection relies on class and id patterns. A page that matches the wrong container yields confident but partial text, and the result silently becomes `None` when filters reject it.
- **Caveat: transparent by default.** The bot-identifying user agent, robots-aware crawler and lack of stealth are good citizenship, but mean protected sites will block it. The default insecure-SSL retry is convenient but worth turning off for sensitive pipelines.
- **Caveat: crawler scope.** The spider is a small single-site link discoverer with in-memory state, not a distributed crawler.

*Sources: code at 6c1977a, deepwiki-open wiki (12 pages), verified Q&A.*

## How adbar/trafilatura answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

**Plain HTTP(S) only, no headless browser.** Pages are fetched via one of two backends:

- **urllib3** (default): `_send_urllib_request` in `downloads.py:217` uses a `urllib3.PoolManager` with configurable retry strategy and decompression. The response body is streamed in 128KB chunks and capped at `MAX_FILE_SIZE` (20MB default) via `_capped()`. 
- **pycurl** (optional, imported at `downloads.py:42`): `_send_pycurl_request` at `downloads.py:477` provides the same functionality via libcurl with shared DNS/SSL session caches across requests. Both backends support SOCKS proxy via the `http_proxy` env var and enforce an SSRF protection layer (`_SafePoolManager`) that rejects non-global IP addresses.

**No JavaScript rendering.** The tool does not use a headless browser, Selenium, or Playwright. It receives raw HTML only. There is no waiting strategy for dynamic content beyond HTTP-level retries on transient status codes (429, 5xx). Both backends follow redirects (configurable via `MAX_REDIRECTS`). Headers are stored from the response; the `User-Agent` string defaults to `trafilatura/<version>` or rotates through user-supplied agents from the config file (`downloads.py:173`).

**Content types:** The system accepts HTML and any text-based response. There is no built-in support for downloading or extracting PDFs, images, or binary files; `IMAGE_EXTENSION` at `utils.py:123` is used only to detect whether a URL points to an image, which is then filtered rather than rendered.


Citations: [trafilatura/downloads.py:217-243](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L217-L243) · [trafilatura/downloads.py:63-68](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L63-L68) · [trafilatura/downloads.py:86-89](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L86-L89) · [trafilatura/downloads.py:169-179](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L169-L179) · [trafilatura/settings.cfg:6-7](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/settings.cfg#L6-L7) · [trafilatura/utils.py:142-179](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/utils.py#L142-L179)

### How is content extracted or converted? (answered)

**Rule-based cascading extractor with no ML component.** The core pipeline lives in `core.py:trafilatura_sequence()` (line 165) and runs five stages:

1. **Tag conversion** (`htmlprocessing.py:408`): HTML elements are converted to a simplified XML vocabulary — headings become `<head>`, lists become `<list>/<item>`, code blocks become `<code>`, quotes become `<quote>`, inline formatting becomes `<hi>`, `<br>` → `<lb>`, `<img>` → `<graphic>`, etc. XPath-based content-area detection via `BODY_XPATH` (`xpaths.py:62`) locates the main content frame by matching common article ID/class patterns (`entry-content`, `post-body`, `articleBody`).

2. **Boilerplate removal** (`htmlprocessing.py:81` `tree_cleaning`): Unwanted elements (`nav`, `footer`, `iframe`, `script`, `aside`, `form`, etc.) are deleted. XPath expressions (`OVERALL_DISCARD_XPATH` in `xpaths.py:274`) prune sidebar, breadcrumb, share-button, paywall, ad, and cookie-consent sections by ID/class heuristics. Link-density tests (`htmlprocessing.py:177`) remove sections rich in links and low in text.

3. **Main extractor** (`main_extractor.py:686` `extract_content`): Iterates `BODY_XPATH` expressions, extracts matching subtrees, and processes child elements via `handle_textelem()` which dispatches to paragraph/list/table/code/quote handlers.

4. **Cascade**: If short extraction → `recover_wild_text()` scavenges orphan `<p>`/`<code>`/`<div>` elements. Then `compare_extraction()` (`external.py:84`) evaluates readability-lxml and justext as fallbacks. A `baseline()` rescue (`baseline.py:167`) tries JSON-LD embedded content, `<article>` tags, paragraphs, and finally the full body. A recall escalation retries in high-recall mode if output covers <20% of page length.

5. **Output conversion** (`core.py:80` `determine_returnstring`): The XML body is serialized to TXT/Markdown (with YAML metadata header), CSV, JSON, HTML, XML, or XML-TEI. Metadata (title, author, date, categories, tags) is extracted via `metadata.py:1` using XPath selectors and htmldate.


Citations: [trafilatura/core.py:165-267](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L165-L267) · [trafilatura/xpaths.py:62-107](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/xpaths.py#L62-L107) · [trafilatura/htmlprocessing.py:81-124](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/htmlprocessing.py#L81-L124) · [trafilatura/external.py:84-138](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/external.py#L84-L138) · [trafilatura/baseline.py:167-237](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/baseline.py#L167-L237)

### How are LLMs used, if at all? (not applicable)

This repository does not use LLMs at all. A full case-insensitive grep across the entire source for `openai`, `anthropic`, `claude`, `gpt`, `gemini`, `langchain`, `llm`, and `huggingface` returned zero matches. The project relies entirely on deterministic rule-based extraction: XPath heuristics, tag conversion, link-density analysis, and three classical algorithmic extractors (readability-lxml, jusText, and its own cascading rule engine). No prompting, chunking, structured output generation via LLM, provider integration, or cost controls exist.



### How are anti-bot measures, proxies and fingerprinting handled? (answered)

**Minimal anti-bot evasion.** The project does not implement stealth patches, CAPTCHA handling, or browser fingerprint spoofing. Its anti-detection measures are:

- **User-agent rotation** (`downloads.py:169-174`): The `_determine_headers()` function accepts a multi-line `USER_AGENTS` config value, picks one at random per request. The default agent identifies itself as `trafilatura/<version>` (transparent, non-stealthy).
- **SOCKS proxy** (`downloads.py:37`): If `http_proxy` env var is set, all requests route through a `SOCKSProxyManager` (urllib3) or `PRE_PROXY` (pycurl). No built-in proxy rotation or proxy list management.
- **Rate limiting**: The crawler sleeps `SLEEP_TIME` seconds (default 5.0, configurable in `settings.cfg:11`) between requests, derived from the domain's crawl delay (`spider.py:339`).
- **robots.txt** (`spider.py:152-170`): The crawler fetches and parses `robots.txt` via `urllib.robotparser.RobotFileParser` and honours `can_fetch()` rules before visiting links.
- **SSRF protection** (`downloads.py:118-133`): By default, connections to non-public IP addresses (loopback, private, link-local) are blocked via a custom `_SafePoolManager` with per-connection peer vetting.

There is no proxy rotation, no CAPTCHA-solving integration, no TLS fingerprint spoofing, and no referer/header randomization apart from the user-agent draw.


Citations: [trafilatura/downloads.py:169-179](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L169-L179) · [trafilatura/downloads.py:63-68](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L63-L68) · [trafilatura/spider.py:152-170](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L152-L170) · [trafilatura/spider.py:339-339](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L339-L339) · [trafilatura/downloads.py:118-166](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L118-L166) · [trafilatura/settings.cfg:11-11](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/settings.cfg#L11-L11)

### How is crawling at scale implemented? (answered)

**A focused, single-domain crawler with URL management via the `courlan` library.** The primary entry point is `focused_crawler()` in `spider.py:308`. 

- **URL queue & dedup**: The `UrlStore` class (from `courlan`) manages a priority queue of discovered URLs with domain-aware deduplication. URLs are tracked in three states: known, visited, and unvisited. Navigation pages (indices, archives) get priority placement in the queue (`spider.py:224-229`).
- **Limits**: `max_seen_urls` defaults to 10; `max_known_urls` defaults to 100,000 (`spider.py:39`). The primary loop at `spider.py:342` stops when either limit is reached or the domain URL store is exhausted.
- **robots.txt & politeness**: Each domain's `robots.txt` is fetched and parsed on init (`spider.py:60`). `is_valid_link()` (`spider.py:88-90`) checks `rules.can_fetch("*", link)`, that the link stays within the reference domain, and that the URL is crawlable. A configurable `SLEEP_TIME` (5s default) separates requests (`spider.py:339`).
- **Discovery modes**: URLs can be seeded from sitemaps (`sitemaps.py` supports XML sitemaps, sitemap indexes, and robots.txt sitemap directives) and feeds (`feeds.py` supports ATOM, JSON, RSS). The `--explore` CLI option combines both.
- **Concurrency**: Crawling itself is single-threaded per domain. For batch URL processing, `buffered_downloads()` in `downloads.py:438` uses a `ThreadPoolExecutor` (default up to 16 threads) to parallelize fetches. No distributed worker architecture exists.

There is no support for multiple concurrent crawl workers, distributed crawling, or queue persistence across restarts.


Citations: [trafilatura/spider.py:42-91](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L42-L91) · [trafilatura/spider.py:189-229](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/spider.py#L189-L229) · [trafilatura/downloads.py:438-447](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/downloads.py#L438-L447) · [trafilatura/sitemaps.py:1-48](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/sitemaps.py#L1-L48) · [trafilatura/feeds.py:1-49](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/feeds.py#L1-L49)

### What is the developer interface? (answered)

**Library API, CLI — no REST service, MCP server, or web UI.**

**Python API** (exposed in `__init__.py:15-32`):
- `fetch_url(url)` → downloads a page and returns a decoded string.
- `extract(html, ...)` → the main extraction function, returns a string in the chosen format.
- `extract_with_metadata(...)` → returns a `Document` object containing both text and metadata.
- `bare_extraction(...)` → returns a `Document` with the lxml body tree (unsafe for serialization).
- `fetch_response(url)` → returns the raw `Response` object.
- All functions accept 20+ keyword arguments for fine-grained control (output format, language filtering, deduplication, precision/recall, formatting, images, links, tables, comments, pruning XPath, custom config).

**CLI** (`cli.py:49-168`): Invoked as `trafilatura` (registered in `pyproject.toml:83`). Input options: `--input-file`, `--input-dir`, `--URL`, stdin. Navigation options: `--feed`, `--sitemap`, `--crawl`, `--explore`, `--probe`. Extraction options: `--fast`, `--precision`, `--recall`, `--with-metadata`, `--target-language`, `--deduplicate`. Parallel processing via `--parallel`.

**Output formats** (`settings.py:29`): `txt` (plain text with optional YAML metadata header), `markdown`, `csv`, `json`, `html`, `xml`, `xmltei` (TEI-conformant XML).

**Configuration**: A `settings.cfg` file (`settings.py:41-49`) controls defaults for download timeout, file sizes, sleep time, user agents, SSRF protection, date search, and extraction thresholds. Users can supply overrides via `--config-file` or a `ConfigParser` object in the API.

There is no REST web service, no MCP server integration, no graphical user interface. Language bindings are Python-only.


Citations: [trafilatura/cli.py:49-168](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/cli.py#L49-L168) · [trafilatura/core.py:269-338](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/core.py#L269-L338) · [trafilatura/settings.py:29-52](https://github.com/adbar/trafilatura/blob/6c1977a00b4e82ebb7e6fad1192467adaf24d430/trafilatura/settings.py#L29-L52)
