# mixmark-io/turndown

> JavaScript library that converts HTML strings or DOM nodes to Markdown via pluggable rules; it does not fetch or extract content.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/mixmark-io/turndown (reviewed at commit `aa84dfa3e2361edea8c43acbfc2b7a9363494bfd`, 2026-09-03)
- Stars: 11463 · Language: HTML · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/turndown/

## Overview

Turndown is a small JavaScript library that converts HTML to Markdown. It is not a scraper and has no AI features: it does not fetch pages, run JavaScript, crawl, or call a model. You give it an HTML string or a DOM node, and it returns a Markdown string. The whole engine is about 1,000 lines in `src/`, with one runtime dependency, `@mixmark-io/domino`, which parses HTML in Node. In the browser it uses the native DOM parser.

**Where it fits in an AI scraping stack:** it is the conversion step between "I have the HTML" and "I have text an LLM can read". Many JavaScript crawlers and agent tools use it for this. A typical pipeline is: fetch or render the page (Playwright, a stealth driver, a crawler), cut out the main content (a readability library or your own selectors), run Turndown, then chunk the result and send it to a model. Turndown does only the third step. It converts *everything* you pass in, including navigation, footers, and the text of `<script>` and `<style>` tags, unless you tell it not to.

The project is mature (version 7.2.4 at the pinned commit from September 2026). It is maintained mainly for correct CommonMark output and safe escaping, not for scraping features.

## Architecture

```mermaid
flowchart LR
  IN["HTML string or DOM node"] --> RN["RootNode"]
  RN --> PARSE["HTMLParser: DOMParser or domino"]
  RN --> CW["collapseWhitespace"]
  CW --> PROC["process(): walk child nodes"]
  PROC --> NODE["Node(): isBlock, isCode, isBlank, flanking whitespace"]
  NODE --> RULES["Rules.forNode()"]
  RULES --> REP["rule.replacement(content, node, options)"]
  REP --> JOIN["join() newline handling"]
  JOIN --> POST["postProcess(): rule.append, trim"]
  POST --> OUT["Markdown string"]
```

| Component | Path | Role |
|---|---|---|
| Service | `src/turndown.js` | `TurndownService`: options, `turndown()`, `use`, `addRule`, `keep`, `remove`, `escape`; the recursive `process` walk |
| Rules | `src/rules.js` | Holds added, CommonMark, keep and remove rules; picks one per node |
| CommonMark rules | `src/commonmark-rules.js` | Paragraphs, headings, lists, code blocks, links, images, emphasis and so on |
| Parser | `src/html-parser.js` | Native `DOMParser` in browsers, `domino.createDocument` in Node |
| Root node | `src/root-node.js` | Wraps input in `<x-turndown>`, parses it, collapses whitespace |
| Node helpers | `src/node.js` | Marks block, code and blank nodes; computes leading/trailing whitespace |
| Utilities | `src/utilities.js` | Block/void element lists, Markdown escaping regexes |
| Whitespace | `src/collapse-whitespace.js` | Collapses HTML whitespace the way a browser renders it, except inside `<pre>` |

## How a request flows

Take `new TurndownService({ headingStyle: 'atx' }).turndown(html)`:

1. **Options.** The constructor merges your options over the defaults: setext headings, `* * *` rules, `*` bullets, indented code blocks, `_` and `**` delimiters, inline links. It then builds a `Rules` object from `COMMONMARK_RULES` ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L8-L36)).
2. **Validate.** `turndown()` accepts a string, an element, a document or a fragment, and throws `TypeError` for anything else ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L47-L58), [L221-L230](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L221-L230)).
3. **Parse.** `RootNode` wraps a string in `<x-turndown id="turndown-root">`, parses it, and takes that element as the root. A DOM node is deep-cloned, so your document is not changed. Then `collapseWhitespace` runs on the tree ([root-node.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/root-node.js#L5-L27)). In Node the parser is `domino.createDocument`. In a browser it is the native `DOMParser`, or `document.implementation.createHTMLDocument` as a fallback ([html-parser.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/html-parser.js#L11-L68)).
4. **Walk.** `process` reduces over child nodes. Text nodes are escaped with `escape()`, except inside code. Element nodes go to `replacementForNode`. That function picks a rule, converts the children first (depth-first), trims the content when the node has flanking whitespace, and calls `rule.replacement(content, node, options)` ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L140-L193)).
5. **Join.** `join` merges each replacement into the output and keeps at most two newlines between blocks ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L204-L211)).
6. **Finish.** `postProcess` calls every rule's `append` hook (reference-style links use it to add the link list at the end) and trims leading and trailing whitespace ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L164-L173)).

## Key components

### Rule selection

`Rules.forNode` applies a fixed order. A blank node always gets the blank rule. After that come the rules in `array` (added rules are put first, then the CommonMark ones), then `keep` filters, then `remove` filters, and finally the default rule ([rules.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/rules.js#L24-L59)). A filter can be a tag name, an array of tag names, or a function `(node, options) => boolean` ([rules.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/rules.js#L61-L80)). One consequence: `remove('a')` does nothing, because the `inlineLink` rule matches first. To override a built-in element, use `addRule`.

### Built-in rules

`commonmark-rules.js` has 13 rules: paragraph, line break, headings (setext for h1–h2 unless `atx`), blockquote, lists, list items (with `start` for `<ol>`), indented and fenced code blocks (the fence grows when the code contains backticks, and the `language-*` class sets the info string), horizontal rule, inline and reference links, emphasis, strong, inline code and images ([commonmark-rules.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/commonmark-rules.js#L1-L271)). Tables, strikethrough and task lists are not included. They come from the separate `turndown-plugin-gfm` package.

### The default rule

An element with no matching rule keeps its converted children, with blank lines around it if it is a block element ([turndown.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L30-L32)). For scraped pages this is the most important behaviour. Because the default rules keep the text of any element they do not match, `<script>` and `<style>` bodies end up as plain text, a `<table>` flattens to one paragraph per cell, and `<nav>` links are kept. Plan to `remove(['script', 'style', 'noscript'])` and either add the GFM table plugin or `keep('table')`.

### Escaping

`escapeMarkdown` runs 13 regexes over every text node. It backslash-escapes `\ * _ \` [ ]` everywhere and `- + = # > ~~~ 1.` at the start of a line ([utilities.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/utilities.js#L82-L102)). This keeps the output correct when rendered, but it adds many backslashes to prose with underscores or asterisks. For LLM input, which is never rendered, you can override `TurndownService.prototype.escape` with a lighter function.

## Extending it

- **`addRule(key, { filter, replacement })`.** Adds a rule ahead of the built-ins. The `replacement` function gets the converted content, the DOM node and the options. An `append` hook can add text at the end of the document.
- **`keep(filter)` / `remove(filter)`.** `keep` outputs a node as raw HTML. `remove` drops it with its children. Both apply only to nodes no other rule matched.
- **`use(plugin)`.** A plugin is a function that receives the service and calls the methods above. `turndown-plugin-gfm` is the usual one.
- **Options.** `blankReplacement`, `keepReplacement` and `defaultReplacement` change the three special rules. `preformattedCode` keeps whitespace inside `<code>`.

## Running it

- `npm install turndown`. Node 18 or newer, according to `engines`. The npm package has CJS, ES and UMD builds, plus browser builds that leave out domino through the `browser` field in `package.json`. An IIFE file in `dist/` works with a `<script>` tag.
- No services, network or configuration are needed. Conversion is synchronous and runs in the calling thread. Wrap it in a worker if you convert large pages on a busy server.
- Security: in Node, domino does not run scripts or load external resources. In a browser, the native parser decides. `SECURITY.md` advises sanitizing untrusted HTML first, because the parser's side effects happen before Turndown runs.

## Strengths and caveats

- **Strength: small and predictable.** One dependency, synchronous, deterministic output, and a short rule list that is easy to read and override.
- **Strength: correct Markdown.** Fence lengths adapt to the code, link destinations with spaces or brackets are escaped, and whitespace handling copies browser rendering.
- **Strength: works everywhere.** The same API runs in Node, browsers, extensions and edge runtimes.
- **Caveat: no content extraction.** It converts the whole input, including boilerplate and script text. You need a readability step or selectors before it.
- **Caveat: CommonMark only by default.** Tables are flattened unless you add a plugin. This matters a lot for data-heavy pages fed to an LLM.
- **Caveat: aggressive escaping.** The backslashes cost tokens and can confuse a model. Override `escape` if the output is only for machines.
- **Caveat: shared state in reference links.** The `referenceLink` rule stores pending references on the shared rule object and clears them in `append` ([commonmark-rules.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/commonmark-rules.js#L160-L205)). This is safe in synchronous use, but it ties together all instances that use `linkStyle: 'referenced'`.

*Sources: code at aa84dfa, OpenDeepWiki wiki (7 pages), verified Q&A.*

## How mixmark-io/turndown answers the AI web scraping questions

### How are pages fetched and rendered? (not applicable)

Turndown does not fetch or render pages. It takes an HTML string (or a pre-parsed DOM node) as input and converts it to Markdown. There is no HTTP client, no headless browser integration, no JavaScript rendering, no waiting strategy, and no support for PDF or image content types. The library's entry point, `turndown()`, accepts either a string or an HTMLElement/Document/DocumentFragment node (`src/turndown.js:47-58`). When a string is provided, it is parsed into a DOM tree using either the native browser parser or, in Node.js, the domino library — a minimal DOM implementation that does not execute scripts or fetch external resources (`src/html-parser.js:49-53`). The SECURITY.md explicitly warns that Turndown itself does not fetch anything and that any external resource loading would be the DOM parser's behavior, not Turndown's. In short, the project is an offline format converter, not a scraper.


Citations: [src/turndown.js:47-58](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L47-L58) · [src/html-parser.js:49-53](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/html-parser.js#L49-L53)

### How is content extracted or converted? (not applicable)

Turndown converts entire HTML documents or fragments to Markdown; it does not perform content extraction, readability-style boilerplate removal, or schema-based extraction. The conversion is driven by a rule system where each HTML element type has a corresponding replacement function that produces Markdown. Rules match by tag name, array of tag names, or a custom filter function (`src/rules.js:69-80`). The built-in CommonMark rules cover paragraphs, headings, blockquotes, lists, code blocks, horizontal rules, links (inline and reference styles), images, emphasis, strong, and inline code (`src/commonmark-rules.js:1-271`). Users can extend or override rules via `addRule()`, `keep()`, or `remove()` (`src/turndown.js:88-117`). A pre-processing step collapses extraneous whitespace before conversion (`src/collapse-whitespace.js:33-98`). There is no concept of selecting a subset of content by CSS selector (for extraction purposes) or inferring a page's main content — every element in the input is either converted, kept as HTML, or removed, based on the rules configured.

> **Editor's note.** Correction: elements without a matching rule fall through to the default rule, which keeps their text content, so `<script>` and `<style>` bodies appear in the Markdown unless you call `remove(['script', 'style'])`; tables are flattened to cell text unless a plugin such as turndown-plugin-gfm is added. For scraped pages this means you must strip boilerplate before or during conversion.

Citations: [src/rules.js:69-80](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/rules.js#L69-L80) · [src/commonmark-rules.js:1-271](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/commonmark-rules.js#L1-L271) · [src/turndown.js:88-117](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L88-L117) · [src/collapse-whitespace.js:33-98](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/collapse-whitespace.js#L33-L98)

### How are LLMs used, if at all? (not applicable)

Turndown contains no LLM usage of any kind. There are no prompts, no chunking of large pages, no structured output schemas, no LLM provider integrations (OpenAI, Anthropic, or others), and no cost controls. The project is a deterministic HTML-to-Markdown converter implemented entirely in JavaScript with a rule-based approach — every replacement decision is made by matching HTML elements against predefined filters and running the corresponding replacement functions. There are no dependencies on any AI or ML libraries in `package.json`; the sole runtime dependency is `@mixmark-io/domino` for DOM parsing. The library's only output format is Markdown text, and its output is purely a function of its input rules and the HTML it receives.



### How are anti-bot measures, proxies and fingerprinting handled? (not applicable)

Turndown implements no anti-bot measures whatsoever. There are no stealth patches, no fingerprint spoofing, no proxy rotation, no CAPTCHA handling, and no rate limiting. The library does not make network requests of any kind — it is a pure format conversion library operating on in-memory HTML strings or DOM nodes. The `src/` directory contains no references to proxies, headers, IP addresses, user-agent strings, or any networking concepts. While the `package-lock.json` file lists `http-proxy-agent` and `https-proxy-agent` as transitive dependencies (from the test runner or build tools), these are not imported or used anywhere in Turndown's own source code. The project has no need for anti-bot measures because it never contacts external servers.



### How is crawling at scale implemented? (not applicable)

Turndown provides no crawling infrastructure. There are no URL queues, no concurrency/threading models, no URL deduplication, no depth or limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. The library is invoked synchronously on a single HTML input at a time — it converts whatever is passed to it and returns the Markdown result. There is no concept of following links, discovering pages, or managing a crawl frontier. The entirety of Turndown's public API is the `TurndownService` class with its `turndown()`, `addRule()`, `keep()`, `remove()`, and `use()` methods (`src/turndown.js:38-130`).


Citations: [src/turndown.js:38-130](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L38-L130)

### What is the developer interface? (answered)

Turndown exposes a library API only — there is no CLI binary (no `bin` entry in `package.json`), no REST service, no MCP server, and no user-facing UI beyond an HTML demo page (`index.html`). The public interface is the `TurndownService` class, instantiated via `new TurndownService(options)` or the factory shorthand `TurndownService(options)` (`src/turndown.js:8-10`). The primary method is `turndown(input)` which accepts either an HTML string or a DOM node and returns a Markdown string (`src/turndown.js:47-58`). Customization happens through `addRule(key, rule)` for adding conversion rules, `keep(filter)` to preserve elements as HTML, `remove(filter)` to strip elements, `use(plugin)` for plugin-based extensions, and `escape(string)` for escaping Markdown syntax (`src/turndown.js:68-129`). Options are set at construction time and include heading style, HR marker, list bullet, code block style, fence character, emphasis/strong delimiters, link style, and reference style (`src/turndown.js:11-22`). The library ships as multiple build formats (CJS, ES module, UMD, IIFE) for broad compatibility, and a browser-specific variant replaces the domino DOM parser with the native browser DOM (`package.json:6-14`). The only output format is Markdown text. Language bindings are JavaScript-only — there are no Python, Rust, or other language ports in this repository.


Citations: [src/turndown.js:8-10](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L8-L10) · [src/turndown.js:47-58](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L47-L58) · [src/turndown.js:11-22](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/turndown.js#L11-L22) · [package.json:6-14](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/package.json#L6-L14)
