mixmark-io/turndown
JavaScript library that converts HTML strings or DOM nodes to Markdown via pluggable rules; it does not fetch or extract content.
Overview
Turndown is a small JavaScript library that converts HTML to Markdown. It is not a scraper and has no AI features: it does not fetch pages, run JavaScript, crawl, or call a model. You give it an HTML string or a DOM node, and it returns a Markdown string. The whole engine is about 1,000 lines in src/, with one runtime dependency, @mixmark-io/domino, which parses HTML in Node. In the browser it uses the native DOM parser.
Where it fits in an AI scraping stack: it is the conversion step between “I have the HTML” and “I have text an LLM can read”. Many JavaScript crawlers and agent tools use it for this. A typical pipeline is: fetch or render the page (Playwright, a stealth driver, a crawler), cut out the main content (a readability library or your own selectors), run Turndown, then chunk the result and send it to a model. Turndown does only the third step. It converts everything you pass in, including navigation, footers, and the text of <script> and <style> tags, unless you tell it not to.
The project is mature (version 7.2.4 at the pinned commit from September 2026). It is maintained mainly for correct CommonMark output and safe escaping, not for scraping features.
Architecture
flowchart LR
IN["HTML string or DOM node"] --> RN["RootNode"]
RN --> PARSE["HTMLParser: DOMParser or domino"]
RN --> CW["collapseWhitespace"]
CW --> PROC["process(): walk child nodes"]
PROC --> NODE["Node(): isBlock, isCode, isBlank, flanking whitespace"]
NODE --> RULES["Rules.forNode()"]
RULES --> REP["rule.replacement(content, node, options)"]
REP --> JOIN["join() newline handling"]
JOIN --> POST["postProcess(): rule.append, trim"]
POST --> OUT["Markdown string"]
| Component | Path | Role |
|---|---|---|
| Service | src/turndown.js |
TurndownService: options, turndown(), use, addRule, keep, remove, escape; the recursive process walk |
| Rules | src/rules.js |
Holds added, CommonMark, keep and remove rules; picks one per node |
| CommonMark rules | src/commonmark-rules.js |
Paragraphs, headings, lists, code blocks, links, images, emphasis and so on |
| Parser | src/html-parser.js |
Native DOMParser in browsers, domino.createDocument in Node |
| Root node | src/root-node.js |
Wraps input in <x-turndown>, parses it, collapses whitespace |
| Node helpers | src/node.js |
Marks block, code and blank nodes; computes leading/trailing whitespace |
| Utilities | src/utilities.js |
Block/void element lists, Markdown escaping regexes |
| Whitespace | src/collapse-whitespace.js |
Collapses HTML whitespace the way a browser renders it, except inside <pre> |
How a request flows
Take new TurndownService({ headingStyle: 'atx' }).turndown(html):
- Options. The constructor merges your options over the defaults: setext headings,
* * *rules,*bullets, indented code blocks,_and**delimiters, inline links. It then builds aRulesobject fromCOMMONMARK_RULES(turndown.js). - Validate.
turndown()accepts a string, an element, a document or a fragment, and throwsTypeErrorfor anything else (turndown.js, L221-L230). - Parse.
RootNodewraps a string in<x-turndown id="turndown-root">, parses it, and takes that element as the root. A DOM node is deep-cloned, so your document is not changed. ThencollapseWhitespaceruns on the tree (root-node.js). In Node the parser isdomino.createDocument. In a browser it is the nativeDOMParser, ordocument.implementation.createHTMLDocumentas a fallback (html-parser.js). - Walk.
processreduces over child nodes. Text nodes are escaped withescape(), except inside code. Element nodes go toreplacementForNode. That function picks a rule, converts the children first (depth-first), trims the content when the node has flanking whitespace, and callsrule.replacement(content, node, options)(turndown.js). - Join.
joinmerges each replacement into the output and keeps at most two newlines between blocks (turndown.js). - Finish.
postProcesscalls every rule’sappendhook (reference-style links use it to add the link list at the end) and trims leading and trailing whitespace (turndown.js).
Key components
Rule selection
Rules.forNode applies a fixed order. A blank node always gets the blank rule. After that come the rules in array (added rules are put first, then the CommonMark ones), then keep filters, then remove filters, and finally the default rule (rules.js). A filter can be a tag name, an array of tag names, or a function (node, options) => boolean (rules.js). One consequence: remove('a') does nothing, because the inlineLink rule matches first. To override a built-in element, use addRule.
Built-in rules
commonmark-rules.js has 13 rules: paragraph, line break, headings (setext for h1–h2 unless atx), blockquote, lists, list items (with start for <ol>), indented and fenced code blocks (the fence grows when the code contains backticks, and the language-* class sets the info string), horizontal rule, inline and reference links, emphasis, strong, inline code and images (commonmark-rules.js). Tables, strikethrough and task lists are not included. They come from the separate turndown-plugin-gfm package.
The default rule
An element with no matching rule keeps its converted children, with blank lines around it if it is a block element (turndown.js). For scraped pages this is the most important behaviour. Because the default rules keep the text of any element they do not match, <script> and <style> bodies end up as plain text, a <table> flattens to one paragraph per cell, and <nav> links are kept. Plan to remove(['script', 'style', 'noscript']) and either add the GFM table plugin or keep('table').
Escaping
escapeMarkdown runs 13 regexes over every text node. It backslash-escapes \ * _ \ [ ]everywhere and- + = # > ~~~ 1.at the start of a line ([utilities.js](https://github.com/mixmark-io/turndown/blob/aa84dfa3e2361edea8c43acbfc2b7a9363494bfd/src/utilities.js#L82-L102)). This keeps the output correct when rendered, but it adds many backslashes to prose with underscores or asterisks. For LLM input, which is never rendered, you can overrideTurndownService.prototype.escape` with a lighter function.
Extending it
addRule(key, { filter, replacement }). Adds a rule ahead of the built-ins. Thereplacementfunction gets the converted content, the DOM node and the options. Anappendhook can add text at the end of the document.keep(filter)/remove(filter).keepoutputs a node as raw HTML.removedrops it with its children. Both apply only to nodes no other rule matched.use(plugin). A plugin is a function that receives the service and calls the methods above.turndown-plugin-gfmis the usual one.- Options.
blankReplacement,keepReplacementanddefaultReplacementchange the three special rules.preformattedCodekeeps whitespace inside<code>.
Running it
npm install turndown. Node 18 or newer, according toengines. The npm package has CJS, ES and UMD builds, plus browser builds that leave out domino through thebrowserfield inpackage.json. An IIFE file indist/works with a<script>tag.- No services, network or configuration are needed. Conversion is synchronous and runs in the calling thread. Wrap it in a worker if you convert large pages on a busy server.
- Security: in Node, domino does not run scripts or load external resources. In a browser, the native parser decides.
SECURITY.mdadvises sanitizing untrusted HTML first, because the parser’s side effects happen before Turndown runs.
Strengths and caveats
- Strength: small and predictable. One dependency, synchronous, deterministic output, and a short rule list that is easy to read and override.
- Strength: correct Markdown. Fence lengths adapt to the code, link destinations with spaces or brackets are escaped, and whitespace handling copies browser rendering.
- Strength: works everywhere. The same API runs in Node, browsers, extensions and edge runtimes.
- Caveat: no content extraction. It converts the whole input, including boilerplate and script text. You need a readability step or selectors before it.
- Caveat: CommonMark only by default. Tables are flattened unless you add a plugin. This matters a lot for data-heavy pages fed to an LLM.
- Caveat: aggressive escaping. The backslashes cost tokens and can confuse a model. Override
escapeif the output is only for machines. - Caveat: shared state in reference links. The
referenceLinkrule stores pending references on the shared rule object and clears them inappend(commonmark-rules.js). This is safe in synchronous use, but it ties together all instances that uselinkStyle: 'referenced'.
Sources: code at aa84dfa, OpenDeepWiki wiki (7 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit aa84dfa. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
not applicableTurndown does not fetch or render pages. It takes an HTML string (or a pre-parsed DOM node) as input and converts it to Markdown. There is no HTTP client, no headless browser integration, no JavaScript rendering, no waiting strategy, and no support for PDF or image content types. The library's entry point, turndown(), accepts either a string or an HTMLElement/Document/DocumentFragment node (src/turndown.js:47-58). When a string is provided, it is parsed into a DOM tree using either the native browser parser or, in Node.js, the domino library — a minimal DOM implementation that does not execute scripts or fetch external resources (src/html-parser.js:49-53). The SECURITY.md explicitly warns that Turndown itself does not fetch anything and that any external resource loading would be the DOM parser's behavior, not Turndown's. In short, the project is an offline format converter, not a scraper.
How is content extracted or converted?
not applicableTurndown converts entire HTML documents or fragments to Markdown; it does not perform content extraction, readability-style boilerplate removal, or schema-based extraction. The conversion is driven by a rule system where each HTML element type has a corresponding replacement function that produces Markdown. Rules match by tag name, array of tag names, or a custom filter function (src/rules.js:69-80). The built-in CommonMark rules cover paragraphs, headings, blockquotes, lists, code blocks, horizontal rules, links (inline and reference styles), images, emphasis, strong, and inline code (src/commonmark-rules.js:1-271). Users can extend or override rules via addRule(), keep(), or remove() (src/turndown.js:88-117). A pre-processing step collapses extraneous whitespace before conversion (src/collapse-whitespace.js:33-98). There is no concept of selecting a subset of content by CSS selector (for extraction purposes) or inferring a page's main content — every element in the input is either converted, kept as HTML, or removed, based on the rules configured.
<script> and <style> bodies appear in the Markdown unless you call remove(['script', 'style']); tables are flattened to cell text unless a plugin such as turndown-plugin-gfm is added. For scraped pages this means you must strip boilerplate before or during conversion.How are LLMs used, if at all?
not applicableTurndown contains no LLM usage of any kind. There are no prompts, no chunking of large pages, no structured output schemas, no LLM provider integrations (OpenAI, Anthropic, or others), and no cost controls. The project is a deterministic HTML-to-Markdown converter implemented entirely in JavaScript with a rule-based approach — every replacement decision is made by matching HTML elements against predefined filters and running the corresponding replacement functions. There are no dependencies on any AI or ML libraries in package.json; the sole runtime dependency is @mixmark-io/domino for DOM parsing. The library's only output format is Markdown text, and its output is purely a function of its input rules and the HTML it receives.
How are anti-bot measures, proxies and fingerprinting handled?
not applicableTurndown implements no anti-bot measures whatsoever. There are no stealth patches, no fingerprint spoofing, no proxy rotation, no CAPTCHA handling, and no rate limiting. The library does not make network requests of any kind — it is a pure format conversion library operating on in-memory HTML strings or DOM nodes. The src/ directory contains no references to proxies, headers, IP addresses, user-agent strings, or any networking concepts. While the package-lock.json file lists http-proxy-agent and https-proxy-agent as transitive dependencies (from the test runner or build tools), these are not imported or used anywhere in Turndown's own source code. The project has no need for anti-bot measures because it never contacts external servers.
How is crawling at scale implemented?
not applicableTurndown provides no crawling infrastructure. There are no URL queues, no concurrency/threading models, no URL deduplication, no depth or limit controls, no robots.txt parsing, no politeness delays, and no distributed worker support. The library is invoked synchronously on a single HTML input at a time — it converts whatever is passed to it and returns the Markdown result. There is no concept of following links, discovering pages, or managing a crawl frontier. The entirety of Turndown's public API is the TurndownService class with its turndown(), addRule(), keep(), remove(), and use() methods (src/turndown.js:38-130).
What is the developer interface?
answeredTurndown exposes a library API only — there is no CLI binary (no bin entry in package.json), no REST service, no MCP server, and no user-facing UI beyond an HTML demo page (index.html). The public interface is the TurndownService class, instantiated via new TurndownService(options) or the factory shorthand TurndownService(options) (src/turndown.js:8-10). The primary method is turndown(input) which accepts either an HTML string or a DOM node and returns a Markdown string (src/turndown.js:47-58). Customization happens through addRule(key, rule) for adding conversion rules, keep(filter) to preserve elements as HTML, remove(filter) to strip elements, use(plugin) for plugin-based extensions, and escape(string) for escaping Markdown syntax (src/turndown.js:68-129). Options are set at construction time and include heading style, HR marker, list bullet, code block style, fence character, emphasis/strong delimiters, link style, and reference style (src/turndown.js:11-22). The library ships as multiple build formats (CJS, ES module, UMD, IIFE) for broad compatibility, and a browser-specific variant replaces the domino DOM parser with the native browser DOM (package.json:6-14). The only output format is Markdown text. Language bindings are JavaScript-only — there are no Python, Rust, or other language ports in this repository.