LLMs Technical Reviews
Home / AI web scraping / browserless

browserless/browserless

Self-hosted Node server that runs Chromium, Firefox and WebKit in Docker behind Puppeteer/Playwright WebSockets and REST APIs.

GitHub ↗★ 14kTypeScriptSSPL-1.0 OR commercialcommit f7edf0f · 2026-10-06homepage ↗

Overview

Browserless is a Node.js server that runs real browsers in Docker and rents them out over the network. A client connects with stock Puppeteer or Playwright to a WebSocket URL such as ws://host/chromium?token=…, and Browserless launches a fresh browser, pipes the DevTools or Playwright protocol through, and kills the browser when the client disconnects. For callers who do not want a browser library, it also exposes REST endpoints (/content, /scrape, /screenshot, /pdf, /function, /download, /performance) that open a page, run a fixed recipe, and return HTML, JSON or binary.

The value is operational: a concurrency limiter with a queue, per-request timeouts, CPU and memory health gates, token auth, metrics, webhooks for queue, reject and timeout events, OpenAPI docs, a live debugger, and tidy cleanup of profiles and temp directories. It supports Chromium, Chrome, Edge (CDP and Playwright), Firefox and WebKit (Playwright only).

There is no AI in this repository. No LLM calls, no Markdown conversion, no crawler. The README’s LLM-ready crawl, BrowserQL, CAPTCHA solving and residential proxies belong to the hosted product. In an AI scraping stack, Browserless is the browser tier that an agent or extraction library drives. Licensing matters here: the code is SSPL-1.0 or a commercial licence, so closed-source commercial use needs a paid licence.

Architecture

flowchart LR
  C["Puppeteer, Playwright or HTTP client"] --> S["HTTPServer (server.ts)"]
  S --> T["Token auth + AJV schema check"]
  T --> R["Router: route match"]
  R --> L["Limiter: concurrency + queue + timeout"]
  L --> BM["BrowserManager.getBrowserForRequest"]
  BM --> B["ChromiumCDP / *Playwright launch"]
  R --> H["Route handler"]
  H -->|"WebSocket"| P["http-proxy ws pipe"]
  H -->|"REST"| PG["browser.newPage recipe"]
  P --> BR["Browser process"]
  PG --> BR
  BM --> CL["complete(): close + cleanup"]
Component Path Role
Entry and composition src/index.ts, src/browserless.ts Builds config, limiter, router, browser manager; dynamically imports route files; starts the server
HTTP server src/server.ts Request and upgrade handling, CORS, token check, body and query validation
Router src/router.ts Path matching, wraps handlers with the limiter and browser acquisition
Limiter src/limiter.ts queue-based admission control, health checks, timeouts, metrics and webhooks
Browser manager src/browsers/index.ts Launch options parsing, session registry, reconnects, close and cleanup
Browser classes src/browsers/browsers.cdp.ts, browsers.playwright.ts Launch Chromium-family over CDP (optionally with stealth) or Playwright browser servers; proxy WebSockets
Shared routes src/shared/*.ts WebSocket and REST handlers reused by each browser’s src/routes/<browser>/ stubs
Network guard src/network-security.ts Blocks file://, private ranges and cloud metadata hosts
Docker docker/ Base image plus one image per browser, a multi-browser image and an SDK image

How a request flows

Take puppeteer.connect({ browserWSEndpoint: "ws://host:3000/chromium?token=T&launch={\"stealth\":true}" }):

  1. Upgrade. The server sees Upgrade: websocket, moves the token from the query to a header, runs the before hook, parses the URL and finds the WebSocket route (server.ts).
  2. Auth and validate. If the route has auth (or STRICT_TOKEN_USE is on), Token.isAuthorized runs, then query parameters are validated against the JSON schema generated from the route’s TypeScript QuerySchema (server.ts).
  3. Admit. Routes with concurrency = true are wrapped by Limiter.limit, which rejects with 429 when health checks fail or CONCURRENT + QUEUED is exceeded, and otherwise runs or queues the job under the global timeout (router.ts, limiter.ts).
  4. Launch. getBrowserForRequest decodes launch (JSON or base64), merges route defaults, folds --proxy-server into args (or into Playwright’s proxy option), creates a user-data dir and a per-session TMPDIR, and calls browser.launch (browsers/index.ts). ChromiumCDP.launch picks a free port, adds --remote-debugging-port, --no-sandbox and the uBlock extension when blockAds is set, and uses puppeteer-extra with the stealth plugin when stealth is true (browsers.cdp.ts).
  5. Pipe. The route handler is one line, browser.proxyWebSocket(req, socket, head) (chromium.ws.ts), which uses http-proxy to splice the client socket onto the browser’s own DevTools endpoint and resolves when either side closes (browsers.cdp.ts).
  6. Clean up. When the handler resolves, the router calls browserManager.complete, which decrements connected clients and closes the browser unless another client is attached or a keep-until timer is set (router.ts, browsers/index.ts).

A REST call follows steps 2 to 4 and 6, but the handler opens a page itself. /content applies cookies, viewport, user agent, headers and request interception, navigates (or setContent for raw HTML), runs the optional waits, sets X-Response-* headers from the navigation response, and returns page.content() (content.http.ts).

Key components

Route system

Every endpoint is a class with path, method, auth, concurrency, an optional browser, and a handler. At startup Browserless.start imports every compiled route file plus any the SDK user added, loads sibling *.body.json and *.query.json schemas generated from the TypeScript types, skips routes whose browser cannot run on arm64 Linux, refuses to start if a route needs a browser binary that is not installed, and registers the rest (browserless.ts). Per-browser folders under src/routes/ mostly re-export the shared handlers in src/shared/.

Validation

Bodies and query strings are validated with AJV in allErrors mode with no type coercion and no removal of extra properties, after a MAX_PAYLOAD_SIZE cap (10 MB by default) (schema-validator.ts, server.ts). Unknown or mistyped fields fail with a 400 that lists every error.

Limiter

Limiter extends the queue package with concurrency = CONCURRENT (default 10), an extra QUEUED allowance (default 10) and a TIMEOUT (default 30 s) (config.ts, limiter.ts). Success, timeout and error events feed metrics, after hooks and optional alert webhooks. A route can override the timeout with ?timeout= or bypass limits through a bypassLimits predicate.

REST recipes

/scrape waits for each selector (polling every 100 ms) and returns, for every match, innerText, innerHTML, attributes and geometry (scrape.http.ts). /function runs a user-supplied Puppeteer function against a blank page and infers the response content type from the returned value or its file signature (function.http.ts). These are thin, predictable wrappers, not extractors.

Stealth and network policy

The only anti-detection in the open-source code is @zorilla/puppeteer-extra-plugin-stealth, enabled per session with launch.stealth on CDP routes (browsers.cdp.ts). Proxies are plain Chrome flags you supply. In the other direction, every CDP page gets a Network.setBlockedURLs guard built from blocked URL patterns and private network ranges, so a client cannot steer the browser at the host’s internal network (browsers.cdp.ts).

Extending it

  • SDK mode. The npm package exports Browserless and its parts. Pass your own Config, Hooks, Limiter, Router, Token or BrowserManager to the constructor, and add routes with addHTTPRoute or addWebSocketRoute (browserless.ts). A docker/sdk image builds such a project.
  • New endpoint. Write a BrowserHTTPRoute subclass with exported BodySchema/QuerySchema types; the build turns them into JSON schemas and OpenAPI entries automatically.
  • Hooks. before, after, browser and page hooks let you add auth, logging, billing or page setup without forking.
  • Disable features. disableRoutes(...names) removes built-in routes, for example /function on a multi-tenant deployment.

Running it

  • Docker. Pick an image per browser (chromium, chrome, edge, firefox, webkit) or multi. The base image is Ubuntu 26.04 and listens on PORT=3000. Typical settings are TOKEN, CONCURRENT, QUEUED, TIMEOUT, HEALTH=true, CORS, and STRICT_TOKEN_USE.
  • From source. Node with npm install, npm run install:browsers (Playwright downloads), npm run build, then npm start. Docs are served at /docs; the debugger at /debugger when installed.
  • Clients. Puppeteer connects to / or /chromium; Playwright connects to /<browser>/playwright. The Playwright client version is read from the User-Agent header to choose a matching Playwright server build.

Strengths and caveats

  • Strength: drop-in for existing scripts. Change launch() to connect() and a Puppeteer or Playwright script runs remotely with queueing and limits.
  • Strength: production hygiene. Strict schemas, health-gated admission, per-session temp dirs, orphan-browser cleanup, SSRF-style network blocking and metrics are already done.
  • Strength: many engines. CDP and Playwright protocols across five browser families from one server.
  • Caveat: one browser per connection. Every new session launches a fresh browser process. That is clean and isolated, but cold-start costs are paid each time; the keep-until timer exists but nothing in this repo sets it.
  • Caveat: no data layer. No Markdown, no readability, no LLM extraction, no crawler, no proxy rotation and no CAPTCHA solving. Those are cloud features or your job.
  • Caveat: small REST quirks. /content applies setJavaScriptEnabled only when it is truthy, so passing false does not disable JavaScript (content.http.ts).
  • Caveat: licence. SSPL-1.0 or a commercial licence. Check whether SSPL’s terms fit your use before embedding it in a product or a hosted service.

Sources: code at f7edf0f, deepwiki-open wiki (10 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit f7edf0f. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched through a headless browser — no plain-HTTP fetch path exists. The system launches real browser processes (Chromium CDP, Firefox Playwright, WebKit Playwright) inside Docker, then proxies WebSocket traffic between the client's Puppeteer/Playwright library and the browser's CDP or Playwright JSON-RPC interface (src/browsers/browsers.cdp.ts:378-484, src/browsers/browsers.playwright.ts:212-253). JS rendering is on by default (Chromium runs JavaScript), and can be disabled per-request via setJavaScriptEnabled: false on the REST APIs. Waiting strategies are configurable: clients can set gotoOptions (waitUntil events like domcontentloaded), waitForTimeout (millisecond sleep), waitForSelector (element visible/hidden), waitForFunction (JS expression in the page), or waitForEvent (DOM/window event) — all available on the /content, /scrape, /screenshot, and /pdf endpoints (src/shared/content.http.ts:230-251). Webrtc, the content types served are HTML (from /content), application/pdf (via Chromium's built-in PDF renderer in /pdf), PNG/JPEG images (from /screenshot), and arbitrary binary downloads (from /download). The REST API route handlers each open a new browser page (browser.newPage()) and close it in a finally block to prevent leaks (src/shared/content.http.ts:140,272-275). The WebSocket routes (/chromium, /firefox/playwright, etc.) proxy the raw browser protocol stream bidirectionally, letting clients use their own Puppeteer/Playwright scripts (src/shared/chromium.ws.ts:18-35).

Editor's note. Correction: on /content the code runs if (setJavaScriptEnabled) await page.setJavaScriptEnabled(...), so passing false is ignored and JavaScript cannot be turned off this way; every new session also launches a fresh browser process rather than reusing one.

How is content extracted or converted?

answered

Content extraction is provided by three REST endpoints all using Puppeteer's Page API. /content returns the full serialized HTML via getPageContent() (src/shared/content.http.ts:267), which wraps page.content() with retry logic for in-flight navigation teardown (src/utils.ts:909-932). /scrape evaluates a document.querySelectorAll script in the browser context (src/shared/scrape.http.ts:163-223), returning each matching element's innerText, innerHTML, attributes, and bounding-box geometry (top/left/width/height) — configurable by CSS selectors with per-selector timeouts and a bestAttempt mode that tolerates missing elements. There is no built-in HTML→Markdown conversion or readability-style boilerplate removal in the open-source code (the README's "LLM-ready data" claim refers to the premium /crawl cloud API). Schema-based extraction is supported through TypeScript BodySchema / ResponseSchema interfaces that compile to JSON Schema for request validation (src/types.ts:124-138) — validated via AJV at runtime in src/server.ts:340-384. The /function endpoint (src/shared/function.http.ts:59-97) runs arbitrary Puppeteer JavaScript code in the browser context, returning any value (HTML, JSON, binary) back to the caller. No selectors or extraction heuristics are applied there — it is a raw code-execution sandbox.

How are LLMs used, if at all?

not applicable

The open-source repository contains no LLM integration code — no calls to OpenAI, Anthropic, or any other LLM provider; no prompting, chunking, structured output schemas, or cost controls. The README advertises "LLM-ready data" as a feature of the premium /crawl API (cloud/enterprise only) and mentions an MCP Server for AI assistant connectivity, but neither implementation exists in the src/ directory. The only LLM-adjacent artifact is a static/docs/docs.js file that mentions "markdown" and "LLM" in its documentation strings for the cloud API. The open-source project is a pure browser-automation service — LLM features are entirely in the proprietary cloud/enterprise tier.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-detection in the open-source code centers on the @zorilla/puppeteer-extra-plugin-stealth integration. When the stealth: true launch parameter is passed (via query param or CDPLaunchOptions), ChromiumCDP.launch() swaps from plain puppeteer.launch to puppeteerStealth.launch, which applies puppeteer-extra's stealth patches — spoofing browser fingerprint headers, TCP behavior, and WebSocket framing to evade bot detection (src/browsers/browsers.cdp.ts:32-36, 447-448). The legacy shim.ts lifts stealth out of top-level query params into the launch object (src/shim.ts:90-92). Proxy support is available through --proxy-server=URL and --proxy-bypass-list launch arguments, automatically translated to Playwright's proxy config object on Playwright paths (src/browsers/index.ts:707-717, 789-796). The CDP browser classes create a local http-proxy server per instance for WebSocket forwarding (src/browsers/browsers.cdp.ts:73, 575-595). Fingerprint spoofing beyond the stealth plugin is not built in. Navigation security is enforced via the Config.getBlockedURLPatterns() and Config.getBlockedNetworkRanges() system (src/config.ts:416-430), which uses Network.setBlockedURLs in CDP (src/browsers/browsers.cdp.ts:155-195) and per-frame JSON-RPC inspection in Playwright (src/browsers/browsers.playwright.ts:424-485) to block file:// schemes and private-network ranges (loopback, link-local, cloud metadata). The premium anti-bot features — BrowserQL, residential proxy rotation, CAPTCHA solving, fingerprint randomization — are only in the cloud/enterprise tier (README claims).

How is crawling at scale implemented?

not applicable

There is no crawling framework in the open-source codebase. The README describes /crawl, /map (sitemap discovery), and /search APIs as premium cloud/enterprise features. The open-source project handles only individual page requests through REST endpoints and WebSocket browser connections. Concurrency management is handled by the Limiter class (src/limiter.ts:29-298), which uses the queue library to enforce configurable CONCURRENT and QUEUED limits with timeout, health-check gating (CPU/memory), and webhook alerts — but this is a per-request admission queue, not a crawling pipeline. No URL deduplication, depth limiting, politeness delays, robots.txt parsing, or distributed worker coordination exists.

What is the developer interface?

answered

The developer interface is a self-hosted Docker service with multiple access patterns. WebSocket protocol — clients connect with standard Puppeteer (puppeteer-core) at ws://host/ or Playwright at ws://host/{browser}/playwright, and the server proxies traffic to a real browser process (src/shared/chromium.ws.ts:18-35; src/routes/chromium/ws/playwright.ts). REST APIs — seven POST endpoints (/content, /scrape, /screenshot, /pdf, /function, /download, /performance) plus management GET endpoints (/sessions, /config, /pressure, /metrics, /active, /kill), all JSON-validated via AJV on schema derived from TypeScript interfaces (src/server.ts:340-384). Output formats: HTML, JSON, PDF binary, PNG/JPEG binary, or any content type returned from custom Puppeteer functions. OpenAPI docs at /docs and the Debug Viewer at /debugger for inspecting live sessions (src/config.ts:263-264). Authentication is token-based (TOKEN env var), optionally strict-mode requiring it on every route (src/config.ts:250-251). CLI entry at bin/browserless.js (packaged as browserless binary). SDK extension via @browserless.io/browserless allows overriding config, hooks, limiter, metrics, and routes; the Browserless class accepts injected modules in its constructor (src/browserless.ts:83-138). CORS is configurable for cross-origin requests (src/config.ts:268-283, 895-911). The MCP Server mentioned in the README is not part of the open-source code — it is a cloud-only integration.