browserless/browserless
Self-hosted Node server that runs Chromium, Firefox and WebKit in Docker behind Puppeteer/Playwright WebSockets and REST APIs.
Overview
Browserless is a Node.js server that runs real browsers in Docker and rents them out over the network. A client connects with stock Puppeteer or Playwright to a WebSocket URL such as ws://host/chromium?token=…, and Browserless launches a fresh browser, pipes the DevTools or Playwright protocol through, and kills the browser when the client disconnects. For callers who do not want a browser library, it also exposes REST endpoints (/content, /scrape, /screenshot, /pdf, /function, /download, /performance) that open a page, run a fixed recipe, and return HTML, JSON or binary.
The value is operational: a concurrency limiter with a queue, per-request timeouts, CPU and memory health gates, token auth, metrics, webhooks for queue, reject and timeout events, OpenAPI docs, a live debugger, and tidy cleanup of profiles and temp directories. It supports Chromium, Chrome, Edge (CDP and Playwright), Firefox and WebKit (Playwright only).
There is no AI in this repository. No LLM calls, no Markdown conversion, no crawler. The README’s LLM-ready crawl, BrowserQL, CAPTCHA solving and residential proxies belong to the hosted product. In an AI scraping stack, Browserless is the browser tier that an agent or extraction library drives. Licensing matters here: the code is SSPL-1.0 or a commercial licence, so closed-source commercial use needs a paid licence.
Architecture
flowchart LR
C["Puppeteer, Playwright or HTTP client"] --> S["HTTPServer (server.ts)"]
S --> T["Token auth + AJV schema check"]
T --> R["Router: route match"]
R --> L["Limiter: concurrency + queue + timeout"]
L --> BM["BrowserManager.getBrowserForRequest"]
BM --> B["ChromiumCDP / *Playwright launch"]
R --> H["Route handler"]
H -->|"WebSocket"| P["http-proxy ws pipe"]
H -->|"REST"| PG["browser.newPage recipe"]
P --> BR["Browser process"]
PG --> BR
BM --> CL["complete(): close + cleanup"]
| Component | Path | Role |
|---|---|---|
| Entry and composition | src/index.ts, src/browserless.ts |
Builds config, limiter, router, browser manager; dynamically imports route files; starts the server |
| HTTP server | src/server.ts |
Request and upgrade handling, CORS, token check, body and query validation |
| Router | src/router.ts |
Path matching, wraps handlers with the limiter and browser acquisition |
| Limiter | src/limiter.ts |
queue-based admission control, health checks, timeouts, metrics and webhooks |
| Browser manager | src/browsers/index.ts |
Launch options parsing, session registry, reconnects, close and cleanup |
| Browser classes | src/browsers/browsers.cdp.ts, browsers.playwright.ts |
Launch Chromium-family over CDP (optionally with stealth) or Playwright browser servers; proxy WebSockets |
| Shared routes | src/shared/*.ts |
WebSocket and REST handlers reused by each browser’s src/routes/<browser>/ stubs |
| Network guard | src/network-security.ts |
Blocks file://, private ranges and cloud metadata hosts |
| Docker | docker/ |
Base image plus one image per browser, a multi-browser image and an SDK image |
How a request flows
Take puppeteer.connect({ browserWSEndpoint: "ws://host:3000/chromium?token=T&launch={\"stealth\":true}" }):
- Upgrade. The server sees
Upgrade: websocket, moves the token from the query to a header, runs thebeforehook, parses the URL and finds the WebSocket route (server.ts). - Auth and validate. If the route has
auth(orSTRICT_TOKEN_USEis on),Token.isAuthorizedruns, then query parameters are validated against the JSON schema generated from the route’s TypeScriptQuerySchema(server.ts). - Admit. Routes with
concurrency = trueare wrapped byLimiter.limit, which rejects with 429 when health checks fail orCONCURRENT + QUEUEDis exceeded, and otherwise runs or queues the job under the global timeout (router.ts, limiter.ts). - Launch.
getBrowserForRequestdecodeslaunch(JSON or base64), merges route defaults, folds--proxy-serverinto args (or into Playwright’sproxyoption), creates a user-data dir and a per-sessionTMPDIR, and callsbrowser.launch(browsers/index.ts).ChromiumCDP.launchpicks a free port, adds--remote-debugging-port,--no-sandboxand the uBlock extension whenblockAdsis set, and usespuppeteer-extrawith the stealth plugin whenstealthis true (browsers.cdp.ts). - Pipe. The route handler is one line,
browser.proxyWebSocket(req, socket, head)(chromium.ws.ts), which useshttp-proxyto splice the client socket onto the browser’s own DevTools endpoint and resolves when either side closes (browsers.cdp.ts). - Clean up. When the handler resolves, the router calls
browserManager.complete, which decrements connected clients and closes the browser unless another client is attached or a keep-until timer is set (router.ts, browsers/index.ts).
A REST call follows steps 2 to 4 and 6, but the handler opens a page itself. /content applies cookies, viewport, user agent, headers and request interception, navigates (or setContent for raw HTML), runs the optional waits, sets X-Response-* headers from the navigation response, and returns page.content() (content.http.ts).
Key components
Route system
Every endpoint is a class with path, method, auth, concurrency, an optional browser, and a handler. At startup Browserless.start imports every compiled route file plus any the SDK user added, loads sibling *.body.json and *.query.json schemas generated from the TypeScript types, skips routes whose browser cannot run on arm64 Linux, refuses to start if a route needs a browser binary that is not installed, and registers the rest (browserless.ts). Per-browser folders under src/routes/ mostly re-export the shared handlers in src/shared/.
Validation
Bodies and query strings are validated with AJV in allErrors mode with no type coercion and no removal of extra properties, after a MAX_PAYLOAD_SIZE cap (10 MB by default) (schema-validator.ts, server.ts). Unknown or mistyped fields fail with a 400 that lists every error.
Limiter
Limiter extends the queue package with concurrency = CONCURRENT (default 10), an extra QUEUED allowance (default 10) and a TIMEOUT (default 30 s) (config.ts, limiter.ts). Success, timeout and error events feed metrics, after hooks and optional alert webhooks. A route can override the timeout with ?timeout= or bypass limits through a bypassLimits predicate.
REST recipes
/scrape waits for each selector (polling every 100 ms) and returns, for every match, innerText, innerHTML, attributes and geometry (scrape.http.ts). /function runs a user-supplied Puppeteer function against a blank page and infers the response content type from the returned value or its file signature (function.http.ts). These are thin, predictable wrappers, not extractors.
Stealth and network policy
The only anti-detection in the open-source code is @zorilla/puppeteer-extra-plugin-stealth, enabled per session with launch.stealth on CDP routes (browsers.cdp.ts). Proxies are plain Chrome flags you supply. In the other direction, every CDP page gets a Network.setBlockedURLs guard built from blocked URL patterns and private network ranges, so a client cannot steer the browser at the host’s internal network (browsers.cdp.ts).
Extending it
- SDK mode. The npm package exports
Browserlessand its parts. Pass your ownConfig,Hooks,Limiter,Router,TokenorBrowserManagerto the constructor, and add routes withaddHTTPRouteoraddWebSocketRoute(browserless.ts). Adocker/sdkimage builds such a project. - New endpoint. Write a
BrowserHTTPRoutesubclass with exportedBodySchema/QuerySchematypes; the build turns them into JSON schemas and OpenAPI entries automatically. - Hooks.
before,after,browserandpagehooks let you add auth, logging, billing or page setup without forking. - Disable features.
disableRoutes(...names)removes built-in routes, for example/functionon a multi-tenant deployment.
Running it
- Docker. Pick an image per browser (
chromium,chrome,edge,firefox,webkit) ormulti. The base image is Ubuntu 26.04 and listens onPORT=3000. Typical settings areTOKEN,CONCURRENT,QUEUED,TIMEOUT,HEALTH=true,CORS, andSTRICT_TOKEN_USE. - From source. Node with
npm install,npm run install:browsers(Playwright downloads),npm run build, thennpm start. Docs are served at/docs; the debugger at/debuggerwhen installed. - Clients. Puppeteer connects to
/or/chromium; Playwright connects to/<browser>/playwright. The Playwright client version is read from theUser-Agentheader to choose a matching Playwright server build.
Strengths and caveats
- Strength: drop-in for existing scripts. Change
launch()toconnect()and a Puppeteer or Playwright script runs remotely with queueing and limits. - Strength: production hygiene. Strict schemas, health-gated admission, per-session temp dirs, orphan-browser cleanup, SSRF-style network blocking and metrics are already done.
- Strength: many engines. CDP and Playwright protocols across five browser families from one server.
- Caveat: one browser per connection. Every new session launches a fresh browser process. That is clean and isolated, but cold-start costs are paid each time; the keep-until timer exists but nothing in this repo sets it.
- Caveat: no data layer. No Markdown, no readability, no LLM extraction, no crawler, no proxy rotation and no CAPTCHA solving. Those are cloud features or your job.
- Caveat: small REST quirks.
/contentappliessetJavaScriptEnabledonly when it is truthy, so passingfalsedoes not disable JavaScript (content.http.ts). - Caveat: licence. SSPL-1.0 or a commercial licence. Check whether SSPL’s terms fit your use before embedding it in a product or a hosted service.
Sources: code at f7edf0f, deepwiki-open wiki (10 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit f7edf0f. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPages are fetched through a headless browser — no plain-HTTP fetch path exists. The system launches real browser processes (Chromium CDP, Firefox Playwright, WebKit Playwright) inside Docker, then proxies WebSocket traffic between the client's Puppeteer/Playwright library and the browser's CDP or Playwright JSON-RPC interface (src/browsers/browsers.cdp.ts:378-484, src/browsers/browsers.playwright.ts:212-253). JS rendering is on by default (Chromium runs JavaScript), and can be disabled per-request via setJavaScriptEnabled: false on the REST APIs. Waiting strategies are configurable: clients can set gotoOptions (waitUntil events like domcontentloaded), waitForTimeout (millisecond sleep), waitForSelector (element visible/hidden), waitForFunction (JS expression in the page), or waitForEvent (DOM/window event) — all available on the /content, /scrape, /screenshot, and /pdf endpoints (src/shared/content.http.ts:230-251). Webrtc, the content types served are HTML (from /content), application/pdf (via Chromium's built-in PDF renderer in /pdf), PNG/JPEG images (from /screenshot), and arbitrary binary downloads (from /download). The REST API route handlers each open a new browser page (browser.newPage()) and close it in a finally block to prevent leaks (src/shared/content.http.ts:140,272-275). The WebSocket routes (/chromium, /firefox/playwright, etc.) proxy the raw browser protocol stream bidirectionally, letting clients use their own Puppeteer/Playwright scripts (src/shared/chromium.ws.ts:18-35).
if (setJavaScriptEnabled) await page.setJavaScriptEnabled(...), so passing false is ignored and JavaScript cannot be turned off this way; every new session also launches a fresh browser process rather than reusing one.How is content extracted or converted?
answeredContent extraction is provided by three REST endpoints all using Puppeteer's Page API. /content returns the full serialized HTML via getPageContent() (src/shared/content.http.ts:267), which wraps page.content() with retry logic for in-flight navigation teardown (src/utils.ts:909-932). /scrape evaluates a document.querySelectorAll script in the browser context (src/shared/scrape.http.ts:163-223), returning each matching element's innerText, innerHTML, attributes, and bounding-box geometry (top/left/width/height) — configurable by CSS selectors with per-selector timeouts and a bestAttempt mode that tolerates missing elements. There is no built-in HTML→Markdown conversion or readability-style boilerplate removal in the open-source code (the README's "LLM-ready data" claim refers to the premium /crawl cloud API). Schema-based extraction is supported through TypeScript BodySchema / ResponseSchema interfaces that compile to JSON Schema for request validation (src/types.ts:124-138) — validated via AJV at runtime in src/server.ts:340-384. The /function endpoint (src/shared/function.http.ts:59-97) runs arbitrary Puppeteer JavaScript code in the browser context, returning any value (HTML, JSON, binary) back to the caller. No selectors or extraction heuristics are applied there — it is a raw code-execution sandbox.
How are LLMs used, if at all?
not applicableThe open-source repository contains no LLM integration code — no calls to OpenAI, Anthropic, or any other LLM provider; no prompting, chunking, structured output schemas, or cost controls. The README advertises "LLM-ready data" as a feature of the premium /crawl API (cloud/enterprise only) and mentions an MCP Server for AI assistant connectivity, but neither implementation exists in the src/ directory. The only LLM-adjacent artifact is a static/docs/docs.js file that mentions "markdown" and "LLM" in its documentation strings for the cloud API. The open-source project is a pure browser-automation service — LLM features are entirely in the proprietary cloud/enterprise tier.
How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-detection in the open-source code centers on the @zorilla/puppeteer-extra-plugin-stealth integration. When the stealth: true launch parameter is passed (via query param or CDPLaunchOptions), ChromiumCDP.launch() swaps from plain puppeteer.launch to puppeteerStealth.launch, which applies puppeteer-extra's stealth patches — spoofing browser fingerprint headers, TCP behavior, and WebSocket framing to evade bot detection (src/browsers/browsers.cdp.ts:32-36, 447-448). The legacy shim.ts lifts stealth out of top-level query params into the launch object (src/shim.ts:90-92). Proxy support is available through --proxy-server=URL and --proxy-bypass-list launch arguments, automatically translated to Playwright's proxy config object on Playwright paths (src/browsers/index.ts:707-717, 789-796). The CDP browser classes create a local http-proxy server per instance for WebSocket forwarding (src/browsers/browsers.cdp.ts:73, 575-595). Fingerprint spoofing beyond the stealth plugin is not built in. Navigation security is enforced via the Config.getBlockedURLPatterns() and Config.getBlockedNetworkRanges() system (src/config.ts:416-430), which uses Network.setBlockedURLs in CDP (src/browsers/browsers.cdp.ts:155-195) and per-frame JSON-RPC inspection in Playwright (src/browsers/browsers.playwright.ts:424-485) to block file:// schemes and private-network ranges (loopback, link-local, cloud metadata). The premium anti-bot features — BrowserQL, residential proxy rotation, CAPTCHA solving, fingerprint randomization — are only in the cloud/enterprise tier (README claims).
How is crawling at scale implemented?
not applicableThere is no crawling framework in the open-source codebase. The README describes /crawl, /map (sitemap discovery), and /search APIs as premium cloud/enterprise features. The open-source project handles only individual page requests through REST endpoints and WebSocket browser connections. Concurrency management is handled by the Limiter class (src/limiter.ts:29-298), which uses the queue library to enforce configurable CONCURRENT and QUEUED limits with timeout, health-check gating (CPU/memory), and webhook alerts — but this is a per-request admission queue, not a crawling pipeline. No URL deduplication, depth limiting, politeness delays, robots.txt parsing, or distributed worker coordination exists.
What is the developer interface?
answeredThe developer interface is a self-hosted Docker service with multiple access patterns. WebSocket protocol — clients connect with standard Puppeteer (puppeteer-core) at ws://host/ or Playwright at ws://host/{browser}/playwright, and the server proxies traffic to a real browser process (src/shared/chromium.ws.ts:18-35; src/routes/chromium/ws/playwright.ts). REST APIs — seven POST endpoints (/content, /scrape, /screenshot, /pdf, /function, /download, /performance) plus management GET endpoints (/sessions, /config, /pressure, /metrics, /active, /kill), all JSON-validated via AJV on schema derived from TypeScript interfaces (src/server.ts:340-384). Output formats: HTML, JSON, PDF binary, PNG/JPEG binary, or any content type returned from custom Puppeteer functions. OpenAPI docs at /docs and the Debug Viewer at /debugger for inspecting live sessions (src/config.ts:263-264). Authentication is token-based (TOKEN env var), optionally strict-mode requiring it on every route (src/config.ts:250-251). CLI entry at bin/browserless.js (packaged as browserless binary). SDK extension via @browserless.io/browserless allows overriding config, hooks, limiter, metrics, and routes; the Browserless class accepts injected modules in its constructor (src/browserless.ts:83-138). CORS is configurable for cross-origin requests (src/config.ts:268-283, 895-911). The MCP Server mentioned in the README is not part of the open-source code — it is a cloud-only integration.