LLMs Technical Reviews
Home / Browser & computer control / browser-harness

browser-use/browser-harness

Python CDP daemon and pre-imported helpers that let a coding agent drive your real Chrome from heredoc scripts or MCP tools.

GitHub ↗★ 18kPythonMITcommit afbcc38 · 2026-09-07homepage ↗

Overview

browser-harness is a thin layer between a coding agent (Claude Code, Codex, Cursor and similar) and a Chromium browser. It contains no LLM, no planner and no page serializer. The agent writes a short Python script, pipes it into the browser-harness CLI as a heredoc, and the script calls pre-imported helpers such as new_tab, click_at_xy, js and cdp. A long-lived daemon holds the Chrome DevTools Protocol (CDP) WebSocket open between invocations, so each script starts on the same tab the previous one left.

The design bet is that a capable coding agent needs raw CDP and a few sharp helpers, not an abstraction. Element finding is left to the agent: the bundled SKILL.md tells it to filter Accessibility.getFullAXTree, read a box with DOM.getBoxModel, and click the centre with click_at_xy. The “self-healing” in the description means the agent adds its own helpers to an agent_helpers.py file in its workspace. The core package stays fixed.

The default target is the user’s own running Chrome, Edge or Brave, with its logged-in sessions. The same daemon can also attach to any CDP endpoint, or to a Browser Use Cloud browser created through start_remote_daemon(). Most of the code is not browser driving. It handles connection lifecycle: Chrome 144+ “Allow remote debugging” prompts, stale sessions, PID reuse, Windows IPC, cloud billing cleanup, and an optional recording-to-video pipeline.

Architecture

flowchart LR
  A["Coding agent"] -->|"heredoc script"| CLI["run.py: main / exec"]
  A -->|"MCP stdio"| MCP["mcp_server.py tools"]
  CLI --> H["helpers.py"]
  MCP --> H
  CLI --> ADM["admin.py: ensure_daemon"]
  MCP --> ADM
  ADM -->|"spawn"| D["daemon.py: Daemon"]
  H -->|"JSON line over socket"| IPC["_ipc.py"]
  IPC --> D
  D -->|"CDP WebSocket"| B["Chrome / CDP endpoint / cloud"]
  CLI --> REC["recorder.py"]
  CLI --> TEL["telemetry.py"]
  ADM -->|"POST /browsers"| BU["Browser Use Cloud API"]
Component Path Role
CLI runner src/browser_harness/run.py Subcommands, daemon bootstrap, helper tracing, exec of the piped script
Helpers src/browser_harness/helpers.py cdp, input, tabs, js, waits, screenshots, http_get; loads agent_helpers.py
Daemon src/browser_harness/daemon.py Finds the browser, holds the CDP WebSocket, relays requests, buffers events, recovers stale sessions
IPC src/browser_harness/_ipc.py Unix socket (mode 0600) on POSIX; loopback TCP plus a random token on Windows
Admin src/browser_harness/admin.py ensure_daemon, restart and stop, Chrome launch, cloud browsers, doctor, update
MCP server src/mcp_server.py 23 browser_* tools wrapping the same helpers
Recorder and video recorder.py, video.py, video_render.py Per-action JPEG frames and events.jsonl; brief-driven video export
Auth and telemetry auth.py, telemetry.py Browser Use Cloud login (PKCE, device code, API key); opt-out PostHog events
Skills SKILL.md, interaction-skills/, agent-workspace/domain-skills/ Markdown instructions for the agent: 18 interaction topics, about 100 site folders

How a request flows

Take a heredoc script that runs new_tab("https://example.com") and then print(page_info()):

  1. Read the script. main() wraps stdout and stderr to keep their tails, then calls _run. _run reads stdin and, unless a daemon already answers, may auto-start a cloud browser. That needs BU_AUTOSPAWN, cloud auth, no local Chrome on 9222/9223 and no explicit BU_CDP_* endpoint. Otherwise it calls ensure_daemon() (run.py).
  2. Ensure the daemon. If a daemon is alive, ensure_daemon sends a real Target.getTargets through it and returns on a result. A stale cloud daemon is stopped, which also stops its billable browser. Otherwise a detached python -m browser_harness.daemon is spawned under a file lock, and the function watches the daemon log for handshake-wait, “Chrome not running” and permission errors (admin.py).
  3. Connect. get_ws_url() takes BU_CDP_WS or resolves BU_CDP_URL via /json/version. Without those, it reads DevToolsActivePort from every known Chrome, Edge, Brave, Arc and Comet profile directory. It falls back to probing ports 9222 and 9223 (daemon.py). For local Chrome, _PatientCDPClient opens the WebSocket with no handshake timeout, so the “Allow remote debugging?” sheet stays up until the user answers (daemon.py).
  4. Attach. attach_first_page picks the first real page, then a blank tab, then the New Tab Page, and creates about:blank only as a last resort. A named daemon (BU_NAME other than default) gets its own background tab. It enables Page, DOM, Runtime and Network on the session (daemon.py).
  5. Execute. _install_helper_trace wraps every public helper so each call is timed, appended to a 500-entry trace and passed to recorder.observe. Then exec(code, globals()) runs the script with the helpers in scope (run.py, L408-L409).
  6. Relay. Each helper ends in cdp() → _send(), which opens a fresh socket and sends one newline-terminated JSON object (method, params, session_id) with a 5 s response timeout. Screenshots get 60 s (helpers.py, _ipc.py). The daemon’s handle strips the session for Target.* calls and forwards everything else on the attached session with send_raw (daemon.py).
  7. Report. On exit, main() sends a cli_event to telemetry with the outcome, then the process ends. The daemon and its tab stay up for the next script.

Key components

Daemon and IPC

The daemon is one asyncio process per BU_NAME. Besides relaying CDP calls, it answers meta requests: ping (with its PID, so restart_daemon can verify identity before signalling), drain_events, current_tab, set_session, pending_dialog and shutdown. It taps the CDP event registry to keep the last 500 events and to track open JavaScript dialogs (daemon.py). If a call fails with “Session with given id not found”, it re-attaches under a lock and retries once. A session-replacement map stops a delayed request from landing on a tab the user switched to meanwhile (daemon.py). On POSIX the socket is created under umask 077, so it is 0600 from the start. Windows has no equivalent, so the daemon writes a 32-byte hex token to the port file and rejects requests without it (_ipc.py).

Helpers

helpers.py is deliberately small. click_at_xy sends a bare mousePressed/mouseReleased pair with no move or scroll-into-view (helpers.py). type_text is Input.insertText. fill_input focuses by CSS selector, selects all with a raw key event, types key by key through press_key, then fires input and change for React and Vue (helpers.py). press_key maps characters to US-layout code and virtual key values and adds Shift where needed (helpers.py). js evaluates with returnByValue and awaitPromise, and retries wrapped in a function when Chrome rejects a top-level return (helpers.py). wait_for_network_idle is built on drain_events and counts only events from the active session (helpers.py). http_get skips the browser and goes through Browser Use’s fetch-use proxy when BROWSER_USE_API_KEY is set (helpers.py).

Connection self-repair

Most of admin.py deals with the ways a local Chrome connection fails. If the log says Chrome is not running, ensure_daemon launches it, preferring a profile whose remote-debugging toggle is already on (admin.py). If remote debugging is off, it opens chrome://inspect/#remote-debugging at most once every 180 s and raises an instruction for the agent to relay (admin.py). It never opens a second connection while an Allow prompt is pending. The comments explain why: a retry used to create another prompt, so one approval turned into an endless loop. Errors are written as messages for the calling agent (permission-blocked: ..., chrome-not-running: ...), not as stack traces.

MCP server

mcp_server.py is an optional extra (browser-harness[mcp]). It registers 23 tools, browser_new_tab through browser_stop_recording. Each runs ensure_daemon(), sends helper stdout to stderr so the stdio JSON-RPC stream stays clean, and turns any exception into an MCP ToolError (mcp_server.py). Screenshots come back as a file path plus dimensions, not image bytes (mcp_server.py). That suits a local coding agent that can open files, and it does not suit a remote client.

Recording and telemetry

Recording is off on fresh installs. When it is on, every state-changing helper in ACTIONS writes a JPEG frame and an events.jsonl line. URLs are scrubbed of token-like query parameters, and text typed into password fields is masked (recorder.py). Telemetry is on by default and is disabled with BH_TELEMETRY=0 or browser-harness telemetry (telemetry.py). The generic capture() path filters property names, but capture_cli_event sends the script text (up to 20,000 characters), the stdout tail, and each helper call’s arguments (up to 300 characters) under a persistent install id (telemetry.py). Anyone scripting against private pages should know this.

Extending it

  • Agent helpers. _load_agent_helpers executes agent_helpers.py from the workspace directory and copies its public names into the helper namespace. The workspace is $BH_AGENT_WORKSPACE, or ~/.config/browser-harness/agent-workspace by default (helpers.py). This is where the agent puts new primitives.
  • Domain skills. With BH_DOMAIN_SKILLS=1, goto_url returns up to 10 Markdown filenames from domain-skills/<site>/ in the workspace (helpers.py). The repo ships about 100 site folders written by agents, but they are only used if the workspace points at them.
  • Raw CDP. cdp("Domain.method", **params) reaches any protocol method, and drain_events() exposes the daemon’s event buffer.
  • Orchestrators. BU_NAME runs isolated daemons. BU_CDP_WS and BU_CDP_URL pin an endpoint, and BH_REQUIRE_EXISTING_DAEMON=1 makes the CLI fail closed instead of discovering another Chrome.

Running it

  • Install. uv tool install --python 3.12 browser-harness, then register the output of browser-harness skill as an agent skill. The package needs Python 3.11 or later and four pinned dependencies: cdp-use, fetch-use, pillow and websockets.
  • Local browser. Start Chrome (or Edge, Brave, Arc, Comet) and tick “Allow remote debugging” on chrome://inspect/#remote-debugging. Each new connection on Chrome 144+ shows an Allow prompt. On macOS, browser-harness mac-approve can click it. browser-harness doctor diagnoses install, daemon and browser state.
  • Dedicated or remote browser. Set BU_CDP_URL (HTTP DevTools endpoint) or BU_CDP_WS.
  • Cloud. browser-harness auth login, then start_remote_daemon("name", profileName=..., proxyCountryCode=...) inside a script. It POSTs to the Browser Use API, starts a daemon on the returned cdpUrl, and stops the cloud browser if startup fails (admin.py).
  • MCP. uvx --from 'browser-harness[mcp]' browser-harness-mcp.

Strengths and caveats

  • Strength: real browser, real sessions. Attaching to the user’s running Chrome gives the agent existing logins without cookie export. The connection code handles the newer Chrome permission model carefully.
  • Strength: a persistent daemon makes scripts cheap. Each heredoc call is a short Python process talking to a warm CDP connection, so the agent can work in small, inspectable steps.
  • Strength: honest engineering on the hard parts. PID-reuse checks, stale-session reattachment and cloud-stop retries have dedicated unit tests (tests/unit/test_daemon.py and test_admin.py together run to about 2,300 lines).
  • Caveat: no perception layer. There is no DOM snapshot, element index or screenshot annotation. Results depend on how well the calling agent uses the AX tree and coordinates.
  • Caveat: exec of arbitrary code. The harness runs whatever script the agent sends, in-process, with the user’s logged-in browser attached. That is the design, but the sandbox is the agent’s, not the harness’s.
  • Caveat: one mutable tab per daemon. SKILL.md warns that two agents sharing the default daemon can act on each other’s tabs. For real isolation you need named daemons, each of which triggers another Allow prompt, or cloud browsers.
  • Caveat: telemetry by default. CLI events include script text and output tails unless you opt out.
  • Caveat: the cloud features are a service. Stealth, proxies and CAPTCHA handling are Browser Use Cloud features, not code in this repository.

Sources: code at afbcc38, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (13 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit afbcc38. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

The page is represented to the model primarily through screenshots and accessibility tree data, not through DOM serialization or set-of-marks indexes. There is no built-in DOM-to-text serialization, element indexing scheme, or visual marking overlay.

  • Screenshots: capture_screenshot() in helpers.py (line 331) calls Page.captureScreenshot via CDP and returns a PNG. It supports max_dim param to downscale large screenshots (e.g. max_dim=1800 to stay under a 2000px-per-side limit some image-aware LLMs enforce). Screenshots are served to the MCP layer as a path + dimensions object from mcp_server.py:192-200. The recorder also captures JPEG frames (quality 80) after each action via _capture() in recorder.py:263-310.

  • Accessibility tree: SKILL.md documents the recommended element-finding workflow: use cdp("Accessibility.getFullAXTree")["nodes"] which returns every element's role, name, and backendDOMNodeId. Filter this in Python before printing since it runs thousands of nodes. Click coordinates come from cdp("DOM.getBoxModel", backendNodeId=n)["model"]["content"] — average the four corners for viewport-space click coordinates.

  • Raw HTML via JS: SKILL.md recommends falling back to js(...) for DOM inspection when the AX tree lacks an element (canvas, exotic widgets). There is no automatic DOM pruning or serialization logic in the harness itself — the agent decides what to fetch and how much.

  • Page metadata: page_info() in helpers.py:159-169 returns url, title, viewport dimensions, scroll position, and page scroll dimensions as JSON via Runtime.evaluate. Native dialogs override this with a {dialog: ...} object.

There is no built-in "set-of-marks" overlay, element index system, or automatic size-limiting/pruning of DOM text. The harness lets the agent fetch whatever page representation it needs, keeping the harness itself representation-agnostic.

How are actions executed and how are elements targeted?

answered

Actions are executed exclusively through CDP (Chrome DevTools Protocol) over a persistent WebSocket connection. There is no Playwright or OS-level input automation.

  • Architecture: The daemon in daemon.py holds a single CDP WebSocket connection (_PatientCDPClient for local Chrome, standard CDPClient for remote) and an IPC relay over Unix socket (or TCP loopback on Windows). All helper functions in helpers.py communicate with the daemon via _send() which serializes JSON requests over the IPC socket, and the daemon relays them to CDP.

  • Element targeting: Elements are targeted primarily by viewport coordinates via click_at_xy(x, y) (helpers.py:174-194). The SKILL.md-recommended workflow is: accessibility node → DOM.getBoxModel to get bounding box → average corners for center coordinates → click_at_xy. Raw CSS selectors are used for filling inputs (fill_input at line 208), waiting (wait_for_element at line 496), and file uploads (upload_file at line 621). upload_file uses DOM.querySelector + DOM.setFileInputFiles. There are no zero-indexed element indexes or set-of-marks identifiers.

  • Typing: type_text() (helpers.py:196) uses Input.insertText which inserts text directly. fill_input() (helpers.py:208) is a more robust alternative that focuses the element, clears with select-all (Ctrl/Cmd+A via raw CDP key events), types per-character via press_key(), then fires synthetic input + change DOM events to wake framework listeners.

  • Scrolling: CDP Input.dispatchMouseEvent type mouseWheel via scroll() (helpers.py:326-327).

  • Key presses: press_key() (helpers.py:295-324) dispatches keyDown, char, keyUp CDP events for each key, with physical keyboard codes for US layout.

  • Tabs: Managed via CDP Target.* methods. new_tab() creates via Target.createTarget, switch_tab() attaches via Target.attachToTarget, close_tab() via Target.closeTarget. The daemon tracks the active session ID.

  • MCP layer: In mcp_server.py, each browser action (click, type, scroll, goto, screenshot, etc.) is exposed as a named MCP tool over stdio, returning JSON.

How is the agent loop / planning implemented?

answered

browser-harness does not implement an agent loop, planner, or executor. It is a browser-control harness, not an agent framework.

  • What it is: The README states "Connect an LLM directly to your real browser through one editable CDP websocket." The project provides the transport layer — CDP helpers, a daemon for persistent connections, IPC for CLI and MCP access — but the agent loop itself is external. The calling agent (Claude Code, Codex, or any MCP client) is responsible for planning, tool-call orchestration, stop conditions, and inter-step memory.

  • MCP server as tool interface: The mcp_server.py file (lines 138-291) exposes each browser helper as a named MCP tool (browser_click, browser_type, browser_goto, browser_screenshot, etc.). An MCP client calls these tools sequentially in its own loop. The daemon preserves the attached tab and CDP session across calls, providing continuity without the harness tracking steps.

  • Helper tracing (observation only): run.py:145-178 implements _traced() which wraps every helper call to record: the helper name, truncated args (300 chars), duration in seconds, and any error message. This trace (_helper_trace) is capped at 500 entries (_MAX_TRACED_STEPS at line 145) and is used purely for telemetry, not for agent memory or loop control. The recorder (recorder.py) captures screenshots and event traces for human review and video production, not for agent planning.

  • No planner/executor split: There is no planner vs executor distinction, no tool-call schema (beyond the MCP tool definitions), and no stop-condition logic. The caller decides when to stop. The harness provides wait(), wait_for_load(), wait_for_element(), and wait_for_network_idle() as building blocks the caller can use in its own loop.

  • CLI entry point: run.py:240-302 (main()) reads a heredoc script, auto-bootstraps the daemon, and exec()s the script. Each invocation is one-shot — there's no built-in iterative loop.

Editor's note. Correction: the _traced() wrapper is not telemetry-only; after every successful helper call it also invokes recorder.observe, which writes a frame and an events.jsonl line when recording is enabled (run.py L174).

How are failures, retries and self-healing handled?

answered

Reliability is handled through daemon self-healing, connection retry with escalation, and IPC-level error propagation — not through retries of failed actions, caching of successful actions, or replanning.

  • Daemon self-healing: ensure_daemon() in admin.py:525-708 is the core reliability function. It is idempotent and handles: stale daemon (detected via CDP health check — if Target.getTargets returns no "result", the daemon is recycled); Chrome not running (launches the browser automatically via _launch_browser() at line 1274); remote debugging not enabled (opens chrome://inspect/#remote-debugging with a 180s reopen throttle at line 1387); and the Chrome Allow popup (prints instructions, raises "permission-blocked" without retrying). It uses a _spawn_lock (line 256-316) with file locking to prevent concurrent daemon creation.

  • Connection retry: ensure_daemon() loops up to 3 times (line 561) with escalating strategies: each iteration can launch Chrome, open the inspect page, or wait for approval. If a local daemon dies during handshake, the PID file is cleaned under lock and a new attempt is made — but if the popup was denied, it raises "permission-blocked" and stops.

  • Error propagation: Every helper wraps CDP calls via _send() which can raise _IPCResponseTimeout (a TimeoutError subclass, line 48) and RuntimeError for CDP errors (line 68-69). These propagate to the caller (the agent or MCP client) as exceptions. The MCP server catches all exceptions with a bare except Exception (mcp_server.py:132-133) and converts them to ToolError.

  • CDP session recovery: In daemon.py:767-808, when a CDP request hits "Session with given id not found", the daemon automatically re-attaches to the current page and retries the request once against the replacement session. This is the only built-in retry of a failed CDP operation.

  • IPC timeouts: Default IPC response timeout is 5s (DEFAULT_IPC_RESPONSE_TIMEOUT_SECONDS, helpers.py:44). Screenshots get a longer 60s timeout (SCREENSHOT_IPC_RESPONSE_TIMEOUT_SECONDS, line 45). The IPC connect timeout is 5s (IPC_CONNECT_TIMEOUT_SECONDS, line 41).

  • No action caching or replanning: The harness does not cache successful actions or replan on failure. Those are the calling agent's responsibilities.

Which models are supported and how are they called?

answered

browser-harness is model-agnostic — it does not call any LLM, implement provider integrations, enforce vision requirements, or define structured output schemas. The repository contains no model-specific code whatsoever.

  • No LLM calls in the harness: There is no OpenAI, Anthropic, Gemini, or any other provider client in the dependency list or source code. The pyproject.toml (dependencies) includes cdp-use for CDP WebSocket connections, Pillow for image processing, and mcp for the MCP server — no ML/LLM packages.

  • How the agent connects: The README says "Connect an LLM directly to your real browser through one editable CDP websocket." The harness is the transport layer that the LLM agent (Claude Code, Codex, Cursor, etc.) uses as a tool. The agent decides which model to use, how to call it, and how to structure the interaction.

  • MCP as the model interface: The MCP server (mcp_server.py) exposes helpers as MCP tools over stdio. This is the intended model-facing API — any MCP-capable client can consume these tools. The tool schema is derived from each helper function's Python annotations (via functools.wraps and the @_tool decorator at line 116-135). There is no vision requirement at the harness level, though the SKILL.md recommends screenshots as one of several page-representation methods.

  • No structured output / tool calling: The harness does not enforce structured output or tool-call schemas. Its helpers return plain Python values (dicts, strings, None), and the MCP wrapper serializes them to JSON text. The calling agent implements its own tool-calling loop.

  • No small or specialized models: There are no lightweight model alternatives, no model routing, and no fallback logic. The harness treats every caller identically via the MCP tool interface or the Python CLI.

In summary, browser-harness is entirely model-agnostic. Model support is determined by the external agent that consumes the harness's tools, not by the harness itself.

How are browser sessions, profiles, auth and anti-bot handled?

answered

Browser sessions are handled through three modes: local, CDP URL/WS, and cloud.

  • Local browser: The default mode. The daemon discovers a running Chrome/Chromium/Edge/Brave instance by scanning known profile directories (PROFILES in daemon.py:112-113 which aggregates platform-specific paths from profile_dirs(), lines 102-109). It reads DevToolsActivePort from the profile directory to find the WebSocket endpoint (get_ws_url(), lines 264-342). Daemon identity verification uses process-start fingerprinting (_process_start_time() in admin.py:17-105) to prevent PID-reuse bugs during restart.

  • Remote CDP endpoint: Set via BU_CDP_WS (direct WebSocket URL) or BU_CDP_URL (HTTP DevTools endpoint, auto-resolved to WS via /json/version). Used for dedicated automation Chrome instances that avoid the per-connection Allow dialog (daemon.py:270-290). BU_CDP_URL also blocks cloud auto-bootstrap (run.py:99-100).

  • Cloud browsers (Browser Use Cloud): Provisioned via start_remote_daemon() in admin.py:997-1036. This calls POST /api/v3/browsers to spin up a managed cloud Chrome, then connects the daemon via the returned cdpUrl. Cloud browsers support: profile-based cookie sync (profileName/profileId, resolved from list_cloud_profiles()), proxy country selection (proxyCountryCode), custom proxies (customProxy), configurable viewport, and auto-recording. Auto-bootstrap to cloud is opt-in via BU_AUTOSPAWN env var (run.py:400-401).

  • Profiles and cookies: Local profiles are detected via list_local_profiles() (shells out to profile-use list --json). sync_local_profile() uses profile-use sync to push local Chrome cookies to a cloud profile for logged-in sessions. The Browser Use Cloud API lists profiles with cookieDomains arrays as a login-state summary (admin.py:957-984).

  • Anti-bot: No built-in stealth or CAPTCHA resolution. The README states cloud browsers "run with clean managed IPs and stealth settings" and CAPTCHA solving (README claim — managed by Browser Use Cloud infrastructure, not in this repo). The http_get() helper routes through fetch-use proxy when BROWSER_USE_API_KEY is set (helpers.py:631-638).

  • Auth: PKCE OAuth flow or device-code flow to Browser Use Cloud via auth.py. Auth persisted in a chmod 600 JSON file at the config dir. Also supports BROWSER_USE_API_KEY environment variable.