# browser-use/browser-use

> Async Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/browser-use/browser-use (reviewed at commit `7be96ed8bafa8dfe1eef228b59cf5c884b8b2431`, 2026-10-02)
- Stars: 117268 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/browser-use/

## Overview

browser-use is a Python library for an autonomous browser agent. You give an `Agent` a task in plain language and a chat model. It then runs a step loop. Each step reads the page as compact indexed DOM text plus a screenshot, asks the model for a structured `AgentOutput` (reasoning fields and a list of actions), and executes those actions in Chrome. The loop stops when the model calls `done`, the step budget runs out, or too many steps fail.

The library drives Chromium directly over the Chrome DevTools Protocol through the `cdp-use` client. Playwright is not a package dependency at this commit; it is only invoked as a subprocess to install Chromium when no local browser is found. Browser work is organised around a `bubus` event bus: `BrowserSession` owns the CDP connection, and a set of "watchdogs" handle clicks, DOM capture, downloads, popups, security rules, storage state and CAPTCHA events. The library is aimed at Python developers who want a ready-made agent loop with many provider adapters, and who are fine with an LLM call on every step.

The repo also ships more around the library: an MCP server, a `@sandbox` decorator for running code on Browser Use Cloud, a cloud browser client and history replay. Note that the `browser-use` console script now mostly delegates to the separate `browser-harness` package; the agent described here is the Python API.

## Architecture

```mermaid
flowchart LR
  T["Task + LLM"] --> A["Agent.run / step"]
  A --> MM["MessageManager"]
  A --> LLM["BaseChatModel.ainvoke"]
  LLM --> OUT["AgentOutput (actions)"]
  OUT --> MA["Agent.multi_act"]
  MA --> TO["Tools.act -> Registry"]
  TO --> EB["bubus EventBus"]
  EB --> WD["Watchdogs"]
  WD --> CDP["cdp-use client"]
  A --> BS["BrowserSession state summary"]
  BS --> DOM["DomService + DOMTreeSerializer"]
  DOM --> CDP
  CDP --> CH["Chrome (local or cloud)"]
```

| Component | Path | Role |
|---|---|---|
| Agent | `browser_use/agent/service.py` | Step loop, retries, planning, loop detection, history, rerun |
| Agent schemas | `browser_use/agent/views.py` | `AgentSettings`, `AgentOutput`, `ActionLoopDetector`, history models |
| Message manager | `browser_use/agent/message_manager/` | Builds per-step prompts, compacts old history |
| Tools and registry | `browser_use/tools/service.py`, `tools/registry/service.py` | Built-in actions, `@tools.action` decorator, dynamic action union |
| Browser session | `browser_use/browser/session.py`, `profile.py` | CDP connection, targets, reconnection, launch and profile settings |
| Watchdogs | `browser_use/browser/watchdogs/` | Event handlers for actions, DOM, screenshots, downloads, popups, security, CAPTCHA |
| DOM pipeline | `browser_use/dom/service.py`, `dom/serializer/` | Merge DOMSnapshot + DOM + AX tree, filter, index, serialize |
| LLM adapters | `browser_use/llm/` | `BaseChatModel` protocol and one package per provider |
| Cloud and extras | `browser/cloud/`, `sandbox/`, `mcp/`, `skills/` | Cloud browsers, remote execution, MCP server, hosted skills |

## How a request flows

1. **Construct.** `Agent.__init__` falls back to `ChatBrowserUse` when no LLM is passed and turns on flash mode for it. It auto-sets screenshot size for Claude Sonnet and an LLM timeout per model family ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L225-L277)). It also enables coordinate clicking for some model names ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L326-L333)).
2. **Run.** `Agent.run(max_steps=500)` starts the session and then calls `_execute_step` until done. Each step runs under `asyncio.wait_for(step_timeout)`.
3. **Step.** `step()` first waits if a CAPTCHA solver is active. It then runs `_prepare_context`, `_get_next_action`, `_execute_actions` and `_post_process`, with one `_handle_step_error` for any exception ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1035-L1086)).
4. **Perceive.** `_prepare_context` calls `BrowserSession.get_browser_state_summary(include_screenshot=True)`. It filters actions by URL, builds the state message, and injects budget, replan, exploration and loop-detection nudges ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1087-L1160)). The summary request is a `BrowserStateRequestEvent` on the event bus ([session.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L1597-L1640)).
5. **Decide.** `get_model_output` calls `llm.ainvoke(messages, output_format=AgentOutput)`, trims to `max_actions_per_step`, and switches to `fallback_llm` on 401/402/429/5xx or truncated output ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1945-L1981), [L1983-L2025](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1983-L2025)).
6. **Act.** `multi_act` runs the actions in order. It stops early on `done`, on an error, on an action marked `terminates_sequence`, or when the URL or focused tab changed ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2730-L2848)). Each action goes through `Tools.act`, which calls `registry.execute_action` under a per-action timeout of 180 s by default ([tools/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L2178-L2240)).
7. **Execute in the browser.** For `click`, `_click_by_index` looks up the node in the selector map and dispatches `ClickElementEvent` ([tools/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L704-L734)). `DefaultActionWatchdog` refuses `<select>` and file inputs, checks occlusion, and sends `Input.dispatchMouseEvent` moved/pressed/released, with a JS click fallback ([default_action_watchdog.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/watchdogs/default_action_watchdog.py#L703-L720), [L904-L950](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/watchdogs/default_action_watchdog.py#L904-L950)).
8. **Record.** `_post_process` updates the plan and loop detector and counts consecutive failures, but only for single-action steps ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1219-L1256)). `_finalize` appends an `AgentHistory` item.

## Key components

### DOM pipeline and element indexes

`DomService` fires `DOMSnapshot.captureSnapshot` (with paint order and rects), `DOM.getDocument(depth=-1, pierce=True)` and per-frame `Accessibility.getFullAXTree` in parallel, with a 10 s wait and retries ([dom/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/service.py#L569-L600)). `DOMTreeSerializer.serialize_accessible_elements` then simplifies the tree, removes elements hidden by paint order, drops redundant parents, filters children inside clickable parents by bounding box, and assigns indexes ([serializer.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L114-L165)). The index the model sees is the CDP `backend_node_id` when it is unique, with a synthetic id only on collision ([serializer.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L637-L658), [L747-L757](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L747-L757)). So an element usually keeps its number across steps. It is not a 1..N counter.

### Agent output and planning

`AgentOutput` holds `thinking`, `evaluation_previous_goal`, `memory`, `next_goal`, optional plan fields and a non-empty `action` list ([views.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/views.py#L388-L400)). Its `action` type is a discriminated union built from the registry at runtime, so custom actions appear in the schema automatically. Flash mode removes the reasoning and plan fields to save tokens.

### Tools registry

`Registry.action(description, param_model, domains, terminates_sequence)` is the decorator behind every built-in and custom action ([registry/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/registry/service.py#L291-L327)). `domains` hides an action outside matching URLs, and `terminates_sequence` tells `multi_act` to stop the queue after it. `Controller` is just an alias for `Tools` ([tools/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L2327)).

### BrowserSession and watchdogs

`BrowserSession` is a Pydantic model that owns the root CDP client, per-target sessions and the event bus. `attach_all_watchdogs` wires the handlers: downloads, storage state, local browser launch, security (allowed/prohibited domains), about:blank, popups, permissions, default actions, screenshots, DOM, recording, HAR and CAPTCHA. The crash watchdog is commented out ([session.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L1690-L1712)). Dropped WebSockets get three reconnect attempts, about 54 s in total ([session.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L585-L594)).

### LLM layer

Every provider implements the `BaseChatModel` protocol: a `model` string, a `provider` name and an overloaded `ainvoke(messages, output_format)` that returns plain text or a validated Pydantic object ([llm/base.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/llm/base.py#L32-L59)). Adapters cover OpenAI, Anthropic, Google, Azure, AWS Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, OrcaRouter, Cerebras, Vercel, LiteLLM, OCI and Browser Use's own hosted model.

## Extending it

- **Custom actions.** Decorate an async function with `@tools.action("description", param_model=...)` and pass `tools=` to `Agent`. Parameters can include injected objects such as `browser_session`, `page_extraction_llm` or `file_system`.
- **Exclude or restrict actions.** Use `Tools(exclude_actions=[...])` or `tools.exclude_action(name)`, or set `domains=` on registration.
- **New model provider.** Write any class that satisfies `BaseChatModel`. No base-class inheritance is required.
- **Hooks.** `run(on_step_start=..., on_step_end=...)` gets the agent object each step. History can be saved and replayed with `rerun_history`, which retries each step and re-matches element indexes against the live DOM ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L3103-L3133)).
- **Other surfaces.** `browser_use/mcp/server.py` exposes `browser_navigate`, `browser_click`, `browser_get_state` and similar tools over MCP. `@sandbox` ships a function to Browser Use Cloud with cloudpickle and injects a remote browser.

## Running it

- `uv add browser-use` (Python 3.11 or later), set a provider key such as `OPENAI_API_KEY`, then `await Agent(task=..., llm=ChatOpenAI(...)).run()`.
- **Local browser.** `LocalBrowserWatchdog` launches an installed Chrome/Chromium with flags from `BrowserProfile`, and falls back to installing Chromium through a Playwright subprocess. `user_data_dir` and `storage_state` keep logins between runs, and `cdp_url` attaches to a browser that is already running.
- **Cloud browser.** `Browser(use_cloud=True)` plus `BROWSER_USE_API_KEY` requests a remote browser with proxies and fingerprinting, then connects over the same CDP path.
- No database or server is needed for the library itself. Anonymous PostHog telemetry is on by default (`ANONYMIZED_TELEMETRY=false` turns it off), and `BROWSER_USE_CLOUD_SYNC` follows that setting unless set explicitly.

## Strengths and caveats

- **Strength: complete loop out of the box.** Planning fields, loop detection with escalating nudges, message compaction, fallback LLM, per-step and per-action timeouts, and a final recovery call after `max_failures` are already implemented.
- **Strength: stable element references.** Using backend node ids as indexes, and stopping the action queue when the page changes, cuts down stale-index clicks.
- **Strength: broad provider support** through one small protocol, with vision switched off automatically for models that lack it.
- **Caveat: cost and latency.** Every step captures DOM and a screenshot and calls the LLM. Nothing caches decisions between runs, except an explicit `rerun_history`.
- **Caveat: large, fast-moving core.** `agent/service.py` and `browser/session.py` are each over 4,000 lines, and defaults such as model-name pattern lists change often.
- **Caveat: hosted tilt.** The default LLM (`ChatBrowserUse`), CAPTCHA solving, stealth and skills depend on Browser Use Cloud. A local run gets only Chrome flags such as `--disable-blink-features=AutomationControlled`.

*Sources: code at 7be96ed, deepwiki-open wiki (10 pages), verified Q&A.*

## How browser-use/browser-use answers the Browser & computer control questions

### How is the page represented to the model? (answered)

**DOM Serialization.** The agent represents pages through `SerializedDOMState` (`browser_use/dom/views.py:935`), which wraps a tree of `SimplifiedNode` objects built from three CDP sources: `DOMSnapshot.captureSnapshot`, the full accessibility tree via `accessibility.GetFullAXTree`, and the DOM document structure (`browser_use/dom/service.py:6-25`). The `DomService` class orchestrates this, combining snapshot bounds, AX node properties, and DOM nodes into `EnhancedDOMTreeNode` objects with sibling/shadow-root/iframe support. Cross-origin iframes with dimensions ≥ 10px each side are included (`browser_use/dom/service.py:40-42`).

**Serialization to text.** `DOMTreeSerializer.serialize_accessible_elements()` (`browser_use/dom/serializer/serializer.py:114-165`) runs a pipeline: (1) create a simplified tree by filtering non-content elements (style, script, SVG decorative children) and detecting interactive elements; (2) paint-order filtering to skip elements visually occluded; (3) tree optimization (remove unnecessary parents); (4) bounding-box filtering that excludes children contained within propagating parents (`a`, `button`, `div[role=button]`) at 99% containment threshold; (5) assign numeric interactive indices to clickable/scrollable elements. The final `serialize_tree()` (`browser_use/dom/serializer/serializer.py:989`) renders a compact text format: `[index]<tagname attr="val" />` for interactive elements, ordinary `<tagname>` for containers, `|scroll element|` markers for scrollable containers, `|SHADOW(open/closed)|` for shadow-DOM hosts, text nodes inline, and iframes with hidden-element hints showing scroll distances. Compound controls (date pickers, range sliders, file inputs, select dropdowns) get synthetic `compound_components=...` attributes with sub-role information (`browser_use/dom/serializer/serializer.py:1059-1100`).

**Element indices (set-of-marks).** Interactive elements get sequentially numbered indices stored in a `DOMSelectorMap` (dict of `int → EnhancedDOMTreeNode`). The LLM sees `[1]<button type="submit">` and uses that index to click. New elements since the last step are marked with a `*` prefix. `max_clickable_elements_length` caps serialized text at 40000 chars (`browser_use/agent/views.py:92`).

**Screenshots.** A PNG screenshot is always captured every step via `BrowserSession.get_browser_state_summary(include_screenshot=True)` (`browser_use/agent/service.py:1096-1099`). When `use_vision` is enabled, the screenshot is sent as an image content part to the LLM. For Claude Sonnet models, screenshots are auto-resized to 1400×850 to fit context windows (`browser_use/agent/service.py:247-250`). Configurable via `llm_screenshot_size` parameter.

**Accessibility tree.** The AX tree provides `name`, `role`, `description`, `properties` (valuemin/max/now, expanded, pressed, invalid, etc.) per element. These feed into `ClickableElementDetector.is_interactive()` (`browser_use/dom/serializer/clickable_elements.py:6-52`) which heuristically scores elements as interactive based on tag, role, JS click listeners, ARIA attributes, pointer cursor, and form-control heuristics.

> **Editor's note.** Correction: element indexes are the CDP `backend_node_id` when that is unique, not a sequential counter, so they mostly stay stable between steps.

Citations: [browser_use/dom/service.py:1-72](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/service.py#L1-L72) · [browser_use/dom/views.py:935-955](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/views.py#L935-L955) · [browser_use/dom/serializer/serializer.py:114-165](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L114-L165) · [browser_use/dom/serializer/serializer.py:989-1100](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L989-L1100) · [browser_use/agent/service.py:245-252](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L245-L252) · [browser_use/dom/serializer/clickable_elements.py:1-52](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/clickable_elements.py#L1-L52)

### How are actions executed and how are elements targeted? (answered)

**Architecture.** Actions are dispatched through the `Tools` (`browser_use/tools/service.py`) → `Registry.execute_action()` → event bus → handler chain. `Tools.act()` (`browser_use/tools/service.py:2178-2236`) receives an `ActionModel` (a Pydantic union discriminated by action name), extracts the action name and params, and calls `registry.execute_action()` under an `asyncio.wait_for` timeout (env `BROWSER_USE_ACTION_TIMEOUT_S`, default 180s). The registry maps action names to registered handler functions, which typically dispatch events on the browser session's event bus.

**Element targeting (index-based).** Most actions reference elements by a numeric `index` from the `selector_map` — the same indices shown in the DOM serialization. The handler looks up the `EnhancedDOMTreeNode` and resolves it to a CDP `backend_node_id` + `frame_id`. This avoids fragile XPath or CSS selectors. For models that support it (Claude Sonnet 4+, Opus 4+, Gemini 2.5 Pro), coordinate-based clicking is available via `coordinate_x`/`coordinate_y` in `ClickElementAction` (`browser_use/tools/views.py:69-75`). Enabled automatically at agent init (`browser_use/agent/service.py:327-333`).

**Available action types** (`browser_use/tools/views.py`): `navigate` (with `new_tab` option), `click` (by index or coordinates), `input_text` (with `clear` flag to append), `scroll` (direction + page fraction + optional element index), `send_keys` (keyboard shortcuts like `Control+o`), `upload_file` (index + file path), `switch_tab` / `close_tab` (4-char tab ID), `search` (Google/Bing/DuckDuckGo), `search_page` (regex/literal text search in page), `find_elements` (CSS selector), `extract` (LLM-based extraction from page markdown), `screenshot` (save to file), `save_as_pdf`, `get_dropdown_options`, `done` (terminate with result).

**Execution flow.** `Agent.multi_act()` (`browser_use/agent/service.py:2728-2848`) runs sequential actions from a single LLM response. It enforces two page-change guards: (1) static `terminates_sequence=True` on actions like navigate/search/go_back/switch; (2) runtime detection comparing pre/post-action URL and focused target ID. Any page change aborts remaining queued actions. A `wait_between_actions` delay (from `BrowserProfile`) is applied between steps. `send_keys` dispatches keyboard events via CDP. File upload uses CDP's `input.dispatchFile` or sets the file input value. Tab management uses CDP `Target.attachToTarget`/`Target.detachFromTarget`. Scroll is implemented via CDP JavaScript `window.scrollBy()` or `element.scrollIntoView()`.

**Multi-step actions per LLM call.** Controlled by `max_actions_per_step` (default 5). The `AgentOutput` schema allows multiple actions per response. Each action is dispatched sequentially; if one errors, the remaining actions are still attempted (error is recorded, not fatal).

> **Editor's note.** Correction: `multi_act` stops at the first failing action (`tools/service.py` ~L2809). Remaining actions are not run. The coordinate-click model check matches `gemini-3-pro`, not Gemini 2.5 Pro.

Citations: [browser_use/tools/service.py:2178-2236](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L2178-L2236) · [browser_use/tools/views.py:69-75](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/views.py#L69-L75) · [browser_use/tools/views.py:122-143](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/views.py#L122-L143) · [browser_use/agent/service.py:2728-2848](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2728-L2848) · [browser_use/agent/service.py:327-333](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L327-L333) · [browser_use/agent/service.py:2766-2828](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2766-L2828)

### How is the agent loop / planning implemented? (answered)

**Main loop.** `Agent.run(max_steps=500)` (`browser_use/agent/service.py:2501-2727`) is the entry point. It starts the browser session, dispatches session/task events, runs initial actions (URL extraction from task), then enters a `while self.state.n_steps <= max_steps` loop calling `_execute_step()` which internally calls `step()`.

**Step loop.** `step()` (`browser_use/agent/service.py:1033-1085`) has four phases:
1. **CAPTCHA wait**: checks if a CAPTCHA solver is in progress and blocks if needed (`browser_use/agent/service.py:1044-1061`).
2. **Context preparation** (`_prepare_context`): captures browser state (DOM + screenshot), checks downloads, checks stop/pause signals, updates action models for the current page, builds plan description, creates LLM messages with the state, runs optional message compaction, injects loop-detection/replan/exploration/budget nudges.
3. **LLM call + execution** (`_get_next_action` + `_execute_actions`): calls LLM with the prepared messages, parses structured `AgentOutput`, then calls `multi_act()` to run all returned actions sequentially.
4. **Post-processing** (`_post_process`): updates plan state, records actions for loop detection, increments consecutive failures counter on error.

**Tool-call schema.** LLM output is parsed into `AgentOutput` (`browser_use/agent/views.py:388-398`) containing `thinking`, `evaluation_previous_goal`, `memory`, `next_goal`, `current_plan_item`, `plan_update`, and `action` (list of `ActionModel`). The `ActionModel` is a discriminated union of all registered action types. In flash mode, planning fields are stripped. The output goes through `AgentOutput.type_with_custom_actions()` which dynamically creates a Pydantic model with the registry's current actions as a discriminated union.

**Planning.** When `enable_planning=True`, the LLM can output `plan_update` (list of step strings) to create/replace a plan stored as `PlanItem` objects with status tracking (pending/current/done/skipped). `_render_plan_description()` formats it for re-injection. Replan nudges trigger after 3 consecutive failures (`planning_replan_on_stall`). An exploration nudge suggests creating a plan after 5 steps without one.

**Memory between steps.** The `MessageManager` maintains message history. `_maybe_compact_messages()` summarizes old history every 25 steps (configurable) when the prompt exceeds ~10K tokens, using a compaction LLM. `memory` field in `AgentOutput` allows the LLM to persist structured notes. `ActionResult.long_term_memory` lets actions inject persistent context.

**Stop conditions.** `max_steps` reached (forces `done` action with last-step warning), `max_failures` exceeded (default 5, with optional final recovery attempt), manual stop/pause via signal handler (SIGINT toggles pause/resume; second SIGINT force-exits). The `_force_done_after_last_step()` method replaces the action model with a done-only schema on the final step.

**History.** Each step records an `AgentHistory` item with model output, results, browser state (URL, title, tabs, screenshot path). After completion, an optional LLM judge evaluates the trace.


Citations: [browser_use/agent/service.py:2501-2660](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2501-L2660) · [browser_use/agent/service.py:1033-1085](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1033-L1085) · [browser_use/agent/views.py:388-398](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/views.py#L388-L398) · [browser_use/agent/service.py:1162-1173](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1162-L1173) · [browser_use/agent/service.py:1464-1506](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1464-L1506) · [browser_use/agent/service.py:1417-1462](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1417-L1462)

### How are failures, retries and self-healing handled? (answered)

**Error classification.** `Agent._handle_step_error()` (`browser_use/agent/service.py:1258-1314`) categorizes exceptions into: `InterruptedError` (user pause/stop — silent), `ConnectionError`/WebSocket-closed errors (browser disconnect — triggers reconnection), and all other exceptions (counted as consecutive failures). `_is_connection_like_error()` checks error strings for patterns like 'websocket connection closed', 'browser has been closed'. Browser-closed errors set `self.state.stopped = True` and exit the loop.

**Retries for empty actions.** `_get_model_output_with_retry()` (`browser_use/agent/service.py:1670-1704`) wraps the LLM call: if the model returns an empty or all-empty action list, it sends a clarification message and retries once. If the retry also returns empty, it inserts a synthetic `done(success=false)` action.

**Consecutive failure counting.** Single-action steps that error increment `consecutive_failures` (capped at `max_failures`, default 5). Multi-action steps with errors are handled by loop detection instead (not counted as failures). When `consecutive_failures >= max_failures`, a final recovery response is attempted (if `final_response_after_failure=True`), then the agent stops.

**Step-level timeout.** `_execute_step()` (`browser_use/agent/service.py:2438-2498`) wraps the entire `step()` in `asyncio.wait_for(timeout=self.settings.step_timeout)` (default 180s). On timeout, it increments failures and continues. LLM calls have their own `llm_timeout` (auto-detected per model: 30s for Groq, 90s for Claude/DeepSeek/o3, 75s default).

**Action timeout.** `Tools.act()` applies `asyncio.wait_for(timeout=action_timeout)` (env `BROWSER_USE_ACTION_TIMEOUT_S`, default 180s) to prevent hung CDP calls. The timeout is above the 120s extract action LLM cap (`browser_use/tools/service.py:99-110`).

**Fallback LLM.** If the primary LLM fails with a rate limit (429), auth error (401/402), server error (500-504), or truncated output, and a `fallback_llm` is configured, the agent switches providers mid-run (`browser_use/agent/service.py:1983-2025`). This is a one-way switch; once on fallback, the agent stays there.

**Connection recovery.** `BrowserSession` has WebSocket reconnection logic (`browser_use/browser/session.py:585-593`) with 3 attempts, delays (1+2+4s), 54s total wait timeout. When `_handle_step_error` detects a connection error during reconnection, it waits for the reconnect event. If reconnection succeeds, the step is retried; otherwise, it's treated as a terminal browser-closed error.

**Loop detection.** `ActionLoopDetector` (`browser_use/agent/views.py:157-248`) tracks the last 20 action hashes for repetition and page fingerprints for stagnation. At thresholds (5/8/12 repeated actions, 5 stagnant pages), escalating nudges are injected into the LLM context. Wait/done/go_back actions are exempt from tracking. This is advisory only — never blocks actions.

**CAPTCHA handling.** `CaptchaWatchdog` (`browser_use/browser/watchdogs/captcha_watchdog.py`) listens for CDP `captchaSolverStarted/Finished` events from the browser proxy. The agent's `step()` checks `wait_if_captcha_solving()` at the start and blocks until resolution or timeout. The result (success/failed/timeout) is injected into the action result for the LLM. Cloud browsers include proxy rotation and stealth fingerprinting to reduce CAPTCHA incidence.

**Message compaction.** Old conversation history is summarized by an LLM every 25 steps when the prompt exceeds ~10K tokens (`MessageCompactionSettings`). This prevents context windows from overflowing while preserving key information.


Citations: [browser_use/agent/service.py:1258-1314](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1258-L1314) · [browser_use/agent/service.py:1670-1704](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1670-L1704) · [browser_use/agent/service.py:1983-2025](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1983-L2025) · [browser_use/tools/service.py:99-142](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L99-L142) · [browser_use/agent/views.py:157-248](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/views.py#L157-L248) · [browser_use/browser/watchdogs/captcha_watchdog.py:1-100](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/watchdogs/captcha_watchdog.py#L1-L100)

### Which models are supported and how are they called? (answered)

**Supported providers.** The library supports 14+ LLM providers, each in its own `browser_use/llm/<provider>/` package: OpenAI (`openai/chat.py`), Anthropic (`anthropic/chat.py`), Google Gemini (`google/chat.py`), Azure OpenAI (`azure/chat.py`), AWS Bedrock (both Anthropic via `aws/chat_anthropic.py` and native via `aws/chat_bedrock.py`), DeepSeek (`deepseek/chat.py`), Groq (`groq/chat.py`), Mistral (`mistral/chat.py`), Ollama (`ollama/chat.py`), OpenRouter (`openrouter/chat.py`), OrcaRouter (`orcarouter/chat.py`), Cerebras (`cerebras/chat.py`), Vercel (`vercel/chat.py`), lite llm (`litellm/chat.py`), OCI Raw (`oci_raw/chat.py`), and a custom fine-tuned browser-use model (`browser_use/llm/browser_use/chat.py`). Lazy imports via `__getattr__` keep startup fast (`browser_use/llm/__init__.py:82-98`).

**Base protocol.** All LLMs implement the `BaseChatModel` Protocol (`browser_use/llm/base.py:32-75`) which requires `model` attribute, `provider` property, and `ainvoke(messages, output_format=None, **kwargs) -> ChatInvokeCompletion`. The `ainvoke` method is overloaded: without `output_format` it returns a string; with a Pydantic model class it returns a validated structured output. The `Agent` always calls with `output_format=AgentOutput` (a Pydantic model with `action` as a discriminated union).

**Vision support.** Controlled by `use_vision` parameter (default `True` or `'auto'`). When enabled, screenshots are sent as image content parts alongside the DOM text. Some models auto-disable vision: DeepSeek (doesn't support it, `browser_use/agent/service.py:476-478`), older Grok variants (grok-3/grok-code, line 483-485). `vision_detail_level` (`'auto' | 'low' | 'high'`) controls image quality. Claude Sonnet models auto-configure 1400×850 screenshot resizing to fit vision context windows.

**Fine-tuned model.** `ChatBrowserUse` (`browser_use/llm/browser_use/chat.py`) is a first-party fine-tuned model that defaults when no LLM is specified. When used, `flash_mode=True` is auto-set, stripping planning fields from the output and disabling `use_thinking` — yielding faster, more token-efficient inference (`browser_use/agent/service.py:238-243`).

**Structured output / tool calling.** All providers return Pydantic-validated structured outputs. The mechanism differs per provider: OpenAI uses JSON mode / response_format; Anthropic uses tool_use with a single tool; Google uses response_schema; fine-tuned models return raw JSON matching the schema. The `output_format` parameter is passed through `ainvoke()`. For the agent loop, `AgentOutput` is dynamically generated via `AgentOutput.type_with_custom_actions(ActionModel)` which creates a discriminated union of all registered action types.

**Model-to-timeout mapping.** LLM timeouts are auto-configured per model family (`browser_use/agent/service.py:261-277`): Gemini 3 Pro gets 90s, Gemini other gets 75s, Groq gets 30s (fast inference), o3/Claude/DeepSeek get 90s.

**Pre-configured model instances.** `browser_use/llm/models.py` provides shorthands like `openai_gpt_4o`, `google_gemini_2_5_pro`, `anthropic_claude_sonnet_4` etc., each constructing the appropriate Chat class with model string and optional API key from environment.


Citations: [browser_use/llm/__init__.py:82-134](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/llm/__init__.py#L82-L134) · [browser_use/agent/service.py:237-278](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L237-L278) · [browser_use/agent/service.py:476-486](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L476-L486) · [browser_use/agent/service.py:1943-1981](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1943-L1981) · [browser_use/agent/service.py:327-333](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L327-L333)

### How are browser sessions, profiles, auth and anti-bot handled? (answered)

**BrowserSession model.** `BrowserSession` (`browser_use/browser/session.py:134`) is the central session manager, owning CDP WebSocket connections, target/session maps, and an event bus coordinating 12+ watchdogs. Configuration is stored in a `BrowserProfile` (`browser_use/browser/profile.py`) which controls browser launch args, proxy, extensions, display config, domain filtering, etc. The session supports both Pydantic model construction and direct parameter overloading from a comprehensible API surface.

**Local vs Cloud browsers.** Two modes: **local** launches a Chrome/Chromium process via CDP, providing temporary or persistent profiles, and **cloud** creates remote browser instances through `cloud.browser-use.com` API (`browser_use/browser/cloud/cloud.py`). Cloud mode authenticates via `BROWSER_USE_API_KEY` env var or auth config file, calls the cloud API's `/v{2,3,4}/browsers` endpoint to get a CDP URL, then connects normally. Proxy country codes are configurable for geo-targeting. The agent auto-detects which mode via `BrowserProfile.use_cloud` or explicit `cloud_browser_params`.

**Persistent profiles and cookies.** `BrowserProfile` supports `user_data_dir` for persistent profiles, and `storage_state` for saved cookies/session data from Playwright's storage format. The `StorageStateWatchdog` saves/restores cookies and local storage on browser connect/disconnect. Cookies are loaded from a file specified in `BrowserProfile.storage_state` or from a user-data-dir profile. The `DownloadStorageStateEvent` is dispatched when saving state.

**Stealth and anti-detection.** Chrome is launched with automation-control flags disabled (`'--disable-blink-features=AutomationControlled'`, `browser_use/browser/profile.py:201`). Multiple fingerprint-reducing Chrome components are disabled (`CHROME_DISABLED_COMPONENTS` list, `browser_use/browser/profile.py:45-93`), including `AutomationControlled`, `HeavyAdPrivacyMitigations`, `CalculateNativeWinOcclusion`, `OverscrollHistoryNavigation`. Cloud browsers add proxy rotation and stealth fingerprinting on the server side.

**Proxies.** `ProxySettings` (in `BrowserProfile`) supports HTTP/SOCKS5 proxies with authentication. Cloud browsers offer country-code proxy selection (`ProxyCountryCode`). The `SecurityWatchdog` enforces `allowed_domains`/`prohibited_domains` lists, intercepting navigation attempts to blocked domains at the CDP level.

**Watchdog architecture.** `BrowserSession` uses an event-bus pattern with specialized watchdogs (`browser_use/browser/watchdogs/`): `LocalBrowserWatchdog` (Chrome process launch), `DownloadsWatchdog` (PDF auto-download with filename sanitization), `PopupsWatchdog` (JS alerts/dialogs), `SecurityWatchdog` (domain filtering, IP blocking, sensitive data), `DOMWatchdog` (DOM snapshot + element highlighting), `ScreenshotWatchdog` (screenshot capture + optional resize), `CaptchaWatchdog` (proxy CAPTCHA solver events), `AboutBlankWatchdog` (empty page redirects), `CrashWatchdog` (browser crash detection), `StorageStateWatchdog` (cookie/profile persistence), `PermissionsWatchdog` (browser permission management), `RecordingWatchdog`/`HARRecordingWatchdog`. Each watchdog attaches handlers to the session's event bus and is started/stopped with the browser lifecycle.

**CDP connection management.** `BrowserSession.connect()` establishes a WebSocket to the CDP URL, initializes a `CDPClient`, creates a `SessionManager` for target/session tracking, enables page monitoring (lifecycle events, accessibility), and dispatches `BrowserConnectedEvent`. WebSocket reconnection handles transient disconnects with exponential backoff (3 attempts, 1/2/4s delays, 54s total timeout). Connections carry an intentional-stopped flag so cleanup handlers don't cascade.

**CAPTCHA.** The `CaptchaWatchdog` listens for CDP `BrowserUse.captchaSolverStarted/Finished` events from the browser proxy (e.g., the cloud browser's integrated solver or a third-party solving service). It exposes `wait_if_captcha_solving()` which the agent loop calls. Cloud browsers' proxy-based CAPTCHA solver is claimed to reduce CAPTCHA encounters via stealth fingerprinting and IP rotation.

> **Editor's note.** Correction: CrashWatchdog is commented out in `session.py` (~L1700/L1715) and is not active.

Citations: [browser_use/browser/session.py:134-340](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L134-L340) · [browser_use/browser/session.py:567-594](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L567-L594) · [browser_use/browser/session.py:786-830](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L786-L830) · [browser_use/browser/profile.py:45-230](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/profile.py#L45-L230) · [browser_use/browser/cloud/cloud.py:1-100](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/cloud/cloud.py#L1-L100) · [browser_use/browser/watchdogs/captcha_watchdog.py:44-100](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/watchdogs/captcha_watchdog.py#L44-L100)
