# Browser & computer control: comparison

> Frameworks that let an LLM drive a browser (or a whole desktop/phone) to complete tasks.

Canonical page: https://llms-technical-reviews.com/compare/browser-control/

## How is the page represented to the model?

Both projects send the model a text rendering of the page built from CDP, not raw HTML. They differ in how much they add on top.

[browser-use](/p/browser-use/) merges `DOMSnapshot.captureSnapshot`, the full DOM and the accessibility tree. Its `DOMTreeSerializer` drops elements hidden by paint order and redundant wrappers, then prints interactive nodes as `[index]<tag attr=...>`. The index is the CDP `backend_node_id` where possible, so it stays stable across steps. New elements are starred. A screenshot is taken every step and sent when `use_vision` is on. Serialized text is capped at 40,000 characters.

[Stagehand](/p/stagehand/) builds a "hybrid snapshot": an accessibility outline merged with DOM data, with iframes stitched in and shadow DOM pierced. Each line carries a `frameOrdinal-backendNodeId` id that maps to an XPath. Callers can scope it to a locator or exclude parts. There are no screenshots in `act` or `observe`. `extract` can attach a plain viewport PNG on request, with no set-of-marks overlay.

Pick browser-use when pages are visual or ambiguous and a vision model is worth the extra tokens on every step. Pick Stagehand for cheaper, text-only prompts scoped to part of a page, especially for extraction.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/page-perception/index.md

## How are actions executed and how are elements targeted?

Both projects keep the model away from raw coordinates by default. The model names an element id, and the library performs the input with CDP.

[browser-use](/p/browser-use/) exposes a fixed action catalogue: `click`, `input_text`, `scroll`, `send_keys`, `upload_file`, `switch_tab`, `navigate`, `extract` and more. Elements are targeted by the numeric index from its DOM dump. Clicks go through an event bus to a watchdog that checks occlusion and sends `Input.dispatchMouseEvent`, with a JS-click fallback. For some Claude, Gemini 3 Pro and Browser Use models, coordinate clicking is also enabled. One model reply may contain up to five actions. The queue stops when the URL or tab changes, or after an action marked `terminates_sequence`.

[Stagehand](/p/stagehand/) asks for exactly one `{elementId, method, arguments}` per inference. The method comes from an enum (click, fill, type, press, scroll, selectOption, hover, drag and similar). The id becomes an XPath, and its CDP-based `Locator` resolves it, scrolls it into view, reads the box model and clicks at the centre. A `twoStep` flag covers custom dropdowns. File inputs are filled with in-page `File` objects.

Use browser-use when the model should run whole tasks with many action types and tabs. Use Stagehand when you want individual, auditable actions inside your own script.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/action-execution/index.md

## How is the agent loop / planning implemented?

This is where the two designs differ most.

[browser-use](/p/browser-use/) is an agent. `Agent.run` loops through `step()`: capture state, call the LLM for an `AgentOutput` (thinking, evaluation of the last goal, memory, next goal, optional plan update, action list), execute, then post-process. Planning is optional. It keeps a plan list with statuses, and nudges the model to replan after repeated failures or to make a plan after five steps without one. Memory is the message history, compacted by an LLM summary on long runs, plus a free-text `memory` field. The run stops on `done`, at `max_steps` (where the schema is forced to `done`), after `max_failures` consecutive failures, or on a user stop.

[Stagehand](/p/stagehand/) has no loop at this commit, and its SDK has no `agent()` method. `act`, `observe` and `extract` each make one or two structured LLM calls and keep no memory between calls. Multi-step behaviour is left to your own code, or to an agent harness that uses Stagehand's integration tools (`run`, `snapshot`, `screenshot`) over MCP or natively.

Choose browser-use for "give it a goal and let it go" tasks. Choose Stagehand if you already have an agent framework or prefer scripted control flow with AI only at the fuzzy steps.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/agent-loop/index.md

## How are failures, retries and self-healing handled?

Both projects bound each operation with timeouts. They differ in where recovery happens.

[browser-use](/p/browser-use/) recovers inside the loop. Steps time out after 180 s and actions after 180 s by default. LLM calls get a per-model timeout. An empty action list gets one clarification retry. Rate-limit, auth, 5xx and truncation errors can switch the run once to a `fallback_llm`. Dropped CDP WebSockets get three reconnect attempts. An `ActionLoopDetector` watches repeated actions and unchanged pages and adds escalating hints to the prompt. After `max_failures` consecutive failed steps, one final recovery call is made. Nothing is cached between runs. `rerun_history` can replay a saved run and re-match elements.

[Stagehand](/p/stagehand/) recovers per call. With `selfHeal`, a failed action triggers a fresh snapshot and a new inference. `TimeoutError` is always passed to the caller. Before each snapshot it waits for DOM and network quiet. Its main reliability tool is a server-side action cache keyed on the accessibility tree. A hit replays the stored actions without the LLM and falls back to inference if replay fails. The cache needs a Browserbase API key and session.

Use browser-use for long unattended runs that must keep going. Use Stagehand's cache when the same flows repeat often on Browserbase.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/reliability/index.md

## Which models are supported and how are they called?

Both projects require structured output. They differ in how many providers they reach and how.

[browser-use](/p/browser-use/) defines a small `BaseChatModel` protocol (`ainvoke(messages, output_format)`) with native adapters for about fifteen backends. These include OpenAI, Anthropic, Google, Azure, Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, Cerebras, LiteLLM and its own hosted `ChatBrowserUse`, which is the default when no model is given. The agent always asks for a Pydantic `AgentOutput` whose action field is a discriminated union. Vision is on by default and switched off for DeepSeek and some Grok models. Timeouts, screenshot size and coordinate clicking are tuned by model name.

[Stagehand](/p/stagehand/) calls models from inside its browser extension through the Vercel AI SDK. Five providers are supported (OpenAI, Anthropic, Google, Groq, Cerebras), with allow-listed model ids. Every call uses JSON Schema structured output generated from Zod. Without a provider key, calls go through the Browserbase Model Gateway. A client-side `generate` hook lets you route inference through any model you control. Vision is only used for optional `extract` screenshots.

Pick browser-use for the widest provider choice, including local Ollama. Pick Stagehand when a short list of frontier providers, or your own `generate` function, is enough.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/models/index.md

## How are browser sessions, profiles, auth and anti-bot handled?

Both projects run local Chrome over CDP or a vendor cloud browser. Anti-bot features are mostly in the paid clouds.

[browser-use](/p/browser-use/) configures everything in `BrowserProfile`. Settings include `user_data_dir` for persistent profiles, `storage_state` files for cookies, HTTP/SOCKS proxies and allowed or prohibited domain lists, which a security watchdog enforces. Local Chrome starts with a list of flags such as `--disable-blink-features=AutomationControlled`. `use_cloud=True` requests a Browser Use Cloud browser with proxies, country selection and fingerprinting. When that cloud's solver is running, the agent pauses on CAPTCHA events and reports the result to the model.

[Stagehand](/p/stagehand/) offers two factories. `localBrowser.launch` starts Chrome with a temp or fixed `userDataDir`, a proxy flag and headless mode, then loads its extension. `browserbase.launch` creates a Browserbase session. Its options include advanced stealth, fingerprint settings, managed or external proxies, regions, CAPTCHA solving with selector hints, persistent contexts, recording and "verified" sessions. Cookies and extra headers are set through context methods. Any browser it drives must accept the Stagehand extension.

For self-hosted runs, both offer only profiles and proxies. Choose by cloud: Browser Use Cloud with browser-use, or Browserbase with Stagehand. Note that Stagehand needs a browser that can load extensions.

More projects in this category are being researched.

Per-project answers: https://llms-technical-reviews.com/browser-control/q/sessions/index.md

## Projects

- [browser-use/browser-use](https://llms-technical-reviews.com/p/browser-use/index.md) — Async Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP.
- [browserbase/stagehand](https://llms-technical-reviews.com/p/stagehand/index.md) — Browser automation SDK (TS, Python, Go) whose act/observe/extract run inside a Chrome extension that drives pages over CDP.