LLMs Technical Reviews

Browser & computer control: browser-use vs stagehand

Frameworks that let an LLM drive a browser (or a whole desktop/phone) to complete tasks. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.

At a glance

● answered from code · — not applicable (the project does not do this) · ? insufficient evidence

How is the page represented to the model?

Both projects send the model a text rendering of the page built from CDP, not raw HTML. They differ in how much they add on top.

browser-use merges DOMSnapshot.captureSnapshot, the full DOM and the accessibility tree. Its DOMTreeSerializer drops elements hidden by paint order and redundant wrappers, then prints interactive nodes as [index]<tag attr=...>. The index is the CDP backend_node_id where possible, so it stays stable across steps. New elements are starred. A screenshot is taken every step and sent when use_vision is on. Serialized text is capped at 40,000 characters.

Stagehand builds a "hybrid snapshot": an accessibility outline merged with DOM data, with iframes stitched in and shadow DOM pierced. Each line carries a frameOrdinal-backendNodeId id that maps to an XPath. Callers can scope it to a locator or exclude parts. There are no screenshots in act or observe. extract can attach a plain viewport PNG on request, with no set-of-marks overlay.

Pick browser-use when pages are visual or ambiguous and a vision model is worth the extra tokens on every step. Pick Stagehand for cheaper, text-only prompts scoped to part of a page, especially for extraction.

More projects in this category are being researched.

How are actions executed and how are elements targeted?

Both projects keep the model away from raw coordinates by default. The model names an element id, and the library performs the input with CDP.

browser-use exposes a fixed action catalogue: click, input_text, scroll, send_keys, upload_file, switch_tab, navigate, extract and more. Elements are targeted by the numeric index from its DOM dump. Clicks go through an event bus to a watchdog that checks occlusion and sends Input.dispatchMouseEvent, with a JS-click fallback. For some Claude, Gemini 3 Pro and Browser Use models, coordinate clicking is also enabled. One model reply may contain up to five actions. The queue stops when the URL or tab changes, or after an action marked terminates_sequence.

Stagehand asks for exactly one {elementId, method, arguments} per inference. The method comes from an enum (click, fill, type, press, scroll, selectOption, hover, drag and similar). The id becomes an XPath, and its CDP-based Locator resolves it, scrolls it into view, reads the box model and clicks at the centre. A twoStep flag covers custom dropdowns. File inputs are filled with in-page File objects.

Use browser-use when the model should run whole tasks with many action types and tabs. Use Stagehand when you want individual, auditable actions inside your own script.

More projects in this category are being researched.

How is the agent loop / planning implemented?

This is where the two designs differ most.

browser-use is an agent. Agent.run loops through step(): capture state, call the LLM for an AgentOutput (thinking, evaluation of the last goal, memory, next goal, optional plan update, action list), execute, then post-process. Planning is optional. It keeps a plan list with statuses, and nudges the model to replan after repeated failures or to make a plan after five steps without one. Memory is the message history, compacted by an LLM summary on long runs, plus a free-text memory field. The run stops on done, at max_steps (where the schema is forced to done), after max_failures consecutive failures, or on a user stop.

Stagehand has no loop at this commit, and its SDK has no agent() method. act, observe and extract each make one or two structured LLM calls and keep no memory between calls. Multi-step behaviour is left to your own code, or to an agent harness that uses Stagehand's integration tools (run, snapshot, screenshot) over MCP or natively.

Choose browser-use for "give it a goal and let it go" tasks. Choose Stagehand if you already have an agent framework or prefer scripted control flow with AI only at the fuzzy steps.

More projects in this category are being researched.

How are failures, retries and self-healing handled?

Both projects bound each operation with timeouts. They differ in where recovery happens.

browser-use recovers inside the loop. Steps time out after 180 s and actions after 180 s by default. LLM calls get a per-model timeout. An empty action list gets one clarification retry. Rate-limit, auth, 5xx and truncation errors can switch the run once to a fallback_llm. Dropped CDP WebSockets get three reconnect attempts. An ActionLoopDetector watches repeated actions and unchanged pages and adds escalating hints to the prompt. After max_failures consecutive failed steps, one final recovery call is made. Nothing is cached between runs. rerun_history can replay a saved run and re-match elements.

Stagehand recovers per call. With selfHeal, a failed action triggers a fresh snapshot and a new inference. TimeoutError is always passed to the caller. Before each snapshot it waits for DOM and network quiet. Its main reliability tool is a server-side action cache keyed on the accessibility tree. A hit replays the stored actions without the LLM and falls back to inference if replay fails. The cache needs a Browserbase API key and session.

Use browser-use for long unattended runs that must keep going. Use Stagehand's cache when the same flows repeat often on Browserbase.

More projects in this category are being researched.

Which models are supported and how are they called?

Both projects require structured output. They differ in how many providers they reach and how.

browser-use defines a small BaseChatModel protocol (ainvoke(messages, output_format)) with native adapters for about fifteen backends. These include OpenAI, Anthropic, Google, Azure, Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, Cerebras, LiteLLM and its own hosted ChatBrowserUse, which is the default when no model is given. The agent always asks for a Pydantic AgentOutput whose action field is a discriminated union. Vision is on by default and switched off for DeepSeek and some Grok models. Timeouts, screenshot size and coordinate clicking are tuned by model name.

Stagehand calls models from inside its browser extension through the Vercel AI SDK. Five providers are supported (OpenAI, Anthropic, Google, Groq, Cerebras), with allow-listed model ids. Every call uses JSON Schema structured output generated from Zod. Without a provider key, calls go through the Browserbase Model Gateway. A client-side generate hook lets you route inference through any model you control. Vision is only used for optional extract screenshots.

Pick browser-use for the widest provider choice, including local Ollama. Pick Stagehand when a short list of frontier providers, or your own generate function, is enough.

More projects in this category are being researched.

How are browser sessions, profiles, auth and anti-bot handled?

Both projects run local Chrome over CDP or a vendor cloud browser. Anti-bot features are mostly in the paid clouds.

browser-use configures everything in BrowserProfile. Settings include user_data_dir for persistent profiles, storage_state files for cookies, HTTP/SOCKS proxies and allowed or prohibited domain lists, which a security watchdog enforces. Local Chrome starts with a list of flags such as --disable-blink-features=AutomationControlled. use_cloud=True requests a Browser Use Cloud browser with proxies, country selection and fingerprinting. When that cloud's solver is running, the agent pauses on CAPTCHA events and reports the result to the model.

Stagehand offers two factories. localBrowser.launch starts Chrome with a temp or fixed userDataDir, a proxy flag and headless mode, then loads its extension. browserbase.launch creates a Browserbase session. Its options include advanced stealth, fingerprint settings, managed or external proxies, regions, CAPTCHA solving with selector hints, persistent contexts, recording and "verified" sessions. Cookies and extra headers are set through context methods. Any browser it drives must accept the Stagehand extension.

For self-hosted runs, both offer only profiles and proxies. Choose by cloud: Browser Use Cloud with browser-use, or Browserbase with Stagehand. Note that Stagehand needs a browser that can load extensions.

More projects in this category are being researched.

The projects