# Browser & computer control

> Frameworks that let an LLM drive a browser (or a whole desktop/phone) to complete tasks.

These projects let a language model operate a real browser, or in some cases a whole desktop or phone. The model gets a view of the screen, picks an action, and the framework carries it out with clicks, typing and navigation. Designs range from full autonomous agents with their own step loop, to SDKs that add a single AI-driven `act` or `extract` call to an ordinary script.

When choosing, check five things. How is the page shown to the model: DOM text, accessibility tree, screenshots, or a mix? That decides token cost and which models work. Who owns the loop: the framework, or your code or agent harness? What happens when an action fails? Look for retries, self-healing and caching of known-good actions. Which model providers are supported? And are stealth, proxies and CAPTCHA handling in the open-source code, or only in a paid cloud browser?


## Projects

- [browser-use/browser-use](https://llms-technical-reviews.com/p/browser-use/) — Async Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP. (★117268, Python)
- [browserbase/stagehand](https://llms-technical-reviews.com/p/stagehand/) — Browser automation SDK (TS, Python, Go) whose act/observe/extract run inside a Chrome extension that drives pages over CDP. (★25549, TypeScript)

In the research queue (not yet published): browser-use/browser-harness, browser-use/workflow-use, browser-use/jev-ultrafast, Skyvern-AI/skyvern, hyperbrowserai/HyperAgent, trycua/cua, lahfir/agent-desktop, awlevin/typesafe-computer-use, omdsh-dev/dsh-browser, jkudish/jev-browser, droidrun/mobile-jev, ekzhang/openjev-sglang, savka777/jev-use, razaanstha/ulka.

## Comparison questions

- [How is the page represented to the model?](https://llms-technical-reviews.com/browser-control/q/page-perception/) — DOM serialization, accessibility tree, screenshots, set-of-marks/element indexes; size limits and pruning.
- [How are actions executed and how are elements targeted?](https://llms-technical-reviews.com/browser-control/q/action-execution/) — CDP / Playwright / OS-level input; selectors vs indexes vs coordinates; typing, scrolling, file upload, tabs.
- [How is the agent loop / planning implemented?](https://llms-technical-reviews.com/browser-control/q/agent-loop/) — Step loop; planner vs executor; tool-call schema; stop conditions; memory between steps.
- [How are failures, retries and self-healing handled?](https://llms-technical-reviews.com/browser-control/q/reliability/) — Error classes caught; retries; replanning; caching of successful actions or workflows; timeouts.
- [Which models are supported and how are they called?](https://llms-technical-reviews.com/browser-control/q/models/) — Providers; vision requirement; structured output / tool calling; small or specialised models.
- [How are browser sessions, profiles, auth and anti-bot handled?](https://llms-technical-reviews.com/browser-control/q/sessions/) — Local vs remote/cloud browsers; persistent profiles and cookies; stealth; proxies; CAPTCHA handling.

Full comparison: https://llms-technical-reviews.com/compare/browser-control/