LLMs Technical Reviews
Home / Browser & computer control

Browser & computer control

Frameworks that let an LLM drive a browser (or a whole desktop/phone) to complete tasks.

These projects let a language model operate a real browser, or in some cases a whole desktop or phone. The model gets a view of the screen, picks an action, and the framework carries it out with clicks, typing and navigation. Designs range from full autonomous agents with their own step loop, to SDKs that add a single AI-driven act or extract call to an ordinary script.

When choosing, check five things. How is the page shown to the model: DOM text, accessibility tree, screenshots, or a mix? That decides token cost and which models work. Who owns the loop: the framework, or your code or agent harness? What happens when an action fails? Look for retries, self-healing and caching of known-good actions. Which model providers are supported? And are stealth, proxies and CAPTCHA handling in the open-source code, or only in a paid cloud browser?

Projects (2)

ProjectStarsLanguageLicense
browser-use/browser-useAsync Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP.★ 117kPythonMIT
browserbase/stagehandBrowser automation SDK (TS, Python, Go) whose act/observe/extract run inside a Chrome extension that drives pages over CDP.★ 26kTypeScriptMIT

In the research queue: browser-use/browser-harness, browser-use/workflow-use, browser-use/jev-ultrafast, Skyvern-AI/skyvern, hyperbrowserai/HyperAgent, trycua/cua, lahfir/agent-desktop, awlevin/typesafe-computer-use, omdsh-dev/dsh-browser, jkudish/jev-browser, droidrun/mobile-jev, ekzhang/openjev-sglang, savka777/jev-use, razaanstha/ulka.

Comparison questions

Each question is answered separately for every project in this category, from that project's source code.

All verdicts on one page →

  1. How is the page represented to the model?DOM serialization, accessibility tree, screenshots, set-of-marks/element indexes; size limits and pruning.
  2. How are actions executed and how are elements targeted?CDP / Playwright / OS-level input; selectors vs indexes vs coordinates; typing, scrolling, file upload, tabs.
  3. How is the agent loop / planning implemented?Step loop; planner vs executor; tool-call schema; stop conditions; memory between steps.
  4. How are failures, retries and self-healing handled?Error classes caught; retries; replanning; caching of successful actions or workflows; timeouts.
  5. Which models are supported and how are they called?Providers; vision requirement; structured output / tool calling; small or specialised models.
  6. How are browser sessions, profiles, auth and anti-bot handled?Local vs remote/cloud browsers; persistent profiles and cookies; stealth; proxies; CAPTCHA handling.