How are actions executed and how are elements targeted?
CDP / Playwright / OS-level input; selectors vs indexes vs coordinates; typing, scrolling, file upload, tabs.
Verdict
Both projects keep the model away from raw coordinates by default. The model names an element id, and the library performs the input with CDP.
browser-use exposes a fixed action catalogue: click, input_text, scroll, send_keys, upload_file, switch_tab, navigate, extract and more. Elements are targeted by the numeric index from its DOM dump. Clicks go through an event bus to a watchdog that checks occlusion and sends Input.dispatchMouseEvent, with a JS-click fallback. For some Claude, Gemini 3 Pro and Browser Use models, coordinate clicking is also enabled. One model reply may contain up to five actions. The queue stops when the URL or tab changes, or after an action marked terminates_sequence.
Stagehand asks for exactly one {elementId, method, arguments} per inference. The method comes from an enum (click, fill, type, press, scroll, selectOption, hover, drag and similar). The id becomes an XPath, and its CDP-based Locator resolves it, scrolls it into view, reads the box model and clicks at the centre. A twoStep flag covers custom dropdowns. File inputs are filled with in-page File objects.
Use browser-use when the model should run whole tasks with many action types and tabs. Use Stagehand when you want individual, auditable actions inside your own script.
More projects in this category are being researched.
Per-project answers
browser-use/browser-use
answeredArchitecture. Actions are dispatched through the Tools (browser_use/tools/service.py) → Registry.execute_action() → event bus → handler chain. Tools.act() (browser_use/tools/service.py:2178-2236) receives an ActionModel (a Pydantic union discriminated by action name), extracts the action name and params, and calls registry.execute_action() under an asyncio.wait_for timeout (env BROWSER_USE_ACTION_TIMEOUT_S, default 180s). The registry maps action names to registered handler functions, which typically dispatch events on the browser session's event bus.
Element targeting (index-based). Most actions reference elements by a numeric index from the selector_map — the same indices shown in the DOM serialization. The handler looks up the EnhancedDOMTreeNode and resolves it to a CDP backend_node_id + frame_id. This avoids fragile XPath or CSS selectors. For models that support it (Claude Sonnet 4+, Opus 4+, Gemini 2.5 Pro), coordinate-based clicking is available via coordinate_x/coordinate_y in ClickElementAction (browser_use/tools/views.py:69-75). Enabled automatically at agent init (browser_use/agent/service.py:327-333).
Available action types (browser_use/tools/views.py): navigate (with new_tab option), click (by index or coordinates), input_text (with clear flag to append), scroll (direction + page fraction + optional element index), send_keys (keyboard shortcuts like Control+o), upload_file (index + file path), switch_tab / close_tab (4-char tab ID), search (Google/Bing/DuckDuckGo), search_page (regex/literal text search in page), find_elements (CSS selector), extract (LLM-based extraction from page markdown), screenshot (save to file), save_as_pdf, get_dropdown_options, done (terminate with result).
Execution flow. Agent.multi_act() (browser_use/agent/service.py:2728-2848) runs sequential actions from a single LLM response. It enforces two page-change guards: (1) static terminates_sequence=True on actions like navigate/search/go_back/switch; (2) runtime detection comparing pre/post-action URL and focused target ID. Any page change aborts remaining queued actions. A wait_between_actions delay (from BrowserProfile) is applied between steps. send_keys dispatches keyboard events via CDP. File upload uses CDP's input.dispatchFile or sets the file input value. Tab management uses CDP Target.attachToTarget/Target.detachFromTarget. Scroll is implemented via CDP JavaScript window.scrollBy() or element.scrollIntoView().
Multi-step actions per LLM call. Controlled by max_actions_per_step (default 5). The AgentOutput schema allows multiple actions per response. Each action is dispatched sequentially; if one errors, the remaining actions are still attempted (error is recorded, not fatal).
multi_act stops at the first failing action (tools/service.py ~L2809). Remaining actions are not run. The coordinate-click model check matches gemini-3-pro, not Gemini 2.5 Pro.browserbase/stagehand
answeredActions are never executed by the LLM directly — the LLM returns a structured decision (elementId, method, arguments, twoStep flag), and actService.ts executes it deterministically via CDP primitives. The twoStep flag enables two-phase actions (e.g., clicking to expand a non-<select> dropdown, then choosing the option).
From decision to execution: The LLM returns an ActInferenceSchema object with an elementId (like "0-18372"). normalizeActInferenceElement looks up the ID in the combinedXpathMap to get an XPath, then wraps it as xpath=/.... The takeDeterministicAction function calls performUnderstudyMethod (actHandlerUtils.ts), which resolves the XPath to a Locator object (via resolveLocatorWithHops for cross-iframe support) and dispatches to the appropriate method handler.
Method handlers are mapped in METHOD_HANDLER_MAP and include: click, doubleClick, fill, type, press (key), scrollTo, scrollIntoView, mouse.wheel, nextChunk/prevChunk (scroll by element height), selectOption, hover, dragAndDrop.
CDP-level execution: The Locator class (understudy/locator.ts) resolves the selector to an objectId inside an isolated world (Page.createIsolatedWorld), then uses CDP commands: DOM.scrollIntoViewIfNeeded, DOM.getBoxModel (to find the element's center coordinates), and Input.dispatchMouseEvent for clicks (with mouseMoved/mousePressed/mouseReleased events). Typing uses Input.insertText (efficient) or per-character Input.dispatchKeyEvent. Filling uses a JavaScript function injected via Runtime.callFunctionOn that sets the element's value; if the element needs IME/character-level input, it falls back to Input.insertText. File uploads construct File objects in-page via assignFilePayloadsToInputElement and assign them to <input type="file">. Scrolling uses Runtime.callFunctionOn with JavaScript that scrolls the element/window by its height.
Selectors: The LLM returns encoded element IDs, which are resolved to xpath=... selectors. The Locator class supports CSS selectors, XPath expressions, and >>-delimited iframe hops (via deepLocator.ts). Coordinates are never sent to the LLM — they are computed server-side from DOM.getBoxModel.
Tabs: page.goto, page.goBack, page.goForward use CDP Navigation.goto/Navigation.goBack. The BrowserContext manages pages as top-level CDP targets.
← How is the page represented to the model? · How is the agent loop / planning implemented? →