# How are actions executed and how are elements targeted?

> Browser & computer control — a good answer covers: CDP / Playwright / OS-level input; selectors vs indexes vs coordinates; typing, scrolling, file upload, tabs.

Canonical page: https://llms-technical-reviews.com/browser-control/q/action-execution/

## Verdict

Both projects keep the model away from raw coordinates by default. The model names an element id, and the library performs the input with CDP.

[browser-use](/p/browser-use/) exposes a fixed action catalogue: `click`, `input_text`, `scroll`, `send_keys`, `upload_file`, `switch_tab`, `navigate`, `extract` and more. Elements are targeted by the numeric index from its DOM dump. Clicks go through an event bus to a watchdog that checks occlusion and sends `Input.dispatchMouseEvent`, with a JS-click fallback. For some Claude, Gemini 3 Pro and Browser Use models, coordinate clicking is also enabled. One model reply may contain up to five actions. The queue stops when the URL or tab changes, or after an action marked `terminates_sequence`.

[Stagehand](/p/stagehand/) asks for exactly one `{elementId, method, arguments}` per inference. The method comes from an enum (click, fill, type, press, scroll, selectOption, hover, drag and similar). The id becomes an XPath, and its CDP-based `Locator` resolves it, scrolls it into view, reads the box model and clicks at the centre. A `twoStep` flag covers custom dropdowns. File inputs are filled with in-page `File` objects.

Use browser-use when the model should run whole tasks with many action types and tabs. Use Stagehand when you want individual, auditable actions inside your own script.

More projects in this category are being researched.

## Per-project answers

### browser-use/browser-use (answered)

**Architecture.** Actions are dispatched through the `Tools` (`browser_use/tools/service.py`) → `Registry.execute_action()` → event bus → handler chain. `Tools.act()` (`browser_use/tools/service.py:2178-2236`) receives an `ActionModel` (a Pydantic union discriminated by action name), extracts the action name and params, and calls `registry.execute_action()` under an `asyncio.wait_for` timeout (env `BROWSER_USE_ACTION_TIMEOUT_S`, default 180s). The registry maps action names to registered handler functions, which typically dispatch events on the browser session's event bus.

**Element targeting (index-based).** Most actions reference elements by a numeric `index` from the `selector_map` — the same indices shown in the DOM serialization. The handler looks up the `EnhancedDOMTreeNode` and resolves it to a CDP `backend_node_id` + `frame_id`. This avoids fragile XPath or CSS selectors. For models that support it (Claude Sonnet 4+, Opus 4+, Gemini 2.5 Pro), coordinate-based clicking is available via `coordinate_x`/`coordinate_y` in `ClickElementAction` (`browser_use/tools/views.py:69-75`). Enabled automatically at agent init (`browser_use/agent/service.py:327-333`).

**Available action types** (`browser_use/tools/views.py`): `navigate` (with `new_tab` option), `click` (by index or coordinates), `input_text` (with `clear` flag to append), `scroll` (direction + page fraction + optional element index), `send_keys` (keyboard shortcuts like `Control+o`), `upload_file` (index + file path), `switch_tab` / `close_tab` (4-char tab ID), `search` (Google/Bing/DuckDuckGo), `search_page` (regex/literal text search in page), `find_elements` (CSS selector), `extract` (LLM-based extraction from page markdown), `screenshot` (save to file), `save_as_pdf`, `get_dropdown_options`, `done` (terminate with result).

**Execution flow.** `Agent.multi_act()` (`browser_use/agent/service.py:2728-2848`) runs sequential actions from a single LLM response. It enforces two page-change guards: (1) static `terminates_sequence=True` on actions like navigate/search/go_back/switch; (2) runtime detection comparing pre/post-action URL and focused target ID. Any page change aborts remaining queued actions. A `wait_between_actions` delay (from `BrowserProfile`) is applied between steps. `send_keys` dispatches keyboard events via CDP. File upload uses CDP's `input.dispatchFile` or sets the file input value. Tab management uses CDP `Target.attachToTarget`/`Target.detachFromTarget`. Scroll is implemented via CDP JavaScript `window.scrollBy()` or `element.scrollIntoView()`.

**Multi-step actions per LLM call.** Controlled by `max_actions_per_step` (default 5). The `AgentOutput` schema allows multiple actions per response. Each action is dispatched sequentially; if one errors, the remaining actions are still attempted (error is recorded, not fatal).

> **Editor's note.** Correction: `multi_act` stops at the first failing action (`tools/service.py` ~L2809). Remaining actions are not run. The coordinate-click model check matches `gemini-3-pro`, not Gemini 2.5 Pro.

Citations: [browser_use/tools/service.py:2178-2236](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L2178-L2236) · [browser_use/tools/views.py:69-75](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/views.py#L69-L75) · [browser_use/tools/views.py:122-143](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/views.py#L122-L143) · [browser_use/agent/service.py:2728-2848](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2728-L2848) · [browser_use/agent/service.py:327-333](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L327-L333) · [browser_use/agent/service.py:2766-2828](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2766-L2828)

### browserbase/stagehand (answered)

Actions are **never executed by the LLM directly** — the LLM returns a structured decision (`elementId`, `method`, `arguments`, `twoStep` flag), and `actService.ts` executes it deterministically via CDP primitives. The `twoStep` flag enables two-phase actions (e.g., clicking to expand a non-`<select>` dropdown, then choosing the option).

**From decision to execution**: The LLM returns an `ActInferenceSchema` object with an `elementId` (like `"0-18372"`). `normalizeActInferenceElement` looks up the ID in the `combinedXpathMap` to get an XPath, then wraps it as `xpath=/...`. The `takeDeterministicAction` function calls `performUnderstudyMethod` (`actHandlerUtils.ts`), which resolves the XPath to a `Locator` object (via `resolveLocatorWithHops` for cross-iframe support) and dispatches to the appropriate method handler.

**Method handlers** are mapped in `METHOD_HANDLER_MAP` and include: `click`, `doubleClick`, `fill`, `type`, `press` (key), `scrollTo`, `scrollIntoView`, `mouse.wheel`, `nextChunk`/`prevChunk` (scroll by element height), `selectOption`, `hover`, `dragAndDrop`.

**CDP-level execution**: The `Locator` class (`understudy/locator.ts`) resolves the selector to an `objectId` inside an isolated world (`Page.createIsolatedWorld`), then uses CDP commands: `DOM.scrollIntoViewIfNeeded`, `DOM.getBoxModel` (to find the element's center coordinates), and `Input.dispatchMouseEvent` for clicks (with mouseMoved/mousePressed/mouseReleased events). Typing uses `Input.insertText` (efficient) or per-character `Input.dispatchKeyEvent`. Filling uses a JavaScript function injected via `Runtime.callFunctionOn` that sets the element's value; if the element needs IME/character-level input, it falls back to `Input.insertText`. File uploads construct `File` objects in-page via `assignFilePayloadsToInputElement` and assign them to `<input type="file">`. Scrolling uses `Runtime.callFunctionOn` with JavaScript that scrolls the element/window by its height.

**Selectors**: The LLM returns encoded element IDs, which are resolved to `xpath=...` selectors. The `Locator` class supports CSS selectors, XPath expressions, and `>>`-delimited iframe hops (via `deepLocator.ts`). Coordinates are never sent to the LLM — they are computed server-side from `DOM.getBoxModel`.

**Tabs**: `page.goto`, `page.goBack`, `page.goForward` use CDP `Navigation.goto`/`Navigation.goBack`. The `BrowserContext` manages pages as top-level CDP targets.


Citations: [packages/extension/services/actService.ts:286-328](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/services/actService.ts#L286-L328) · [packages/extension/handlers/handlerUtils/actHandlerUtils.ts:47-116](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/handlers/handlerUtils/actHandlerUtils.ts#L47-L116) · [packages/extension/understudy/locator.ts:393-464](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/locator.ts#L393-L464) · [packages/extension/understudy/locator.ts:680-727](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/locator.ts#L680-L727) · [packages/extension/types/private/handlers.ts:1-14](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/types/private/handlers.ts#L1-L14) · [packages/extension/handlers/handlerUtils/actHandlerUtils.ts:120-137](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/handlers/handlerUtils/actHandlerUtils.ts#L120-L137)
