# How is the page represented to the model?

> Browser & computer control — a good answer covers: DOM serialization, accessibility tree, screenshots, set-of-marks/element indexes; size limits and pruning.

Canonical page: https://llms-technical-reviews.com/browser-control/q/page-perception/

## Verdict

Both projects send the model a text rendering of the page built from CDP, not raw HTML. They differ in how much they add on top.

[browser-use](/p/browser-use/) merges `DOMSnapshot.captureSnapshot`, the full DOM and the accessibility tree. Its `DOMTreeSerializer` drops elements hidden by paint order and redundant wrappers, then prints interactive nodes as `[index]<tag attr=...>`. The index is the CDP `backend_node_id` where possible, so it stays stable across steps. New elements are starred. A screenshot is taken every step and sent when `use_vision` is on. Serialized text is capped at 40,000 characters.

[Stagehand](/p/stagehand/) builds a "hybrid snapshot": an accessibility outline merged with DOM data, with iframes stitched in and shadow DOM pierced. Each line carries a `frameOrdinal-backendNodeId` id that maps to an XPath. Callers can scope it to a locator or exclude parts. There are no screenshots in `act` or `observe`. `extract` can attach a plain viewport PNG on request, with no set-of-marks overlay.

Pick browser-use when pages are visual or ambiguous and a vision model is worth the extra tokens on every step. Pick Stagehand for cheaper, text-only prompts scoped to part of a page, especially for extraction.

More projects in this category are being researched.

## Per-project answers

### browser-use/browser-use (answered)

**DOM Serialization.** The agent represents pages through `SerializedDOMState` (`browser_use/dom/views.py:935`), which wraps a tree of `SimplifiedNode` objects built from three CDP sources: `DOMSnapshot.captureSnapshot`, the full accessibility tree via `accessibility.GetFullAXTree`, and the DOM document structure (`browser_use/dom/service.py:6-25`). The `DomService` class orchestrates this, combining snapshot bounds, AX node properties, and DOM nodes into `EnhancedDOMTreeNode` objects with sibling/shadow-root/iframe support. Cross-origin iframes with dimensions ≥ 10px each side are included (`browser_use/dom/service.py:40-42`).

**Serialization to text.** `DOMTreeSerializer.serialize_accessible_elements()` (`browser_use/dom/serializer/serializer.py:114-165`) runs a pipeline: (1) create a simplified tree by filtering non-content elements (style, script, SVG decorative children) and detecting interactive elements; (2) paint-order filtering to skip elements visually occluded; (3) tree optimization (remove unnecessary parents); (4) bounding-box filtering that excludes children contained within propagating parents (`a`, `button`, `div[role=button]`) at 99% containment threshold; (5) assign numeric interactive indices to clickable/scrollable elements. The final `serialize_tree()` (`browser_use/dom/serializer/serializer.py:989`) renders a compact text format: `[index]<tagname attr="val" />` for interactive elements, ordinary `<tagname>` for containers, `|scroll element|` markers for scrollable containers, `|SHADOW(open/closed)|` for shadow-DOM hosts, text nodes inline, and iframes with hidden-element hints showing scroll distances. Compound controls (date pickers, range sliders, file inputs, select dropdowns) get synthetic `compound_components=...` attributes with sub-role information (`browser_use/dom/serializer/serializer.py:1059-1100`).

**Element indices (set-of-marks).** Interactive elements get sequentially numbered indices stored in a `DOMSelectorMap` (dict of `int → EnhancedDOMTreeNode`). The LLM sees `[1]<button type="submit">` and uses that index to click. New elements since the last step are marked with a `*` prefix. `max_clickable_elements_length` caps serialized text at 40000 chars (`browser_use/agent/views.py:92`).

**Screenshots.** A PNG screenshot is always captured every step via `BrowserSession.get_browser_state_summary(include_screenshot=True)` (`browser_use/agent/service.py:1096-1099`). When `use_vision` is enabled, the screenshot is sent as an image content part to the LLM. For Claude Sonnet models, screenshots are auto-resized to 1400×850 to fit context windows (`browser_use/agent/service.py:247-250`). Configurable via `llm_screenshot_size` parameter.

**Accessibility tree.** The AX tree provides `name`, `role`, `description`, `properties` (valuemin/max/now, expanded, pressed, invalid, etc.) per element. These feed into `ClickableElementDetector.is_interactive()` (`browser_use/dom/serializer/clickable_elements.py:6-52`) which heuristically scores elements as interactive based on tag, role, JS click listeners, ARIA attributes, pointer cursor, and form-control heuristics.

> **Editor's note.** Correction: element indexes are the CDP `backend_node_id` when that is unique, not a sequential counter, so they mostly stay stable between steps.

Citations: [browser_use/dom/service.py:1-72](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/service.py#L1-L72) · [browser_use/dom/views.py:935-955](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/views.py#L935-L955) · [browser_use/dom/serializer/serializer.py:114-165](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L114-L165) · [browser_use/dom/serializer/serializer.py:989-1100](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/serializer.py#L989-L1100) · [browser_use/agent/service.py:245-252](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L245-L252) · [browser_use/dom/serializer/clickable_elements.py:1-52](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/dom/serializer/clickable_elements.py#L1-L52)

### browserbase/stagehand (answered)

The page is represented to the model as a **hybrid text snapshot** combining the Chrome accessibility tree and the DOM tree. The core mechanism lives in `capture.ts`, which calls `Accessibility.getFullAXTree` via CDP for each frame's a11y tree and `DOM.getDocument` for the DOM structure. These are merged into a single hierarchical text outline where each element is rendered as a line like `[0-18372] button: Submit` — the encoded ID (`frameOrdinal-backendNodeId`) serves as the element index that the LLM returns when choosing an action. An xpathMap (a `Record<string, string>`) maps each encoded ID to an absolute XPath, which the execution layer uses to locate the element for CDP interaction.

Iframes are handled via a multi-step pipeline: `buildSessionIndexes` calls `DOM.getDocument` once per CDP session; per-frame maps are sliced from the shared index; and `computeFramePrefixes` walks the frame tree to compute absolute XPath prefixes for child frames. The per-frame outlines are stitched into a single `combinedTree` via `injectSubtrees` in `treeFormatUtils.ts`. Shadow DOM is pierced by default (`pierceShadow: true`).

**Size limits and pruning**: The DOM tree retrieval adaptively retries with shallower depths when CDP's CBOR encoder stack overflows — `DOM_DEPTH_ATTEMPTS` in `domTree.ts` tries `[-1, 256, 128, 64, 32, 16, 8, 4, 2, 1]`, and each truncated node is hydrated individually via `DOM.describeNode` with its own depth fallback. Structural AX roles (`generic`, `none`, `InlineTextBox`) are pruned from the outline. Nodes can also be excluded via `ignoreLocators`.

**Screenshots** are optional and only used in `extract()` when `options.screenshot: true`. They are captured as raw PNG data (`page.screenshot` → `Page.screenshot`) and sent alongside the DOM text as an image content block. The screenshot is *not* annotated with bounding boxes or set-of-marks overlays — the `AnnotatedScreenshotText` constant in `LLMClient.ts` describes a planned annotation feature but the `screenshotScripts/index.ts` only exports `resolveMaskRect`.


Citations: [packages/extension/understudy/a11y/snapshot/capture.ts:60-90](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/a11y/snapshot/capture.ts#L60-L90) · [packages/extension/understudy/a11y/snapshot/domTree.ts:7-15](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/a11y/snapshot/domTree.ts#L7-L15) · [packages/extension/understudy/a11y/snapshot/a11yTree.ts:20-55](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/a11y/snapshot/a11yTree.ts#L20-L55) · [packages/extension/understudy/a11y/snapshot/treeFormatUtils.ts:8-15](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/understudy/a11y/snapshot/treeFormatUtils.ts#L8-L15) · [packages/extension/services/extractService.ts:107-139](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/services/extractService.ts#L107-L139) · [packages/extension/llm/LLMClient.ts:46-48](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/llm/LLMClient.ts#L46-L48)
