How is the page represented to the model?
DOM serialization, accessibility tree, screenshots, set-of-marks/element indexes; size limits and pruning.
Verdict
Both projects send the model a text rendering of the page built from CDP, not raw HTML. They differ in how much they add on top.
browser-use merges DOMSnapshot.captureSnapshot, the full DOM and the accessibility tree. Its DOMTreeSerializer drops elements hidden by paint order and redundant wrappers, then prints interactive nodes as [index]<tag attr=...>. The index is the CDP backend_node_id where possible, so it stays stable across steps. New elements are starred. A screenshot is taken every step and sent when use_vision is on. Serialized text is capped at 40,000 characters.
Stagehand builds a “hybrid snapshot”: an accessibility outline merged with DOM data, with iframes stitched in and shadow DOM pierced. Each line carries a frameOrdinal-backendNodeId id that maps to an XPath. Callers can scope it to a locator or exclude parts. There are no screenshots in act or observe. extract can attach a plain viewport PNG on request, with no set-of-marks overlay.
Pick browser-use when pages are visual or ambiguous and a vision model is worth the extra tokens on every step. Pick Stagehand for cheaper, text-only prompts scoped to part of a page, especially for extraction.
More projects in this category are being researched.
Per-project answers
browser-use/browser-use
answeredDOM Serialization. The agent represents pages through SerializedDOMState (browser_use/dom/views.py:935), which wraps a tree of SimplifiedNode objects built from three CDP sources: DOMSnapshot.captureSnapshot, the full accessibility tree via accessibility.GetFullAXTree, and the DOM document structure (browser_use/dom/service.py:6-25). The DomService class orchestrates this, combining snapshot bounds, AX node properties, and DOM nodes into EnhancedDOMTreeNode objects with sibling/shadow-root/iframe support. Cross-origin iframes with dimensions ≥ 10px each side are included (browser_use/dom/service.py:40-42).
Serialization to text. DOMTreeSerializer.serialize_accessible_elements() (browser_use/dom/serializer/serializer.py:114-165) runs a pipeline: (1) create a simplified tree by filtering non-content elements (style, script, SVG decorative children) and detecting interactive elements; (2) paint-order filtering to skip elements visually occluded; (3) tree optimization (remove unnecessary parents); (4) bounding-box filtering that excludes children contained within propagating parents (a, button, div[role=button]) at 99% containment threshold; (5) assign numeric interactive indices to clickable/scrollable elements. The final serialize_tree() (browser_use/dom/serializer/serializer.py:989) renders a compact text format: [index]<tagname attr="val" /> for interactive elements, ordinary <tagname> for containers, |scroll element| markers for scrollable containers, |SHADOW(open/closed)| for shadow-DOM hosts, text nodes inline, and iframes with hidden-element hints showing scroll distances. Compound controls (date pickers, range sliders, file inputs, select dropdowns) get synthetic compound_components=... attributes with sub-role information (browser_use/dom/serializer/serializer.py:1059-1100).
Element indices (set-of-marks). Interactive elements get sequentially numbered indices stored in a DOMSelectorMap (dict of int → EnhancedDOMTreeNode). The LLM sees [1]<button type="submit"> and uses that index to click. New elements since the last step are marked with a * prefix. max_clickable_elements_length caps serialized text at 40000 chars (browser_use/agent/views.py:92).
Screenshots. A PNG screenshot is always captured every step via BrowserSession.get_browser_state_summary(include_screenshot=True) (browser_use/agent/service.py:1096-1099). When use_vision is enabled, the screenshot is sent as an image content part to the LLM. For Claude Sonnet models, screenshots are auto-resized to 1400×850 to fit context windows (browser_use/agent/service.py:247-250). Configurable via llm_screenshot_size parameter.
Accessibility tree. The AX tree provides name, role, description, properties (valuemin/max/now, expanded, pressed, invalid, etc.) per element. These feed into ClickableElementDetector.is_interactive() (browser_use/dom/serializer/clickable_elements.py:6-52) which heuristically scores elements as interactive based on tag, role, JS click listeners, ARIA attributes, pointer cursor, and form-control heuristics.
backend_node_id when that is unique, not a sequential counter, so they mostly stay stable between steps.browserbase/stagehand
answeredThe page is represented to the model as a hybrid text snapshot combining the Chrome accessibility tree and the DOM tree. The core mechanism lives in capture.ts, which calls Accessibility.getFullAXTree via CDP for each frame's a11y tree and DOM.getDocument for the DOM structure. These are merged into a single hierarchical text outline where each element is rendered as a line like [0-18372] button: Submit — the encoded ID (frameOrdinal-backendNodeId) serves as the element index that the LLM returns when choosing an action. An xpathMap (a Record<string, string>) maps each encoded ID to an absolute XPath, which the execution layer uses to locate the element for CDP interaction.
Iframes are handled via a multi-step pipeline: buildSessionIndexes calls DOM.getDocument once per CDP session; per-frame maps are sliced from the shared index; and computeFramePrefixes walks the frame tree to compute absolute XPath prefixes for child frames. The per-frame outlines are stitched into a single combinedTree via injectSubtrees in treeFormatUtils.ts. Shadow DOM is pierced by default (pierceShadow: true).
Size limits and pruning: The DOM tree retrieval adaptively retries with shallower depths when CDP's CBOR encoder stack overflows — DOM_DEPTH_ATTEMPTS in domTree.ts tries [-1, 256, 128, 64, 32, 16, 8, 4, 2, 1], and each truncated node is hydrated individually via DOM.describeNode with its own depth fallback. Structural AX roles (generic, none, InlineTextBox) are pruned from the outline. Nodes can also be excluded via ignoreLocators.
Screenshots are optional and only used in extract() when options.screenshot: true. They are captured as raw PNG data (page.screenshot → Page.screenshot) and sent alongside the DOM text as an image content block. The screenshot is not annotated with bounding boxes or set-of-marks overlays — the AnnotatedScreenshotText constant in LLMClient.ts describes a planned annotation feature but the screenshotScripts/index.ts only exports resolveMaskRect.