# hyperbrowserai/HyperAgent

> TypeScript SDK that adds LLM-driven ai(), perform() and extract() to Playwright pages, picking elements from a CDP accessibility tree.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/hyperbrowserai/HyperAgent (reviewed at commit `a7ec1f45e2f7b99ea505218b0516d156da7597b5`, 2026-05-11)
- Stars: 1589 · Language: TypeScript · License: AGPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/hyperagent/

## Overview

HyperAgent (`@hyperbrowser/agent`) is a TypeScript library from Hyperbrowser that adds LLM-driven methods to an ordinary Playwright `Page`. You create a `HyperAgent`, get a page from it, and call `page.ai(task)` for a multi-step goal, `page.perform(instruction)` for one element action, or `page.extract(task, zodSchema)` for data. Between those calls the page is still a Playwright page, so normal `goto`, `locator` and `evaluate` code works as usual.

The model sees the page as text. Every step serialises the Chrome accessibility tree into lines tagged with an encoded id (`frameIndex-backendNodeId`, e.g. `0-5125`) that covers iframes and out-of-process frames. The model answers with structured JSON that names an id, a method such as `click` or `fill`, and arguments. HyperAgent then resolves the id to a live DOM node and sends raw CDP input events. Screenshots are opt-in: `enableVisualMode` defaults to false, so `page.ai()` is text-only unless you ask for the overlay screenshot.

It sits between a pure tool layer and a full agent framework. `page.ai()` is a real single-loop agent with memory and stop rules. `page.perform()` is a one-shot "find and act" call. Successful runs also produce an action cache: a list of XPaths and methods that can be replayed later without the LLM, or turned into a plain script.

## Architecture

```mermaid
flowchart LR
  U["Your code"] --> HA["HyperAgent"]
  HA --> HP["HyperPage: ai / perform / extract"]
  HP --> LOOP["runAgentTask loop"]
  HP --> FIND["findElementWithInstruction"]
  LOOP --> DOM["captureDOMState / getA11yDOM"]
  FIND --> DOM
  LOOP --> LLM["HyperAgentLLM.invokeStructured"]
  FIND --> EX["examineDom (LLM)"]
  EX --> LLM
  LOOP --> ACT["Action registry"]
  ACT --> PERF["performAction"]
  FIND --> PERF
  PERF --> CDP["dispatchCDPAction (CDP Input)"]
  PERF --> PW["Playwright locator path"]
  ACT --> MCP["MCP tools"]
  HA --> BP["Browser provider: Local / Hyperbrowser"]
  BP --> BR["Chromium"]
  CDP --> BR
  DOM --> BR
```

| Component | Path | Role |
|---|---|---|
| Agent class | `src/agent/index.ts` | `HyperAgent`: LLM and browser setup, `executeTask`, `executeSingleAction`, cache replay, MCP, and the `HyperPage` wrapper |
| Agent loop | `src/agent/tools/agent.ts` | `runAgentTask`: snapshot, prompt, structured LLM call, one action per step, stop rules |
| Prompt builder | `src/agent/messages/` | System prompt and per-step message assembly |
| Actions | `src/agent/actions/` | Built-in actions (`actElement`, `goToUrl`, `refreshPage`, `extract`, `wait`, `complete`, optional `pdf`) |
| Element finder | `src/agent/examine-dom/`, `src/agent/shared/find-element.ts` | One LLM call that ranks tree elements for a single instruction (`perform`) |
| DOM provider | `src/context-providers/a11y-dom/` | Accessibility-tree capture across frames, id/XPath maps, optional visual overlay, snapshot cache |
| CDP layer | `src/cdp/` | Session pooling over Playwright CDP sessions, frame graph, element resolution, input dispatch |
| LLM adapters | `src/llm/providers/` | OpenAI, Anthropic, Gemini, DeepSeek behind one `HyperAgentLLM` interface |
| Browser providers | `src/browser-providers/` | Local Chrome launch or a Hyperbrowser cloud session over `connectOverCDP` |
| MCP client | `src/agent/mcp/client.ts` | Connects stdio or SSE MCP servers and registers their tools as actions |
| CLI | `src/cli/index.ts` | `hyperagent-cli` interactive runner with a user-question action |

## How a request flows

Take `await page.ai("find the cheapest flight to Lisbon")`:

1. **Setup.** The constructor builds the LLM client from an `LLMConfig` (or defaults to OpenAI `gpt-4o` when only `OPENAI_API_KEY` is set) and picks the browser provider; `cdpActions` defaults to true ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L106-L144)). `initBrowser` starts the provider and opens a context with `viewport: null` ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L151-L180)).
2. **Page wrapper.** `setupHyperPage` keeps a stack of tabs opened from the current page, so the agent follows popups. `page.ai` calls `executeTask(task, params, activePage)` ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1473-L1572)).
3. **Task state.** `executeTask` creates a task id and `TaskState`, then calls `runAgentTask` with the action list, which ends with `complete` or a schema-typed `complete` when `outputSchema` is set ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L459-L517), [index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L188-L200)).
4. **Snapshot.** Each step waits for the DOM to settle, then calls `captureDOMState` ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L316-L348)). `getA11yDOM` builds backend-id and XPath maps with `DOM.getDocument`, attaches OOPIF sessions, and fetches `Accessibility.getFullAXTree` per frame ([a11y-dom/index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/context-providers/a11y-dom/index.ts#L908-L1000)). With `enableVisualMode`, an overlay of boxes and ids is composited onto a CDP screenshot with Jimp ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L65-L98)).
5. **Prompt.** `buildAgentStepMessages` sends the goal, URL, variable names, every previous step's thoughts, memory, action and result, then the element tree and the optional screenshot ([builder.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/messages/builder.ts#L9-L92)).
6. **Decide.** `invokeStructured` is called with `{thoughts, memory, action}`, where `action` is a Zod union of all registered actions ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L100-L125)). A response that fails validation is retried up to three times with the Zod errors fed back, and the last three schema errors are carried into later steps ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L399-L508)).
7. **Act.** For `actElement`, `performAction` uses CDP when it is enabled and a backend-node map exists: `resolveElement` turns the encoded id into a node in the right frame session, and `dispatchCDPAction` runs the method ([perform-action.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/shared/perform-action.ts#L64-L114)). A click scrolls into view, takes the box centre and sends `Input.dispatchMouseEvent` moved/pressed/released ([interactions.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/cdp/interactions.ts#L347-L396)).
8. **Loop control.** Five consecutive failed actions or `wait` actions fail the task as stuck. `maxSteps` cancels it. `complete` sets the output ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L521-L625)).
9. **Return.** The result contains every step and an `actionCache` with one entry per step; `onStep` and `onComplete` hooks fire along the way ([agent.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L669-L702)).

`page.perform(instruction)` skips the loop. `executeSingleAction` calls `findElementWithInstruction`, which captures a fresh tree and asks `examineDom` for ranked matches, retrying up to 10 times. It then runs the top match through the same `performAction` ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1079-L1170), [find-element.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/shared/find-element.ts#L60-L160)).

## Key components

### Accessibility-tree DOM

The page representation is the accessibility tree, not raw HTML. Ids join a frame ordinal and the CDP `backendNodeId`, and the snapshot keeps `xpathMap`, `backendNodeMap` and `frameMap` beside the text so any id can be acted on later. Same-origin iframes and OOPIFs are merged, and ad or tracking frames are filtered in `src/cdp/frame-filters.ts`. An optional snapshot cache (`useDomCache`) is invalidated on navigation events and after any non-read-only action.

### Action system

Every action is an `AgentActionDefinition`: a `type`, a Zod `actionParams` schema and a `run` function. The default list is `goToUrl`, `refreshPage`, `actElement`, `extract` and `wait`, plus `pdf` when `GEMINI_API_KEY` is set ([actions/index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/index.ts#L29-L47)). `pageBack`, `pageForward` and `thinking` exist in the folder but are commented out. `actElement` limits methods to twelve: click, fill, type, press, selectOptionFromDropdown, check, uncheck, hover and four scroll variants ([act-element.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/act-element.ts#L6-L39)).

The `extract` action is different from the others. It converts the whole page HTML to Markdown, trims it to the token budget, and sends it with a full-page screenshot to a plain `invoke` call ([extract.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/extract.ts#L24-L70)). That means extraction needs a vision-capable model even when visual mode is off.

### `page.extract()`

`page.extract()` is not a separate pipeline. It runs `executeTask` with a fixed extraction prompt, `maxSteps: 2` by default and the schema as `outputSchema`, then `JSON.parse`s the `complete` output ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1590-L1638)). Expect at most two agent steps, usually an `extract` action followed by `complete`.

### LLM adapters

`createLLMClient` switches over four providers ([providers/index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/index.ts#L18-L57)). OpenAI and DeepSeek use `response_format` JSON schema through the OpenAI SDK, and `baseURL` lets you point them at a compatible endpoint ([openai.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/openai.ts#L150-L180)). Anthropic has no JSON mode here, so it turns the actions into tools and forces a tool call ([anthropic.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/anthropic.ts#L93-L114)). Gemini uses its response schema. You can also pass any object that implements `HyperAgentLLM`.

### Action cache and replay

Each step is recorded with method, arguments, frame index and XPath. `runFromActionCache` replays a cache through `perform*` helpers such as `performClick(xpath)`. Each cached step is tried up to `maxSteps` times. When a `performInstruction` is supplied and the XPath fails, it falls back to an LLM `perform` call ([run-cached-action.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/shared/run-cached-action.ts#L104-L200), [action-cache-exec.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/shared/action-cache-exec.ts#L65-L80)). `createScriptFromActionCache` prints the cache as a TypeScript script of those helper calls.

### Browser providers

`LocalBrowserProvider` launches Playwright Chromium on the `chrome` channel, always headed, with `--disable-blink-features=AutomationControlled` ([local.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/browser-providers/local.ts#L11-L21)). `HyperbrowserProvider` creates a cloud session with `@hyperbrowser/sdk` and connects with `connectOverCDP`; stealth, proxies and similar features belong to that session config ([hyperbrowser.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/browser-providers/hyperbrowser.ts#L33-L56)).

## Extending it

- **Custom actions.** Pass `customActions` to the constructor. Each one needs a unique `type` and a Zod schema; `complete` is reserved ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1296-L1313)). The CLI's `UserInteractionAction` (ask the human a question) is a working example, and `examples/custom-tool` has search and Wikipedia tools.
- **MCP servers.** `initializeMCPClient({ servers })` connects stdio or SSE servers and registers each of their tools as an agent action ([index.ts](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1319-L1355)). `examples/mcp` covers Google Sheets, Notion and weather.
- **Own model.** Use `LLMConfig` with `baseURL`, or supply a `HyperAgentLLM` implementation.
- **Hooks and output.** `TaskParams` takes `outputSchema`, `maxSteps`, `onStep`, `onComplete`, `debugOnAgentOutput`, `enableVisualMode` and `useDomCache`. `aiAsync` returns a task handle that can be paused, resumed or cancelled.
- **Mixed code.** Because `HyperPage` is a Playwright `Page`, scripted steps and `perform` calls can be interleaved freely.

## Running it

- **Install.** `npm install @hyperbrowser/agent`, plus Chrome on the machine for local runs (the provider uses the `chrome` channel).
- **Keys.** Pick a provider in `llm`; its key comes from `apiKey` or `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY` (or `GOOGLE_API_KEY`) or `DEEPSEEK_API_KEY`. Without an `llm` config the agent only starts if `OPENAI_API_KEY` is set.
- **Cloud.** Set `browserProvider: "Hyperbrowser"` and `HYPERBROWSER_API_KEY` to run in a Hyperbrowser session.
- **CLI.** `npx @hyperbrowser/agent -c "task"` runs one task interactively; `--hyperbrowser` switches to the cloud and `-m` loads an MCP config.
- **Debug.** `debug: true` writes per-step folders with the element tree, messages, screenshots, frame graph and timing JSON.

## Strengths and caveats

- **Strength: CDP-native execution.** Element ids map straight to backend nodes, with XPath recovery when nodes go stale, and inputs are real CDP mouse and key events. Iframes and OOPIFs are handled in the core path, not as an afterthought.
- **Strength: cheap by default.** The loop is text-only unless you enable visual mode, and `perform` costs one LLM call per action.
- **Strength: replayable runs.** The action cache and script generator give a clear way to move a working AI flow onto a deterministic one, with an LLM fallback per step.
- **Caveat: shallow agent.** One action per step, no planner, and history grows with every step. The only stop rules are `maxSteps`, `complete` and five failures or waits in a row.
- **Caveat: `extract` is heavy.** It sends the whole page as Markdown plus a screenshot, so it needs a vision model and can be expensive on long pages.
- **Caveat: local browser is fixed.** The local provider always runs headed Chrome, and the only built-in stealth is one Blink flag. Headless runs, stronger stealth and proxies mean using the Hyperbrowser cloud or your own provider.
- **Caveat: variables are hints only.** Variables are listed to the model as `<<key>>` with a description, but at this commit no code substitutes their values back into actions.
- **Caveat: licence.** The repository's LICENSE file and `package.json` are AGPL-3.0, which matters if you embed it in a hosted service.

*Sources: code at a7ec1f4, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (16 pages), verified Q&A.*

## How hyperbrowserai/HyperAgent answers the Browser & computer control questions

### How is the page represented to the model? (answered)

The page is represented to the model primarily through the **CDP accessibility tree**, serialized as a hierarchical text tree with **encoded element IDs** in the format `frameIndex-backendNodeId` (e.g., `"0-5125"`). The main entry point is `getA11yDOM()` (`src/context-providers/a11y-dom/index.ts:908`), which calls `Accessibility.getFullAXTree` via CDP for the main frame and all iframes — same-origin iframes via `contentDocumentBackendNodeId`, out-of-process iframes (OOPIFs) via separate CDP sessions discovered through `FrameContextManager.captureOOPIFs` (`src/context-providers/a11y-dom/index.ts:993`). Raw `AXNode` arrays are converted to `AccessibilityNode` objects by `buildHierarchicalTree()` (`src/context-providers/a11y-dom/build-tree.ts:71`), then serialized into a simplified text tree via `formatSimplifiedTree` — this text is what the LLM sees in the `=== Elements ===` section of the prompt (`src/agent/messages/builder.ts:63`). 

When `enableVisualMode` is true, **bounding boxes** are collected in batch via `batchCollectBoundingBoxesWithFailures` and a **visual overlay** is rendered by `renderA11yOverlay()` (`src/context-providers/a11y-dom/index.ts:1187`), which composites bounding boxes with encoded IDs over a CDP screenshot using `Jimp`. The overlay is composited with the raw screenshot via `compositeScreenshot()` (`src/agent/tools/agent.ts:65`). The composited screenshot is passed to the LLM as a base64 PNG alongside the text tree (`src/agent/messages/builder.ts:69`). 

**DOM caching** is handled by `domSnapshotCache` — a snapshot is cached when `useCache` is true and visual mode is off, and invalidated on navigation/frame events via `markDomSnapshotDirty()` (`src/context-providers/a11y-dom/dom-cache.ts`). The DOM capture has 3 retry attempts for transient CDP errors (`src/agent/shared/dom-capture.ts:11`), with recovery for "Execution context was destroyed" / "Target closed" errors. **Streaming** mode (`enableDomStreaming`) sends frame chunks progressively via `FrameChunkEvent` with ordered assembly in `DomChunkAggregator` (`src/agent/shared/dom-capture.ts:28`).


Citations: [src/context-providers/a11y-dom/index.ts:908-1254](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/context-providers/a11y-dom/index.ts#L908-L1254) · [src/context-providers/a11y-dom/build-tree.ts:71-100](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/context-providers/a11y-dom/build-tree.ts#L71-L100) · [src/agent/messages/builder.ts:9-92](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/messages/builder.ts#L9-L92) · [src/agent/tools/agent.ts:316-370](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L316-L370) · [src/agent/shared/dom-capture.ts:88-163](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/shared/dom-capture.ts#L88-L163)

### How are actions executed and how are elements targeted? (answered)

Actions are executed **CDP-first**, falling back to Playwright when CDP is unavailable. The central dispatch is `performAction()` (`src/agent/actions/shared/perform-action.ts:20`). When `ctx.cdpActions` is true and a `backendNodeMap` exists, it uses CDP: `resolveElement()` (`src/cdp/element-resolver.ts:35`) resolves an encoded ID (`frameIndex-backendNodeId`) to a `{session, frameId, backendNodeId, objectId}` via `DOM.resolveNode`, with **XPath-based recovery** when backendNodeIds go stale (calling `DOM.describeNode` after evaluating an XPath in the correct execution context). The resolved element is then dispatched by `dispatchCDPAction()` (`src/cdp/interactions.ts`) which uses CDP `Input.dispatchMouseEvent` / `Input.dispatchKeyEvent` / `Input.dispatchDragEvent` for clicks, typing, and fills, and JavaScript evaluation for scroll actions. 

**CDP methods** (defined in `src/cdp/interactions.ts:17-31`) include: `click`, `hover`, `fill`, `type`, `press`, `check`, `uncheck`, `selectOptionFromDropdown`, `scrollToElement`, `scrollToPercentage`, `nextChunk`, `prevChunk`. When CDP is unavailable, the system builds a **Playwright locator** via `getElementLocator()` (`src/agent/shared/element-locator.ts`) using XPaths from `xpathMap`, resolving frame indices to Playwright frames, then calls `executePlaywrightMethod()` (`src/agent/shared/execute-playwright-method.ts`) with a 3500ms click timeout. 

**Element targeting** uses the encoded ID scheme `frameIndex-backendNodeId`. The `actElement` action schema requires `elementId`, `method` (one of 12 CDP methods), and `arguments` (string array) (`src/agent/actions/act-element.ts:12-36`). **Scrolling** also has a specialized `ScrollAction` (`src/agent/actions/scroll.ts:4`) for up/down/left/right by viewport height via `window.scrollBy`. **File upload**, **tab management**, and page navigation actions (goToUrl, pageBack, pageForward, refreshPage) are defined as separate action definitions (`src/agent/actions/index.ts:29-43`). Tab-following is handled by `setupHyperPage()` which maintains a page stack — when the active page opens a new tab, the stack auto-pushes; when a tab closes it auto-splices (`src/agent/index.ts:1482-1558`).

> **Editor's note.** Correction: there is no Input.dispatchDragEvent and no file-upload action in src/. The default action list is goToUrl, refreshPage, actElement, extract and wait (plus pdf with GEMINI_API_KEY); ScrollAction, pageBack and pageForward exist as files but are not registered. Playwright is not a runtime fallback when CDP fails; performAction picks one path up front.

Citations: [src/agent/actions/shared/perform-action.ts:20-154](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/shared/perform-action.ts#L20-L154) · [src/cdp/element-resolver.ts:35-120](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/cdp/element-resolver.ts#L35-L120) · [src/cdp/interactions.ts:1-31](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/cdp/interactions.ts#L1-L31) · [src/agent/actions/act-element.ts:1-55](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/actions/act-element.ts#L1-L55) · [src/agent/index.ts:1473-1558](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1473-L1558)

### How is the agent loop / planning implemented? (answered)

The core agent loop is `runAgentTask()` in `src/agent/tools/agent.ts:208`. It is a **single-step loop** with no separate planner/executor distinction — each iteration captures DOM state, builds messages, invokes the LLM for one action, executes it, and repeats. Each LLM call returns a **structured output** validated via Zod: `{thoughts: string, memory: string, action: union-schema}` (`src/types/agent/types.ts:6-21`). The `thoughts` field captures the model's reasoning, `memory` accumulates state changes (e.g., "Clicked login button → login form appeared"), and `action` selects from available action types via a Zod discriminated union built by `getActionSchema()` (`src/agent/tools/agent.ts:100`).

**Messages** are constructed by `buildAgentStepMessages()` (`src/agent/messages/builder.ts:9`) which appends: the final goal, current URL, variables, previous action history (assistant thought/memory/action → user result), the DOM tree, and optionally a composited screenshot with scroll info. The **system prompt** (`src/agent/messages/system-prompt.ts`) describes input/output format, available actions, and guidelines (one action per step, use wait when unsure, avoid loops).

**Stop conditions** are checked at each iteration: `TaskStatus.PAUSED` sleeps 100ms and continues; `endTaskStatuses` (COMPLETED, FAILED, CANCELLED) breaks; `maxSteps` threshold cancels the task (`src/agent/tools/agent.ts:303-306`); and a **consecutive failure/waits counter** of 5 (`MAX_CONSECUTIVE_FAILURES_OR_WAITS`) fails the task if exceeded (`src/agent/tools/agent.ts:267`). The loop also tracks schema validation errors — up to 3 are fed back into the next step's messages for cross-step learning (`src/agent/tools/agent.ts:400-413`). **LLM structured output** has up to 3 retry attempts per step with Zod error feedback appended as user messages (`src/agent/tools/agent.ts:424-508`).


Citations: [src/agent/tools/agent.ts:208-510](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L208-L510) · [src/types/agent/types.ts:1-21](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/types/agent/types.ts#L1-L21) · [src/agent/messages/builder.ts:9-92](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/messages/builder.ts#L9-L92) · [src/agent/messages/system-prompt.ts:1-78](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/messages/system-prompt.ts#L1-L78) · [src/agent/tools/agent.ts:562-650](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L562-L650)

### How are failures, retries and self-healing handled? (answered)

Reliability is handled through **multiple layers of retries**, **error classification**, and **action caching**. The general-purpose `retry()` utility (`src/utils/retry.ts:2`) provides exponential backoff (2^i * 1000ms) with up to 3 attempts — used for LLM invocations, DOM state capture, and scroll info retrieval. **DOM capture** has its own retry loop (`src/agent/shared/dom-capture.ts:105`) with 3 attempts for transient CDP errors ("Execution context was destroyed", "Cannot find context", "Target closed"), plus a placeholder-snapshot detector via `isPlaceholderSnapshot()` that retries on "Error: Could not extract accessibility tree". 

**LLM structured output** failures get 3 attempts per step with Zod validation errors fed back as corrective user messages (`src/agent/tools/agent.ts:424-508`). Schema errors accumulate across steps (last 3 tracked) and are injected into future step messages to guide the model (`src/agent/tools/agent.ts:400-413`). The **consecutive failure/waits counter** terminates tasks after 5 consecutive waits or action failures, with a clear error message (`src/agent/tools/agent.ts:267, 575-624`). 

**Action caching** records each step's XPath, frame index, method, and arguments via `buildActionCacheEntry()` (`src/agent/shared/action-cache.ts:95`). Successful tasks produce an `ActionCacheOutput` that can be replayed via `runFromActionCache()` (`src/agent/index.ts:519-834`), which replays cached actions by dispatching through `perform*` helpers with XPath retry support and fallback to instruction-based execution. **Element resolution** has its own recovery: if `DOM.resolveNode` fails with "Could not find node with given id", the system falls back to XPath-based recovery via `recoverBackendNodeId()` (`src/cdp/element-resolver.ts:228`) which evaluates the XPath in the correct frame's execution context, then calls `DOM.describeNode` to get a fresh backendNodeId. **Page context switching** is detected and triggers retry with a 500ms delay (`src/agent/index.ts:1537-1558`).


Citations: [src/utils/retry.ts:1-25](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/utils/retry.ts#L1-L25) · [src/agent/tools/agent.ts:560-625](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/tools/agent.ts#L560-L625) · [src/cdp/element-resolver.ts:228-317](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/cdp/element-resolver.ts#L228-L317) · [src/agent/index.ts:519-834](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L519-L834) · [src/agent/index.ts:1525-1558](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L1525-L1558)

### Which models are supported and how are they called? (answered)

The project supports **four LLM providers**: OpenAI, Anthropic, Gemini, and DeepSeek, all implementing the common `HyperAgentLLM` interface (`src/llm/types.ts:76`) with three key methods: `invoke()`, `invokeStructured()`, `getCapabilities()`. Provider selection happens in `createLLMClient()` (`src/llm/providers/index.ts:18`) which routes `LLMConfig.provider` to the correct client constructor. Default is OpenAI (`gpt-4o`) if no LLM is configured but `OPENAI_API_KEY` is set (`src/agent/index.ts:108-113`). 

**Vision requirement**: All providers must support multimodal (vision) input for the screenshot-in-prompt flow — `getCapabilities().multimodal` returns `true` for all four (`src/llm/providers/openai.ts:218`, `src/llm/providers/anthropic.ts:127`, etc.). 

**Structured output** varies by provider:
- **OpenAI** uses `response_format: {type: "json_schema"}` with Zod-jsonSchema conversion via `convertToOpenAIJsonSchema()` (`src/llm/providers/openai.ts:153-206`). Temperature is omitted for GPT-5 models at temperature 1. 
- **Anthropic** uses **tool calling** with two paths: `invokeStructuredViaTools()` converts action definitions to Anthropic tool schemas and forces tool use (`tool_choice: "any"`) (`src/llm/providers/anthropic.ts:133-230`); `invokeStructuredViaSimpleTool()` wraps arbitrary Zod schemas as a single "structured_output" tool (`src/llm/providers/anthropic.ts:236-283`). 
- **Gemini** uses `responseMimeType: "application/json"` with `responseSchema` for JSON mode (`src/llm/providers/gemini.ts`). 
- **DeepSeek** reuses the OpenAI SDK (`src/llm/providers/deepseek.ts:23`) with `baseURL: "https://api.deepseek.com"` and supports `response_format` for JSON mode. 

Custom `baseURL` is supported for OpenAI and DeepSeek, enabling **proxy/compatible endpoints**. The `getCapabilities()` method reports `multimodal`, `toolCalling`, and `jsonMode` flags — the agent loop uses `invokeStructured()` exclusively for deterministic action parsing with Zod validation.

> **Editor's note.** Correction: vision is not required for the agent loop, because screenshots are opt-in (enableVisualMode defaults to false). A vision model is needed for visual mode and for the built-in extract action, which always sends a page screenshot with the Markdown.

Citations: [src/llm/providers/index.ts:1-70](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/index.ts#L1-L70) · [src/llm/types.ts:52-95](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/types.ts#L52-L95) · [src/llm/providers/openai.ts:71-227](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/openai.ts#L71-L227) · [src/llm/providers/anthropic.ts:133-283](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/anthropic.ts#L133-L283) · [src/llm/providers/deepseek.ts:1-55](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/llm/providers/deepseek.ts#L1-L55)

### How are browser sessions, profiles, auth and anti-bot handled? (answered)

Browser sessions are managed through the abstract `BrowserProvider<T>` base class (`src/types/browser-providers/types.ts:3`), with two concrete implementations: **`LocalBrowserProvider`** (`src/browser-providers/local.ts`) and **`HyperbrowserProvider`** (`src/browser-providers/hyperbrowser.ts`). Selection is controlled by `HyperAgentConfig.browserProvider` (default: `"Local"`) (`src/agent/index.ts:81-83, 124`). 

**LocalBrowserProvider** launches Playwright's Chromium with `channel: "chrome"`, `headless: false`, and the stealth flag `--disable-blink-features=AutomationControlled` (`src/browser-providers/local.ts:13-18`). This flag helps evade anti-bot detection by suppressing the "Chrome is being controlled by automated software" banner. Additional launch options can be passed via `localConfig`. 

**HyperbrowserProvider** (`src/browser-providers/hyperbrowser.ts`) connects to **cloud browser sessions** via the `@hyperbrowser/sdk`. It creates a remote session through the Hyperbrowser API (`client.sessions.create()`), then connects Playwright to it via `chromium.connectOverCDP(session.wsEndpoint)` (`src/browser-providers/hyperbrowser.ts:35-41`). Session details (live URL, session ID, info URL) are logged in debug mode. On close, it both closes the Playwright browser and calls `client.sessions.stop()` to terminate the cloud session (`src/browser-providers/hyperbrowser.ts:58-63`). 

**Profiles, cookies, and auth** are managed through raw Playwright `BrowserContext` — the agent creates one context per browser session with `viewport: null` (`src/agent/index.ts:160-162`). The context is not configured with a persistent profile directory by default; users can pass a Playwright-provided context via `initPage` into `executeTask()`/`executeTaskAsync()` for session reuse. **Anti-bot** beyond the `AutomationControlled` flag is not implemented — the system prompt instructs the agent to "Accept or close" cookie banners and to try refreshing or alternative approaches for CAPTCHAs (`src/agent/messages/system-prompt.ts:74-77`). **Proxies** are not configured in the agent layer but can be passed through Playwright launch options via `localConfig` or `hyperbrowserConfig`, which are passed through to the underlying `chromium.launch()` / `chromium.connectOverCDP()` calls.


Citations: [src/types/browser-providers/types.ts:1-9](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/types/browser-providers/types.ts#L1-L9) · [src/browser-providers/local.ts:1-31](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/browser-providers/local.ts#L1-L31) · [src/browser-providers/hyperbrowser.ts:1-71](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/browser-providers/hyperbrowser.ts#L1-L71) · [src/agent/index.ts:151-180](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/index.ts#L151-L180) · [src/agent/messages/system-prompt.ts:72-78](https://github.com/hyperbrowserai/HyperAgent/blob/a7ec1f45e2f7b99ea505218b0516d156da7597b5/src/agent/messages/system-prompt.ts#L72-L78)
