hyperbrowserai/HyperAgent
TypeScript SDK that adds LLM-driven ai(), perform() and extract() to Playwright pages, picking elements from a CDP accessibility tree.
Overview
HyperAgent (@hyperbrowser/agent) is a TypeScript library from Hyperbrowser that adds LLM-driven methods to an ordinary Playwright Page. You create a HyperAgent, get a page from it, and call page.ai(task) for a multi-step goal, page.perform(instruction) for one element action, or page.extract(task, zodSchema) for data. Between those calls the page is still a Playwright page, so normal goto, locator and evaluate code works as usual.
The model sees the page as text. Every step serialises the Chrome accessibility tree into lines tagged with an encoded id (frameIndex-backendNodeId, e.g. 0-5125) that covers iframes and out-of-process frames. The model answers with structured JSON that names an id, a method such as click or fill, and arguments. HyperAgent then resolves the id to a live DOM node and sends raw CDP input events. Screenshots are opt-in: enableVisualMode defaults to false, so page.ai() is text-only unless you ask for the overlay screenshot.
It sits between a pure tool layer and a full agent framework. page.ai() is a real single-loop agent with memory and stop rules. page.perform() is a one-shot “find and act” call. Successful runs also produce an action cache: a list of XPaths and methods that can be replayed later without the LLM, or turned into a plain script.
Architecture
flowchart LR
U["Your code"] --> HA["HyperAgent"]
HA --> HP["HyperPage: ai / perform / extract"]
HP --> LOOP["runAgentTask loop"]
HP --> FIND["findElementWithInstruction"]
LOOP --> DOM["captureDOMState / getA11yDOM"]
FIND --> DOM
LOOP --> LLM["HyperAgentLLM.invokeStructured"]
FIND --> EX["examineDom (LLM)"]
EX --> LLM
LOOP --> ACT["Action registry"]
ACT --> PERF["performAction"]
FIND --> PERF
PERF --> CDP["dispatchCDPAction (CDP Input)"]
PERF --> PW["Playwright locator path"]
ACT --> MCP["MCP tools"]
HA --> BP["Browser provider: Local / Hyperbrowser"]
BP --> BR["Chromium"]
CDP --> BR
DOM --> BR
| Component | Path | Role |
|---|---|---|
| Agent class | src/agent/index.ts |
HyperAgent: LLM and browser setup, executeTask, executeSingleAction, cache replay, MCP, and the HyperPage wrapper |
| Agent loop | src/agent/tools/agent.ts |
runAgentTask: snapshot, prompt, structured LLM call, one action per step, stop rules |
| Prompt builder | src/agent/messages/ |
System prompt and per-step message assembly |
| Actions | src/agent/actions/ |
Built-in actions (actElement, goToUrl, refreshPage, extract, wait, complete, optional pdf) |
| Element finder | src/agent/examine-dom/, src/agent/shared/find-element.ts |
One LLM call that ranks tree elements for a single instruction (perform) |
| DOM provider | src/context-providers/a11y-dom/ |
Accessibility-tree capture across frames, id/XPath maps, optional visual overlay, snapshot cache |
| CDP layer | src/cdp/ |
Session pooling over Playwright CDP sessions, frame graph, element resolution, input dispatch |
| LLM adapters | src/llm/providers/ |
OpenAI, Anthropic, Gemini, DeepSeek behind one HyperAgentLLM interface |
| Browser providers | src/browser-providers/ |
Local Chrome launch or a Hyperbrowser cloud session over connectOverCDP |
| MCP client | src/agent/mcp/client.ts |
Connects stdio or SSE MCP servers and registers their tools as actions |
| CLI | src/cli/index.ts |
hyperagent-cli interactive runner with a user-question action |
How a request flows
Take await page.ai("find the cheapest flight to Lisbon"):
- Setup. The constructor builds the LLM client from an
LLMConfig(or defaults to OpenAIgpt-4owhen onlyOPENAI_API_KEYis set) and picks the browser provider;cdpActionsdefaults to true (index.ts).initBrowserstarts the provider and opens a context withviewport: null(index.ts). - Page wrapper.
setupHyperPagekeeps a stack of tabs opened from the current page, so the agent follows popups.page.aicallsexecuteTask(task, params, activePage)(index.ts). - Task state.
executeTaskcreates a task id andTaskState, then callsrunAgentTaskwith the action list, which ends withcompleteor a schema-typedcompletewhenoutputSchemais set (index.ts, index.ts). - Snapshot. Each step waits for the DOM to settle, then calls
captureDOMState(agent.ts).getA11yDOMbuilds backend-id and XPath maps withDOM.getDocument, attaches OOPIF sessions, and fetchesAccessibility.getFullAXTreeper frame (a11y-dom/index.ts). WithenableVisualMode, an overlay of boxes and ids is composited onto a CDP screenshot with Jimp (agent.ts). - Prompt.
buildAgentStepMessagessends the goal, URL, variable names, every previous step’s thoughts, memory, action and result, then the element tree and the optional screenshot (builder.ts). - Decide.
invokeStructuredis called with{thoughts, memory, action}, whereactionis a Zod union of all registered actions (agent.ts). A response that fails validation is retried up to three times with the Zod errors fed back, and the last three schema errors are carried into later steps (agent.ts). - Act. For
actElement,performActionuses CDP when it is enabled and a backend-node map exists:resolveElementturns the encoded id into a node in the right frame session, anddispatchCDPActionruns the method (perform-action.ts). A click scrolls into view, takes the box centre and sendsInput.dispatchMouseEventmoved/pressed/released (interactions.ts). - Loop control. Five consecutive failed actions or
waitactions fail the task as stuck.maxStepscancels it.completesets the output (agent.ts). - Return. The result contains every step and an
actionCachewith one entry per step;onStepandonCompletehooks fire along the way (agent.ts).
page.perform(instruction) skips the loop. executeSingleAction calls findElementWithInstruction, which captures a fresh tree and asks examineDom for ranked matches, retrying up to 10 times. It then runs the top match through the same performAction (index.ts, find-element.ts).
Key components
Accessibility-tree DOM
The page representation is the accessibility tree, not raw HTML. Ids join a frame ordinal and the CDP backendNodeId, and the snapshot keeps xpathMap, backendNodeMap and frameMap beside the text so any id can be acted on later. Same-origin iframes and OOPIFs are merged, and ad or tracking frames are filtered in src/cdp/frame-filters.ts. An optional snapshot cache (useDomCache) is invalidated on navigation events and after any non-read-only action.
Action system
Every action is an AgentActionDefinition: a type, a Zod actionParams schema and a run function. The default list is goToUrl, refreshPage, actElement, extract and wait, plus pdf when GEMINI_API_KEY is set (actions/index.ts). pageBack, pageForward and thinking exist in the folder but are commented out. actElement limits methods to twelve: click, fill, type, press, selectOptionFromDropdown, check, uncheck, hover and four scroll variants (act-element.ts).
The extract action is different from the others. It converts the whole page HTML to Markdown, trims it to the token budget, and sends it with a full-page screenshot to a plain invoke call (extract.ts). That means extraction needs a vision-capable model even when visual mode is off.
page.extract()
page.extract() is not a separate pipeline. It runs executeTask with a fixed extraction prompt, maxSteps: 2 by default and the schema as outputSchema, then JSON.parses the complete output (index.ts). Expect at most two agent steps, usually an extract action followed by complete.
LLM adapters
createLLMClient switches over four providers (providers/index.ts). OpenAI and DeepSeek use response_format JSON schema through the OpenAI SDK, and baseURL lets you point them at a compatible endpoint (openai.ts). Anthropic has no JSON mode here, so it turns the actions into tools and forces a tool call (anthropic.ts). Gemini uses its response schema. You can also pass any object that implements HyperAgentLLM.
Action cache and replay
Each step is recorded with method, arguments, frame index and XPath. runFromActionCache replays a cache through perform* helpers such as performClick(xpath). Each cached step is tried up to maxSteps times. When a performInstruction is supplied and the XPath fails, it falls back to an LLM perform call (run-cached-action.ts, action-cache-exec.ts). createScriptFromActionCache prints the cache as a TypeScript script of those helper calls.
Browser providers
LocalBrowserProvider launches Playwright Chromium on the chrome channel, always headed, with --disable-blink-features=AutomationControlled (local.ts). HyperbrowserProvider creates a cloud session with @hyperbrowser/sdk and connects with connectOverCDP; stealth, proxies and similar features belong to that session config (hyperbrowser.ts).
Extending it
- Custom actions. Pass
customActionsto the constructor. Each one needs a uniquetypeand a Zod schema;completeis reserved (index.ts). The CLI’sUserInteractionAction(ask the human a question) is a working example, andexamples/custom-toolhas search and Wikipedia tools. - MCP servers.
initializeMCPClient({ servers })connects stdio or SSE servers and registers each of their tools as an agent action (index.ts).examples/mcpcovers Google Sheets, Notion and weather. - Own model. Use
LLMConfigwithbaseURL, or supply aHyperAgentLLMimplementation. - Hooks and output.
TaskParamstakesoutputSchema,maxSteps,onStep,onComplete,debugOnAgentOutput,enableVisualModeanduseDomCache.aiAsyncreturns a task handle that can be paused, resumed or cancelled. - Mixed code. Because
HyperPageis a PlaywrightPage, scripted steps andperformcalls can be interleaved freely.
Running it
- Install.
npm install @hyperbrowser/agent, plus Chrome on the machine for local runs (the provider uses thechromechannel). - Keys. Pick a provider in
llm; its key comes fromapiKeyorOPENAI_API_KEY,ANTHROPIC_API_KEY,GEMINI_API_KEY(orGOOGLE_API_KEY) orDEEPSEEK_API_KEY. Without anllmconfig the agent only starts ifOPENAI_API_KEYis set. - Cloud. Set
browserProvider: "Hyperbrowser"andHYPERBROWSER_API_KEYto run in a Hyperbrowser session. - CLI.
npx @hyperbrowser/agent -c "task"runs one task interactively;--hyperbrowserswitches to the cloud and-mloads an MCP config. - Debug.
debug: truewrites per-step folders with the element tree, messages, screenshots, frame graph and timing JSON.
Strengths and caveats
- Strength: CDP-native execution. Element ids map straight to backend nodes, with XPath recovery when nodes go stale, and inputs are real CDP mouse and key events. Iframes and OOPIFs are handled in the core path, not as an afterthought.
- Strength: cheap by default. The loop is text-only unless you enable visual mode, and
performcosts one LLM call per action. - Strength: replayable runs. The action cache and script generator give a clear way to move a working AI flow onto a deterministic one, with an LLM fallback per step.
- Caveat: shallow agent. One action per step, no planner, and history grows with every step. The only stop rules are
maxSteps,completeand five failures or waits in a row. - Caveat:
extractis heavy. It sends the whole page as Markdown plus a screenshot, so it needs a vision model and can be expensive on long pages. - Caveat: local browser is fixed. The local provider always runs headed Chrome, and the only built-in stealth is one Blink flag. Headless runs, stronger stealth and proxies mean using the Hyperbrowser cloud or your own provider.
- Caveat: variables are hints only. Variables are listed to the model as
<<key>>with a description, but at this commit no code substitutes their values back into actions. - Caveat: licence. The repository’s LICENSE file and
package.jsonare AGPL-3.0, which matters if you embed it in a hosted service.
Sources: code at a7ec1f4, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (16 pages), verified Q&A.
How it answers the Browser & computer control questions
Each answer was drafted by a code-reading agent at commit a7ec1f4. Its citations were checked mechanically. Compare with the other browser & computer control →
How is the page represented to the model?
answeredThe page is represented to the model primarily through the CDP accessibility tree, serialized as a hierarchical text tree with encoded element IDs in the format frameIndex-backendNodeId (e.g., "0-5125"). The main entry point is getA11yDOM() (src/context-providers/a11y-dom/index.ts:908), which calls Accessibility.getFullAXTree via CDP for the main frame and all iframes — same-origin iframes via contentDocumentBackendNodeId, out-of-process iframes (OOPIFs) via separate CDP sessions discovered through FrameContextManager.captureOOPIFs (src/context-providers/a11y-dom/index.ts:993). Raw AXNode arrays are converted to AccessibilityNode objects by buildHierarchicalTree() (src/context-providers/a11y-dom/build-tree.ts:71), then serialized into a simplified text tree via formatSimplifiedTree — this text is what the LLM sees in the === Elements === section of the prompt (src/agent/messages/builder.ts:63).
When enableVisualMode is true, bounding boxes are collected in batch via batchCollectBoundingBoxesWithFailures and a visual overlay is rendered by renderA11yOverlay() (src/context-providers/a11y-dom/index.ts:1187), which composites bounding boxes with encoded IDs over a CDP screenshot using Jimp. The overlay is composited with the raw screenshot via compositeScreenshot() (src/agent/tools/agent.ts:65). The composited screenshot is passed to the LLM as a base64 PNG alongside the text tree (src/agent/messages/builder.ts:69).
DOM caching is handled by domSnapshotCache — a snapshot is cached when useCache is true and visual mode is off, and invalidated on navigation/frame events via markDomSnapshotDirty() (src/context-providers/a11y-dom/dom-cache.ts). The DOM capture has 3 retry attempts for transient CDP errors (src/agent/shared/dom-capture.ts:11), with recovery for "Execution context was destroyed" / "Target closed" errors. Streaming mode (enableDomStreaming) sends frame chunks progressively via FrameChunkEvent with ordered assembly in DomChunkAggregator (src/agent/shared/dom-capture.ts:28).
How are actions executed and how are elements targeted?
answeredActions are executed CDP-first, falling back to Playwright when CDP is unavailable. The central dispatch is performAction() (src/agent/actions/shared/perform-action.ts:20). When ctx.cdpActions is true and a backendNodeMap exists, it uses CDP: resolveElement() (src/cdp/element-resolver.ts:35) resolves an encoded ID (frameIndex-backendNodeId) to a {session, frameId, backendNodeId, objectId} via DOM.resolveNode, with XPath-based recovery when backendNodeIds go stale (calling DOM.describeNode after evaluating an XPath in the correct execution context). The resolved element is then dispatched by dispatchCDPAction() (src/cdp/interactions.ts) which uses CDP Input.dispatchMouseEvent / Input.dispatchKeyEvent / Input.dispatchDragEvent for clicks, typing, and fills, and JavaScript evaluation for scroll actions.
CDP methods (defined in src/cdp/interactions.ts:17-31) include: click, hover, fill, type, press, check, uncheck, selectOptionFromDropdown, scrollToElement, scrollToPercentage, nextChunk, prevChunk. When CDP is unavailable, the system builds a Playwright locator via getElementLocator() (src/agent/shared/element-locator.ts) using XPaths from xpathMap, resolving frame indices to Playwright frames, then calls executePlaywrightMethod() (src/agent/shared/execute-playwright-method.ts) with a 3500ms click timeout.
Element targeting uses the encoded ID scheme frameIndex-backendNodeId. The actElement action schema requires elementId, method (one of 12 CDP methods), and arguments (string array) (src/agent/actions/act-element.ts:12-36). Scrolling also has a specialized ScrollAction (src/agent/actions/scroll.ts:4) for up/down/left/right by viewport height via window.scrollBy. File upload, tab management, and page navigation actions (goToUrl, pageBack, pageForward, refreshPage) are defined as separate action definitions (src/agent/actions/index.ts:29-43). Tab-following is handled by setupHyperPage() which maintains a page stack — when the active page opens a new tab, the stack auto-pushes; when a tab closes it auto-splices (src/agent/index.ts:1482-1558).
How is the agent loop / planning implemented?
answeredThe core agent loop is runAgentTask() in src/agent/tools/agent.ts:208. It is a single-step loop with no separate planner/executor distinction — each iteration captures DOM state, builds messages, invokes the LLM for one action, executes it, and repeats. Each LLM call returns a structured output validated via Zod: {thoughts: string, memory: string, action: union-schema} (src/types/agent/types.ts:6-21). The thoughts field captures the model's reasoning, memory accumulates state changes (e.g., "Clicked login button → login form appeared"), and action selects from available action types via a Zod discriminated union built by getActionSchema() (src/agent/tools/agent.ts:100).
Messages are constructed by buildAgentStepMessages() (src/agent/messages/builder.ts:9) which appends: the final goal, current URL, variables, previous action history (assistant thought/memory/action → user result), the DOM tree, and optionally a composited screenshot with scroll info. The system prompt (src/agent/messages/system-prompt.ts) describes input/output format, available actions, and guidelines (one action per step, use wait when unsure, avoid loops).
Stop conditions are checked at each iteration: TaskStatus.PAUSED sleeps 100ms and continues; endTaskStatuses (COMPLETED, FAILED, CANCELLED) breaks; maxSteps threshold cancels the task (src/agent/tools/agent.ts:303-306); and a consecutive failure/waits counter of 5 (MAX_CONSECUTIVE_FAILURES_OR_WAITS) fails the task if exceeded (src/agent/tools/agent.ts:267). The loop also tracks schema validation errors — up to 3 are fed back into the next step's messages for cross-step learning (src/agent/tools/agent.ts:400-413). LLM structured output has up to 3 retry attempts per step with Zod error feedback appended as user messages (src/agent/tools/agent.ts:424-508).
How are failures, retries and self-healing handled?
answeredReliability is handled through multiple layers of retries, error classification, and action caching. The general-purpose retry() utility (src/utils/retry.ts:2) provides exponential backoff (2^i * 1000ms) with up to 3 attempts — used for LLM invocations, DOM state capture, and scroll info retrieval. DOM capture has its own retry loop (src/agent/shared/dom-capture.ts:105) with 3 attempts for transient CDP errors ("Execution context was destroyed", "Cannot find context", "Target closed"), plus a placeholder-snapshot detector via isPlaceholderSnapshot() that retries on "Error: Could not extract accessibility tree".
LLM structured output failures get 3 attempts per step with Zod validation errors fed back as corrective user messages (src/agent/tools/agent.ts:424-508). Schema errors accumulate across steps (last 3 tracked) and are injected into future step messages to guide the model (src/agent/tools/agent.ts:400-413). The consecutive failure/waits counter terminates tasks after 5 consecutive waits or action failures, with a clear error message (src/agent/tools/agent.ts:267, 575-624).
Action caching records each step's XPath, frame index, method, and arguments via buildActionCacheEntry() (src/agent/shared/action-cache.ts:95). Successful tasks produce an ActionCacheOutput that can be replayed via runFromActionCache() (src/agent/index.ts:519-834), which replays cached actions by dispatching through perform* helpers with XPath retry support and fallback to instruction-based execution. Element resolution has its own recovery: if DOM.resolveNode fails with "Could not find node with given id", the system falls back to XPath-based recovery via recoverBackendNodeId() (src/cdp/element-resolver.ts:228) which evaluates the XPath in the correct frame's execution context, then calls DOM.describeNode to get a fresh backendNodeId. Page context switching is detected and triggers retry with a 500ms delay (src/agent/index.ts:1537-1558).
Which models are supported and how are they called?
answeredThe project supports four LLM providers: OpenAI, Anthropic, Gemini, and DeepSeek, all implementing the common HyperAgentLLM interface (src/llm/types.ts:76) with three key methods: invoke(), invokeStructured(), getCapabilities(). Provider selection happens in createLLMClient() (src/llm/providers/index.ts:18) which routes LLMConfig.provider to the correct client constructor. Default is OpenAI (gpt-4o) if no LLM is configured but OPENAI_API_KEY is set (src/agent/index.ts:108-113).
Vision requirement: All providers must support multimodal (vision) input for the screenshot-in-prompt flow — getCapabilities().multimodal returns true for all four (src/llm/providers/openai.ts:218, src/llm/providers/anthropic.ts:127, etc.).
Structured output varies by provider:
- OpenAI uses
response_format: {type: "json_schema"}with Zod-jsonSchema conversion viaconvertToOpenAIJsonSchema()(src/llm/providers/openai.ts:153-206). Temperature is omitted for GPT-5 models at temperature 1. - Anthropic uses tool calling with two paths:
invokeStructuredViaTools()converts action definitions to Anthropic tool schemas and forces tool use (tool_choice: "any") (src/llm/providers/anthropic.ts:133-230);invokeStructuredViaSimpleTool()wraps arbitrary Zod schemas as a single "structured_output" tool (src/llm/providers/anthropic.ts:236-283). - Gemini uses
responseMimeType: "application/json"withresponseSchemafor JSON mode (src/llm/providers/gemini.ts). - DeepSeek reuses the OpenAI SDK (
src/llm/providers/deepseek.ts:23) withbaseURL: "https://api.deepseek.com"and supportsresponse_formatfor JSON mode.
Custom baseURL is supported for OpenAI and DeepSeek, enabling proxy/compatible endpoints. The getCapabilities() method reports multimodal, toolCalling, and jsonMode flags — the agent loop uses invokeStructured() exclusively for deterministic action parsing with Zod validation.
How are browser sessions, profiles, auth and anti-bot handled?
answeredBrowser sessions are managed through the abstract BrowserProvider<T> base class (src/types/browser-providers/types.ts:3), with two concrete implementations: LocalBrowserProvider (src/browser-providers/local.ts) and HyperbrowserProvider (src/browser-providers/hyperbrowser.ts). Selection is controlled by HyperAgentConfig.browserProvider (default: "Local") (src/agent/index.ts:81-83, 124).
LocalBrowserProvider launches Playwright's Chromium with channel: "chrome", headless: false, and the stealth flag --disable-blink-features=AutomationControlled (src/browser-providers/local.ts:13-18). This flag helps evade anti-bot detection by suppressing the "Chrome is being controlled by automated software" banner. Additional launch options can be passed via localConfig.
HyperbrowserProvider (src/browser-providers/hyperbrowser.ts) connects to cloud browser sessions via the @hyperbrowser/sdk. It creates a remote session through the Hyperbrowser API (client.sessions.create()), then connects Playwright to it via chromium.connectOverCDP(session.wsEndpoint) (src/browser-providers/hyperbrowser.ts:35-41). Session details (live URL, session ID, info URL) are logged in debug mode. On close, it both closes the Playwright browser and calls client.sessions.stop() to terminate the cloud session (src/browser-providers/hyperbrowser.ts:58-63).
Profiles, cookies, and auth are managed through raw Playwright BrowserContext — the agent creates one context per browser session with viewport: null (src/agent/index.ts:160-162). The context is not configured with a persistent profile directory by default; users can pass a Playwright-provided context via initPage into executeTask()/executeTaskAsync() for session reuse. Anti-bot beyond the AutomationControlled flag is not implemented — the system prompt instructs the agent to "Accept or close" cookie banners and to try refreshing or alternative approaches for CAPTCHAs (src/agent/messages/system-prompt.ts:74-77). Proxies are not configured in the agent layer but can be passed through Playwright launch options via localConfig or hyperbrowserConfig, which are passed through to the underlying chromium.launch() / chromium.connectOverCDP() calls.