razaanstha/ulka
Chromium side-panel agent where an FX planner hands bounded subgoals to Jev, which picks actions from an accessibility-tree snapshot over CDP.
Overview
Ulka is an experimental Manifest V3 extension that adds a chat side panel to a Chromium browser (developed in Helium). You type a task, and it drives the current tab through the chrome.debugger API. There is no backend: the extension calls Vercel AI Gateway directly with the user’s own key.
It splits the work across three kinds of model call. An FX agent (the libfx WebAssembly runtime) running deepseek/deepseek-v4.1-flash is the planner. It never touches the page directly, only through a dozen host tools such as observe_browser, browser_subgoal and complete_task. Jev (typesafe-ai/jev, through Gateway’s evaluation API) does action selection: given the observed controls, it picks one operation and one target from closed lists. A plain language model writes field values, reviews them, and verifies whether the task is done. Only that last check can end a task with “done”.
Ulka borrows Stagehand’s observe/act/extract and action-caching ideas but runs its own small CDP driver. The code is young, heavily guarded, and full of limits added after specific failures, which the README documents openly.
Architecture
flowchart LR
SP["Side panel chat"] --> BG["background.ts fxChat"]
BG --> FX["runFxBrowser (libfx WASM planner)"]
FX -->|"browser_subgoal"| RJ["runJevSubgoal"]
FX -->|"observe / read / extract"| OBS["CdpObserver"]
FX -->|"complete_task"| VER["OutcomeVerifier"]
RJ --> AR["AgentRunner loop"]
AR --> OBS
AR --> JEV["VercelJevDecisionEngine"]
AR --> TXT["TextGenerator + review"]
AR --> EX["BrowserExecutor"]
OBS --> AX["Accessibility.getFullAXTree + DOM"]
EX --> CDP["chrome.debugger CDP"]
JEV --> GW["Vercel AI Gateway"]
FX --> GW
VER --> GW
| Component | Path | Role |
|---|---|---|
| Service worker | apps/extension/src/background.ts |
Message hub; fxChat builds the FX host (tools, approvals, tab session); runJevSubgoal wires one AgentRunner |
| FX planner | apps/extension/src/agent/fx-agent.ts |
runFxBrowser: tool definitions, tool budget, safety stops, final-verification recovery loop |
| Subgoal loop | apps/extension/src/agent/agent-runner.ts |
Observe, decide, gate, execute, settle; loop, scroll and wait detectors |
| Jev decisions | apps/extension/src/agent/vercel-jev.ts, action-space.ts |
Builds state and choice questions, validates the returned distribution |
| Observation | apps/extension/src/agent/observer.ts, accessibility.ts |
AX tree to DOM mapping, in-page snapshot, settling |
| Execution | apps/extension/src/agent/executor.ts, freshness.ts |
CDP input events with last-moment target validation |
| Text and checks | text-generator.ts, outcome-verifier.ts |
Field values with a reviewer pass; structured completion verdicts |
| Safety | approvals.ts, url-provenance.ts, navigation.ts |
Keyword risk levels, navigation only to user-supplied or observed URLs |
| Protocol | packages/protocol/src/ |
PageSnapshot, AgentOperation, decision types |
How a request flows
- Start. The side panel sends the chat.
chat()reads the Gateway key fromchrome.storage.localand callsfxChatwith the active tab. That sets a shared task budget of 30 actions and 60 model steps (background.ts). - Plan.
runFxBrowsercreates the FX agent with the browsing skill prompt and the host tools. Tool calls run one at a time and stop after 24 calls (36 or 48 onceplan_tasksets milestones) (fx-agent.ts, L257-L263). - Delegate. For interaction, FX calls
browser_subgoal(goal). The host runsrunJevSubgoal, which builds anAgentRunnerinsubgoalmode with its own cap of 8 actions and 16 model calls on top of the shared budget (background.ts). - Observe.
CdpObserver.observefetches the AX tree for the main frame and child frames, maps backend node ids to DOM elements, then runs the in-page observer to buildPageSnapshot.elementswith ids, roles, labels, states and allowed operations (observer.ts, accessibility.ts). - Decide. Unless a cached CLICK/SELECT matches,
VercelJevDecisionEngine.decideOncesends one evaluation: anoperationchoice plus one<op>_targetchoice per operation with several targets. It first removes actions that were exhausted, ineffective twice, or would close a cycle (vercel-jev.ts). - Gate. For
TYPE_TEXT,TextGeneratorwrites the value, and a second call reviews it.classifyActionthen marks the action safe, confirm or blocked. Confirm actions wait for a side-panel approval (agent-runner.ts). - Execute and settle.
BrowserExecutor.executere-validates the snapshot and target, then sends CDP input.waitForChangepolls for a content change and two quiet samples (3 s, or 10 s for WAIT) (agent-runner.ts). - Checkpoint. In subgoal mode the runner hands control back on a URL change, after typing, after Escape, or when Jev says DONE. It returns an unverified checkpoint with the fresh observation, not a success (agent-runner.ts).
- Verify. FX calls
complete_task, or its turn ends.OutcomeVerifierjudges the original request against all host-captured evidence. If it says no, FX gets the feedback and up to two more turns (fx-agent.ts).
Key components
Observation
Accessibility is the primary source. AccessibilitySource.prepare calls Accessibility.getFullAXTree, resolves each node with DOM.resolveNode, and registers the element in a page-global cache (window.__ulkaAgent). The in-page observer then assigns numeric node ids, which the executor and the freshness validator use to find the live element again. If the AX tree is unavailable, the snapshot is built from a DOM selector scan and diagnostics report dom-fallback. The model sees no screenshots. Offscreen controls are kept as context up to 250 rows, with an omission count. While settling, the observer polls a cheaper DOM-only snapshot and re-reads AX once at the end (observer.ts).
Jev decisions
The state carries page text, an element table, action_targets, tabs, scroll position and the last ten actions. Extra rules for comboboxes, dates and form context are added only when such controls are present. validateGatewayChoice rejects any answer whose probabilities are not exactly the offered keys, do not sum to about 1, or do not put the choice on top. A malformed answer gets one retry. A 5xx from the AI SDK route gets one raw POST /v1/evaluate attempt, and after that the run stops (vercel-jev.ts, L278-L295).
Execution
Everything goes through CDP on the attached tab: Input.dispatchMouseEvent for clicks, hover and wheel; Input.dispatchKeyEvent with real key codes for Enter, Escape and arrows; Input.insertText after selecting the field’s full contents; Page.navigateToHistoryEntry and Page.reload. Before typing, it checks that the clicked element still owns focus, so text never lands in a popup that stole it (executor.ts).
Guards and caching
Approvals are keyword rules on the target label: “send”, “submit”, “buy”, “delete” need confirmation, and “delete account” or “transfer money” are blocked. Sensitive-looking fields such as card or SSN need confirmation before typing (approvals.ts). Navigation URLs must come from the user or from a browser result. Successful CLICK/SELECT actions are cached by goal, URL, page state, role, label and ordinal, and replayed only on an exact match (action-cache.ts).
Extending it
- Planner behaviour. Edit
apps/extension/src/agent/skills/web-browsing/SKILL.md, which is bundled into FX’s instructions. Rebuild and reload. - Models.
models.tsholds the planner/text model id. The Jev id is hard-coded invercel-jev.ts. Both route through Gateway only. - Site tools. Pages that register WebMCP tools (
document.modelContext) are discovered at observe time and callable throughwebmcp_call, each call gated by approval. - New operations or tools. Add an
AgentOperationinpackages/protocol, a branch inBrowserExecutor, a risk rule inapprovals.ts, or a newtool(...)inrunFxBrowser.
Running it
- Build. Bun 1.4.2:
bun install --frozen-lockfile && bun run build, then loadapps/extension/distunpacked. - Browser. A Chromium build with side panels, the debugger API and WebAssembly JSPI (FX refuses to start without JSPI). Only Helium on macOS is tested.
- Key. Paste a Vercel AI Gateway key in Settings. It is stored unencrypted in
chrome.storage.local. Every Gateway request explicitly setszeroDataRetention: false(gateway-policy.ts). - Tests.
bun run checkruns typecheck, build and the Bun test suite.bun run smokeruns a headless-Chrome accessibility smoke test.
Strengths and caveats
- Strength: completion is checked, not assumed. Subgoals return unverified checkpoints, and only an independent verifier over host-captured evidence can mark the task done.
- Strength: constrained execution. Jev chooses among observed ids. Targets are re-validated (connected, visible, unoccluded, same label) right before CDP input.
- Strength: transparent operations. Per-reply token and cost estimates, detailed diagnostic events, and a model connection test in the panel.
- Caveat: limits disagree with the README. The README says the eight-action subgoal boundary was removed, but
runJevSubgoalstill passesmaxActions: 8, and the runner enforces it alongside the shared 30/60 budget. - Caveat: heuristic safety. Approval is keyword matching on labels, and the README says plainly that it is not a security boundary. It runs inside your signed-in browser profile.
- Caveat: narrow platform. It needs one Gateway account and a JSPI-capable Chromium, has been tested only in Helium, and has limited support for iframes, shadow DOM and complex widgets.
- Caveat: many model calls. A typed field costs a generation plus a review call, and planning, Jev and verification are separate requests. Latency and cost add up on long tasks.
Sources: code at bb04433, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (31 pages), verified Q&A.
How it answers the Browser & computer control questions
Each answer was drafted by a code-reading agent at commit bb04433. Its citations were checked mechanically. Compare with the other browser & computer control →
How is the page represented to the model?
answeredThe page is represented primarily through a DOM-based serialization with an accessibility-tree fallback, executed as a JavaScript expression injected into the browser page via CDP's Runtime.evaluate. The OBSERVER_EXPRESSION in observer.ts:10-192 runs in-page: it walks visible interactive elements from document.body (or the topmost aria-modal dialog), extracting each control's role, label, href, value, checked, expanded, selected, focused, and other ARIA properties. The result is a PageSnapshot (defined in snapshot.ts:36-51) containing a url, title, concatenated body text (up to 6000 chars, observer.ts:184-188), a scroll state, an elements array of PageElement objects, a guards map (for later target validation), and diagnostics about omitted controls. No screenshots or set-of-marks visual overlays are sent to the model — the representation is purely structured text. Each PageElement carries an id (e.g. "e1"), a role, a label, and an operations array listing which agent actions are valid on it (observer.ts:102-107). Labels come from aria-labelledby, aria-label, <label> elements, alt text, or inner text — resolved in the in-page name() function (observer.ts:22-31). Accessibility fallback: if the DOM scan reveals contenteditable elements or focused controls not in the standard selector set, the observer switches to the AX path (accessibility.ts). AccessibilitySource.prepare() fetches Chrome's full accessibility tree via Accessibility.getFullAXTree, resolves AX backend node IDs to DOM handles, and enriches the snapshot with AX-derived roles and labels (observer.ts:299-310). Size limits and pruning: visible interactive elements get the full observation budget; offscreen controls are capped at 250 rows (observer.ts:98). The candidate pool is sorted so available controls are listed before unavailable ones (observer.ts:82). Rejected and unavailable controls are counted in diagnostics.rejected and diagnostics.unavailable. The page text field is capped at 6000 characters from a tree-walker traversal (observer.ts:184-188). Context elements (form/dialog regions) are limited to 3 per control and labels to 500 chars. A diagnostics.modalScoped boolean tells the model when interaction is scoped to a modal dialog.
Accessibility.getFullAXTree (main and child frames), maps AX nodes to DOM elements, and builds candidates from those AX records; the DOM selector scan is used only when no AX tree is returned (diagnostics dom-fallback).How are actions executed and how are elements targeted?
answeredAll actions execute through Chrome's DevTools Protocol (CDP) — not Playwright, OS-level input, or Selenium. The BrowserExecutor class (executor.ts:13-125) receives a PageSnapshot, an AgentDecision (operation + target), and optional text, then dispatches the appropriate CDP commands. Element targeting: elements are identified by their DOM nodeId (an auto-incrementing integer stored in a page-level WeakMap/Map called __ulkaAgent.nodes); these IDs survive page mutations within a session. Before execution, the FreshnessValidator.validateTarget() method (freshness.ts:15-29) re-checks the element — verifying it is still connected, visible, enabled, not occluded, and that its role/label/href match — and returns a live {x, y} click point from interactionPoint(). If any check fails, a StaleDecisionError is thrown. Clicking and right-clicking: Input.dispatchMouseEvent with mousePressed/mouseReleased and the resolved coordinates (executor.ts:54-58,86-89). Typing: for native date/time inputs, value is set directly via element.value = ... plus input/change events (executor.ts:101-103). For all other text fields, the content is selected (select() or Range selection) and then Input.insertText inserts the full value atomically (executor.ts:105-121). Keyboard actions (ArrowDown, ArrowUp, Enter, Escape): the element is focused first, then Input.dispatchKeyEvent with the correct windowsVirtualKeyCode and key string, including Enter's \r character event for native form submission (executor.ts:65-76). Scrolling: Input.dispatchMouseEvent with mouseWheel and a deltaY value of ±420 or ±560 (executor.ts:38-48,62-63). If the snapshot has a scrollTarget, the scroll container is validated live. Page navigation: GO_BACK/GO_FORWARD use Page.getNavigationHistory + Page.navigateToHistoryEntry (executor.ts:32-36). RELOAD uses Page.reload (executor.ts:31). Tabs are managed via the Chrome extension API (NativeTabs, native-tabs.ts): open (chrome.tabs.create), switch (chrome.tabs.update), close (chrome.tabs.remove). The TabController interface at executor.ts:7-11 is implemented in background.ts:286-303 with background-mode safety checks. Downloads are handled through Chrome's downloads API (read-only listing with downloads.ts). No file upload mechanism was found beyond native date typing.
How is the agent loop / planning implemented?
answeredThe agent runs a two-tier loop: an outer FX (Helium) planner and an inner Jev subgoal executor. Outer FX loop (fx-agent.ts): runFxBrowser() creates an FX agent (powered by libfx WASM runtime) with a fixed set of tools: observe_browser, navigate_browser, browser_subgoal, search_web, plan_task, complete_task, read_page, extract_page, act_action, native_tabs, webmcp_call, ask_user, list_downloads. The FX agent is an LLM that plans and calls tools in a loop — it can plan_task to track milestones, browser_subgoal to delegate atomic work, and complete_task to verify. The loop runs up to 3 recovery attempts if verification fails (fx-agent.ts:280-330). Tool calls are capped at 24 (expanded to 36-48 when a plan with 3+ milestones exists) (fx-agent.ts:75,260). Inner Jev loop (agent-runner.ts): AgentRunner.run() implements a step loop for atomic subgoals. Each iteration: (1) observe page via CdpObserver.observe() or waitForChange(), (2) ask Jev for a decision via VercelJevDecisionEngine.decide(), (3) generate text if TYPE_TEXT, (4) check action approval policy, (5) execute via BrowserExecutor.execute(), (6) wait for page change via waitForChange(), (7) repeat until DONE or blocked. Jev decision-making (vercel-jev.ts): The VercelJevDecisionEngine.buildActionSpace() method constructs an ActionSpace from the snapshot — a map of valid operations to human-readable descriptions, and operation-specific target lists. This space is sent to the typesafe-ai/jev evaluation model via Gateway's HTTP API (SDK experimental_evaluate or raw HTTP fallback). Jev applies decisionRules() context (vercel-jev.ts:22-29) that adds combobox, date, and context rules when relevant. The Jev response must produce a valid choice with well-formed probabilities (summing to ~1, highest selected) (vercel-jev.ts:278-295). Stop conditions include: step limits (30 actions, 60 model calls), action repetition detection (hasIneffectiveRepetition, hasActionCycle in history.ts), excessive scroll-to-now-content (3 empty scrolls), page-state revisits (3 visits to same state), repeated stale targets (3 consecutive StaleDecisionError), and tool error cascades. Memory between steps is tracked via TaskMemory (agent-runner.ts:28-36): an action history array, attempts map (action key → count), stateVisits map (progress state → count), and an actionCache for CLICK/SELECT replay across subgoals.
browser_subgoal runs an AgentRunner capped at 8 actions and 16 model calls (runJevSubgoal in background.ts), and in subgoal mode the runner returns an unverified checkpoint after a URL change, typing, Escape or Jev's DONE.How are failures, retries and self-healing handled?
answeredError classes caught: The runner catches StaleDecisionError (element changed, detached, or occluded since snapshot), TextTargetMismatchError (wrong field for text), TextGenerationUnavailableError, and VerificationUnavailableError. The FX agent additionally catches gateway-specific errors: JevGatewayUnavailableError, JevGatewayRejectedError, JevGatewayInvalidResponseError, tab-closed errors, and debugger-conflict errors (fx-agent.ts:182-203). Retries: The Jev decision engine retries malformed evaluation output once (vercel-jev.ts:144-151). Gateway SDK failures (status 5xx, retryable gateway errors) trigger a single HTTP fallback attempt (vercel-jev.ts:116-133). The outcome verifier retries once on transient service errors or malformed responses, with increased maxOutputTokens on schema failures (outcome-verifier.ts:82-119). Text generation repairs malformed or rejected content once via a second model call (text-generator.ts:114-149). The observer tries up to 20 times (50ms apart) for a document to be ready (observer.ts:298). The FX outer loop has 3 recovery attempts if verification fails (fx-agent.ts:281). Replanning: Stale targets trigger feedback with evidence text guiding the Jev model to re-observe and choose differently (agent-runner.ts:189-198). Failed verification sets excludeDone=true and passes evidence as feedback for the next decision. Blocked subgoals are counted — after 2 consecutive blocks the FX planner stops (fx-agent.ts:157-158). Action caching (action-cache.ts): Successful CLICK and SELECT actions that produced a page change are cached by semantic key (goal + URL + operation + role + label + ordinal + page precondition state). On fresh runs, findCachedAction() replays a matching cached action without a model call (agent-runner.ts:88,94). Timeouts: Page settling during waitForChange() uses a 3-second budget (10s for explicit WAIT) (agent-runner.ts:215). Navigation waits up to 30s (navigation.ts:28-68). Model connection tests timeout at 10s. The observer's observe() retries for up to ~1 second. Consecutive-error stops: 3 consecutive stale targets, 3 consecutive ineffective actions (no page change), 3 consecutive tool errors, or 2 consecutive blocked subgoals all abort the run.
Which models are supported and how are they called?
answeredPrimary model: deepseek/deepseek-v4.1-flash is set as the LANGUAGE_MODEL in models.ts:1. This drives both the FX planner/agent (via libfx WASM) and serves as the TEXT_MODEL for field content generation. The decision model is entirely different: action selection uses typesafe-ai/jev — a specialized evaluation/judge model accessed through Vercel's AI Gateway via experimental_evaluate (the ai SDK) with an HTTP fallback (vercel-jev.ts:92,198,248). Jev receives structured state + multiple-choice questions and returns a chosen operation with probability-weighted confidence. Vision: there is no screenshot or vision input anywhere in the codebase. The model receives only structured text: the DOM serialization as PageSnapshot, a PageExtractor that extracts typed fields from page content, and ReadPageExpression for reading offscreen text. Structured output / tool calling: the FX agent uses libfx's createFxAgent() which provides native tool calling — tools are defined with Zod schemas, converted to JSON Schema via z.toJSONSchema(), and executed by the agent runtime (fx-agent.ts:99-215). The Jev evaluator uses the Gateway's evaluation API with choice-type questions. Text generation and verification use ai SDK's Output.object() with Zod schemas for structured JSON responses — the text generator returns {suitable, reason, text}, the date generator returns {suitable, reason, date}, and the verifier returns {satisfied, evidence}. Small/specialised models: Jev (typesafe-ai/jev) is itself a specialised model for evaluation/choice rather than free-form generation. The text generator (TEXT_MODEL) reuses the same model as the planner. Providers: all model calls go through the Vercel AI Gateway (apiKey stored in Chrome storage, accessed via createGateway() and createFxGatewayFetch()). The gateway applies a configured zeroDataRetention: false policy (gateway-policy.ts:8). SDK: Uses the ai SDK (version ^7.0.116) for streaming text generation, structured output, and evaluation. Model pricing for cost tracking is loaded from https://ai-gateway.vercel.sh/v1/models in model-pricing.ts and reported via ModelUsageLedger.
How are browser sessions, profiles, auth and anti-bot handled?
answeredLocal browser only: Ulka operates as a Chrome extension (MV3) and controls the user's local Chrome browser — it attaches to existing tabs via CDP, not remote/cloud browsers. The AttachedBrowser class (browser.ts:7-22) wraps a CDP debugger session: attach() sends api.attach(target, "1.3"), detach() releases it, and switchTo() detaches, updates the tab ID, and re-attaches. TaskBrowserSession (browser.ts:24-47) manages reuse of a single AttachedBrowser instance across a task, ensuring only one CDP connection at a time. The extension stores API keys in Chrome's storage.local. Persistent profiles and cookies: there is no auth/profile/cookie management in the codebase — no cookie injection, no saved session restoration, and no login profile handling. The extension uses whichever Chrome profile the user is already signed into. Stealth and anti-bot: there is no stealth feature — no user-agent manipulation, no request header stripping, no fingerprint masking, and no anti-detection measures. The CDP attachment is standard Chrome Debugger Protocol without modifications. Proxies: there is no proxy configuration. CAPTCHA handling: there is no CAPTCHA detection or solving mechanism. When a CAPTCHA appears, the system has no special handling — it would just observe whatever interactive controls the page exposes. The BLOCKED operation (snapshot.ts:8, agent-runner.ts:118) is a catch-all escape for when the model decides it cannot progress, but nothing specific to CAPTCHAs exists. Background mode: the extension supports a background mode (storage.local ulkaBackground) where tab switches happen without activating the tab, allowing the user to continue browsing while the agent works (background.ts:98,105). Navigation guard: UrlProvenance (url-provenance.ts traced in fx-agent.ts as urls.assertObserved()) ensures the agent only navigates to URLs that were either supplied by the user or observed in a previous browser result — preventing navigation to unverified or hallucinated URLs. validateNavigationUrl() (navigation.ts:5-13) restricts navigation to https:// URLs or internal chrome:// pages (extensions, history, downloads, bookmarks, settings, newtab, version only). Internal chrome:// pages return a stub PageSnapshot since their content is unavailable via CDP (navigation.ts:17-26).