# razaanstha/ulka

> Chromium side-panel agent where an FX planner hands bounded subgoals to Jev, which picks actions from an accessibility-tree snapshot over CDP.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/razaanstha/ulka (reviewed at commit `bb04433dcebb75b1a93eb5a6dc7d25909930cc65`, 2026-09-26)
- Stars: 22 · Language: TypeScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/ulka/

## Overview

Ulka is an experimental Manifest V3 extension that adds a chat side panel to a Chromium browser (developed in Helium). You type a task, and it drives the current tab through the `chrome.debugger` API. There is no backend: the extension calls Vercel AI Gateway directly with the user's own key.

It splits the work across three kinds of model call. An **FX** agent (the `libfx` WebAssembly runtime) running `deepseek/deepseek-v4.1-flash` is the planner. It never touches the page directly, only through a dozen host tools such as `observe_browser`, `browser_subgoal` and `complete_task`. **Jev** (`typesafe-ai/jev`, through Gateway's evaluation API) does action selection: given the observed controls, it picks one operation and one target from closed lists. A plain **language model** writes field values, reviews them, and verifies whether the task is done. Only that last check can end a task with "done".

Ulka borrows Stagehand's observe/act/extract and action-caching ideas but runs its own small CDP driver. The code is young, heavily guarded, and full of limits added after specific failures, which the README documents openly.

## Architecture

```mermaid
flowchart LR
  SP["Side panel chat"] --> BG["background.ts fxChat"]
  BG --> FX["runFxBrowser (libfx WASM planner)"]
  FX -->|"browser_subgoal"| RJ["runJevSubgoal"]
  FX -->|"observe / read / extract"| OBS["CdpObserver"]
  FX -->|"complete_task"| VER["OutcomeVerifier"]
  RJ --> AR["AgentRunner loop"]
  AR --> OBS
  AR --> JEV["VercelJevDecisionEngine"]
  AR --> TXT["TextGenerator + review"]
  AR --> EX["BrowserExecutor"]
  OBS --> AX["Accessibility.getFullAXTree + DOM"]
  EX --> CDP["chrome.debugger CDP"]
  JEV --> GW["Vercel AI Gateway"]
  FX --> GW
  VER --> GW
```

| Component | Path | Role |
|---|---|---|
| Service worker | `apps/extension/src/background.ts` | Message hub; `fxChat` builds the FX host (tools, approvals, tab session); `runJevSubgoal` wires one `AgentRunner` |
| FX planner | `apps/extension/src/agent/fx-agent.ts` | `runFxBrowser`: tool definitions, tool budget, safety stops, final-verification recovery loop |
| Subgoal loop | `apps/extension/src/agent/agent-runner.ts` | Observe, decide, gate, execute, settle; loop, scroll and wait detectors |
| Jev decisions | `apps/extension/src/agent/vercel-jev.ts`, `action-space.ts` | Builds state and choice questions, validates the returned distribution |
| Observation | `apps/extension/src/agent/observer.ts`, `accessibility.ts` | AX tree to DOM mapping, in-page snapshot, settling |
| Execution | `apps/extension/src/agent/executor.ts`, `freshness.ts` | CDP input events with last-moment target validation |
| Text and checks | `text-generator.ts`, `outcome-verifier.ts` | Field values with a reviewer pass; structured completion verdicts |
| Safety | `approvals.ts`, `url-provenance.ts`, `navigation.ts` | Keyword risk levels, navigation only to user-supplied or observed URLs |
| Protocol | `packages/protocol/src/` | `PageSnapshot`, `AgentOperation`, decision types |

## How a request flows

1. **Start.** The side panel sends the chat. `chat()` reads the Gateway key from `chrome.storage.local` and calls `fxChat` with the active tab. That sets a shared task budget of 30 actions and 60 model steps ([background.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/background.ts#L93-L110)).
2. **Plan.** `runFxBrowser` creates the FX agent with the browsing skill prompt and the host tools. Tool calls run one at a time and stop after 24 calls (36 or 48 once `plan_task` sets milestones) ([fx-agent.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L99-L120), [L257-L263](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L257-L263)).
3. **Delegate.** For interaction, FX calls `browser_subgoal(goal)`. The host runs `runJevSubgoal`, which builds an `AgentRunner` in `subgoal` mode with its own cap of 8 actions and 16 model calls on top of the shared budget ([background.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/background.ts#L316-L357)).
4. **Observe.** `CdpObserver.observe` fetches the AX tree for the main frame and child frames, maps backend node ids to DOM elements, then runs the in-page observer to build `PageSnapshot.elements` with ids, roles, labels, states and allowed operations ([observer.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/observer.ts#L297-L316), [accessibility.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/accessibility.ts#L70-L130)).
5. **Decide.** Unless a cached CLICK/SELECT matches, `VercelJevDecisionEngine.decideOnce` sends one evaluation: an `operation` choice plus one `<op>_target` choice per operation with several targets. It first removes actions that were exhausted, ineffective twice, or would close a cycle ([vercel-jev.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/vercel-jev.ts#L154-L275)).
6. **Gate.** For `TYPE_TEXT`, `TextGenerator` writes the value, and a second call reviews it. `classifyAction` then marks the action safe, confirm or blocked. Confirm actions wait for a side-panel approval ([agent-runner.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/agent-runner.ts#L143-L177)).
7. **Execute and settle.** `BrowserExecutor.execute` re-validates the snapshot and target, then sends CDP input. `waitForChange` polls for a content change and two quiet samples (3 s, or 10 s for WAIT) ([agent-runner.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/agent-runner.ts#L179-L221)).
8. **Checkpoint.** In subgoal mode the runner hands control back on a URL change, after typing, after Escape, or when Jev says DONE. It returns an *unverified checkpoint* with the fresh observation, not a success ([agent-runner.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/agent-runner.ts#L222-L270)).
9. **Verify.** FX calls `complete_task`, or its turn ends. `OutcomeVerifier` judges the original request against all host-captured evidence. If it says no, FX gets the feedback and up to two more turns ([fx-agent.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L278-L331)).

## Key components

### Observation

Accessibility is the primary source. `AccessibilitySource.prepare` calls `Accessibility.getFullAXTree`, resolves each node with `DOM.resolveNode`, and registers the element in a page-global cache (`window.__ulkaAgent`). The in-page observer then assigns numeric node ids, which the executor and the freshness validator use to find the live element again. If the AX tree is unavailable, the snapshot is built from a DOM selector scan and diagnostics report `dom-fallback`. The model sees no screenshots. Offscreen controls are kept as context up to 250 rows, with an omission count. While settling, the observer polls a cheaper DOM-only snapshot and re-reads AX once at the end ([observer.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/observer.ts#L242-L292)).

### Jev decisions

The state carries page text, an element table, `action_targets`, tabs, scroll position and the last ten actions. Extra rules for comboboxes, dates and form context are added only when such controls are present. `validateGatewayChoice` rejects any answer whose probabilities are not exactly the offered keys, do not sum to about 1, or do not put the choice on top. A malformed answer gets one retry. A 5xx from the AI SDK route gets one raw `POST /v1/evaluate` attempt, and after that the run stops ([vercel-jev.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/vercel-jev.ts#L107-L133), [L278-L295](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/vercel-jev.ts#L278-L295)).

### Execution

Everything goes through CDP on the attached tab: `Input.dispatchMouseEvent` for clicks, hover and wheel; `Input.dispatchKeyEvent` with real key codes for Enter, Escape and arrows; `Input.insertText` after selecting the field's full contents; `Page.navigateToHistoryEntry` and `Page.reload`. Before typing, it checks that the clicked element still owns focus, so text never lands in a popup that stole it ([executor.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/executor.ts#L21-L124)).

### Guards and caching

Approvals are keyword rules on the target label: "send", "submit", "buy", "delete" need confirmation, and "delete account" or "transfer money" are blocked. Sensitive-looking fields such as card or SSN need confirmation before typing ([approvals.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/approvals.ts#L1-L28)). Navigation URLs must come from the user or from a browser result. Successful CLICK/SELECT actions are cached by goal, URL, page state, role, label and ordinal, and replayed only on an exact match ([action-cache.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/action-cache.ts#L50-L67)).

## Extending it

- **Planner behaviour.** Edit `apps/extension/src/agent/skills/web-browsing/SKILL.md`, which is bundled into FX's instructions. Rebuild and reload.
- **Models.** `models.ts` holds the planner/text model id. The Jev id is hard-coded in `vercel-jev.ts`. Both route through Gateway only.
- **Site tools.** Pages that register WebMCP tools (`document.modelContext`) are discovered at observe time and callable through `webmcp_call`, each call gated by approval.
- **New operations or tools.** Add an `AgentOperation` in `packages/protocol`, a branch in `BrowserExecutor`, a risk rule in `approvals.ts`, or a new `tool(...)` in `runFxBrowser`.

## Running it

- **Build.** Bun 1.4.2: `bun install --frozen-lockfile && bun run build`, then load `apps/extension/dist` unpacked.
- **Browser.** A Chromium build with side panels, the debugger API and WebAssembly JSPI (FX refuses to start without JSPI). Only Helium on macOS is tested.
- **Key.** Paste a Vercel AI Gateway key in Settings. It is stored unencrypted in `chrome.storage.local`. Every Gateway request explicitly sets `zeroDataRetention: false` ([gateway-policy.ts](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/gateway-policy.ts#L6-L37)).
- **Tests.** `bun run check` runs typecheck, build and the Bun test suite. `bun run smoke` runs a headless-Chrome accessibility smoke test.

## Strengths and caveats

- **Strength: completion is checked, not assumed.** Subgoals return unverified checkpoints, and only an independent verifier over host-captured evidence can mark the task done.
- **Strength: constrained execution.** Jev chooses among observed ids. Targets are re-validated (connected, visible, unoccluded, same label) right before CDP input.
- **Strength: transparent operations.** Per-reply token and cost estimates, detailed diagnostic events, and a model connection test in the panel.
- **Caveat: limits disagree with the README.** The README says the eight-action subgoal boundary was removed, but `runJevSubgoal` still passes `maxActions: 8`, and the runner enforces it alongside the shared 30/60 budget.
- **Caveat: heuristic safety.** Approval is keyword matching on labels, and the README says plainly that it is not a security boundary. It runs inside your signed-in browser profile.
- **Caveat: narrow platform.** It needs one Gateway account and a JSPI-capable Chromium, has been tested only in Helium, and has limited support for iframes, shadow DOM and complex widgets.
- **Caveat: many model calls.** A typed field costs a generation plus a review call, and planning, Jev and verification are separate requests. Latency and cost add up on long tasks.

*Sources: code at bb04433, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (31 pages), verified Q&A.*

## How razaanstha/ulka answers the Browser & computer control questions

### How is the page represented to the model? (answered)

The page is represented primarily through a **DOM-based serialization** with an *accessibility-tree fallback*, executed as a JavaScript expression injected into the browser page via CDP's `Runtime.evaluate`. The `OBSERVER_EXPRESSION` in `observer.ts:10-192` runs in-page: it walks visible interactive elements from `document.body` (or the topmost `aria-modal` dialog), extracting each control's `role`, `label`, `href`, `value`, `checked`, `expanded`, `selected`, `focused`, and other ARIA properties. The result is a `PageSnapshot` (defined in `snapshot.ts:36-51`) containing a `url`, `title`, concatenated body `text` (up to 6000 chars, `observer.ts:184-188`), a `scroll` state, an `elements` array of `PageElement` objects, a `guards` map (for later target validation), and diagnostics about omitted controls. **No screenshots or set-of-marks visual overlays are sent to the model** — the representation is purely structured text. Each `PageElement` carries an `id` (e.g. `"e1"`), a `role`, a `label`, and an `operations` array listing which agent actions are valid on it (`observer.ts:102-107`). Labels come from `aria-labelledby`, `aria-label`, `<label>` elements, `alt` text, or inner text — resolved in the in-page `name()` function (`observer.ts:22-31`). **Accessibility fallback:** if the DOM scan reveals `contenteditable` elements or focused controls not in the standard selector set, the observer switches to the AX path (`accessibility.ts`). `AccessibilitySource.prepare()` fetches Chrome's full accessibility tree via `Accessibility.getFullAXTree`, resolves AX backend node IDs to DOM handles, and enriches the snapshot with AX-derived roles and labels (`observer.ts:299-310`). **Size limits and pruning:** visible interactive elements get the full observation budget; offscreen controls are capped at 250 rows (`observer.ts:98`). The candidate pool is sorted so available controls are listed before unavailable ones (`observer.ts:82`). Rejected and unavailable controls are counted in `diagnostics.rejected` and `diagnostics.unavailable`. The page `text` field is capped at 6000 characters from a tree-walker traversal (`observer.ts:184-188`). Context elements (form/dialog regions) are limited to 3 per control and labels to 500 chars. A `diagnostics.modalScoped` boolean tells the model when interaction is scoped to a modal dialog.

> **Editor's note.** Correction: accessibility is the primary source, not a fallback. Every observation first calls `Accessibility.getFullAXTree` (main and child frames), maps AX nodes to DOM elements, and builds candidates from those AX records; the DOM selector scan is used only when no AX tree is returned (diagnostics `dom-fallback`).

Citations: [apps/extension/src/agent/observer.ts:10-192](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/observer.ts#L10-L192) · [packages/protocol/src/snapshot.ts:10-51](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/packages/protocol/src/snapshot.ts#L10-L51) · [apps/extension/src/agent/accessibility.ts:68-131](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/accessibility.ts#L68-L131) · [apps/extension/src/agent/model-context.ts:1-9](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/model-context.ts#L1-L9) · [apps/extension/src/agent/model-context.ts:1-9](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/model-context.ts#L1-L9)

### How are actions executed and how are elements targeted? (answered)

All actions execute through Chrome's **DevTools Protocol (CDP)** — not Playwright, OS-level input, or Selenium. The `BrowserExecutor` class (`executor.ts:13-125`) receives a `PageSnapshot`, an `AgentDecision` (operation + target), and optional text, then dispatches the appropriate CDP commands. **Element targeting:** elements are identified by their DOM `nodeId` (an auto-incrementing integer stored in a page-level `WeakMap`/`Map` called `__ulkaAgent.nodes`); these IDs survive page mutations within a session. Before execution, the `FreshnessValidator.validateTarget()` method (`freshness.ts:15-29`) re-checks the element — verifying it is still connected, visible, enabled, not occluded, and that its role/label/href match — and returns a live `{x, y}` click point from `interactionPoint()`. If any check fails, a `StaleDecisionError` is thrown. **Clicking and right-clicking:** `Input.dispatchMouseEvent` with `mousePressed`/`mouseReleased` and the resolved coordinates (`executor.ts:54-58,86-89`). **Typing:** for native date/time inputs, value is set directly via `element.value = ...` plus `input`/`change` events (`executor.ts:101-103`). For all other text fields, the content is selected (`select()` or `Range` selection) and then `Input.insertText` inserts the full value atomically (`executor.ts:105-121`). **Keyboard actions** (ArrowDown, ArrowUp, Enter, Escape): the element is focused first, then `Input.dispatchKeyEvent` with the correct `windowsVirtualKeyCode` and `key` string, including Enter's `\r` character event for native form submission (`executor.ts:65-76`). **Scrolling:** `Input.dispatchMouseEvent` with `mouseWheel` and a `deltaY` value of ±420 or ±560 (`executor.ts:38-48,62-63`). If the snapshot has a `scrollTarget`, the scroll container is validated live. **Page navigation:** `GO_BACK`/`GO_FORWARD` use `Page.getNavigationHistory` + `Page.navigateToHistoryEntry` (`executor.ts:32-36`). `RELOAD` uses `Page.reload` (`executor.ts:31`). **Tabs** are managed via the Chrome extension API (`NativeTabs`, `native-tabs.ts`): open (`chrome.tabs.create`), switch (`chrome.tabs.update`), close (`chrome.tabs.remove`). The `TabController` interface at `executor.ts:7-11` is implemented in `background.ts:286-303` with background-mode safety checks. **Downloads** are handled through Chrome's downloads API (read-only listing with `downloads.ts`). **No file upload** mechanism was found beyond native date typing.


Citations: [apps/extension/src/agent/executor.ts:13-125](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/executor.ts#L13-L125) · [apps/extension/src/agent/freshness.ts:10-29](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/freshness.ts#L10-L29) · [apps/extension/src/agent/native-tabs.ts:11-94](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/native-tabs.ts#L11-L94)

### How is the agent loop / planning implemented? (answered)

The agent runs a **two-tier loop**: an outer FX (Helium) planner and an inner Jev subgoal executor. **Outer FX loop** (`fx-agent.ts`): `runFxBrowser()` creates an FX agent (powered by `libfx` WASM runtime) with a fixed set of tools: `observe_browser`, `navigate_browser`, `browser_subgoal`, `search_web`, `plan_task`, `complete_task`, `read_page`, `extract_page`, `act_action`, `native_tabs`, `webmcp_call`, `ask_user`, `list_downloads`. The FX agent is an LLM that plans and calls tools in a loop — it can `plan_task` to track milestones, `browser_subgoal` to delegate atomic work, and `complete_task` to verify. The loop runs up to 3 recovery attempts if verification fails (`fx-agent.ts:280-330`). Tool calls are capped at 24 (expanded to 36-48 when a plan with 3+ milestones exists) (`fx-agent.ts:75,260`). **Inner Jev loop** (`agent-runner.ts`): `AgentRunner.run()` implements a step loop for atomic subgoals. Each iteration: (1) observe page via `CdpObserver.observe()` or `waitForChange()`, (2) ask Jev for a decision via `VercelJevDecisionEngine.decide()`, (3) generate text if TYPE_TEXT, (4) check action approval policy, (5) execute via `BrowserExecutor.execute()`, (6) wait for page change via `waitForChange()`, (7) repeat until DONE or blocked. **Jev decision-making** (`vercel-jev.ts`): The `VercelJevDecisionEngine.buildActionSpace()` method constructs an `ActionSpace` from the snapshot — a map of valid operations to human-readable descriptions, and operation-specific target lists. This space is sent to the `typesafe-ai/jev` evaluation model via Gateway's HTTP API (SDK `experimental_evaluate` or raw HTTP fallback). Jev applies `decisionRules()` context (`vercel-jev.ts:22-29`) that adds combobox, date, and context rules when relevant. The Jev response must produce a valid choice with well-formed probabilities (summing to ~1, highest selected) (`vercel-jev.ts:278-295`). **Stop conditions** include: step limits (30 actions, 60 model calls), action repetition detection (`hasIneffectiveRepetition`, `hasActionCycle` in `history.ts`), excessive scroll-to-now-content (3 empty scrolls), page-state revisits (3 visits to same state), repeated stale targets (3 consecutive `StaleDecisionError`), and tool error cascades. **Memory between steps** is tracked via `TaskMemory` (`agent-runner.ts:28-36`): an action `history` array, `attempts` map (action key → count), `stateVisits` map (progress state → count), and an `actionCache` for CLICK/SELECT replay across subgoals.

> **Editor's note.** Correction: besides the shared task budget of 30 actions and 60 model steps, each `browser_subgoal` runs an `AgentRunner` capped at 8 actions and 16 model calls (`runJevSubgoal` in background.ts), and in subgoal mode the runner returns an unverified checkpoint after a URL change, typing, Escape or Jev's DONE.

Citations: [apps/extension/src/agent/fx-agent.ts:64-341](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L64-L341) · [apps/extension/src/agent/agent-runner.ts:54-280](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/agent-runner.ts#L54-L280) · [apps/extension/src/agent/vercel-jev.ts:89-275](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/vercel-jev.ts#L89-L275) · [apps/extension/src/agent/action-space.ts:1-78](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/action-space.ts#L1-L78)

### How are failures, retries and self-healing handled? (answered)

**Error classes caught:** The runner catches `StaleDecisionError` (element changed, detached, or occluded since snapshot), `TextTargetMismatchError` (wrong field for text), `TextGenerationUnavailableError`, and `VerificationUnavailableError`. The FX agent additionally catches gateway-specific errors: `JevGatewayUnavailableError`, `JevGatewayRejectedError`, `JevGatewayInvalidResponseError`, tab-closed errors, and debugger-conflict errors (`fx-agent.ts:182-203`). **Retries:** The Jev decision engine retries malformed evaluation output once (`vercel-jev.ts:144-151`). Gateway SDK failures (status 5xx, retryable gateway errors) trigger a single HTTP fallback attempt (`vercel-jev.ts:116-133`). The outcome verifier retries once on transient service errors or malformed responses, with increased `maxOutputTokens` on schema failures (`outcome-verifier.ts:82-119`). Text generation repairs malformed or rejected content once via a second model call (`text-generator.ts:114-149`). The observer tries up to 20 times (50ms apart) for a document to be ready (`observer.ts:298`). The FX outer loop has 3 recovery attempts if verification fails (`fx-agent.ts:281`). **Replanning:** Stale targets trigger `feedback` with evidence text guiding the Jev model to re-observe and choose differently (`agent-runner.ts:189-198`). Failed verification sets `excludeDone=true` and passes evidence as feedback for the next decision. Blocked subgoals are counted — after 2 consecutive blocks the FX planner stops (`fx-agent.ts:157-158`). **Action caching** (`action-cache.ts`): Successful CLICK and SELECT actions that produced a page change are cached by semantic key (goal + URL + operation + role + label + ordinal + page precondition state). On fresh runs, `findCachedAction()` replays a matching cached action without a model call (`agent-runner.ts:88,94`). **Timeouts:** Page settling during `waitForChange()` uses a 3-second budget (10s for explicit WAIT) (`agent-runner.ts:215`). Navigation waits up to 30s (`navigation.ts:28-68`). Model connection tests timeout at 10s. The observer's `observe()` retries for up to ~1 second. **Consecutive-error stops:** 3 consecutive stale targets, 3 consecutive ineffective actions (no page change), 3 consecutive tool errors, or 2 consecutive blocked subgoals all abort the run.


Citations: [apps/extension/src/agent/agent-runner.ts:180-200](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/agent-runner.ts#L180-L200) · [apps/extension/src/agent/action-cache.ts:1-68](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/action-cache.ts#L1-L68) · [apps/extension/src/agent/fx-agent.ts:148-165](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L148-L165) · [apps/extension/src/agent/outcome-verifier.ts:82-120](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/outcome-verifier.ts#L82-L120)

### Which models are supported and how are they called? (answered)

**Primary model:** `deepseek/deepseek-v4.1-flash` is set as the `LANGUAGE_MODEL` in `models.ts:1`. This drives both the FX planner/agent (via `libfx` WASM) and serves as the `TEXT_MODEL` for field content generation. **The decision model is entirely different:** action selection uses `typesafe-ai/jev` — a specialized evaluation/judge model accessed through Vercel's AI Gateway via `experimental_evaluate` (the `ai` SDK) with an HTTP fallback (`vercel-jev.ts:92,198,248`). Jev receives structured state + multiple-choice questions and returns a chosen operation with probability-weighted confidence. **Vision:** there is **no screenshot or vision input** anywhere in the codebase. The model receives only structured text: the DOM serialization as `PageSnapshot`, a `PageExtractor` that extracts typed fields from page content, and `ReadPageExpression` for reading offscreen text. **Structured output / tool calling:** the FX agent uses `libfx`'s `createFxAgent()` which provides native tool calling — tools are defined with Zod schemas, converted to JSON Schema via `z.toJSONSchema()`, and executed by the agent runtime (`fx-agent.ts:99-215`). The Jev evaluator uses the Gateway's evaluation API with choice-type questions. Text generation and verification use `ai` SDK's `Output.object()` with Zod schemas for structured JSON responses — the text generator returns `{suitable, reason, text}`, the date generator returns `{suitable, reason, date}`, and the verifier returns `{satisfied, evidence}`. **Small/specialised models:** Jev (`typesafe-ai/jev`) is itself a specialised model for evaluation/choice rather than free-form generation. The text generator (`TEXT_MODEL`) reuses the same model as the planner. **Providers:** all model calls go through the **Vercel AI Gateway** (`apiKey` stored in Chrome storage, accessed via `createGateway()` and `createFxGatewayFetch()`). The gateway applies a configured `zeroDataRetention: false` policy (`gateway-policy.ts:8`). **SDK:** Uses the `ai` SDK (version `^7.0.116`) for streaming text generation, structured output, and evaluation. **Model pricing** for cost tracking is loaded from `https://ai-gateway.vercel.sh/v1/models` in `model-pricing.ts` and reported via `ModelUsageLedger`.


Citations: [apps/extension/src/agent/models.ts:1-6](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/models.ts#L1-L6) · [apps/extension/src/agent/vercel-jev.ts:143-250](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/vercel-jev.ts#L143-L250) · [apps/extension/src/agent/fx-agent.ts:217-276](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L217-L276) · [apps/extension/src/agent/structured-generation.ts:1-47](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/structured-generation.ts#L1-L47) · [apps/extension/src/agent/gateway-policy.ts:1-37](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/gateway-policy.ts#L1-L37)

### How are browser sessions, profiles, auth and anti-bot handled? (answered)

**Local browser only:** Ulka operates as a Chrome extension (MV3) and controls the user's **local Chrome** browser — it attaches to existing tabs via CDP, not remote/cloud browsers. The `AttachedBrowser` class (`browser.ts:7-22`) wraps a CDP debugger session: `attach()` sends `api.attach(target, "1.3")`, `detach()` releases it, and `switchTo()` detaches, updates the tab ID, and re-attaches. **TaskBrowserSession** (`browser.ts:24-47`) manages reuse of a single `AttachedBrowser` instance across a task, ensuring only one CDP connection at a time. The extension stores API keys in Chrome's `storage.local`. **Persistent profiles and cookies:** there is **no auth/profile/cookie management** in the codebase — no cookie injection, no saved session restoration, and no login profile handling. The extension uses whichever Chrome profile the user is already signed into. **Stealth and anti-bot:** there is **no stealth** feature — no user-agent manipulation, no request header stripping, no fingerprint masking, and no anti-detection measures. The CDP attachment is standard Chrome Debugger Protocol without modifications. **Proxies:** there is no proxy configuration. **CAPTCHA handling:** there is no CAPTCHA detection or solving mechanism. When a CAPTCHA appears, the system has no special handling — it would just observe whatever interactive controls the page exposes. The `BLOCKED` operation (`snapshot.ts:8`, `agent-runner.ts:118`) is a catch-all escape for when the model decides it cannot progress, but nothing specific to CAPTCHAs exists. **Background mode:** the extension supports a `background` mode (`storage.local ulkaBackground`) where tab switches happen without activating the tab, allowing the user to continue browsing while the agent works (`background.ts:98,105`). **Navigation guard:** `UrlProvenance` (`url-provenance.ts` traced in `fx-agent.ts` as `urls.assertObserved()`) ensures the agent only navigates to URLs that were either supplied by the user or observed in a previous browser result — preventing navigation to unverified or hallucinated URLs. `validateNavigationUrl()` (`navigation.ts:5-13`) restricts navigation to `https://` URLs or internal `chrome://` pages (extensions, history, downloads, bookmarks, settings, newtab, version only). Internal `chrome://` pages return a stub `PageSnapshot` since their content is unavailable via CDP (`navigation.ts:17-26`).


Citations: [apps/extension/src/agent/browser.ts:7-47](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/browser.ts#L7-L47) · [apps/extension/src/agent/navigation.ts:5-26](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/navigation.ts#L5-L26) · [apps/extension/src/agent/fx-agent.ts:249-249](https://github.com/razaanstha/ulka/blob/bb04433dcebb75b1a93eb5a6dc7d25909930cc65/apps/extension/src/agent/fx-agent.ts#L249-L249)
