# awlevin/typesafe-computer-use

> macOS computer-use loop in which a TypeSafe classifier picks each action from OCR and accessibility state; an LLM writes only free text.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/awlevin/typesafe-computer-use (reviewed at commit `44ca11f0935b021b73020825da054b5c92cc1288`, 2026-09-29)
- Stars: 1174 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/typesafe-computer-use/

## Overview

typesafe-computer-use, called "jev" in its own code, is a computer-use agent that drives a real Mac toward a goal typed in plain English. Its one idea is to take the frontier model out of the per-step loop. Each step reads the screen deterministically (Vision OCR plus the accessibility tree), code adds the facts a model would otherwise have to work out, and a small hosted classifier, TypeSafe, picks the next action from a fixed menu. An LLM, called the "writer", runs only when a text field needs free text, when a URL outside a small catalog is needed, or when the classifier stops and someone has to read the screen and say what happened.

The project's own measurement puts a decision at about $0.0002 and 0.13 to 0.38 s of model latency, against roughly $0.03 and 5 s for a frontier model reading a screenshot. The README is frank about the price of that: everything the big model would do for free, such as comparing event dates, has to be rebuilt as deterministic state in code.

It is not a browser library in the Playwright sense. The main path (`clicker`) clicks real pixels and presses real keys through Quartz, and it reaches websites by bringing Chrome to the front with AppleScript. There is a separate, opt-in browser backend that reads the DOM over raw CDP, and an OSWorld adapter that runs the same loop against an Ubuntu VM for benchmarking. Windows support is marked experimental and untested.

## Architecture

```mermaid
flowchart LR
  CLI["clicker CLI"] --> RUN["runner.run step loop"]
  RUN --> CAP["perception.capture"]
  CAP --> AD["platform_adapter.desktop"]
  AD --> MAC["macos.py"]
  AD --> WIN["windows.py"]
  AD --> OSW["osworld/desktop.py"]
  RUN --> PER["perception.perceive: OCR + AX merge"]
  PER --> DEC["decide.decide"]
  DEC --> TS["TypeSafe system_one"]
  RUN --> ACT["actions.perform"]
  ACT --> AD
  ACT --> WR["writer: text and URLs"]
  RUN --> HO["hand_off: answer model"]
  HO --> WR
  BENCH["clicker-bench"] --> BR["browser/ runner over CDP"]
  BR --> TS
```

| Component | Path | Role |
|---|---|---|
| CLI | `typesafe_computer_use/cli.py` | `clicker` (run or dry run) and `clicker-inspect` (capture and dump the exact payload) |
| Step loop | `typesafe_computer_use/runner.py` | Capture, perceive, decide, act; stop rules; hand-off to the writer; run folder |
| Perception | `typesafe_computer_use/perception.py` | Screenshot, cropped OCR with a tile cache, AX items, merge into one numbered item list |
| AX walk | `typesafe_computer_use/ax_walk.py` | Bounded, platform-free accessibility-tree walk with pruning rules |
| Decision | `typesafe_computer_use/decide.py` | State dict, criteria, one multi-`Choice` TypeSafe request, typed-text check |
| Actions | `typesafe_computer_use/actions.py` | One handler per action kind, each returning a history line |
| Writer | `typesafe_computer_use/writer.py`, `openai_writer.py` | Structured free-text calls: field text, URLs, the final answer with focus or question |
| Platform adapters | `platform_adapter.py`, `macos.py`, `windows.py` | The `Desktop` protocol and its macOS and Windows implementations |
| Browser backend | `typesafe_computer_use/browser/` | Opt-in DOM perception and real CDP input events on a throwaway Chrome |
| OSWorld agent | `typesafe_computer_use/osworld/` | `JevAgent` and a `Desktop` that turns inputs into pyautogui code for the VM |
| Tests | `tests/` | Pure-logic tests plus a simulated computer (`world.py`) driven by the real loop |

## How a request flows

Take `uv run clicker "open the Playground" --act`:

1. **Start.** `main` loads `.env`, refuses to run without `TYPESAFE_API_KEY`, refuses `--act` without Accessibility permission, builds the writer, and calls `run` with a `RunConfig` ([cli.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/cli.py#L40-L104)). Without `--act` it is a one-step dry run.
2. **Loop.** `run` opens a `TypeSafeClient`, wraps both model clients in call meters, and calls `run_step` up to `--steps` times (default 100). When a step stops, `hand_off` decides whether the run ends or resumes ([runner.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L99-L145)).
3. **Capture.** `capture` takes a screenshot (`screencapture` on macOS) and asks the adapter for the frontmost app and pid, window bounds, focused field and active tab URL ([perception.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L37-L70)).
4. **Perceive.** `perceive` runs OCR on the frontmost window's crop, walks the accessibility tree, drops AX nodes that are not actually drawn, and merges the two sources into one numbered list. It also fills side tables: AX refs for pressing, items covered by a popup, and labelled controls that exist but are off screen ([perception.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L84-L123)).
5. **Stall check.** `screen_moved` compares a screen signature with the last one. Three actions in a row that changed nothing end the step as `stalled` ([runner.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L372-L391)).
6. **Decide.** `decide` sends one `system_one` request with up to four `Choice` questions: `kind`, `item`, `site` and, when there are hidden controls, `offscreen`. The state carries the goal, the clock, app, URL, focused field, the last 8 actions, the actions already tried on this screen, and every item with its region, date hint, row-mates and popup ([decide.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L127-L162), [L205-L256](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L205-L256)).
7. **Gate.** `resolve` stops on `done` or `none`, or when confidence is under `--min-confidence` (0.4). For a click, confidence is the minimum of the kind and item answers ([runner.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L340-L369), [decide.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L188-L198)).
8. **Act.** `perform` routes the chosen key to `click_item`, `press_offscreen` or a handler table ([actions.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L34-L49)). Every action returns a history line, and `repeating` stops the run after two actions in a row that were already tried on the same screen ([runner.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L399-L417)).
9. **Hand off.** On any stop, `hand_off` asks the answer model to read the screenshot and screen text. It returns `achieved`, `answer`, and optionally a one-move `focus` (the classifier resumes with it) or a `question` for the user. The default limits are 10 hand-offs and 3 questions. A third stall with no new page is `stuck` and final ([runner.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L157-L238)).

## Key components

### Perception: OCR plus accessibility

OCR reads only the frontmost window plus the menu-bar strip above it. A tile cache compares each capture with the previous one at 1/8 scale in 256 px tiles and re-reads only the changed areas, up to four rectangles. The accessibility walk is capped at 4000 nodes or 0.6 s, and it prunes nodes whose frames lie: off-display subtrees, slivers under 4 pt, closed `AXMenu` trees and nameless groups ([ax_walk.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/ax_walk.py#L54-L59)). `merge_with_origins` pairs each AX control with an overlapping OCR block of matching text. It drops symbol-only OCR blocks that sit on a button (an icon read as `←`), cuts the list to 255 items, the TypeSafe `Choice` ceiling, and renumbers it in reading order ([perception.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L636-L662)).

### Decision: facts in code, choice in the classifier

The design rule, stated in CONTRIBUTING, is "the classifier picks; code decides facts". So the state carries computed facts. `dates.py` adds "dated 2026-10-13 (in 27 days)" to any block with a date. `row_mates` tells three "Buy" buttons apart by the text in their rows ([decide.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L72-L88)). The runner lists the actions that already led back to this screen. The action menu is kept mutually exclusive on purpose, because the author found that every stall came from two options meaning the same thing ([decide.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L29-L59)). There is no "click the address bar". A website is reached only through `use_browser` plus the `site` question, with an eight-site catalog and `other` for the writer ([config.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/config.py#L9-L32)).

### Actions and input

`click_item` first tries `AXPress` on the element itself, which works under sticky headers. It falls back to a pixel click, closing a covering popup first with its close button or Escape, never with its other buttons ([actions.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L52-L70)). Text goes in through `AXValue` with a read-back, then keystrokes after a clear ([actions.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L105-L123)). Unsubmitted text is checked by a second TypeSafe `Noul` question. Below 0.5 the original value is restored through the same element ([actions.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L172-L199)). On macOS, input is raw Quartz events, and every click and keystroke first checks the abort corner ([macos.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/macos.py#L37-L80)).

### The writer

All free text goes through `_structured`, which sends a JSON Schema as `output_config` and, for non-Anthropic endpoints, also writes the schema into the system prompt and turns thinking off ([writer.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L74-L112)). Defaults are `claude-haiku-4-5` for field text and URLs and `claude-sonnet-5` for the answer model. `CLICKER_WRITER_BASE_URL` with `CLICKER_WRITER_API=openai` points it at any OpenAI-compatible server ([writer.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L41-L63)). Proposed URLs must be `https` with no whitespace ([writer.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L302-L315)). The answer prompt forbids answers from memory and limits a focus to one move the classifier can actually make ([writer.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L332-L369)).

### Platform adapters

`Desktop` is a `Protocol` with two dozen methods: abort, input, apps and windows, capture/OCR/AX, and element actions ([platform_adapter.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/platform_adapter.py#L36-L71)). `pick_host` loads `windows.py` on Windows and `macos.py` elsewhere. Where that fails, as on Linux, it uses a `NoDesktop` that raises on any call ([platform_adapter.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/platform_adapter.py#L93-L106)). `using(adapter)` swaps in another computer for a block, which is how the OSWorld agent drives a VM without changing the loop ([platform_adapter.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/platform_adapter.py#L143-L161)).

### Browser backend and OSWorld

`browser/` is a second, independent loop. It starts a dedicated Chrome with a temporary profile, loopback CDP and one allowed origin ([cdp.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/cdp.py#L94-L125)), and reads elements with one `Runtime.evaluate`. It clicks with `Input.dispatchMouseEvent` at the element centre ([act.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/act.py#L43-L64)). Its request asks `kind`, `element` and a `satisfied` `Noul` ([browser/decide.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/decide.py#L148-L197)). It is reached only through `clicker-bench`, not `clicker`. `JevAgent` runs the unchanged `runner.run` on a worker thread inside `using(OSWorldDesktop(...))`. Each step's inputs are collected as pyautogui code, and the run always reports `DONE`, never `FAIL` ([osworld/agent.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/osworld/agent.py#L1-L19), [desktop.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/osworld/desktop.py#L162-L178)).

## Extending it

- **A new action.** Add it to `fixed_actions` and to `_HANDLERS` together ([actions.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L224-L234)), and keep it exclusive of the existing ones. The contributing guide asks for a replay note showing the decision before and after.
- **A new site.** Add a key to `SITES`. Anything not in the catalog is proposed by the writer.
- **A new platform.** Implement every `Desktop` method (a test holds adapters to the protocol, parameter for parameter) and add a branch to `pick_host`. The AX walk takes children, attributes and actions as callables, so only those bindings change.
- **Another writer model.** Set `CLICKER_WRITER_*` and `CLICKER_ANSWER_*`. `CLICKER_WRITER_VISION=false` sends a text-only answer model no screenshot.
- **A task the loop cannot do.** Write it as an `xfail` scenario in `tests/test_scenarios.py` on the simulated computer in `tests/world.py`, which runs the real `runner.run` with a policy in place of the classifier.

## Running it

- **Install.** macOS 14+, Python 3.12+, `uv sync`, and a `.env` with `TYPESAFE_API_KEY` (required) and `ANTHROPIC_API_KEY` (needed for typing, writer URLs and the final answer). The terminal needs Screen Recording and Accessibility permissions.
- **Run.** `uv run clicker "goal"` is a dry run of one step. `--act` drives the machine, with `--steps`, `--delay`, `--min-confidence` and `--handoffs`. Ctrl-C or the top-left screen corner aborts.
- **Inspect and replay.** `clicker-inspect` captures after a countdown and opens the annotated screen and exact payload ([cli.py](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/cli.py#L106-L142)). Every run writes `runs/<timestamp>/` with raw captures, payloads and every probability, and `clicker --image <capture> --app ... --url ...` re-decides a step offline.
- **Benchmark.** `scripts/osworld setup` and `run-jev chrome/<task> --ocr rapidocr` need a Linux host with KVM. The committed record holds 24 scored jev task runs, 19 of them passing, mostly Chrome tasks, with the 52 proxy-dependent tasks deferred.

## Strengths and caveats

- **Strength: cost and latency.** One small classifier request per step, with free output tokens and a calibrated confidence to gate on, is a different cost class from screenshot-per-step agents.
- **Strength: debuggability.** Every step saves the capture, the exact payload and the full probability distributions, and a step can be replayed offline. This is better than most agents in this category.
- **Strength: disciplined boundaries.** Facts are computed in code, free text comes only from the writer with a structured reply and a credential guard, and platform calls stay behind one protocol.
- **Caveat: it drives your real machine.** The main path takes over the mouse and keyboard of the logged-in Mac. Using the computer during a run fights it for focus, and only the main display is captured.
- **Caveat: perception limits.** OCR sees only text, and the AX tree only helps in apps that publish one. The README cites Spotify at 0% labelled controls. Identical labels without distinguishing rows split the vote.
- **Caveat: hard-coded reasoning.** Every comparison the classifier cannot make (dates, rows, popups) must be rebuilt in code. New task types tend to need new state, not just a better prompt.
- **Caveat: vendor dependency.** Every decision goes to the hosted TypeSafe API. There is no local or alternative classifier.
- **Caveat: weak browser story for this category.** There are no stealth, proxy, CAPTCHA or session controls. The DOM backend is a benchmark tool, not the default path.

*Sources: code at 44ca11f, deepwiki-open wiki (14 pages), OpenDeepWiki wiki (11 pages), verified Q&A.*

## How awlevin/typesafe-computer-use answers the Browser & computer control questions

### How is the page represented to the model? (answered)

The page is represented to the model as a numbered list of **Items**, each carrying text, screen coordinates, and source (OCR, accessibility tree, or both). Perception is a two-source pipeline:

**OCR path:** `perception.py:capture()` takes a screenshot via `screencapture`, crops to the frontmost window plus an 8pt margin and the menu bar strip, then runs macOS Vision OCR (`ocrmac`'s `recognize_text`). A change-detection system compares the current capture to the previous one at 1/8 scale in 256px tiles; unchanged tiles reuse their OCR lines. Changed tiles are clustered into contiguous blobs, each re-read as a single OCR crop — up to 4 rectangles max, preferring a full re-read above 60% change. Lines below MIN_OCR_CONFIDENCE (0.3) or matching the goal echo are filtered, then merged into blocks by vertical proximity and alignment. (`perception.py:84–123, 258–291, 300–319`)

**Accessibility tree:** `ax_walk.py:walk_actionable()` does a bounded BFS of the frontmost app's AX tree (cap 4000 nodes, 0.6s), collecting labelled on-screen controls. Nodes with real frames wholly off the display are pruned to `screen.offscreen` (cap 120) — reachable via AXPress, never as click targets. The walk prunes `AXMenu` subtrees, nameless `AXGroup` layout boxes, and sub-4pt slivers. (`ax_walk.py:134–212`)

**Merge and budget:** OCR blocks and AX controls are merged by box-overlap (≥50% of the smaller box) and text-match (containment or ≥50% shared words). An OCR block centered on a button/link with no alphanumerics is treated as that control's icon and dropped. The merged list is capped at 255 items (TypeSafe's Choice ceiling) — AX-sourced items are kept first, then OCR items by confidence. (`perception.py:636–701, config.py:10`)

**Facts code adds:** date-distance hints on blocks containing dates (via `dates.py`), row-mates for duplicated labels, popup coverage markers, the action already tried on this screen, the focused field, the URL, and the app name. All is rendered into `decide.py:base_state()` for the classifier. (`decide.py:128–162, 91–105`)

The **browser backend** replaces all of the above with a single CDP `Runtime.evaluate` call — a JS script collects interactive DOM elements, assigns them `data-tscu` attribute indices, checks occlusion via `elementFromPoint`, and returns an ordered list with exact text, role, viewport coordinates, and visible page text blocks. No OCR, no accessibility tree, no screenshot permission needed. (`browser/perceive.py:29–168, docs/browser-backend.md:1–27`)


Citations: [typesafe_computer_use/perception.py:84-123](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L84-L123) · [typesafe_computer_use/perception.py:258-291](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L258-L291) · [typesafe_computer_use/ax_walk.py:134-212](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/ax_walk.py#L134-L212) · [typesafe_computer_use/decide.py:128-162](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L128-L162) · [typesafe_computer_use/perception.py:636-662](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L636-L662) · [typesafe_computer_use/browser/perceive.py:29-168](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/perceive.py#L29-L168)

### How are actions executed and how are elements targeted? (answered)

Actions are dispatched through `actions.py:perform()`, which routes by the classifier's `kind` Choice into deterministic handlers. Elements are targeted by **index** into the item list, not by CSS selectors or raw coordinates — the index maps back to the item's pixel center or accessibility element ref.

**MacOS native:** `macos.py` posts real Quartz events through `CGEventPost` to `kCGHIDEventTap`. A click (`click_at`) moves the pointer then sends mouse down/up events. Keys (`press`) use `CGEventCreateKeyboardEvent`, supporting Command-prefixed keys via `kCGEventFlagMaskCommand`. Typing emits one unicode keystroke per character. Scrolling parks the pointer at the window center first. (`macos.py:74–120`)

**Accessibility-first clicks:** When an item came from the AX tree, `click_item` first tries `ax_press(ref)` — this sends AXPress to the control element itself, landing even if covered by a sticky header. If that fails or there's no AX ref, it falls back to a pixel-coordinate mouse click. If the item sits under a popup (tracked in `screen.covered`), the popup's close button (or Escape) is pressed first. (`actions.py:52–70, 73–86`)

**Text input:** `fill_field` first tries `ax_set_value` on the focused element — one call, no keystrokes. It verifies by reading the value back: if it ends with the typed text, done. Otherwise it `clear_field` (Cmd-A, Delete) then types character by character. After typing, a TypeSafe Noul check (`verify_typed`) scores whether the field holds a sensible value; below 0.5, the original value is restored via AX (`restore_field`). (`actions.py:105–136`)

**Browser (CDP) path:** The browser backend uses real `Input.dispatchMouseEvent` / `Input.insertText` CDP events. Elements are targeted by their `data-tscu="<index>"` attribute, set by the perception JS. A click uses scrollIntoView first if off-screen, then dispatches mouse move/press/release. Typing focuses the element, then calls Input.insertText. (`browser/act.py:43–86`)

**Other actions:** `use_browser` activates the named browser via AppleScript (`tell app to activate`), then opens a URL from the SITES catalog or from the writer. `scroll_down`/`scroll_up` set lines via Quartz scroll wheel events, or `pyautogui.scroll` in OSWorld. `go_back` sends Cmd-[ (or Alt-Left on Windows). `wait` sleeps for 3 seconds, plus the step's own settle delay. (`actions.py:138–234`)

The **OSWorld adapter** translates all inputs into `pyautogui` code strings, accumulated and sent as one step's action list. Clicks become `pyautogui.click(x, y)`, keys become `pyautogui.press(...)`. The text you type reaches the VM only through `repr()`, never as raw code. (`osworld/desktop.py:160–180`)


Citations: [typesafe_computer_use/actions.py:34-49](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L34-L49) · [typesafe_computer_use/macos.py:74-81](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/macos.py#L74-L81) · [typesafe_computer_use/actions.py:52-70](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L52-L70) · [typesafe_computer_use/actions.py:105-136](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L105-L136) · [typesafe_computer_use/browser/act.py:43-64](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/act.py#L43-L64) · [typesafe_computer_use/osworld/desktop.py:160-180](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/osworld/desktop.py#L160-L180)

### How is the agent loop / planning implemented? (answered)

The agent loop lives in `runner.py:run()`. It is a synchronous step loop: one capture + percept + decide + act cycle per iteration, up to `--steps` (default 100).

**Step structure:** `run_step` captures the screen (screenshot, app, window, field, URL), runs `perceive()` to produce the item list, constructs the state dict with goal, date context, previous actions, tried actions, and screen items, then calls `decide()`. That is a **single TypeSafe `system_one` request** with up to four Choice questions: `kind` (which action), `item` (which on-screen item — only for click_item), `site` (which website — only for use_browser), and `offscreen` (which hidden control — only for press_offscreen). The responses carry full probability distributions and calibrated confidence scores. (`runner.py:284–337`, `decide.py:204–256`)

**Stop conditions** are checked in `resolve()` after the decision: (1) `done` or `none` — goal achieved or impossible; (2) confidence below `--min-confidence` (0.4); (3) dry run (`--act` not passed). Then comes the deterministic action via `perform()`. (`runner.py:340–369`)

**Stall detection** uses a screen signature — the tuple of (app, URL, focused field, (text, row) pairs in reading order). Two captures differing by ≤1 line and ≤1-in-10 lines count as the same screen. Three consecutive idle actions (screen unchanged) or two consecutive repeated actions (same action on same screen) stop the run. Three stalls without reaching a new page produce `stuck`, which is final. (`runner.py:372–417`, `models.py:91–117`)

**Writer handoff:** When the classifier stops, `hand_off()` calls the writer model (the stronger `CLICKER_ANSWER_MODEL`, default `claude-sonnet-5`) with the screenshot, all screen text, earlier screens' text, the stop reason, and earlier stops with their focuses. The writer returns `{achieved, answer, focus, question}`. If the goal isn't reached, it can set a **focus** (one concrete next action for the classifier, e.g. "Click the 'Sort by price' button") and the loop resumes with that focus in state. Or it can ask a **question** for the user. Up to `--handoffs` 10 handoffs and 3 questions per run. A focus that produces no action leaves the prior answer standing. (`runner.py:157–220`, `writer.py:332–423`)

**Memory between steps:** The classifier receives the last 8 actions as `previous_actions`, the actions already tried on the current screen as `already_tried_on_this_screen`, and `Guidance` (current focus + user replies from the handoff). The writer gets the same plus earlier stops and earlier screens' text. The OCR cache carries one step's raw lines into the next via `OcrCache`. (`decide.py:128–162`, `runner.py:87–96`)


Citations: [typesafe_computer_use/runner.py:99-145](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L99-L145) · [typesafe_computer_use/runner.py:284-337](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L284-L337) · [typesafe_computer_use/decide.py:204-256](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L204-L256) · [typesafe_computer_use/runner.py:157-220](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L157-L220) · [typesafe_computer_use/runner.py:372-417](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L372-L417) · [typesafe_computer_use/models.py:91-117](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/models.py#L91-L117)

### How are failures, retries and self-healing handled? (answered)

Jev uses deterministic code-side checks and escalation rules rather than action-level retries.

**Stall detection (the primary reliability mechanism):** Every step computes a `signature` — (app, URL, focused field, (text, row) tuples). `same_screen()` allows at most 1 differing text line and that line must be ≤1-in-10, so a clock tick or OCR jitter doesn't mask a stall. After **3 idle actions** (screen unchanged — a refused click, a wait, a scroll at the bottom) or **2 repeated actions** (same action on same screen — a cycle) the run stops with outcome `stalled`. After 3 stalls with no new page between them, the run escalates to `stuck`, which the writer cannot hand back from. (`runner.py:24–27, 372–391, 399–417, models.py:91–117`)

**Error classes caught:** `Missed` is raised when a click's pointer never reached the target — the action line says "click refused: … was not clicked". `Abort` (mouse in top-left corner or Ctrl-C) stops the loop, as does `KeyboardInterrupt`. The `osworld` adapter additionally catches `BaseException` in its worker thread and hands it back to `predict`. (`models.py:36–41`, `macos.py:37–40`, `osworld/agent.py:195–221`)

**Popup auto-close:** When an item sits under a popup (tracked by `ax_walk` as the covering window), `click_item` closes the popup first — using its close button when available, otherwise Escape — then clicks the item. Chrome's "Restore pages?" bubble is identified by name and never closed with its Restore button. Popup tracking only works via the OSWorld accessibility tree (the macOS tree doesn't report covering windows). (`perception.py:114–119`, `actions.py:52–86`)

**Text verification and recovery:** After typing into a field, `verify_typed` runs a TypeSafe Noul asking "did the typing succeed, and is the value sensible?". Below 0.5, `restore_field` sets the original value back through the same AX element, but only if it still holds exactly what was just typed — no keystroke fallback, since focus may have moved. (`actions.py:191–199, 126–136`, `decide.py:259–276`)

**No retries, no caching of successful actions:** There is no retry-on-failure — a refused action is recorded in history and counts toward the idle limit. The OCR cache (`OcrCache`) caches raw OCR lines across unchanged tiles, not action outcomes. The `already_tried_on_this_screen` list prevents re-picking the same fruitless action. The critic is the next capture: nothing in the action description says what came of it, so every action is verified by the screen it produced. (`runner.py:393–396`, `perception.py:164–227`)


Citations: [typesafe_computer_use/runner.py:24-27](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L24-L27) · [typesafe_computer_use/runner.py:372-417](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/runner.py#L372-L417) · [typesafe_computer_use/models.py:91-117](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/models.py#L91-L117) · [typesafe_computer_use/actions.py:126-136](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/actions.py#L126-L136) · [typesafe_computer_use/decide.py:259-276](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L259-L276) · [typesafe_computer_use/perception.py:164-227](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/perception.py#L164-L227)

### Which models are supported and how are they called? (answered)

Jev uses two distinct model tiers for two distinct roles:

**Classifier (every step, 86%+ of model calls):** The TypeSafe SDK (`typesafe.ai`) answers a `Choice` over up to 255 options with a full probability distribution and calibrated confidence. It is a small, fast model: 130–380ms latency, ~$0.0002 per decision, no vision. The classifier never reads pixels or generates text — it receives the structured state dict (goal, items, previous actions, computed facts) and picks one option per Choice. Up to four Choices are answered in a single `system_one` request. (`config.py:10`, `decide.py:204–256`, `calls.py:87–107`)

**Writer (when the classifier needs free text or stops):** By default Claude via the Anthropic Messages API. The per-step writer model (`CLICKER_WRITER_MODEL`, default `claude-haiku-4-5`) generates field text and proposes URLs. The answer model (`CLICKER_ANSWER_MODEL`, default `claude-sonnet-5`) reads the screenshot whenever the classifier stops — a stronger reader for judging goal completion. (`config.py:16–17, 51–53, 99–100`, `writer.py:332–423`)

**Custom endpoints:** `CLICKER_WRITER_BASE_URL` can point at any Anthropic Messages API or (with `CLICKER_WRITER_API=openai`) OpenAI Chat Completions endpoint — LM Studio, Ollama, vLLM, DeepSeek, a LiteLLM proxy. The `OpenAIWriter` adapter steps down response_format (json_schema → json_object → none) and token_limit (max_tokens → max_completion_tokens) when endpoints refuse parameters. Thinking is explicitly disabled for non-Anthropic endpoints. Vision can be toggled off (`CLICKER_WRITER_VISION=false`) for text-only answer models. (`config.py:55–96`, `openai_writer.py:1–70`, `docs/writer-endpoints.md:1–52`)

**Key separation:** `CLICKER_WRITER_API_KEY` only goes to `CLICKER_WRITER_BASE_URL`. `ANTHROPIC_API_KEY` only goes to Anthropic's API. On the Anthropic API the key is sent in both `x-api-key` and `Authorization` for proxy compatibility. (`writer.py:48–63`)

**Structured output:** The writer is called with Anthropic's `output_config/json_schema` or OpenAI's `response_format: json_schema` — every call returns a validated, typed dict. Non-Anthropic endpoints also get the schema spelled out in the prompt as fallback. (`writer.py:74–113`, `openai_writer.py:46–69`)

The intent is that the classifier takes nearly every model call, and the writer only runs when free text is needed or the classifier stops. The `calls:` log line tracks the split. (`runner.py:143`, `calls.py:1–84`)


Citations: [typesafe_computer_use/config.py:14-20](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/config.py#L14-L20) · [typesafe_computer_use/writer.py:37-63](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L37-L63) · [typesafe_computer_use/writer.py:74-113](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/writer.py#L74-L113) · [typesafe_computer_use/openai_writer.py:33-70](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/openai_writer.py#L33-L70) · [typesafe_computer_use/decide.py:204-256](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/decide.py#L204-L256) · [docs/writer-endpoints.md:1-52](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/docs/writer-endpoints.md#L1-L52)

### How are browser sessions, profiles, auth and anti-bot handled? (answered)

The desktop adapter (macOS) drives the **user's real machine** — it activates apps via AppleScript, captures the screen, and sends synthetic input events through Quartz. There is no isolated browser session: the classifier decides to `use_browser`, which brings `CLICKER_BROWSER` (default `Google Chrome`) to the front via AppleScript `activate`, then optionally opens a URL via `open location`. The browser's own cookies, profiles, and extensions are untouched because the loop uses AppleScript's own activation — it never launches a separate Chrome instance. (`macos.py:146–172`, `macos.py:126–128`, `cli.py:30–38`)

**Browser backend (CDP):** `cdp.py:Chrome` starts a **dedicated, ephemeral Chrome instance** with a fresh temporary profile (`tempfile.mkdtemp(prefix="tscu-chrome-")`). The debugging websocket listens on loopback only (`127.0.0.1`) and restricts origins to itself. The profile is removed on `close()`. This never touches the user's profile. Headless mode is available (`--headless=new`) in addition to headed. (`browser/cdp.py:80–87, 94–125, 148–159`, `docs/browser-backend.md:3`)

**OSWorld VM:** JevAgent runs inside an Ubuntu VM via OSWorld's runner. The `OSWorldDesktop` adapter communicates through pyautogui code strings. The browser is started per task by OSWorld — `activate` calls `wmctrl -xa google-chrome`, starting the browser with `subprocess.Popen` if it isn't running. Each task gets a fresh VM snapshot via OSWorld's `reset()`, so no session persists. (`osworld/desktop.py:196–201, 256–281`, `osworld/agent.py:98–106`)

**Windows adapter:** Uses UI Automation's Windows API access — session follows whatever the current Windows desktop has. The browser URL is read from the address bar via UI Automation (Chrome/Edge show it without the scheme). (`windows.py:71–73, 1–10`, `docs/windows.md:1–13`)

**Anti-bot / stealth / CAPTCHA:** None of these are implemented. The README explicitly says passwords are never typed — the user should rely on the browser's password manager or an SSO button the OCR can read. There are no stealth measures, no proxy configuration, no CAPTCHA handling, no user-agent rotation, and no cookie persistence across runs. The CDP chrome starts with `--disable-extensions`, `--disable-sync`, `--disable-background-networking` — these are for determinism in benchmarking, not stealth. (`docs/how-a-step-works.md:172`, `browser/cdp.py:104–112`)


Citations: [typesafe_computer_use/macos.py:146-172](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/macos.py#L146-L172) · [typesafe_computer_use/browser/cdp.py:80-160](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/browser/cdp.py#L80-L160) · [typesafe_computer_use/osworld/desktop.py:196-281](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/typesafe_computer_use/osworld/desktop.py#L196-L281) · [docs/browser-backend.md:1-8](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/docs/browser-backend.md#L1-L8) · [docs/how-a-step-works.md:168-176](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/docs/how-a-step-works.md#L168-L176) · [docs/windows.md:1-22](https://github.com/awlevin/typesafe-computer-use/blob/44ca11f0935b021b73020825da054b5c92cc1288/docs/windows.md#L1-L22)
