LLMs Technical Reviews

lahfir/agent-desktop

Rust CLI and C library that lets agents read and operate macOS apps through accessibility trees, with snapshot refs and verified actions.

GitHub ↗★ 1.8kRustApache-2.0commit 9d7ba42 · 2026-10-04

Overview

agent-desktop is a computer-use tool for native desktop apps, not for web pages. It ships as a single Rust binary (agent-desktop) and as a C-ABI library (libagent_desktop_ffi). An agent runs agent-desktop snapshot --app Finder -i, gets back a JSON accessibility tree with refs such as @s8f3k2p9:e14, and then runs agent-desktop click @s8f3k2p9:e14. Each call is one process that prints one JSON envelope and exits.

The project states plainly that it is not an agent. It has no LLM client, no planner and no memory of its own. The observe-decide-act loop belongs to the caller: Claude Code, a custom harness, or a script. What it does provide is a strict contract around each step. Refs are bound to a process instance and a saved snapshot. Every action re-resolves its target, checks that it is actionable, and reports a delivery disposition that tells the caller whether a retry is safe.

At the pinned commit the only real platform is macOS. It reads the tree through the AX API (AXUIElement) and acts through AX actions first, with synthetic CGEvent input as an opt-in fallback. The windows and linux crates are empty stubs that compile, implement the traits with default methods, and fail closed with PLATFORM_NOT_SUPPORTED. The one browser link is launch --cdp. It starts a Chromium or Electron app with a verified DevTools port so that Playwright or another CDP client can drive the web contents, while agent-desktop keeps the native menus and dialogs.

Architecture

flowchart LR
  AG["Calling agent"] --> CLI["agent-desktop CLI"]
  AG --> FFI["C ABI (libagent_desktop_ffi)"]
  CLI --> POL["command_policy preflight"]
  POL --> DSP["dispatch()"]
  DSP --> CMD["core commands/*"]
  FFI --> CMD
  CMD --> SNAP["snapshot + ref_alloc"]
  CMD --> RA["ref_action: resolve, preflight, verify"]
  SNAP --> STORE["RefStore (refmap.json)"]
  RA --> STORE
  SNAP --> PA["PlatformAdapter trait"]
  RA --> PA
  PA --> MAC["macos crate: AX + CGEvent"]
  PA --> STUB["windows / linux stubs"]
  CMD --> CDP["launch --cdp: DevTools probe"]
Component Path Role
Binary entry src/main.rs Parses the CLI, builds the context, picks the adapter at compile time, prints the JSON envelope
Dispatcher src/dispatch/mod.rs One match arm per command; wraps mutating commands in a scope and the cursor overlay
Command policy src/command_policy/mod.rs Maps each command to the permissions it needs (Accessibility, Screen Recording) and validates ref arguments
Batch src/batch/ Decodes a JSON array into the same Commands enum and runs it through the same dispatch path
Core commands crates/core/src/commands/ One file per command (click.rs, snapshot.rs, launch.rs, …)
Adapter traits crates/core/src/adapter/ ObservationOps, ActionOps, InputOps, SystemOps, combined as PlatformAdapter
Snapshot and refs crates/core/src/snapshot.rs, ref_alloc.rs, refs_store.rs Builds the tree, allocates refs, saves the refmap per snapshot id
Ref actions crates/core/src/ref_action*.rs, post_action.rs Auto-wait, strict resolution, actionability, dispatch, state readback
macOS adapter crates/macos/src/ AX tree reading, strict re-identification, action chains, CGEvent input, clipboard, notifications
FFI crates/ffi/ cdylib with a cbindgen header (include/agent_desktop.h) for Python, Swift, Go, Node, C++
Skills skills/agent-desktop/ Agent-facing docs, compiled into the binary and served by agent-desktop skills

How a request flows

Take agent-desktop click @s8f3k2p9:e14 on macOS:

  1. Parse and build context. run() parses the arguments with clap; a parse error becomes a JSON error with exit code 2. execute_with_deadline resolves the session from --session or AGENT_DESKTOP_SESSION and builds a CommandContext with the trace, headed and agent-id settings (main.rs).
  2. Permission preflight. execute_with_adapter calls build_adapter(), which selects MacOSAdapter with #[cfg(target_os)]. It then asks the adapter for a permission report only when command_policy says the command needs one, and refuses early if it is missing (main.rs, command_policy/mod.rs). Any error raised before dispatch is stamped NotDelivered, so the caller knows nothing happened.
  3. Dispatch. dispatch() opens a mutating-command scope and routes Commands::Click to commands::click::execute (dispatch/mod.rs), which is a one-liner into execute_ref_action_with_context with Action::Click (click.rs).
  4. Load the ref. execute_ref_action_with_context loads the RefEntry for that snapshot from disk and starts the auto-wait loop (helpers.rs). execute_poll_loop resolves and preflights the target repeatedly until the deadline, without holding the lock (ref_action_poll.rs).
  5. Lock and re-resolve. Under the cross-process interaction lease, dispatch_resolved resolves the element again with resolve_element_strict, runs stable_preflight, and scrolls the element into view and retries once if the preflight asks for it (ref_action.rs). On macOS, strict resolution first checks that the saved process instance is still alive, then tries a path fast path and falls back to an identity search (resolve.rs). Zero matches become STALE_REF.
  6. Execute. execute_verified_action calls the adapter’s execute_action (post_action.rs). On macOS, perform_action runs the click chain: a physical CGEvent click, then AXPress, AXOpen, descendant activation, container selection and AXConfirm (dispatch.rs, chain_defs.rs). The physical step returns NotDelivered unless the request is headed (chain_step_exec.rs), so the default run is AX-only.
  7. Verify and reply. For stateful actions (type, set-value, toggle, check and others) the live element is read back and polled for up to 400 ms until the postcondition holds. The result goes back through finish() as one JSON line on stdout. Errors use the same envelope on stdout with exit code 1 (main.rs).

Key components

Snapshots and refs

snapshot::build picks the target window (or a sheet, popover or alert surface), asks the adapter to observe the tree within a 3-second deadline, and passes the raw tree to allocate_refs. Then run_with_context saves the refmap under a new snapshot id and qualifies every ref with it (snapshot.rs, L231-L251). A node gets a ref if its role is interactive or if it advertises an action other than SetFocus, RightClick or ScrollTo. Chromium and Electron put those three actions on almost every node, so counting them would ref the whole page (ref_alloc.rs). Each RefEntry stores pid and process instance, role, name/value/description, native id, a bounds hash and the tree path (ref_alloc.rs). The store keeps the newest 128 snapshots and prunes to 96 (refs_store.rs).

Skeleton traversal

--skeleton caps depth at 3 (commands/snapshot.rs). Truncated containers report a children_count, and named containers at the cut get “anchor” refs. The agent then drills in with snapshot --root @ref. This is the project’s main answer to token cost on dense apps such as Slack, and -i (interactive only) and --compact prune further.

Delivery semantics

Every error carries one of five dispositions: NotDelivered, DeliveryUncertain, DeliveredUnverified, DeliveredVerified or Unknown. Only NotDelivered maps to a safe retry (delivery_semantics.rs). This is the most useful idea in the codebase for agent builders. A caller can tell “the click never happened, re-snapshot and retry” apart from “the click may have happened, look before you act again”.

Headless by default

InteractionPolicy::headless() is the default and disallows focus stealing and cursor movement (interaction_policy.rs). --headed turns on the physical steps. Mutating commands take a file lock under a per-user directory, plus an in-process guard, so two agent-desktop processes cannot type at the same time (interaction_lease.rs). The held-input commands (key-down, key-up, mouse-down, mouse-up) are rejected in stateless mode, because nothing would guarantee the release (input_hold_policy.rs).

CDP hand-off for Chromium apps

launch --cdp rejects conflicting --remote-debugging-* switches and refuses if the app is already running. It then injects --remote-debugging-port bound to 127.0.0.1, launches the app, and polls /json/version until a webSocketDebuggerUrl appears (launch.rs, cdp_endpoint.rs). agent-desktop never sends CDP commands itself. It only returns a verified endpoint.

Sessions, traces and the command host

session start creates a directory with a manifest. Refs, JSONL trace segments and, in full artifacts mode, refmap copies are scoped to it (session/mod.rs). trace show/export reads them back. On macOS there is also an internal command host: a long-lived process on a Unix socket that keeps AX objects retained between calls, with a 15-minute idle timeout (command_host.rs, command_host_client.rs). It is enabled by internal environment variables, not by a documented flag. A lost reply is reported as DeliveryUncertain.

Extending it

  • New platform. Implement the four capability traits in crates/windows or crates/linux. PlatformAdapter is a blanket impl over them (adapter/mod.rs), and every unimplemented method already returns a “not supported” error (adapter/observation.rs). Core must not import a platform crate. CI checks this with cargo tree.
  • New command. Add a file under crates/core/src/commands/, an argument struct under src/cli_args/, a dispatch arm, a command_policy entry and a batch decoder entry (batch/mod.rs). There is no registry, so the compiler shows every place you missed.
  • Other languages. Link the cdylib and call ad_init(AD_ABI_VERSION_MAJOR) first. The header exposes about a hundred ad_* functions: snapshot, ref execution, raw trees, windows, apps, clipboard, notifications and errno-style ad_last_error_* getters. Build it with the release-ffi profile, because the normal release profile aborts on panic and would defeat the FFI’s catch_unwind guards.
  • Agent integration. agent-desktop skills get agent-desktop prints the bundled skill docs, so an agent can learn the command surface from the binary itself. There is no MCP server at this commit; --mcp is listed only as a future phase.

Running it

  • Install. npm install -g agent-desktop downloads a prebuilt binary, or cargo build --release (Rust 1.89+). The README targets macOS 13+.
  • Permissions. Grant Accessibility to the terminal or host process. Screenshots and visual debug also need Screen Recording. agent-desktop permissions and status report what is missing.
  • Basic loop. snapshot --app <App> -i --skeleton → pick a ref → click/type/set-value <ref> → take a fresh snapshot. Add --wait-for <selector> to make an action wait for a post-condition, and --session to keep refs and traces apart per task.
  • No services. There is no daemon to start, no API key and no network access, except the local DevTools probe when you pass --cdp.

Strengths and caveats

  • Strength: honest action results. Strict re-resolution, actionability preflight, state readback and the five-level delivery disposition are more rigorous than most computer-use tools, which report “clicked” and hope.
  • Strength: no cursor hijacking. The headless AX path lets an agent work in background windows while a person uses the machine. Physical input is opt-in.
  • Strength: token discipline. Skeleton snapshots, ref anchors and --root drill-down keep large Electron trees small enough for a model’s context.
  • Strength: two front doors. The CLI suits shell-calling agents, and the C ABI avoids one process per step for embedded use.
  • Caveat: macOS only. Windows and Linux are trait stubs with an empty surface list. Wiki pages that describe UIA or AT-SPI implementations describe plans, not code.
  • Caveat: not a browser tool. Web content inside Chromium apps is only as good as Chromium’s accessibility export. For real DOM work you have to hand the --cdp endpoint to a separate CDP client.
  • Caveat: refs go stale fast. A ref is tied to one process instance and one snapshot, and any UI change can invalidate it. Agents must re-snapshot often and handle STALE_REF. The tool returns the disposition, but retry logic is the caller’s job.
  • Caveat: large and strict codebase. About 62k lines in core and 40k in the macOS crate, split into files of at most 400 lines. That suits the project’s own invariants, but outside contributors face a lot of small files to learn.

Sources: code at 9d7ba42, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (28 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit 9d7ba42. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

agent-desktop does not automate web pages; it automates native desktop applications. The 'page' is a desktop app window, represented to the calling model as a tree of AccessibilityNode objects built from the OS-level accessibility tree (macOS ATK/SPI, via AXUIElementRef and batch attribute reads using AXUIElementCopyMultipleAttributeValues). Each node carries role, NodeIdentity (name, value, description, native_id), NodePresentation (states like focused/checked/disabled/hidden, available_actions, bounds, and hint), and optional ref_id. The tree is serialized as JSON inside the snapshot command's response envelope.

Pruning and size limits:

  • max_depth (default 10) clamps tree depth; --skeleton mode caps depth at 3 and marks truncated containers with children_count.
  • interactive_only removes non-interactive leaf nodes that have no ref and no children (role-based, screened by is_ref_able in ref_alloc.rs:79 which checks interactive roles plus non-ubiquitous actions).
  • compact collapses transparent wrapper nodes (single-child containers with no semantic payload).
  • subtree_truncated flag indicates budget/deadline exhaustion (no room to count children).
  • Budget/deadline capping (via LocatorTraversal.visit in traversal.rs:62-95) can cut the walk short when approaching limits.
  • Refs are allocated in depth-first order to interactive or action-bearing nodes via allocate_refs() (ref_alloc.rs:170). Ubiquitous affordances (SetFocus, RightClick, ScrollTo) do not qualify a node for a ref, preventing 'ref every node' inflation in web-runtime apps.
  • Skeleton anchors: named/described non-interactive containers at truncation boundaries receive refs as drill-down targets.
  • The RefMap has a 1 MB write-side limit (refs.rs:15 MAX_REFMAP_BYTES = 1_048_576).
  • Screenshots are available as a separate command (screenshot.rs), not embedded in the tree.
  • Structural hints are added for splitter rows to label columns.

The calling model does not see a DOM but a pruned, role-annotated accessibility tree with deterministic (per-snapshot) ref identifiers.

Editor's note. Correction: on macOS the tree comes from the AX API (AXUIElementRef, batch reads via AXUIElementCopyMultipleAttributeValues), not ATK/AT-SPI, which is the Linux accessibility stack and is not implemented here.

How are actions executed and how are elements targeted?

answered

Actions are dispatched through a platform-abstracted trait hierarchy — ActionOps.execute_action() receives a NativeHandle (resolved via resolve_element_strict) and an ActionRequest. The macOS adapter implements this via the accessibility API's AXAction/AXUIElementPerformAction, plus OS-level input synthesis for physical actions (mouse clicks via CFArray of CGPoint, keyboard via CX system APIs, clipboard via Cocoa). There is no CDP, Playwright, or web driver.

Elements are targeted by snapshot-qualified refs like @s8f3k2p9:e1. The flow:

  1. RefStore loads the snapshot's RefMap from disk (refs_store.rs).
  2. A RefEntry (pid, role, bounds hash, path, source surface, stable text identity) is looked up.
  3. resolve_element_strict() re-identifies the element on the live platform: matches pid, role, path/source surface, stable text identity, and bounds hash. Returns STALE_REF on 0 matches, AMBIGUOUS_TARGET on 2+ (adapter/observation.rs:79-85).
  4. A live actionability preflight checks visibility, stability (bounds polling until stable), enabled state, and hit-test (ref_action.rs:225-268).
  5. For stateful actions, execute_verified_action (post_action.rs:7-113) runs readback verification: re-reads the element's LiveElement state after execution and checks the postcondition, with up to 400ms of settle time.

Actions: Click, DoubleClick, RightClick, TripleClick, SetValue, SetFocus, Expand, Collapse, Select, Toggle, Check, Uncheck, Scroll, ScrollTo, PressKey, KeyDown, KeyUp, TypeText, Clear, Hover, Drag (27 variants). Batch mode deserializes JSON arrays into the same typed Commands enum and runs the same dispatch path.

Editor's note. Correction: the Action enum has 21 variants, not 27, and physical input is synthesised with CGEvent (only when --headed); the default path is AX actions such as AXPress. Strict resolution with STALE_REF/AMBIGUOUS_TARGET lives in crates/macos/src/tree/resolve.rs, not in the core trait default.

How is the agent loop / planning implemented?

not applicable

This repository does not implement an agent loop, planner, or executor. It is a stateless CLI tool (59 commands in commands/mod.rs) that an external AI agent invokes per interaction—the observation-action loop lives in the calling agent, not in this tool. The binary processes CLI input, executes one command via dispatch() (src/dispatch/mod.rs:48-112), produces one JSON envelope, and exits. It has no tool-call schema for LLMs, no memory between steps (beyond optional filesystem sessions for trace/ref persistence), and no stop conditions. The public API re-export (lib.rs:156-279) confirms all exports are types/traits/commands for tool building, not loop infrastructure. The CLAUDE.md document explicitly states this design: 'agent-desktop is NOT an AI agent. It is a tool that AI agents invoke.'

How are failures, retries and self-healing handled?

answered

agent-desktop has an extensive reliability architecture across multiple layers:

Error hierarchy: AppError (top-level, wraps AdapterError, IO, JSON, internal) → AdapterError (carries ErrorCode, message, suggestion, platform_detail, details, DeliverySemantics, retryability). 16 error codes total (error_code.rs:5-22), including STALE_REF, AMBIGUOUS_TARGET, TIMEOUT, ELEMENT_NOT_FOUND, APP_UNRESPONSIVE, POLICY_DENIED.

Delivery semantics: A 5-level model tracks action disposition — NotDelivered, DeliveryUncertain, DeliveredUnverified, DeliveredVerified, Unknown — each with a mapped RetryDisposition (Safe/Unsafe/Unknown) (delivery_semantics.rs:6-55). Errors with Safe retry (StaleRef, SnapshotNotFound, AppUnresponsive) get a recovery hint in the envelope with strategy, retryability, and required snapshot freshness.

Auto-wait polling: When timeout_ms > 0, ref_action_poll.rs:26-63 polls element resolution + actionability preflight before dispatching. The preflight checks visibility, enabled state, and bounds stability (sampling at 16ms intervals). Deadlines are propagated throughout.

Retry handling: is_retryable_resolution_failure() (adapter_error.rs:59-67) identifies error codes that are retryable (StaleRef, AmbiguousTarget, Timeout, AppUnresponsive), gated by explicit retryable: true in the error's details — a test matrix at adapter_error.rs:236-269 verifies every code. Retry is caller-side; agent-desktop returns the disposition, the calling agent decides.

Post-action verification: For stateful actions (TypeText, SetValue, Check, Uncheck, Toggle, Expand, Collapse, Clear), execute_verified_action (post_action.rs:7-113) re-reads live element state after dispatch and checks the postcondition. It polls for up to 400ms (50ms intervals) for the state to settle before concluding failure.

Stability detection: stable_preflight() (ref_action.rs:225-268) polls element bounds until they stabilize before cursor-driven dispatch. Uses a StabilitySampler — when bounds change mid-wait, it resets and re-stabilizes.

Renderer accessiblity activation: When a renderer has not published its tree yet, core backs off with exponential retry from 25ms to 400ms max (renderer_accessibility.rs:13-14).

Which models are supported and how are they called?

not applicable

This project does not embed, call, or support any language models. It has no provider integrations, no model IDs, no API keys, no vision requirements, no structured output schema for LLMs, and no small/specialised models. The Rust workspace Cargo.toml contains zero AI/ML crate dependencies. The binary processes CLI input, executes one command, produces one JSON response, and exits. The CLAUDE.md document states: 'Does NOT embed or invoke LLMs' under Non-Goals (line 78-80). The calling AI agent (Claude, GPT, Gemini, or any other system) is entirely separate and external.

How are browser sessions, profiles, auth and anti-bot handled?

answered

agent-desktop does not manage browser sessions, profiles, cookies, stealth, proxies, or CAPTCHA — it automates native desktop apps. Its session system is for application state persistence:

Session creation/lifecycle: session start creates a filesystem-backed session under ~/.agent-desktop/sessions/<run-timestamp-pid-counter>/ with a SessionManifest (id, optional name, created_at, ended_at, trace mode, artifacts mode, cursor_overlay config). Sessions are explicitly ended via session end (session/mod.rs:279-327). The session ID is passed via --session or AGENT_DESKTOP_SESSION env var and scopes snapshots/trace files. Sessions are entirely local — no remote/cloud infrastructure.

Trace/observability: When trace is enabled (default for sessions), per-process JSONL trace segments are written under <session>/trace/<pid>-<procTs>.jsonl (session/manifest.rs:38-44). ArtifactsMode::Full additionally copies each snapshot's refmap into the trace directory for post-hoc analysis.

Ref persistence: Snapshot refs are stored by snapshot ID under <state root>/snapshots/<snapshot_id>/refmap.json, with a latest_snapshot_id pointer. With sessions, the store is scoped to <state root>/sessions/<session_id>/. Retention keeps the newest 128 snapshots and prunes to 96 (refs_store.rs:16-17). A 1 MB write-side guard prevents oversized refmap files.

Subagent cursors: Session-aware multi-agent cursor mode (--cursor --multi-agent) provides independent visual cursors per agent, keyed by session ID and global --agent-id. Cursor overlay config can be set per-session and overridden per-agent via <session>/cursor-overlays/<agent_id>.json (session/mod.rs:144-182).

Anti-bot/stealth: Not applicable. The tool respects OS-level accessibility permissions via permission_report() (adapter/system.rs:30-37), requiring macOS accessibility grant. This is a permission, not anti-bot. The state root defaults to ~/.agent-desktop, overridable via AGENT_DESKTOP_HOME.