LLMs Technical Reviews

droidrun/mobile-jev

Android phone agent where Jev picks typed operations from the accessibility tree and the Mobilerun cloud API executes them.

GitHub ↗★ 435JavaScriptMITcommit 395fc22 · 2026-09-17

Overview

Mobile Jev is a small agent for Android phones, not for browsers. You give it a goal such as “Turn on dark theme in Android Settings”. It reads the phone’s accessibility tree through the Mobilerun cloud API, asks TypeSafe’s Jev model for one typed decision, and performs that decision with Mobilerun’s tap, swipe, keyboard and app-launch endpoints. It needs no ADB, emulator or USB connection. The phone is a Mobilerun-managed device that you already have.

The repository is about 2,000 lines of plain JavaScript (ES modules, no build step) in scripts/mobile-agent/, plus a local Next.js “studio” that streams the device screen and shows each decision as it happens. Jev is the only model. It does not write text. It picks an operation and a target from indexed lists, and any text it types is a span copied verbatim from the goal (or from a --text value).

The code is conservative by design. It re-checks the screen before every action, never retries a mutation after an uncertain transport failure, verifies typed text by reading it back, and says plainly that a DONE from the model is not proof of success.

Architecture

flowchart LR
  U["CLI or Studio"] --> L["runAgent loop"]
  L --> O["MobilerunDevice.observe"]
  O --> S["summarizeState: indexed elements"]
  S --> P["TypeSafePolicy.decide"]
  P --> J["Jev API: operation + speculative targets"]
  J --> V["validateChoice, pick branch"]
  V --> F["assertFresh: re-observe and compare"]
  F --> A["Mobilerun REST: tap, swipe, keyboard, apps"]
  A --> C["confirmInput / settle poll"]
  C --> L
Component Path Role
Agent loop scripts/mobile-agent/agent.mjs runAgent: decide, act, observe, stop conditions, timings
Policy scripts/mobile-agent/policy.mjs Builds the operation and target questions, calls Jev, validates distributions
Candidates scripts/mobile-agent/actions.mjs Turns an observation into tap, scroll, key, global and text candidates
Device adapter scripts/mobile-agent/device.mjs Mobilerun client: observe, listApps, act, staleness checks
Transport scripts/mobile-agent/http.mjs Keep-alive HTTPS for Jev; curl with config on stdin for Mobilerun; timing splits
Text and input text.mjs, input-verification.mjs Goal-span text candidates; read-back verification after typing
CLI scripts/mobile-agent/cli.mjs run, observe, screenshot, profile, direct tap/type/back
Studio apps/jev-studio/ Next.js UI, localhost-only API route, SSE events, child-process runs
Demo scripts/demo.mjs, demo-verifiers.mjs Dark-theme demo with independent outcome check

How a request flows

  1. Start. runAgent checks that the device is ready, then fetches the first observation and the installed-app list in parallel (agent.mjs).
  2. Observe. observe() calls /ui-state?filter=false. summarizeState walks a11y_tree, keeps visible nodes with text or interactive flags, gives each one a path id such as ui.0.3, masks password text, and hashes the result into a SHA-256 fingerprint (device.mjs).
  3. Build questions. candidatesFor emits a tap for each enabled clickable or editable node, scroll gestures per scrollable region, back/home, and, when a non-password field is focused, enter and text options (actions.mjs). buildQuestions groups these into one operation Choice (OPEN_APP, TAP, TYPE_TEXT, SCROLL_*, BACK, HOME, ENTER, WAIT, DONE, BLOCKED) and one speculative target Choice per operation family (policy.mjs).
  4. Decide. decide posts the goal, foreground app, indexed elements, visible text, up to 200 apps and the last 8 actions to https://api.typesafe.ai/v1/systemone in one request. It validates only the chosen branch’s answer and maps it back to an action (policy.mjs).
  5. Act. act fetches a fresh observation and runs assertFresh, which checks per action type that the target still means the same thing. Only then does it send the request, for example POST /tap at the centre of the node’s current bounds (device.mjs).
  6. Confirm. The loop logs the action before observing, then polls (every 60 ms, within a 400 ms budget) until the screen changes. After a text entry it calls confirmInput and stops with input_unverified if the field never shows the exact value (agent.mjs).

Key components

One request, speculative targets

Jev’s Choice API answers several independent questions in one call. The policy uses that to ask “which operation?” and, at the same time, “if TAP, which element?”, “if TYPE_TEXT, which span?” and so on. Only the branch that matches the chosen operation is validated and executed. validateChoice rejects distributions that do not cover exactly the offered keys, do not sum to about 1, or whose choice is not the maximum (policy.mjs). This saves a full round trip per step compared with a two-stage design.

Text without generation

textCandidates enumerates every 1–8-word span of the goal and trims punctuation. If more than 254 values result, it gives up and asks for --text (text.mjs). The target question includes a NONE option, which turns into a needs_input status instead of an invented value. This is a clean fit for a model that cannot generate text, but it means any value not written in the goal must come from the caller.

Staleness and stop rules

assertFresh compares device, app, screen size, observation age (30 s) and an action-specific “meaning” of the target, such as the tapped node’s label and children or the focused field’s identity (device.mjs). A stale decision costs a model call but never an input. Three in a row end the run as unstable_screen. The other exits are done, blocked, step_limit, stuck (same fingerprint and action twice), loading_timeout and decision_limit.

Studio

The catch-all API route accepts only localhost hosts and same-origin POSTs. It hands the browser only the device’s stream URL and token, never the account keys (route.ts). Each run is a child process running web-runner.mjs, which emits JSON lines that the store relays over SSE (web-runner.mjs). Runs live in memory, only one run can use the device at a time, and restarting the server clears them.

Extending it

runAgent takes device and policy as plain objects, so another device backend only needs assertReady, observe, act and optionally listApps, and a different policy only needs decide. That seam is real, but the action vocabulary (tap-element, swipe, type, key, global, open-app) and the element shape are Mobilerun-specific. The demo runner shows how to add task-specific verifiers such as darkThemeState, which reads the actual toggle state.

Running it

Node.js 22.16+, pnpm, curl 7.70+, a ready Mobilerun Android device, and MOBILERUN_API_KEY, MOBILERUN_DEVICE_ID and TYPESAFE_API_KEY in .env.local. pnpm devices and pnpm doctor check the setup. pnpm dev serves the studio on 127.0.0.1:3040. pnpm agent run "<goal>" only previews the next decision. Add --execute (and optionally --steps, --trace, --confidence, --text) to act. The CLI defaults to 10 steps and the studio to 30.

Strengths and caveats

  • Strength: careful execution. Pre-action freshness checks, no blind retries, action logging before the next read, and read-back of typed text make failures visible instead of compounding.
  • Strength: fast loop. One Jev call per step, a warm TLS connection, and short polling instead of fixed sleeps. Timings are split into DNS, TLS, wait and download.
  • Caveat: two vendors, both required. Jev is the only supported model and Mobilerun the only device backend. There is no local-device or alternative-LLM path in the code.
  • Caveat: no independent completion check in the loop. Unlike a goal/stuck watcher design, DONE is just one of the operation choices. Only the bundled dark-theme demo verifies the outcome.
  • Caveat: limited text and gestures. No long-press, pinch or drag, and Enter is the only key the model is offered. Typed values must appear in the goal or be passed with --text.
  • Caveat: extra latency per action. Every act does a full UI-state fetch for the freshness check before sending the input.
  • Caveat: single operator. The studio is localhost-only with in-memory state. Shared or remote use needs your own authentication.

Sources: code at 395fc22, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (18 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit 395fc22. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

The page is represented via the Mobilerun accessibility-tree API, not a DOM or screenshot pipeline. MobilerunDevice.observe() (device.mjs:257-260) calls the Mobilerun /ui-state?filter=false endpoint, which returns an accessibility tree of the current Android screen. summarizeState() (device.mjs:88-189) walks the tree recursively via an inner visit() function that reads a11y_tree nodes, filters by isVisibleToUser !== false, clamps bounds to screen dimensions, and builds an element list. Each element carries: id (a dot-separated tree path like ui.0.1), text/label, resourceId, bounds, plus booleans for clickable, editable, scrollable, enabled, focused, password, checkable, checked, selected. Password text is replaced with "[password]" (device.mjs:124).

No screenshots are sent to the model. Jev is a text-only reasoning system; the screenshot method (device.mjs:262-267) exists only for CLI viewing and debug traces. The model receives a set-of-marks via indexed controls: buildQuestions() (policy.mjs:46-65) assigns 1-based string indices ("1", "2", …) to elements. These indices appear in the Jev questions (tap_target, scroll_target, etc.) so the model selects by index, not coordinates.

Size limits: The serialized request body must stay under 150,000 bytes (policy.mjs:217-219). App inventory is capped at 200 apps (policy.mjs:39). Text candidates from the goal are limited to 254 unique values (text.mjs:15). An elements array in the state lists each index with its label, editable, scrollable, checked state and supported operations (policy.mjs:52-63). The observation also includes a fingerprint (SHA-256 of the full content) used for change detection (device.mjs:186). A phone object carries packageName, currentApp, isEditable, keyboardVisible, and focusedElement metadata (device.mjs:165-182).

How are actions executed and how are elements targeted?

answered

Actions are executed via the Mobilerun REST API — no CDP, Playwright, or OS-level input. The MobilerunDevice.act() method (device.mjs:269-376) dispatches typed actions:

  • tap-element: Resolves elementId to its center coordinates from the observation's bounds and sends POST /tap with {x, y} (device.mjs:298-307). Requires the original observation for staleness checking.
  • tap: Raw coordinate tap POST /tap (device.mjs:294-296).
  • swipe: POST /swipe with startX/Y, endX/Y, duration (default 300ms). If a regionId is provided, gestures are re-projected to current bounds (device.mjs:320-336) so scrolling stays within the right region even after UI layout changes.
  • type: POST /keyboard with {text, clear, completionMode} (device.mjs:338-353). Text comes solely from goal spans via textCandidates() (text.mjs:2-19) — Jev selects a span, code copies it. The model never generates arbitrary prose. The completion mode is "accepted" (local polling verification) by default, or "committed" (server waits) for passwords/appends.
  • clear: DELETE /keyboard (device.mjs:354-358).
  • key: PUT /keyboard with key codes from the KEYS map (device.mjs:359-364): back (4), tab (61), enter (66), delete (67), forward_delete (112).
  • global: System navigation (POST /global) with action codes: back=1, home=2, recent=3 (device.mjs:5-6, 365-370).
  • open-app: PUT /apps/:packageName to launch an installed app directly via the API (device.mjs:272-279). Requires prior listApps() discovery.

Element targeting works through the index map built in buildQuestions() (policy.mjs:67-82). Each element ID maps to a string index; the model selects by index, then policy.mjs:237-258 resolves the Jev choice back to the actual action object with the element ID. Direct coordinate actions ({type: "tap", x, y}) are also available for debugging via the CLI.

Typing uses input verification: after accepted completion, the code polls the screen with confirmInput() (input-verification.mjs:46-65) until the field shows the exact text, the field identity is unchanged, and the field isn't a password field. If verification times out, the run stops with input_unverified — it never retypes blindly.

File upload: Not implemented. Tabs: The tab key (code 61) is the only supported keyboard navigation key. There is no browser tab management — this is a mobile phone, not a web browser.

How is the agent loop / planning implemented?

answered

The agent loop lives in runAgent() (agent.mjs:9-190). One agent loop step: observe → policy.decide → validate → (if execute) device.act → observe again → check stop conditions.

Step loop: Runs up to maxSteps iterations (default 10, max 100), with an extended attempt budget of maxSteps × 2 + 4 to accommodate stale-retry re-observations (agent.mjs:76). Each iteration calls policy.decide() (agent.mjs:78-85) which returns a decision with status and optionally an action.

Planner vs executor: Planning is a single-stage monolithic call — there is no separate planner/executor. TypeSafePolicy.decide() (policy.mjs:169-304) builds one JSON body containing state (goal, current app, editable status, visible text, an element index table, and the last 8 actions from history) plus questions (the operation choice + speculative target questions). This is sent as one request to the TypeSafe API (https://api.typesafe.ai/v1/systemone, policy.mjs:222-224).

Tool-call schema: The model receives choice-type questions with criteria maps. The operation question (policy.mjs:84-110) lists candidate operations: OPEN_APP, TAP, SCROLL_DOWN/UP/LEFT/RIGHT, TYPE_TEXT, system nav (BACK, HOME, ENTER), WAIT, DONE, BLOCKED — each with a human-readable description. Compatible targets (tap_target, scroll_target, text_value, app_target) are speculatively answered in the same request but only the selected branch is consumed (policy.mjs:237-258). Response validation via validateChoice() (policy.mjs:8-26) checks probability distributions sum to ≈1 and that the chosen item has the highest probability.

Stop conditions (agent.mjs:89-101): done (model says DONE), blocked (model says BLOCKED), step_limit (actions ≥ maxSteps), stuck (same action+observation fingerprint repeated), loading_timeout (consecutive WAITs exceed waitTimeoutMs, default 15s), unstable_screen (3 consecutive stale observations), decision_limit (budget exhausted), input_unverified (text not confirmed in field).

Memory between steps: The history array (agent.mjs:41) stores the last 8 action entries with operation, label, text, and screenChanged flag, passed into the next decide call (policy.mjs:208-213). The repeated set (agent.mjs:42) tracks observation.fingerprint:action JSON to detect action loops. Actions are logged before the post-action observation (agent.mjs:141-147) so a failed read cannot erase an executed action.

How are failures, retries and self-healing handled?

answered

Error classes: A single custom class — StaleObservationError (device.mjs:8). It is thrown by assertFresh() (device.mjs:31-81) when the current observation doesn't match the one used for decision-making.

Retries: Only StaleObservationError triggers a retry. The agent catches it (agent.mjs:116-119), increments staleRetries, re-observes, and re-decides. After 3 consecutive stale errors, the loop terminates with unstable_screen (agent.mjs:124). Non-stale errors (transport failures, invalid actions, API HTTP errors) are never retried — they propagate immediately (agent.mjs:117), because a device mutation might have partially executed. This is a deliberate design principle: "Never retry uncertain mutations" (agent.mjs:117).

Freshness / self-healing: assertFresh() validates the observation before any action dispatch. It checks: device ID match, observedAt age (<30s default, maxAgeMs), package name, screen dimensions, and action-specific target semantics. For tap-element, it checks the target node's text/label/checked/selected haven't changed (device.mjs:46-49). For type, it checks the focused input has the same identity and meaning (device.mjs:50-55). For global navigation, it checks all clickable/editable controls plus headings haven't changed (device.mjs:56-69). For swipe with regionId, it checks the scrollable region still exists with the same resourceId (device.mjs:70-73). If any check fails, StaleObservationError is thrown and re-observation + re-decision occurs automatically (agent.mjs:123-127).

Input verification: After accepted-mode typing, confirmInput() (input-verification.mjs:46-65) polls the screen every 60ms until the field shows the exact expected text. Timeout defaults to 2500ms. If verification fails, the run stops with input_unverified (agent.mjs:159-171) — it never modifies the field again based on an unverified value.

Caching: Device readiness is cached for 30 seconds (device.mjs:271). Model transport reuses the HTTPS connection via a maxSockets: 2 keepAlive agent (http.mjs:5). Post-action polling is bounded by settleTimeoutMs (default 400ms) with 60ms intervals (agent.mjs:174-183). No workflow-level caching exists.

Timeouts: settleMs (extra fixed wait after actions, default 0), settleTimeoutMs (400ms, polling budget for transition), waitTimeoutMs (15s, consecutive WAIT budget), and inputTimeoutMs (2500ms, text verification timeout). All validated at entry (agent.mjs:24-31).

Which models are supported and how are they called?

answered

One provider: TypeSafe API at https://api.typesafe.ai/v1/systemone (policy.mjs:222-224). Only TypeSafe's proprietary Jev model is supported. There is no support for OpenAI, Anthropic, Gemini, or any other LLM provider in the codebase. The model is configured via TYPESAFE_MODEL env var or the TypeSafePolicy constructor; the default is 'jev-latest' (policy.mjs:157). The README notes that jev-latest is a convenience alias — fixed versions can be pinned for reproducibility (README line 111).

No vision: Jev does not receive screenshots. The entire model input is structured JSON — element indices, labels, text values, app metadata, recent action history, and the goal string. Screenshots are available via device.mjs:262-267 only for CLI debugging, never sent to the model.

Structured output: The model returns JSON with an answers object containing named choice selections. Each choice includes a type ("choice"), choice (the selected option key), confidence (0–1), and probabilities (a map of option → probability summing to ≈1). The response also includes usage (token counts) and model (the actual model serving the request) (policy.mjs:275-285). Validation via validateChoice() (policy.mjs:8-26) enforces that probabilities sum to within 0.025 of 1.0, the selected choice has max probability, and all values are finite 0–1. Invalid distributions cause an immediate hard error.

Single-request operation+target: The operation question and all compatible target questions are sent in one API call (policy.mjs:111-152). Only the selected branch's answer is validated and consumed; unused speculative answers (e.g., tap_target when DONE was chosen) are ignored (policy.mjs:237-258).

No small/specialised fallback models: Jev handles everything. There is no routing, cascading, or fallback between models. The confidence threshold (--confidence, default 0) can surface uncertain status but doesn't switch models.

Transport: Uses Node.js built-in https module with a keepAlive agent (maxSockets: 2) for connection reuse (http.mjs:5). pooledRequest() (http.mjs:9-101) provides detailed per-request timing: DNS, TCP connect, TLS handshake, response wait, and download phases. Timing data is embedded in the response buffer via Object.defineProperty(bytes, 'timing', ...) (http.mjs:56).

How are browser sessions, profiles, auth and anti-bot handled?

answered

This is not a browser automation tool. It controls real Android phones via the Mobilerun cloud API. The MobilerunDevice class (device.mjs:191-377) communicates with https://api.mobilerun.ai/v1 (default) or a staging variant via MOBILERUN_BASE_URL. The device ID is configured through MOBILERUN_DEVICE_ID (.env.example:5). Devices are provisioned and managed outside this project — the code merely connects to one.

Local vs remote/cloud: All device operations go through the Mobilerun cloud API. There is no local ADB, no emulator support, and no direct USB connection. The assertReady() method (device.mjs:249-256) checks the device state via GET /devices/:id and requires state === 'ready'. Device readiness is cached for 30 seconds to avoid redundant checks (device.mjs:271).

Persistent profiles and cookies: Not applicable — this is a mobile phone with its own persistent state (installed apps, accounts, settings). The listApps() method (device.mjs:240-247) queries installed apps with ?includeSystemApps=true so Jev can discover and launch them. The phone retains its state between runs; there is no profile isolation or session reset in the code.

Stealth / anti-bot / proxies: None implemented. The device is a genuine Android phone managed by Mobilerun — there is no browser fingerprint, user-agent, or proxy configuration. The Mobilerun API requires HTTPS (http.mjs:107-111). The TypeSafe API also requires HTTPS (http.mjs:11-13). Only 127.0.0.1 loopback is allowed for testing.

CAPTCHA handling: Not implemented. The README explicitly warns that "Jev's DONE response is not independent proof of success" (README line 103) and recommends checking resulting device state, especially for numeric values and multi-part goals. The demo runner (demo-verifiers.mjs:1-8) performs post-task state verification by checking actual toggle state rather than trusting the model.

Studio sessions: The Next.js studio (apps/jev-studio/) spawns agent runs as isolated child processes running web-runner.mjs (store.mjs:98-103). Each run gets its own process. API credentials live only on the server; the browser receives only device-scoped streaming tokens (route.ts:21-22). The studio is restricted to localhost with same-origin validation (store.mjs:12-24). Runs are held in in-memory storage only, cleared on server restart (store.mjs:178-183). One task owns the configured device at a time — concurrency is rejected with a 409 error (store.mjs:69, 151).