browser-use/jev-ultrafast
Python browser agent where TypeSafe's Jev picks an operation and an indexed DOM element per step, and a small LLM only types text.
Overview
Jev Ultrafast is a small Python browser agent from Browser Use, built around TypeSafe’s hosted Jev model. Its central idea is that the decision model never writes anything. On every step, the page is turned into a numbered table of controls, and Jev answers a set of multiple-choice questions: which operation to run (CLICK, TYPE_TEXT, SELECT, SCROLL_UP, SCROLL_DOWN, WAIT, DONE, BLOCKED) and which element each operation would target. A second, cheap OpenAI-compatible LLM is called only when the chosen operation is TYPE_TEXT, and only to produce one JSON string to type.
The whole package is about 850 lines: a loop (agent.py), a model client (model.py), a CDP executor (browser.py), one in-page JavaScript snapshot (snapshot.js), the prompts (questions.py) and a loopback inspector UI (demo.py). It drives a tab in your own Chrome through the Browser Harness daemon. It is a speed-focused MVP and a reference design, not a general framework. There are no tools, no memory, no vision, no file uploads and no frames or shadow DOM.
It suits readers who want to see how far a “choose, don’t generate” policy can go, or who need a fast loop for simple form-and-search tasks. It is not a drop-in replacement for a full agent framework.
Architecture
flowchart LR
U["Agent.run()"] --> T["command('tick')"]
T --> P["predict: model.choose()"]
P --> AS["action_space(): indexed elements"]
P --> TS["TypeSafe /v1/systemone"]
T --> A["act: Browser.act()"]
A --> FT["field_text(): text LLM"]
A --> FR["fresh(): marker / node guard"]
A --> BO["browser_operation()"]
BO --> CDP["Browser Harness cdp()"]
CDP --> SJ["snapshot.js in page"]
CDP --> IN["Input.dispatch* events"]
| Component | Path | Role |
|---|---|---|
| Agent loop | jev_ultrafast/agent.py |
Agent state machine: tick = predict + act, budgets, history, stop rules |
| Decision client | jev_ultrafast/model.py |
Builds the indexed action space, sends one TypeSafe request, validates heads, calls the text helper |
| Prompts | jev_ultrafast/questions.py |
NEXT_ACTION, TARGET, TEXT_VALUE rules and MAX_STEPS = 60 |
| Browser executor | jev_ultrafast/browser.py |
Owned CDP tab, observation, freshness checks, click/type/select/scroll execution |
| DOM snapshot | jev_ultrafast/snapshot.js |
One Runtime.evaluate that returns controls, visible text, node ids, marker and guards |
| Inspector | jev_ultrafast/demo.py, static/ |
Loopback HTTP UI that shows probabilities and can step the loop |
| Examples | examples/run.py, examples/flights.py |
Generic runner and a Google Flights run with independent verification |
How a request flows
- Start.
Agent(url, goals)joins the goal(s) into one task string, creates aBrowser, and takes a first observation (agent.py).Browser.__init__callsensure_daemon(), opens a backgroundabout:blanktarget, attaches a flattened session, sets a 1120x780 viewport, enables focus emulation, and navigates (browser.py). - Observe.
browser_operation({"operation": "observe"})evaluatessnapshot.jsonce and adds a SHA-256fingerprint(and a JPEG only if screenshots are on) (browser.py, L188-L194). - Tick.
run()keeps yieldingcommand("tick")until the status isdoneorblocked. A tick runspredict, thenact. AStalePagefrom either half is caught, and the page is re-observed for the next tick (agent.py, L163-L165). - Predict.
predictre-observes if the marker changed, enforces the 120-request budget, and callschoose()(agent.py).choose()builds oneoperationquestion plus one<op>_targetquestion per available operation. It posts page URL/title/text, the element table and the last 10 actions tohttps://api.typesafe.ai/v1/systemone, validates the operation head, and then validates only the target head for that operation (model.py). - Act.
actconsumes the decision before doing anything, so it cannot run twice.DONE/BLOCKEDonly set the status, after a freshness check (agent.py). For afillaction it buildsfield_context()and callsfield_text(), unless a cached value exists for an identical context (agent.py). - Execute.
Browser.actchecks freshness again and callsbrowser_operation({"operation": "act"}). That re-resolves the node from the in-page map, rejects disabled, hidden, off-screen or covered targets (elementFromPoint), and then dispatchesmousePressed/mouseReleasedat the centre. A fill adds a select-all key event andInput.insertText. A native<select>is set in JavaScript withinput/changeevents (browser.py, L135-L186). - Log, settle, re-observe. The action is appended to
historybefore the next observation.observe()first waits up to two animation frames (50 ms cap), or up to 200 ms for autocomplete options after typing into a combobox (browser.py). Three non-wait actions in a row that do not change the fingerprint set the status toblocked(agent.py).
Key components
The snapshot
snapshot.js queries links, buttons, inputs, selects, contenteditable nodes and 14 ARIA roles. It skips password, file and hidden inputs, invisible or disabled nodes, and anything outside the viewport (snapshot.js). Names come from a simplified accessible-name routine (aria-labelledby, aria-label, labels, text, title, placeholder; L12-L23). Editable fields produce both a fill and an “Open …” click action. Each <select> option becomes its own select action. Node identity is a page-side WeakMap counter, not a CDP backend node id (L3-L8). Visible on-screen text is capped at 6,000 characters, and actions are capped at 250. Then scroll_down, scroll_up and wait pseudo-actions are appended (L82-L106).
Freshness guards
There are two checks, both plain array equality, not hashes. Clicks and selects compare a narrow page_key (URL, scroll, viewport, form values) plus a per-node guard: role, name, value, ARIA state and the first 6,000 characters of the nearest form/dialog/row (snapshot.js). Everything else compares the full marker, which includes all visible text and action semantics (browser.py). The scoped click guard is the main speed trick: an animation elsewhere on the page no longer forces a new prediction.
Operation and target heads
action_space() gives each DOM node one index and groups targets by operation. Select targets are index:option pairs (model.py). The heads are asked in parallel in one request, so each target prompt has to state the operation it assumes (questions.py). validate_choice() requires a probability for every offered id, values in [0, 1] that sum to 1 within 0.02, and the choice to be the argmax (model.py).
Text helper
field_text() posts the goal, the field, 6,000 characters of page text and the last six actions to <TEXT_MODEL_BASE_URL>/chat/completions with response_format: json_object. The answer must be exactly {"text": "<non-empty, at most 2000 chars>"} (model.py). The code defaults are DeepSeek’s endpoint and deepseek-chat. .env.example switches this to OpenRouter and inception/mercury-2.5 with reasoning off. A {"text": null} answer is treated as an error, not as a skip.
Inspector
uv run jev serves a UI on loopback. It checks the Host header, a per-process token and Origin, and runs one browser step at a time behind a lock (demo.py). The UI can run Google Flights or two local fixture scenarios, and it exposes predict and act separately, so you can see the probabilities before execution.
Extending it
- Drive it step by step.
Agent.command("predict")andcommand("act", {"fingerprint": ...})are public, and the decision dict carries the full request body and raw answers. That is enough for custom UIs or logging. - Change the policy. Edit
NEXT_ACTION/TARGETinquestions.py. There is no prompt-injection hook, so edits are in source. - New operations. Add a
kindtosnapshot.js, map it inaction_space()’soperationsdict, and handle it inbrowser_operation(). Non-element controls (scroll, wait) flow through thecontrolsdict as their own operations. - Swap the text model. Any OpenAI-compatible endpoint works through
TEXT_MODEL_BASE_URL,TEXT_MODELandTEXT_MODEL_REASONING. The decision model is fixed to the TypeSafe API. OnlyTYPESAFE_MODELis configurable. - Verify outcomes yourself.
examples/flights.pyshows the intended pattern: averify()that checks URL, field values and result rows, independent of the model’sDONE(flights.py).
Running it
uv sync, copy.env.exampleto.env, and setTYPESAFE_API_KEYandTEXT_MODEL_API_KEY. Python 3.12+. Dependencies are onlybrowser-harness==0.1.13andhttpx[http2].- Chrome must accept a remote-debugging connection through Browser Harness (
uv run browser-harness --doctor). The agent opens its own background tab in that browser, so it shares your normal profile, cookies and logins. uv run jevstarts the inspector.examples/run.py --url ... --goal ...runs a goal from the terminal. Tests are offline.scripts/check_guards.pyexercises the guards against a local browser without model calls.
Strengths and caveats
- Strength: the model cannot emit code. Every executable target is a code-owned node id from the latest snapshot, and every model answer is validated as a full probability distribution. The text helper is constrained to a single JSON string.
- Strength: low latency by design. One decision round trip per step, no screenshots in the loop, one
Runtime.evaluateper observation, and short event-based settle waits. The project’s own matched comparison (three pairs, one task) reports a median of 7.1 s against 9.5 s for the previous version, and is honest that this is not a benchmark. - Strength: careful mutation semantics. Decisions are consumed before execution, history is written before re-observation, and an interrupted native select raises instead of retrying.
- Caveat: hard dependency on a hosted, proprietary choice model. Without TypeSafe there is no decision engine, and the code has no fallback provider.
- Caveat: brittle error handling. Only
StalePageis recovered. An invalid TypeSafe response, a null text value, a provider HTTP error or a budget stop raises out ofrun()and ends the run. - Caveat: narrow page coverage. No iframes, shadow DOM, canvas, uploads, new tabs, nested scroll containers or keyboard shortcuts. Scroll is a fixed 560 px wheel event at one point.
- Caveat: no session features. There is no profile isolation, persistence, proxy, stealth or CAPTCHA handling in this repo. Whatever the Browser Harness connection provides is what you get.
Sources: code at 1231850, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (8 pages), verified Q&A.
How it answers the Browser & computer control questions
Each answer was drafted by a code-reading agent at commit 1231850. Its citations were checked mechanically. Compare with the other browser & computer control →
How is the page represented to the model?
answeredThe page is represented as a structured, indexed state — never via screenshots in the model's path. The core DOM reader is snapshot.js, a ~107-line JavaScript that runs atomically inside a single Runtime.evaluate call (browser.py:188-190). It walks visible, safe controls (button, link, checkbox, combobox, textbox, searchbox, etc.; snapshot.js:24-27) and produces a numbered action list where each element has a role, accessibility-derived label, value, checked/selected/expanded state, and bounding rectangle. Text is collected from on-screen text nodes within the viewport, capped at 6000 characters (snapshot.js:82-91). The actions array is capped at 250 elements, after which scroll_up/scroll_down/wait controls are appended (snapshot.js:99-104).
A SHA-256 fingerprint over url, text, actions, and scroll detects page changes (browser.py:115-117). Freshness between observations is verified by comparing a marker — a tuple of performance.timeOrigin, URL, scroll position, viewport size, title, text, action semantics, and form control values — against the current DOM (snapshot.js:97-98). Individual elements get an integer node ID via a WeakMap cache (snapshot.js:3-7) and a guard (role, name, value, state, nearby scope text; snapshot.js:47-54) used to detect whether a specific observed node is still valid. No screenshot is sent to the model (README:97); screenshots are only captured for the human inspector (JPEG, quality 72; browser.py:193).
How are actions executed and how are elements targeted?
answeredActions are executed entirely through CDP input dispatch, not Playwright or OS-level automation. All targeting uses integer node IDs assigned by the DOM reader (snapshot.js:4-6), never model-generated selectors, coordinates, or JavaScript (README:105-106). The Browser.act() method first checks freshness via fresh() (browser.py:100-102), then delegates to browser_operation() (browser.py:135-186).
Before any input, the executor revalidates the element via JavaScript: it checks isConnected, disabled state, aria-disabled, inert, visibility and occlusion via elementFromPoint, and read-only for fillable fields (browser.py:144-160). Geometry is resolved at execution time, not cached from observation. If validation fails, StalePage is raised and no CDP input fires.
- click:
Input.dispatchMouseEventwithmousePressed/mouseReleasedat the element's bounding-box center (browser.py:167-168). - fill: same mouse events to focus, then Ctrl+A (Cmd on macOS) via
Input.dispatchKeyEventwithselectAll, followed byInput.insertTextwith the text-helper value (browser.py:169-185). - select: synchronous DOM mutation — sets
e.value, dispatchesinputandchangeevents (browser.py:153-158). - scroll:
Input.dispatchMouseEventwithmouseWheelat a fixed viewport center, sending a delta of ±560 pixels (browser.py:138-139). - wait: plain
time.sleep(0.1)(browser.py:103-104). - DONE/BLOCKED: no browser action; updates status and returns (
agent.py:93-100).
No file-upload or new-tab actions are implemented — acknowledged as outside the MVP (README:126).
How is the agent loop / planning implemented?
answeredThe agent loop lives in Agent (agent.py:12-165). Initialization (__init__, line 12-41) takes a URL and goals, starts a Browser session, observes the first page, and sets up a state dict containing the goal plan, page state, history, and counters. The public run() method (line 163-165) yields states from command("tick") while status is not "done" or "blocked".
Each tick (line 55-64) runs two phases: predict then act. predict (line 65-85) calls model.choose(), which sends a single HTTP POST to https://api.typesafe.ai/v1/systemone (model.py:119) containing: the page URL/title/text, the indexed element table (with roles, values, ARIA states), the goal, and the last 10 history entries (model.py:107-117). The request poses one operation question (CLICK/TYPE_TEXT/SELECT/SCROLL_UP/SCROLL_DOWN/WAIT/DONE/BLOCKED) plus one target question per eligible operation — two decision layers in one round trip (README:34-50). The TypeSafe model returns choice probabilities per head; the chosen target head is validated via validate_choice() and only that head's selection runs (model.py:120-133). If the operation is TYPE_TEXT, a separate call to a small OpenAI-compatible model (field_text(), model.py:160-198) generates the text value.
Act (line 86-161) consumes the decision (setting it to None to prevent double-clicks; line 91), executes via Browser.act() (browser.py:100-107), appends to history (line 121-141), and reobserves the page. Stop conditions: DONE/BLOCKED selection (agent.py:93-98), reaching MAX_STEPS (60; agent.py:102-104 and questions.py:26), or 3 consecutive non-wait actions that caused no page change (agent.py:153-158). The predictor-phase budget is MAX_STEPS * 2 model calls (agent.py:75-76). History persists between steps and the last 10 entries are included in each predict request.
How are failures, retries and self-healing handled?
answeredReliability is enforced at four layers: pre-execution page freshness, model response validation, retry loops for transient failures, and no-progress detection.
Pre-execution freshness. Before any action, Browser.fresh() checks whether the observed DOM still matches the current page using a marker hash (browser.py:88-98). If not, StalePage is raised and the decision is discarded (consumed to prevent double-execution; agent.py:91). The prediction in act() validates that the target element's node ID still maps to the same guard (role, name, value, disabled state, aria-expanded; browser.py:93-97). The decision is consumed before any model call or mutation — a retry cannot double-click (agent.py:91).
Model validation. validate_choice() (model.py:30-45) rejects any TypeSafe response where probabilities sum to other than 1.0 (within 0.02 tolerance), contain NaN/negative values, or assign the choice to a non-maximum element. Invalid text-helper output (wrong JSON shape, empty, >2000 chars) raises ValueError("nothing typed") (model.py:187-193).
Retry loops. HTTP requests to the TypeSafe API retry on 429/503/529 with exponential backoff (0.5s, 1s; model.py:16-27). Observation retries up to 10 times on StalePage with 20ms waits (browser.py:77-86). Stale page during a tick causes automatic reobservation and a fresh prediction without executing the stale decision (agent.py:59-64).
No-progress detection. If 3 consecutive non-wait actions produce no page fingerprint change, status transitions to "blocked" (agent.py:153-158). Generated text for TYPE_TEXT is cached in pending_text and reused only if the entire field-context matches — stale-page retries with the same page state avoid a redundant text helper call (agent.py:109-115). There is no caching of successful actions or workflows beyond this text reuse.
Which models are supported and how are they called?
answeredThe project uses two distinct models for different roles. The primary decision maker is TypeSafe's Jev (default model name "jev-latest", configurable via TYPESAFE_MODEL env var; model.py:108), called at https://api.typesafe.ai/v1/systemone (model.py:119). Jev receives structured questions — operation and target heads with their constraints, the page state (URL, title, text, indexed elements with ARIA roles/values/state), the goal, and recent action history — and returns probabilities per choice plus a confidence value (model.py:119-148). It is a choice model, not a generative model: it picks among offered alternatives rather than generating free text.
The secondary model is a text helper for TYPE_TEXT operations only (model.py:160-198). It calls any OpenAI-compatible chat-completion API at TEXT_MODEL_BASE_URL (default https://api.deepseek.com/v1; model.py:164) using TEXT_MODEL (default deepseek-chat; model.py:165) and requires TEXT_MODEL_API_KEY. It uses response_format: {"type": "json_object"} with max_tokens: 1024 (model.py:174-176). Reasoning is disabled for DeepSeek's base URL and set to "low" effort for others, with a TEXT_MODEL_REASONING=none escape hatch (model.py:166-168). The current demo config uses inception/mercury-2.5 via OpenRouter (README:68). Gemini, GLM, and DeepSeek are also listed as compatible text helpers (README:68).
Vision is not used. Screenshots are captured but never sent to any model; they exist only for the human-facing inspector (README:97). Structured output via response_format: "json_object" is used only for the text helper; the TypeSafe API has its own structured response contract with typed probabilities.
How are browser sessions, profiles, auth and anti-bot handled?
answeredBrowser sessions connect to a local Chrome instance through the Browser Harness daemon (browser.py:21-22). The Browser constructor calls ensure_daemon() to start the harness, then creates a new CDP target via Target.createTarget (opening an about:blank page) and attaches to it via Target.attachToTarget with flatten mode (browser.py:22-24). This gives a single CDP session per Agent instance — there are no per-step subprocesses or new browser instances.
The viewport is set to 1120×780 with deviceScaleFactor: 1 and mobile: false (browser.py:25). Focus emulation is enabled so the owned background tab keeps rendering animations and menus without switching Chrome's visible tab (browser.py:27). The initial URL is navigated with a 15-second deadline polled on readyState (browser.py:28-33).
There are no persistent profile or cookie management features in this codebase. The session inherits whatever Chrome profile the Browser Harness connects to — typically the user's default Chrome profile (via remote debugging; README:66-67). Session IDs (targetId, sessionId) are saved to disk by the example script (examples/flights.py:57-58) but not reused for resumption. On close, Target.closeTarget is called (browser.py:110-111).
No stealth, proxy, or CAPTCHA handling is implemented. The README explicitly acknowledges limits: shadow roots, frames, canvas, uploads, pop-up tabs, and nested scrolling are not handled (README:126-127). Background tabs are kept rendering via CDP focus emulation but pop-up tabs are outside the MVP (README:126). Remote or cloud browsers are not supported — only local CDP.