LLMs Technical Reviews

jkudish/jev-browser

MCP server, CLI and library that drive headless Chromium with TypeSafe's Jev model making one typed choice per step.

GitHub ↗★ 311JavaScriptMITcommit 1f0726c · 2026-10-02

Overview

jev-browser is a small TypeScript package that runs a browser task end to end: you give it a natural-language task and a start URL, and it drives a headless Playwright Chromium until the goal is met, the run is stuck, or a budget runs out. It is used as an MCP server (one tool, jev_navigate), a CLI (jev-browser run), or a library (navigate()).

What sets it apart is the model. Decisions come from TypeSafe’s Jev, a judgment model that returns typed answers instead of text: a Choice (a probability distribution over named options) and Nouls (single probabilities in [0,1]). Each step asks three questions at once: which action to take, whether the goal is already met, and whether the run is stuck. Code owns everything else, including the loop, stop thresholds, recovery and budgets. Jev cannot write text, so typing into fields uses a separate small LLM through the Vercel AI SDK.

The result is a single JSON object: final status, the final page as text, markdown, HTML or an ARIA snapshot, a per-step trace with confidences, console and network errors, token usage with an estimated cost, and an optional final screenshot. A large share of the code is credential hygiene. Password fill and seed cookies are built so that secret values never reach any model, trace or screenshot.

Architecture

flowchart LR
  C["MCP client / CLI / library"] --> N["navigate() loop"]
  N --> PW["Playwright Chromium"]
  PW --> X["extractAndStamp: DOM candidates"]
  X --> AS["buildActionSpace: e1..eN + controls"]
  AS --> J["askJev: action Choice + goal/stuck Nouls"]
  J --> G{"Stop gates"}
  G -->|"continue"| E["Execute via Playwright"]
  E --> TY["Typing model (AI SDK)"]
  E --> N
  G -->|"stop"| R["Result: page, trace, usage, screenshot"]
  N --> BP["detectBotProtection"]
  N --> RD["Redactor (password, cookies)"]
Component Path Role
Navigation loop src/navigate.ts navigate(): browser setup, per-step extraction, Jev call, stop gates, action dispatch, final payload
Action space and helpers src/lib.ts buildActionSpace, buildCriteria, typing-provider selection, bot-protection detector, seed-cookie resolution
Question catalog src/questions.ts stepQuestions (fan-out of three judgments), selectOptionQuestion
Jev transport src/provider.ts Thin wrapper over @jkudish/jev-agent-tools (ask, resolveTransport)
Secrets src/password.ts Origin parsing, handoff-file and env readers, secret validation, multi-encoding redactor
MCP server src/index.ts Registers jev_navigate, resolves secret references, picks stdio or HTTP
HTTP mode src/http.ts Stateless MCP over HTTP with bearer auth and a concurrency cap
CLI / library src/cli.ts, src/library.ts run subcommand; side-effect-free navigate export

How a request flows

A call to jev_navigate with a task and start URL:

  1. MCP entry. The handler validates arguments with Zod, turns password_file/password_env and cookie_file/cookie_env references into values, and calls navigate() with the MCP request’s abort signal. The JSON result and the screenshot come back as text and image content (index.ts).
  2. Guards and budget. Before any timer or browser starts, navigate() refuses unsafe combinations, such as recording on a credential run or cookies on an injected page, validates secrets and builds one redactor for all of them. It then arms a single AbortController for the wall-clock deadline (default 180 s) and caller cancellation (navigate.ts).
  3. Browser. It launches headless Chromium with a 1024x640 viewport, seeds cookies, lowers Playwright’s default timeout to 8 s and navigates to the start URL (navigate.ts).
  4. Per-step observation. Each of up to 24 steps first probes for a Cloudflare wall. Then extractAndStamp scans a[href], button, input, textarea, select and button/link/textbox roles, resolves accessible names, and stamps each candidate with data-jev-id (navigate.ts). buildActionSpace drops noise, dedupes links by destination, assigns e1..eN and a kind (click, type, search, submit, select, fill_password), and caps the list at 240 (lib.ts).
  5. Judgment. The state sent to Jev is the task, URL and title, a 1,500-character excerpt of visible text (or the open modal), the element descriptions and the action history. stepQuestions asks the action Choice plus goal_done and stuck Nouls in one call (questions.ts).
  6. Stop gates first. A done choice, goal_done > 0.85, or stuck > 0.85 after step 2 ends the run before the proposed action runs. If the same action just had no effect, the loop switches to the next-best option in the distribution (navigate.ts).
  7. Execute. Actions map to Playwright calls on [data-jev-id=...]: page.click, page.fill, fill plus Enter for search fields, selectOption by DOM index after a second Jev call, goBack, or scrollBy (navigate.ts).
  8. Settle and record. settle waits for two identical DOM fingerprints, capped at 1.5 s. The loop adopts any newly opened tab, compares URL, title, text length and scroll position to label the outcome, and appends a StepRecord (navigate.ts).
  9. Finish. It extracts the final payload in the requested format, takes a JPEG screenshot unless a credential was used, and passes the whole result through the redactor (navigate.ts).

Key components

Fan-out judgments as watchers

The three questions cannot see each other’s answers, so goal_done is an independent check on the action choice rather than a justification of it. Each step costs one primary Jev call. A select action adds a second call over the option labels. The 240-element cap keeps the action Choice under Jev’s 255-option limit with room for the four controls (scroll_down, scroll_up, back, done) (lib.ts).

Typing model

createTypingGenerator picks OpenAI, OpenRouter, Anthropic, Google or a custom OpenAI-compatible endpoint, based on which key it finds first (navigate.ts). generateTextToType asks for “ONLY the exact text to type”, with tight output caps and reasoning switched off where the provider allows it (navigate.ts). If it fails, ordinary fields stay empty and search boxes fall back to a keyword heuristic. Either way the run is marked degraded with a warning code. That is the right call: a guessed value in a username field looks like progress but is not.

Credential handling

fill_password actions appear only when a password source and origin are configured. The fill runs as one in-page task that checks the input is still a connected password field on the trusted origin before setting the value (navigate.ts). The task string itself is redacted before any model sees it. Capture windows are widened by the longest encoded form of the secret so that an echo cannot be cut in half by truncation. Video recording and Playwright debug output are refused on credential runs.

Bot-protection detection

detectBotProtection uses only page content (title and body markers) to decide whether to stop. A cf-mitigated header is recorded as evidence but never stops a run on its own. A challenge gets up to 8 seconds to clear itself. If it does not, the run ends as blocked with guidance (lib.ts). This covers Cloudflare only. There is no stealth, proxy or CAPTCHA support.

Extending it

  • Custom Jev carrier. Pass transport: { name, ask } to navigate() to route judgments through your own gateway. Built-in carriers (TypeSafe, OpenRouter, Cloudflare, Vercel) come from @jkudish/jev-agent-tools and can be forced with JEV_PROVIDER (provider.ts).
  • Bring your own page. navigate({ page }) reuses a caller’s Playwright page and context and never closes them.
  • Questions. All prompts live in src/questions.ts. Changing a question there changes every consumer.
  • New action kinds need changes in three places: extraction flags, buildActionSpace, and the dispatch chain in navigate(). There is no plugin hook.

Running it

  • MCP (stdio). npx -y @jkudish/jev-browser with TYPESAFE_API_KEY (or another carrier’s credentials). Chromium downloads on install.
  • HTTP. --http or JEV_BROWSER_TRANSPORT=http serves stateless MCP at /mcp. A non-loopback HOST requires JEV_BROWSER_AUTH_TOKEN, checked in constant time. In-flight runs are capped by JEV_BROWSER_MAX_CONCURRENCY (default 4), and each run gets its own Chromium (http.ts).
  • CLI. jev-browser run "<task>" <url> with --format, --max-steps, --screenshot, --record, --cookie-file name=@path and --password-file - for stdin.
  • Requires Node.js 22+, a Jev carrier key, and optionally a typing-provider key.

Strengths and caveats

  • Strength: cheap and bounded. One small structured call per step, with no screenshots in the loop and no free-text generation from the decision model. Every wait is clamped to the remaining deadline.
  • Strength: auditable. Every step records the proposed and executed action, the confidence and the goal/stuck scores, and stop gates fire before the action runs.
  • Strength: unusually careful secret handling for a project this size: reference-only delivery, origin binding, multi-encoding redaction and screenshot suppression.
  • Caveat: proprietary decision model. Jev is a hosted TypeSafe model, and the loop is shaped around its Choice/Noul API. You cannot switch in a general LLM for the judgments.
  • Caveat: shallow perception. The model sees at most 1,500 characters of page text and labels truncated to 60 characters. Tasks that need reading deep into a page rely on scrolling, and the final page content goes to the caller, not to the model.
  • Caveat: narrow action set. There are no hover, drag, file-upload, keyboard-shortcut, iframe-targeting or wait actions, and only one tab is followed at a time.
  • Caveat: approximate cost. est_cost_usd counts only Jev input tokens at a hardcoded jev-1.12 price. Typing-model spend is not included.

Sources: code at 1f0726c, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (10 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit 1f0726c. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

The page is represented to the model as a JSON state object with four parts: current page metadata, a text excerpt, interactive elements, and step history.

DOM extraction (src/navigate.ts:271-384). Elements are extracted from the live DOM — not the accessibility tree — because a11y trees under-report inputs. The CSS selector a[href], button, input, textarea, select, [role=button], [role=link], [role=searchbox], [role=textbox] captures candidates. Each is filtered by visibility (getClientRects, display, visibility), capped at 2000 raw candidates per step. Accessible label resolution follows AccName 1.2 §4.3.2: aria-labelledby, aria-label, native labels (label[for] and wrapping labels), placeholder, title. Each candidate is stamped with a data-jev-id attribute used as a CSS selector for actions.

Action-space pruning (src/lib.ts:69-149). Raw elements are filtered through buildActionSpace(): noise names (jump links, lone punctuation, generic nav text — 43 curated entries in JUNK_NAMES) and noise hrefs (#, javascript:, mailto:) are removed. Hrefs are deduplicated by destination. The cap is 240 elements total (matching Jev's 255-choice limit; src/lib.ts:4). Each surviving element gets a kind (click, type, search, submit, select, fill_password) based on tag, role, and type.

State structure (src/navigate.ts:800-810). Each step sends Jev: task (redacted), current_page (url + title), page_text_excerpt — on credential runs the fixed page-start excerpt; on normal runs the innermost open modal dialog text, or visible viewport text via a tree-walk capped at 5000 nodes — interactive_elements (id + description only), truncation flags, and history. The excerpt is capped at 1500 characters (STATE_EXCERPT_CHARS). Credential runs expand capture windows by the longest secret representation and redact all strings via redactCapped().

No screenshots are sent to the model. Screenshots are captured at the end for user display only (screenshot: final). There are no set-of-marks element indexes — the action space uses stable sequential IDs (e1, e2, etc.) mapped through data-jev-id attributes.

How are actions executed and how are elements targeted?

answered

Actions are executed through Playwright on a real headless Chromium browser. The browser is launched inside navigate() (src/navigate.ts:665), and every action runs against the live page via Playwright locators built from data-jev-id CSS attribute selectors.

Element targeting uses CSS attribute selectors ([data-jev-id="j1"]) created by selectorFor() (src/lib.ts:164-166). Elements are stamped with data-jev-id during DOM extraction and resolved to Playwright locators. There are no screenshot-based element indexes or set-of-marks coordinates — targeting is purely DOM-driven.

Action dispatch happens in the step loop (src/navigate.ts:868-1012):

  • click — page.click(selector) with a timed timeout (src/navigate.ts:1010)
  • type — page.fill(selector, text) without pressing Enter; the typing generator (a small LLM via the Vercel AI SDK) produces the text first (src/navigate.ts:903). No keyboard simulation is used for search fields.
  • search — page.fill() then page.press("Enter") in one step (src/navigate.ts:932-933)
  • submit — either page.click() on the submit button, or page.press("Enter") on a single-line text field (src/navigate.ts:938-944)
  • select — page.selectOption(selector, { index }) by live DOM index, not by label, so redacted labels never become the selection key (src/navigate.ts:964). A second Jev call picks the option.
  • scroll — window.scrollBy() at 80% of viewport height (src/navigate.ts:878-879)
  • back — page.goBack() (src/navigate.ts:874)
  • fill_password — the password value is injected through the native value setter (Object.getOwnPropertyDescriptor(HTMLInputElement.prototype, "value").set) with input/change events, all within a single JS task to prevent interleaving (src/navigate.ts:981-994). The value never reaches the model.

Tabs: when a click opens a new tab, the context-level page event adopts the new page for subsequent steps (src/navigate.ts:1020-1024). File uploads are never offered (src/navigate.ts:92 in buildActionSpace skips file inputs). Typing is separate from Jev — Jev returns only decisions; text generation uses a configurable small model (OpenAI/OpenRouter/Anthropic/Google) or falls back to a keyword heuristic for search fields.

How is the agent loop / planning implemented?

answered

The agent loop is a code-owned step loop in the navigate() function (src/navigate.ts:757-1066). Code controls the loop structure, budgets, and stop gates; the Jev model provides three independent judgments per step.

Step loop: a for loop runs up to maxSteps (default 24, src/navigate.ts:469). Each iteration:

  1. Check remaining time budget (deadline via AbortController, src/navigate.ts:759)
  2. Bot-protection probe — DOM-only detection of Cloudflare interstitials; if found, a settle window (up to 8s) checks if it auto-clears, otherwise the run stops with blocked (src/navigate.ts:769-777)
  3. Extract page — DOM extraction + action-space building (src/navigate.ts:779-794)
  4. Ask Jev — one primary call with three questions via the fan-out pattern (src/navigate.ts:811-823): a Choice (which action), a goal_done Noul (is the goal achieved?), and a stuck Noul (is progress stalled?). The questions are independent — they cannot see each other's answers (src/questions.ts:14-26)
  5. Stop gates — checked BEFORE action execution: done choice, goal_done > 0.85, stuck > 0.85 after step 2 (src/navigate.ts:827-841)
  6. Repeat-no-op recovery — if the same action repeated with no effect, switch to the next-best option from the Choice distribution (src/navigate.ts:846-857)
  7. Execute action — dispatch to Playwright (src/navigate.ts:868-1012)
  8. Settle — wait for DOM stability (two identical fingerprints) with a 1.5s cap (src/navigate.ts:1019)
  9. Tab adoption — if a new tab opened, switch to it (src/navigate.ts:1020-1024)
  10. Compare state — detect page changes for the trace and repeat detection (src/navigate.ts:1027-1051)

Planner vs executor: there is no separate planner. Jev is the planner (one step ahead) and also the evaluator (goal/stuck). There is no memory between steps beyond history (the action/outcome pairs sent in each step's state) and lastRedundant/lastExecuted for repeat detection.

Select actions add a second Jev call for the dropdown option (src/navigate.ts:953-961).

Stop conditions (checked in order): agent chooses done, goal_done > 0.85, stuck > 0.85 and step > 2, max_steps, timeout, or blocked. The done and goal_achieved statuses are independent judgments — agreement between them is a trustworthy finish.

How are failures, retries and self-healing handled?

answered

Reliability is addressed through several layered mechanisms, but there is no explicit retry loop or caching of successful actions. The design philosophy is "code owns control flow; Jev owns the judgments" (src/navigate.ts:1), so failures are handled deterministically.

Error classes caught (src/navigate.ts:1013-1017): Playwright action errors (timeouts, element not found), typing generator failures, and InvalidJevAnswer (malformed second-stage responses). Deadline/timer errors propagate via AbortController.signal (src/navigate.ts:552). InvalidJevAnswer from the primary call is not caught in the action handler — it propagates to the outer try/catch and stops the run. Typing failures produce structured warnings (src/navigate.ts:589-604).

Repeat-no-op recovery (src/navigate.ts:846-857): when the same action repeats with no visible effect (page unchanged, or a fill happened), the code switches to the next-best option from the Choice distribution via pickAlternate() (src/lib.ts:169-180). There is deliberately no low-confidence override — split probability across similar elements is treated as several acceptable alternatives, not uncertainty (README:419-420).

Stop gates run before execution (src/navigate.ts:827-841): done, goal_done > 0.85, and stuck > 0.85 (after step 2) are checked pre-action, so the run stops before spending time on a doomed step.

Timeouts (src/navigate.ts:552-570): a wall-clock deadline timer is set at run start; an AbortController feeds an AbortSignal through Jev calls, typing generation, and every Playwright timeout. Each bounded timeout is capped by remaining() (src/navigate.ts:570), so no single wait can outlive the run budget. Default maxSeconds is 180 (src/navigate.ts:473).

Bot-protection detection (src/lib.ts:378-515 and src/navigate.ts:718-776): DOM-only marker-based Cloudflare interstitial detection. If an interstitial appears, the run gets a short settle window (8s) to auto-pass before being declared blocked. Markers include branded challenge titles, Ray IDs, error codes, and body phrases — requires page-level evidence to stop, never just a response header. Post-action the final page gets the same settle check to catch walls that appeared during judgment.

What is NOT implemented: there is no retry queue for failed actions, no caching of successful action sequences, no replay of prior workflows, no explicit self-healing beyond the repeat-no-op alternate-picking. Credential runs suppress screenshots and refuse video recording so secrets cannot leak through pixels (src/navigate.ts:1113-1115).

Which models are supported and how are they called?

answered

The package uses two separate model systems: Jev for decisions and a small LLM for typing.

Jev judgment model (src/provider.ts and the @jkudish/jev-agent-tools package): Jev is a specialized decision model from TypeSafe that returns only typed answers — Choices (with probabilities over criteria) and Nouls (a scalar in [0,1]). It never generates text. Four built-in carriers are tried in order: TypeSafe (direct, via TYPESAFE_API_KEY), OpenRouter, Cloudflare Workers AI, and Vercel AI Gateway. JEV_PROVIDER forces one. The JEV_BROWSER_MODEL env var pins a version (default jev-latest; on OpenRouter maps to typesafe/jev-1.13). The transport is resolved per run (src/navigate.ts:547), not at import. Library callers can inject a custom JevTransport (src/provider.ts:17-28).

Jev calls use structured questions built from choice() and noul() from the @typesafe-ai/sdk (src/questions.ts:4). Each call returns validated answers before any action executes (src/provider.ts:18-23). Invalid answers raise InvalidJevAnswer and stop the run.

Typing model (src/navigate.ts:151-253): a second, independent model generates text for type and search actions. Configuration is through JEV_BROWSER_TYPE_* vars, separate from JEV_PROVIDER. Supported providers: OpenAI (default: gpt-5.6-luna), OpenRouter (default: google/gemini-2.5-flash-lite), Anthropic (default: claude-haiku-4.5), Google (default: gemini-2.5-flash). Auto-detection picks the first recognized key (src/lib.ts:252-257). A custom JEV_BROWSER_TYPE_BASE_URL can point at any OpenAI-compatible endpoint (Ollama, LM Studio, vLLM). Typing calls use the Vercel AI SDK (generateText) with tight output caps (48-256 tokens depending on provider; reasoning disabled on OpenRouter and non-Pro Gemini). A search action that fails falls back to a keyword heuristic (heuristicQuery() in src/lib.ts:189-196); a type action with no provider types nothing (producing a typing_generator_empty warning).

Vision requirement: Jev is not a vision model — it makes decisions from text-based state (element descriptions, page excerpts, history). Screenshots are taken for the user's result but never sent to the model. There is no vision dependency.

Editor's note. Correction: with no typing provider configured, a type action types nothing and records the warning code typing_fallback_no_provider (navigate.ts L887-L891); typing_generator_empty is reserved for a configured provider that returns empty text.

How are browser sessions, profiles, auth and anti-bot handled?

answered

Each run gets a fresh, isolated headless Chromium browser instance owned by the run (src/navigate.ts:665). The viewport is 1024×640. JEV_BROWSER_HEADED=1 makes it visible. Playwright's default timeout is lowered to 8s (src/navigate.ts:679).

Local vs remote: only local Playwright Chromium is supported. The HTTP transport (src/http.ts) runs the same local browser per request with concurrency capped at JEV_BROWSER_MAX_CONCURRENCY (default 4, src/http.ts:47). Each admitted request gets its own Chromium — no pooled browsers. The HTTP mode is stateless (no MCP sessions persisted). An auth token (JEV_BROWSER_AUTH_TOKEN) is required for non-loopback hosts (src/http.ts:50).

Persistent profiles and cookies: there are no long-lived browser profiles. The package supports seed cookies (src/lib.ts:516-609 and src/navigate.ts:495-508) — cookies injected into the context before the first navigation so the run starts already logged in. Values are delivered by reference (file path or env var name, never in argv/tool args), validated (non-empty, no control characters, 4-4096 bytes), and redacted from all outputs. Default cookie attributes: host-only (exact domain match), path /, httpOnly: true, sameSite: Lax, secure: true on https start URLs. __Host- names enforce these at runtime. Seed cookies are refused on injected pages (would mutate the caller's context).

Auth/password fill (src/password.ts and src/navigate.ts:511-528,967-1008): passwords are code-injected into native password fields, never sent to the model. Delivery channels: one-shot handoff files (validated 600 mode, single link, inside a 700 handoff directory), JEV_PASSWORD_* env vars, or stdin (CLI only). The value is bound to a specific origin (e.g. https://acme.com, http only on localhost). A fill never presses Enter — submission is a separate action. Redaction covers raw, percent-encoded, HTML-entity, markdown-escaped, YAML-quoted, and whitespace-normalized echoes. Video recording and Playwright debug modes are refused on credential runs. Credential runs suppress the final screenshot.

Anti-bot/stealth (src/lib.ts:378-515): Cloudflare interstitial detection is built in. Marker-based detection (titles like "Just a moment...", body phrases like "Ray ID:", "Performance and security by Cloudflare") runs DOM-only. A detected challenge gets an 8s auto-clear window before the run is declared blocked. The cf-mitigated response header is tracked but never alone causes a stop. There is no stealth (no anti-fingerprinting beyond vanilla Playwright headless), no proxy support, and no CAPTCHA-solving. The guidance for blocked runs recommends reusing a clearance-bearing browser session or using the site's API (src/lib.ts:442-446).

Injected pages (src/navigate.ts:666-667): callers can pass an existing Playwright Page to reuse their own browser, context, and session. Jev Browser never closes an injected page, context, or browser.