LLMs Technical Reviews
Home / Browser & computer control / openjev-sglang

ekzhang/openjev-sglang

FastAPI server that clones the Jev decision API on Qwen3.6 + SGLang by reading one-token label logprobs; no agent or browser code.

GitHub ↗★ 336Pythoncommit bf6a53b · 2026-10-01homepage ↗

Overview

OpenJev is a small FastAPI service that reimplements the TypeSafe “Jev” decision API (POST /v1/systemone) on an open model. You send a state (text, JSON, or a chat transcript) and up to 64 named questions of three kinds: Noul (yes/no probability), Choice (pick one of 2-64 options) and Score (expected position on an ordered rubric). The server answers each question by reading the next-token probabilities of answer labels from Qwen3.6-35B-A3B running in SGLang. It never generates more than one token.

It sits in the browser-control category because several Jev-based agents (computer-use and browser agents) call exactly this API to pick their next action. OpenJev itself is not an agent. It has no browser, no screen reading, no action loop and no tools. It is a self-hostable backend you could point such an agent at. The README now says the experiment is superseded: SGLang ships a native decisions endpoint built the same way.

The code is compact (about 1,000 lines in src/openjev/) and carefully defensive: strict request validation, admission limits, cancellation of sibling requests, and a supervisor that kills the container when SGLang dies.

Architecture

flowchart LR
  C["Client (Jev-style request)"] --> G["RequestGuard: auth + body cap"]
  G --> API["FastAPI /v1/systemone"]
  API --> SVC["EvaluationService"]
  SVC --> PC["PromptCompiler: chat template + labels"]
  SVC --> BK["SGLangClient /generate"]
  BK --> SG["SGLang server (localhost)"]
  SG --> M["Qwen3.6-35B-A3B NVFP4"]
  SVC --> SC["scoring: softmax over label logprobs"]
  RT["runtime: launch + watchdog"] --> SG
  MOD["modal_app.py (B200)"] --> RT
Component Path Role
HTTP app src/openjev/api.py create_app, RequestGuard middleware, routes, error mapping, timing headers
Evaluation service src/openjev/service.py Admission control, token budgets, prefix warm-up, parallel branches
Prompt compiler src/openjev/prompts.py Renders the chat template once, splits prefix/suffix, picks single-token labels
Backend client src/openjev/backend.py One-token /generate calls with selected-token logprobs, abort on cancel
Scoring src/openjev/scoring.py Temperature softmax, Noul/Choice/Score answers, entropy confidence
Runtime src/openjev/runtime.py, launch.py Builds the SGLang command line, starts/stops the process group, watchdogs
Settings src/openjev/config.py, profiles.py OPENJEV_* env settings and the single qwen36 profile
Deployment modal_app.py Modal Server on one B200 with a persistent weights volume
Evals evals/ BoolQ and MMLU-Pro comparisons against hosted Jev

How a request flows

  1. Guard. RequestGuard checks the optional Bearer key with a constant-time compare and buffers the body, returning 413 once it passes max_body_bytes, even for chunked uploads (api.py).
  2. Route. The systemone handler starts service.evaluate() and a disconnect watcher side by side. If the client leaves first, it raises 499 and cancels the work (api.py).
  3. Admit. evaluate() accepts the configured model name, the jev-latest alias or any jev-* name. It rejects with 529 when max_concurrent_requests evaluations are active, and turns a deadline overrun into 504 (service.py).
  4. Compile. PromptCompiler.prepare appends a user turn containing a random marker, renders the whole conversation once with thinking disabled, and splits on the marker. The prefix is tokenized once. Each question becomes a suffix with Question:, lettered Options: and Answer: (prompts.py).
  5. Budget. Before any GPU work, _evaluate rejects a branch longer than max_input_tokens (counting the output token) or a total over max_total_input_tokens (service.py).
  6. Warm, then fan out. It sends the bare prefix to SGLang and waits, so the radix cache holds it. Then all branches go out concurrently with asyncio.gather. Any failure cancels the siblings (service.py).
  7. Read logprobs. Each generate() call asks for max_new_tokens=1 and token_ids_logprob for the label tokens only. On cancel or timeout it posts /abort_request (backend.py). parse_generation insists on exactly one output token and finite label values (backend.py).
  8. Score and reply. answer() softmaxes the label logprobs (divided by temperature) and builds the typed answer. The response carries x-openjev-prefix-tokens, an optional x-openjev-cached-tokens header and a Server-Timing split of prepare, prefill and branches.

Key components

Single-token labels

Qwen splits numbers such as 10 into several tokens, so OpenJev does not number options. At startup the compiler walks A-Z, then AA, AB and so on, and keeps only strings that encode to exactly one distinct token and decode back unchanged. It refuses to start unless it finds 64 (prompts.py). Choice keys are hidden from the model unless an option’s description is null, in which case the key is shown as its meaning.

State handling

state_messages treats a list of {role, content} objects, or an object that is exactly {"messages": [...]}, as a real chat and keeps its roles. Anything else is serialized as one user message. Only text content is allowed (prompts.py). The classification instruction tells the model to treat instructions inside the state as material, not commands.

Scoring

normalize is a stable softmax over only the requested labels, so probabilities are conditional on the offered options. Noul returns P(true). Choice returns the argmax plus the full distribution. Score returns the probability-weighted zero-based index plus a legend. confidence is 1 - H(p)/log(n), a concentration measure, not calibration (scoring.py).

Runtime and supervision

backend_command builds the SGLang launch line: localhost only, FP8 KV cache, --mamba-radix-cache-strategy extra_buffer, breakable prefill CUDA graphs. It refuses extra flags that would disable radix caching or CUDA graphs (runtime.py). watch_process calls os._exit(1) if the child dies unexpectedly, so the platform replaces the container instead of serving a dead backend (launch.py). Every request, warm-ups included, asks for at least one token’s logprob, to stay clear of an SGLang crash when mixed batches occur (backend.py).

Extending it

  • Another model. Add a ModelProfile to profiles.py (weights, revision, image, backend flags). The Settings.profile literal currently allows only qwen36, so that type must widen too (config.py). The tokenizer must provide 64 single-token labels and a chat template that keeps the marker intact.
  • Existing backend. openjev serve --connect URL skips process management and talks to a running SGLang that has the same model and selected-token logprobs.
  • Prompt changes. All prompt wording lives in PromptCompiler.prepare. evals/prompt_probe.py exists to compare variants.

Running it

  • Modal. uv run modal deploy modal_app.py creates an unauthenticated Server on one B200 that scales to zero after five idle minutes. Weights and compile caches persist in a Modal Volume (modal_app.py).
  • Own GPU. uv run openjev serve --sglang-python /path/to/sglang/bin/python starts SGLang 0.5.19 from a separate CUDA environment.
  • Check. uv run openjev smoke URL exercises all three answer types, a 64-option question and the 65-option rejection.

Strengths and caveats

  • Strength: efficient by design. N questions cost N+1 one-token calls over a shared, cached prefix. No decoding loop.
  • Strength: honest outputs. Strict response parsing, explicit null for cache counts the backend omits, and documented meaning of confidence.
  • Strength: drop-in. Jev request and response shapes, jev-* model aliases and TypeSafe-style /v1/models mean Jev clients can switch endpoints with little or no change.
  • Caveat: not a browser tool. Nothing here perceives or controls a UI. It only matters to this category as a backend for Jev-driven agents.
  • Caveat: one model, heavy hardware. The only profile is Qwen3.6-35B-A3B NVFP4 tuned for B200-class GPUs. Changing models means code changes.
  • Caveat: open by default. The Modal deployment is unauthenticated unless you set OPENJEV_API_KEY.
  • Caveat: superseded. The authors point users to SGLang’s native decisions endpoint instead.

Sources: code at bf6a53b, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (10 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit bf6a53b. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

not applicable

OpenJev does not represent any web page or browser view to a model. It is a structured classification server: it receives text or chat-message state, appends question instructions and answer options rendered as plain text labels (A: description, B: description, etc.), and sends the result to SGLang for single-token logit readout. The project has no DOM parser, no accessibility tree, no screenshot pipeline, no set-of-marks mechanism, and no element index. The only "pruning" is a configurable input-token cap per branch (default 32,768 tokens) and a total-token budget across all branches (default 262,144 tokens), enforced in service.py:61-67 before any GPU work.

How are actions executed and how are elements targeted?

not applicable

OpenJev executes no actions on any UI or system. It takes no CDP commands, no Playwright calls, and no OS-level input events. The only "actions" are HTTP POST requests to SGLang's /generate endpoint, sending tokenized prompts with max_new_tokens=1 and reading back token-level logprobs. There is no browser element targeting (no selectors, no coordinates, no accessibility-label-based location), no typing, scrolling, file upload, or tab management. If a request disconnects, the server cancels pending sibling SGLang calls and attempts to abort them via POST /abort_request, but this is connection cleanup, not a computer-use action.

How is the agent loop / planning implemented?

not applicable

There is no agent loop, planner, executor, or multi-step reasoning chain in this project. Every request is a single stateless classification: render the chat template once, split into a common prefix and question-specific suffixes, send each branch as an independent one-token generation, and return the renormalized logprobs. The only loop is the concurrent-fan-out pattern in service.py:74-85 where all N branches fire simultaneously via asyncio.gather. There is no memory between requests, no step loop, no planning phase, and no tool-call schema for an LLM to invoke. The system has exactly one endpoint, /v1/systemone, for one-shot evaluation.

How are failures, retries and self-healing handled?

not applicable

OpenJev's error handling is limited to infrastructure-level retries and client disconnect cleanup. SGLang timeouts are raised as BackendError(504), connection failures as BackendError(503), and non-OK HTTP statuses are mapped to error codes with Retry-After: 1 for 429/503/529. The service layer wraps the full evaluation in a top-level asyncio.timeout; exceeded deadlines raise 504. If a client disconnects mid-request, the server cancels all outstanding SGLang tasks and issues an abort request. The only retry is in the health-check polling during startup (runtime.py:101-108, polls every 1 second for up to 1200 seconds) and the smoke test's retry loop (smoke.py:17-28, polls every 2 seconds). There is no retry of failed inference calls for a single request, no self-healing, no caching of successful actions, and no replanning — the system is single-shot, not an agent.

Which models are supported and how are they called?

not applicable

OpenJev targets exactly one model family: Qwen3.6-35B-A3B, configured as the default in defaults.py (nvidia/Qwen3.6-35B-A3B-NVFP4). One profile entry exists in profiles.py. There is no vision requirement — the model receives only token IDs, never images or screenshots. There is no structured output or tool-calling schema for the model itself to emit: the model never generates more than one token, and the answer is derived from logprobs of pre-selected label tokens, not from its generated response text. The server calls SGLang's /generate endpoint with raw input_ids and requests token_ids_logprob for candidate answer tokens. "Supported models" are effectively single-checkpoint deployments; the model is baked into the container at build time, not hot-swappable per request.

Editor's note. Correction: weights are not baked into the container image. They are downloaded on first start into a persistent Modal Volume (openjev-huggingface) and reused by later containers; the model is still fixed per deployment, not per request.

How are browser sessions, profiles, auth and anti-bot handled?

not applicable

OpenJev has no browser sessions, profiles, or auth state. It is a pure inference API with no persistent user context between requests. Each HTTP request is stateless: the backend stores nothing between evaluations. The only "authentication" is an optional shared Bearer token checked against a static OPENJEV_API_KEY environment variable, validated in the ASGI middleware (api.py:36-53). There are no cookies, no stealth measures, no proxy configuration, no CAPTCHA handling, and no persistent browser profiles. SGLang itself has a radix attention cache for KV reuse across requests, but this is an opportunistic performance optimization (not a session) and is explicitly documented as not being a "pinned per-request KV session" (README). The Modal deployment has no auth on the public endpoint by default.