ekzhang/openjev-sglang
FastAPI server that clones the Jev decision API on Qwen3.6 + SGLang by reading one-token label logprobs; no agent or browser code.
Overview
OpenJev is a small FastAPI service that reimplements the TypeSafe “Jev” decision API (POST /v1/systemone) on an open model. You send a state (text, JSON, or a chat transcript) and up to 64 named questions of three kinds: Noul (yes/no probability), Choice (pick one of 2-64 options) and Score (expected position on an ordered rubric). The server answers each question by reading the next-token probabilities of answer labels from Qwen3.6-35B-A3B running in SGLang. It never generates more than one token.
It sits in the browser-control category because several Jev-based agents (computer-use and browser agents) call exactly this API to pick their next action. OpenJev itself is not an agent. It has no browser, no screen reading, no action loop and no tools. It is a self-hostable backend you could point such an agent at. The README now says the experiment is superseded: SGLang ships a native decisions endpoint built the same way.
The code is compact (about 1,000 lines in src/openjev/) and carefully defensive: strict request validation, admission limits, cancellation of sibling requests, and a supervisor that kills the container when SGLang dies.
Architecture
flowchart LR
C["Client (Jev-style request)"] --> G["RequestGuard: auth + body cap"]
G --> API["FastAPI /v1/systemone"]
API --> SVC["EvaluationService"]
SVC --> PC["PromptCompiler: chat template + labels"]
SVC --> BK["SGLangClient /generate"]
BK --> SG["SGLang server (localhost)"]
SG --> M["Qwen3.6-35B-A3B NVFP4"]
SVC --> SC["scoring: softmax over label logprobs"]
RT["runtime: launch + watchdog"] --> SG
MOD["modal_app.py (B200)"] --> RT
| Component | Path | Role |
|---|---|---|
| HTTP app | src/openjev/api.py |
create_app, RequestGuard middleware, routes, error mapping, timing headers |
| Evaluation service | src/openjev/service.py |
Admission control, token budgets, prefix warm-up, parallel branches |
| Prompt compiler | src/openjev/prompts.py |
Renders the chat template once, splits prefix/suffix, picks single-token labels |
| Backend client | src/openjev/backend.py |
One-token /generate calls with selected-token logprobs, abort on cancel |
| Scoring | src/openjev/scoring.py |
Temperature softmax, Noul/Choice/Score answers, entropy confidence |
| Runtime | src/openjev/runtime.py, launch.py |
Builds the SGLang command line, starts/stops the process group, watchdogs |
| Settings | src/openjev/config.py, profiles.py |
OPENJEV_* env settings and the single qwen36 profile |
| Deployment | modal_app.py |
Modal Server on one B200 with a persistent weights volume |
| Evals | evals/ |
BoolQ and MMLU-Pro comparisons against hosted Jev |
How a request flows
- Guard.
RequestGuardchecks the optional Bearer key with a constant-time compare and buffers the body, returning 413 once it passesmax_body_bytes, even for chunked uploads (api.py). - Route. The
systemonehandler startsservice.evaluate()and a disconnect watcher side by side. If the client leaves first, it raises 499 and cancels the work (api.py). - Admit.
evaluate()accepts the configured model name, thejev-latestalias or anyjev-*name. It rejects with 529 whenmax_concurrent_requestsevaluations are active, and turns a deadline overrun into 504 (service.py). - Compile.
PromptCompiler.prepareappends a user turn containing a random marker, renders the whole conversation once with thinking disabled, and splits on the marker. The prefix is tokenized once. Each question becomes a suffix withQuestion:, letteredOptions:andAnswer:(prompts.py). - Budget. Before any GPU work,
_evaluaterejects a branch longer thanmax_input_tokens(counting the output token) or a total overmax_total_input_tokens(service.py). - Warm, then fan out. It sends the bare prefix to SGLang and waits, so the radix cache holds it. Then all branches go out concurrently with
asyncio.gather. Any failure cancels the siblings (service.py). - Read logprobs. Each
generate()call asks formax_new_tokens=1andtoken_ids_logprobfor the label tokens only. On cancel or timeout it posts/abort_request(backend.py).parse_generationinsists on exactly one output token and finite label values (backend.py). - Score and reply.
answer()softmaxes the label logprobs (divided bytemperature) and builds the typed answer. The response carriesx-openjev-prefix-tokens, an optionalx-openjev-cached-tokensheader and aServer-Timingsplit of prepare, prefill and branches.
Key components
Single-token labels
Qwen splits numbers such as 10 into several tokens, so OpenJev does not number options. At startup the compiler walks A-Z, then AA, AB and so on, and keeps only strings that encode to exactly one distinct token and decode back unchanged. It refuses to start unless it finds 64 (prompts.py). Choice keys are hidden from the model unless an option’s description is null, in which case the key is shown as its meaning.
State handling
state_messages treats a list of {role, content} objects, or an object that is exactly {"messages": [...]}, as a real chat and keeps its roles. Anything else is serialized as one user message. Only text content is allowed (prompts.py). The classification instruction tells the model to treat instructions inside the state as material, not commands.
Scoring
normalize is a stable softmax over only the requested labels, so probabilities are conditional on the offered options. Noul returns P(true). Choice returns the argmax plus the full distribution. Score returns the probability-weighted zero-based index plus a legend. confidence is 1 - H(p)/log(n), a concentration measure, not calibration (scoring.py).
Runtime and supervision
backend_command builds the SGLang launch line: localhost only, FP8 KV cache, --mamba-radix-cache-strategy extra_buffer, breakable prefill CUDA graphs. It refuses extra flags that would disable radix caching or CUDA graphs (runtime.py). watch_process calls os._exit(1) if the child dies unexpectedly, so the platform replaces the container instead of serving a dead backend (launch.py). Every request, warm-ups included, asks for at least one token’s logprob, to stay clear of an SGLang crash when mixed batches occur (backend.py).
Extending it
- Another model. Add a
ModelProfiletoprofiles.py(weights, revision, image, backend flags). TheSettings.profileliteral currently allows onlyqwen36, so that type must widen too (config.py). The tokenizer must provide 64 single-token labels and a chat template that keeps the marker intact. - Existing backend.
openjev serve --connect URLskips process management and talks to a running SGLang that has the same model and selected-token logprobs. - Prompt changes. All prompt wording lives in
PromptCompiler.prepare.evals/prompt_probe.pyexists to compare variants.
Running it
- Modal.
uv run modal deploy modal_app.pycreates an unauthenticated Server on one B200 that scales to zero after five idle minutes. Weights and compile caches persist in a Modal Volume (modal_app.py). - Own GPU.
uv run openjev serve --sglang-python /path/to/sglang/bin/pythonstarts SGLang 0.5.19 from a separate CUDA environment. - Check.
uv run openjev smoke URLexercises all three answer types, a 64-option question and the 65-option rejection.
Strengths and caveats
- Strength: efficient by design. N questions cost N+1 one-token calls over a shared, cached prefix. No decoding loop.
- Strength: honest outputs. Strict response parsing, explicit
nullfor cache counts the backend omits, and documented meaning ofconfidence. - Strength: drop-in. Jev request and response shapes,
jev-*model aliases and TypeSafe-style/v1/modelsmean Jev clients can switch endpoints with little or no change. - Caveat: not a browser tool. Nothing here perceives or controls a UI. It only matters to this category as a backend for Jev-driven agents.
- Caveat: one model, heavy hardware. The only profile is Qwen3.6-35B-A3B NVFP4 tuned for B200-class GPUs. Changing models means code changes.
- Caveat: open by default. The Modal deployment is unauthenticated unless you set
OPENJEV_API_KEY. - Caveat: superseded. The authors point users to SGLang’s native decisions endpoint instead.
Sources: code at bf6a53b, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (10 pages), verified Q&A.
How it answers the Browser & computer control questions
Each answer was drafted by a code-reading agent at commit bf6a53b. Its citations were checked mechanically. Compare with the other browser & computer control →
How is the page represented to the model?
not applicableOpenJev does not represent any web page or browser view to a model. It is a structured classification server: it receives text or chat-message state, appends question instructions and answer options rendered as plain text labels (A: description, B: description, etc.), and sends the result to SGLang for single-token logit readout. The project has no DOM parser, no accessibility tree, no screenshot pipeline, no set-of-marks mechanism, and no element index. The only "pruning" is a configurable input-token cap per branch (default 32,768 tokens) and a total-token budget across all branches (default 262,144 tokens), enforced in service.py:61-67 before any GPU work.
How are actions executed and how are elements targeted?
not applicableOpenJev executes no actions on any UI or system. It takes no CDP commands, no Playwright calls, and no OS-level input events. The only "actions" are HTTP POST requests to SGLang's /generate endpoint, sending tokenized prompts with max_new_tokens=1 and reading back token-level logprobs. There is no browser element targeting (no selectors, no coordinates, no accessibility-label-based location), no typing, scrolling, file upload, or tab management. If a request disconnects, the server cancels pending sibling SGLang calls and attempts to abort them via POST /abort_request, but this is connection cleanup, not a computer-use action.
How is the agent loop / planning implemented?
not applicableThere is no agent loop, planner, executor, or multi-step reasoning chain in this project. Every request is a single stateless classification: render the chat template once, split into a common prefix and question-specific suffixes, send each branch as an independent one-token generation, and return the renormalized logprobs. The only loop is the concurrent-fan-out pattern in service.py:74-85 where all N branches fire simultaneously via asyncio.gather. There is no memory between requests, no step loop, no planning phase, and no tool-call schema for an LLM to invoke. The system has exactly one endpoint, /v1/systemone, for one-shot evaluation.
How are failures, retries and self-healing handled?
not applicableOpenJev's error handling is limited to infrastructure-level retries and client disconnect cleanup. SGLang timeouts are raised as BackendError(504), connection failures as BackendError(503), and non-OK HTTP statuses are mapped to error codes with Retry-After: 1 for 429/503/529. The service layer wraps the full evaluation in a top-level asyncio.timeout; exceeded deadlines raise 504. If a client disconnects mid-request, the server cancels all outstanding SGLang tasks and issues an abort request. The only retry is in the health-check polling during startup (runtime.py:101-108, polls every 1 second for up to 1200 seconds) and the smoke test's retry loop (smoke.py:17-28, polls every 2 seconds). There is no retry of failed inference calls for a single request, no self-healing, no caching of successful actions, and no replanning — the system is single-shot, not an agent.
Which models are supported and how are they called?
not applicableOpenJev targets exactly one model family: Qwen3.6-35B-A3B, configured as the default in defaults.py (nvidia/Qwen3.6-35B-A3B-NVFP4). One profile entry exists in profiles.py. There is no vision requirement — the model receives only token IDs, never images or screenshots. There is no structured output or tool-calling schema for the model itself to emit: the model never generates more than one token, and the answer is derived from logprobs of pre-selected label tokens, not from its generated response text. The server calls SGLang's /generate endpoint with raw input_ids and requests token_ids_logprob for candidate answer tokens. "Supported models" are effectively single-checkpoint deployments; the model is baked into the container at build time, not hot-swappable per request.
openjev-huggingface) and reused by later containers; the model is still fixed per deployment, not per request.How are browser sessions, profiles, auth and anti-bot handled?
not applicableOpenJev has no browser sessions, profiles, or auth state. It is a pure inference API with no persistent user context between requests. Each HTTP request is stateless: the backend stores nothing between evaluations. The only "authentication" is an optional shared Bearer token checked against a static OPENJEV_API_KEY environment variable, validated in the ASGI middleware (api.py:36-53). There are no cookies, no stealth measures, no proxy configuration, no CAPTCHA handling, and no persistent browser profiles. SGLang itself has a radix attention cache for KV reuse across requests, but this is an opportunistic performance optimization (not a session) and is explicitly documented as not being a "pinned per-request KV session" (README). The Modal deployment has no auth on the public endpoint by default.