# ekzhang/openjev-sglang

> FastAPI server that clones the Jev decision API on Qwen3.6 + SGLang by reading one-token label logprobs; no agent or browser code.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/ekzhang/openjev-sglang (reviewed at commit `bf6a53bbb1f75040b06de39abff1230524553ab5`, 2026-10-01)
- Stars: 336 · Language: Python · License: n/a
- Canonical page: https://llms-technical-reviews.com/p/openjev-sglang/

## Overview

OpenJev is a small FastAPI service that reimplements the TypeSafe "Jev" decision API (`POST /v1/systemone`) on an open model. You send a `state` (text, JSON, or a chat transcript) and up to 64 named questions of three kinds: **Noul** (yes/no probability), **Choice** (pick one of 2-64 options) and **Score** (expected position on an ordered rubric). The server answers each question by reading the next-token probabilities of answer labels from Qwen3.6-35B-A3B running in SGLang. It never generates more than one token.

It sits in the browser-control category because several Jev-based agents (computer-use and browser agents) call exactly this API to pick their next action. OpenJev itself is **not** an agent. It has no browser, no screen reading, no action loop and no tools. It is a self-hostable backend you could point such an agent at. The README now says the experiment is superseded: SGLang ships a native decisions endpoint built the same way.

The code is compact (about 1,000 lines in `src/openjev/`) and carefully defensive: strict request validation, admission limits, cancellation of sibling requests, and a supervisor that kills the container when SGLang dies.

## Architecture

```mermaid
flowchart LR
  C["Client (Jev-style request)"] --> G["RequestGuard: auth + body cap"]
  G --> API["FastAPI /v1/systemone"]
  API --> SVC["EvaluationService"]
  SVC --> PC["PromptCompiler: chat template + labels"]
  SVC --> BK["SGLangClient /generate"]
  BK --> SG["SGLang server (localhost)"]
  SG --> M["Qwen3.6-35B-A3B NVFP4"]
  SVC --> SC["scoring: softmax over label logprobs"]
  RT["runtime: launch + watchdog"] --> SG
  MOD["modal_app.py (B200)"] --> RT
```

| Component | Path | Role |
|---|---|---|
| HTTP app | `src/openjev/api.py` | `create_app`, `RequestGuard` middleware, routes, error mapping, timing headers |
| Evaluation service | `src/openjev/service.py` | Admission control, token budgets, prefix warm-up, parallel branches |
| Prompt compiler | `src/openjev/prompts.py` | Renders the chat template once, splits prefix/suffix, picks single-token labels |
| Backend client | `src/openjev/backend.py` | One-token `/generate` calls with selected-token logprobs, abort on cancel |
| Scoring | `src/openjev/scoring.py` | Temperature softmax, Noul/Choice/Score answers, entropy confidence |
| Runtime | `src/openjev/runtime.py`, `launch.py` | Builds the SGLang command line, starts/stops the process group, watchdogs |
| Settings | `src/openjev/config.py`, `profiles.py` | `OPENJEV_*` env settings and the single `qwen36` profile |
| Deployment | `modal_app.py` | Modal Server on one B200 with a persistent weights volume |
| Evals | `evals/` | BoolQ and MMLU-Pro comparisons against hosted Jev |

## How a request flows

1. **Guard.** `RequestGuard` checks the optional Bearer key with a constant-time compare and buffers the body, returning 413 once it passes `max_body_bytes`, even for chunked uploads ([api.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L34-L78)).
2. **Route.** The `systemone` handler starts `service.evaluate()` and a disconnect watcher side by side. If the client leaves first, it raises 499 and cancels the work ([api.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L199-L235)).
3. **Admit.** `evaluate()` accepts the configured model name, the `jev-latest` alias or any `jev-*` name. It rejects with 529 when `max_concurrent_requests` evaluations are active, and turns a deadline overrun into 504 ([service.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L35-L51)).
4. **Compile.** `PromptCompiler.prepare` appends a user turn containing a random marker, renders the whole conversation once with thinking disabled, and splits on the marker. The prefix is tokenized once. Each question becomes a suffix with `Question:`, lettered `Options:` and `Answer:` ([prompts.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L97-L152)).
5. **Budget.** Before any GPU work, `_evaluate` rejects a branch longer than `max_input_tokens` (counting the output token) or a total over `max_total_input_tokens` ([service.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L53-L67)).
6. **Warm, then fan out.** It sends the bare prefix to SGLang and waits, so the radix cache holds it. Then all branches go out concurrently with `asyncio.gather`. Any failure cancels the siblings ([service.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L69-L85)).
7. **Read logprobs.** Each `generate()` call asks for `max_new_tokens=1` and `token_ids_logprob` for the label tokens only. On cancel or timeout it posts `/abort_request` ([backend.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L39-L84)). `parse_generation` insists on exactly one output token and finite label values ([backend.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L93-L123)).
8. **Score and reply.** `answer()` softmaxes the label logprobs (divided by `temperature`) and builds the typed answer. The response carries `x-openjev-prefix-tokens`, an optional `x-openjev-cached-tokens` header and a `Server-Timing` split of prepare, prefill and branches.

## Key components

### Single-token labels

Qwen splits numbers such as `10` into several tokens, so OpenJev does not number options. At startup the compiler walks `A`-`Z`, then `AA`, `AB` and so on, and keeps only strings that encode to exactly one distinct token and decode back unchanged. It refuses to start unless it finds 64 ([prompts.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L76-L95)). Choice keys are hidden from the model unless an option's description is `null`, in which case the key is shown as its meaning.

### State handling

`state_messages` treats a list of `{role, content}` objects, or an object that is exactly `{"messages": [...]}`, as a real chat and keeps its roles. Anything else is serialized as one user message. Only text content is allowed ([prompts.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L22-L48)). The classification instruction tells the model to treat instructions inside the state as material, not commands.

### Scoring

`normalize` is a stable softmax over only the requested labels, so probabilities are conditional on the offered options. Noul returns `P(true)`. Choice returns the argmax plus the full distribution. Score returns the probability-weighted zero-based index plus a legend. `confidence` is `1 - H(p)/log(n)`, a concentration measure, not calibration ([scoring.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/scoring.py#L7-L40)).

### Runtime and supervision

`backend_command` builds the SGLang launch line: localhost only, FP8 KV cache, `--mamba-radix-cache-strategy extra_buffer`, breakable prefill CUDA graphs. It refuses extra flags that would disable radix caching or CUDA graphs ([runtime.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/runtime.py#L18-L57)). `watch_process` calls `os._exit(1)` if the child dies unexpectedly, so the platform replaces the container instead of serving a dead backend ([launch.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/launch.py#L13-L33)). Every request, warm-ups included, asks for at least one token's logprob, to stay clear of an SGLang crash when mixed batches occur ([backend.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L52-L60)).

## Extending it

- **Another model.** Add a `ModelProfile` to `profiles.py` (weights, revision, image, backend flags). The `Settings.profile` literal currently allows only `qwen36`, so that type must widen too ([config.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/config.py#L12-L35)). The tokenizer must provide 64 single-token labels and a chat template that keeps the marker intact.
- **Existing backend.** `openjev serve --connect URL` skips process management and talks to a running SGLang that has the same model and selected-token logprobs.
- **Prompt changes.** All prompt wording lives in `PromptCompiler.prepare`. `evals/prompt_probe.py` exists to compare variants.

## Running it

- **Modal.** `uv run modal deploy modal_app.py` creates an unauthenticated Server on one B200 that scales to zero after five idle minutes. Weights and compile caches persist in a Modal Volume ([modal_app.py](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/modal_app.py#L47-L75)).
- **Own GPU.** `uv run openjev serve --sglang-python /path/to/sglang/bin/python` starts SGLang 0.5.19 from a separate CUDA environment.
- **Check.** `uv run openjev smoke URL` exercises all three answer types, a 64-option question and the 65-option rejection.

## Strengths and caveats

- **Strength: efficient by design.** N questions cost N+1 one-token calls over a shared, cached prefix. No decoding loop.
- **Strength: honest outputs.** Strict response parsing, explicit `null` for cache counts the backend omits, and documented meaning of `confidence`.
- **Strength: drop-in.** Jev request and response shapes, `jev-*` model aliases and TypeSafe-style `/v1/models` mean Jev clients can switch endpoints with little or no change.
- **Caveat: not a browser tool.** Nothing here perceives or controls a UI. It only matters to this category as a backend for Jev-driven agents.
- **Caveat: one model, heavy hardware.** The only profile is Qwen3.6-35B-A3B NVFP4 tuned for B200-class GPUs. Changing models means code changes.
- **Caveat: open by default.** The Modal deployment is unauthenticated unless you set `OPENJEV_API_KEY`.
- **Caveat: superseded.** The authors point users to SGLang's native decisions endpoint instead.

*Sources: code at bf6a53b, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (10 pages), verified Q&A.*

## How ekzhang/openjev-sglang answers the Browser & computer control questions

### How is the page represented to the model? (not applicable)

OpenJev does not represent any web page or browser view to a model. It is a structured classification server: it receives text or chat-message state, appends question instructions and answer options rendered as plain text labels (`A: description`, `B: description`, etc.), and sends the result to SGLang for single-token logit readout. The project has no DOM parser, no accessibility tree, no screenshot pipeline, no set-of-marks mechanism, and no element index. The only "pruning" is a configurable input-token cap per branch (default 32,768 tokens) and a total-token budget across all branches (default 262,144 tokens), enforced in `service.py:61-67` before any GPU work.


Citations: [src/openjev/prompts.py:97-152](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L97-L152) · [src/openjev/service.py:53-108](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L53-L108) · [src/openjev/models.py:66-106](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/models.py#L66-L106)

### How are actions executed and how are elements targeted? (not applicable)

OpenJev executes no actions on any UI or system. It takes no CDP commands, no Playwright calls, and no OS-level input events. The only "actions" are HTTP POST requests to SGLang's `/generate` endpoint, sending tokenized prompts with `max_new_tokens=1` and reading back token-level logprobs. There is no browser element targeting (no selectors, no coordinates, no accessibility-label-based location), no typing, scrolling, file upload, or tab management. If a request disconnects, the server cancels pending sibling SGLang calls and attempts to abort them via `POST /abort_request`, but this is connection cleanup, not a computer-use action.


Citations: [src/openjev/backend.py:27-91](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L27-L91) · [src/openjev/api.py:199-235](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L199-L235) · [src/openjev/runtime.py:60-81](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/runtime.py#L60-L81)

### How is the agent loop / planning implemented? (not applicable)

There is no agent loop, planner, executor, or multi-step reasoning chain in this project. Every request is a single stateless classification: render the chat template once, split into a common prefix and question-specific suffixes, send each branch as an independent one-token generation, and return the renormalized logprobs. The only loop is the concurrent-fan-out pattern in `service.py:74-85` where all N branches fire simultaneously via `asyncio.gather`. There is no memory between requests, no step loop, no planning phase, and no tool-call schema for an LLM to invoke. The system has exactly one endpoint, `/v1/systemone`, for one-shot evaluation.


Citations: [src/openjev/service.py:35-108](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L35-L108) · [src/openjev/prompts.py:97-152](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L97-L152) · [src/openjev/api.py:199-235](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L199-L235)

### How are failures, retries and self-healing handled? (not applicable)

OpenJev's error handling is limited to infrastructure-level retries and client disconnect cleanup. SGLang timeouts are raised as `BackendError(504)`, connection failures as `BackendError(503)`, and non-OK HTTP statuses are mapped to error codes with `Retry-After: 1` for 429/503/529. The service layer wraps the full evaluation in a top-level `asyncio.timeout`; exceeded deadlines raise 504. If a client disconnects mid-request, the server cancels all outstanding SGLang tasks and issues an abort request. The only retry is in the health-check polling during startup (`runtime.py:101-108`, polls every 1 second for up to 1200 seconds) and the smoke test's retry loop (`smoke.py:17-28`, polls every 2 seconds). There is no retry of failed inference calls for a single request, no self-healing, no caching of successful actions, and no replanning — the system is single-shot, not an agent.


Citations: [src/openjev/backend.py:39-91](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L39-L91) · [src/openjev/service.py:35-51](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/service.py#L35-L51) · [src/openjev/api.py:149-157](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L149-L157) · [src/openjev/runtime.py:101-109](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/runtime.py#L101-L109) · [src/openjev/smoke.py:17-28](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/smoke.py#L17-L28)

### Which models are supported and how are they called? (not applicable)

OpenJev targets exactly one model family: Qwen3.6-35B-A3B, configured as the default in `defaults.py` (`nvidia/Qwen3.6-35B-A3B-NVFP4`). One profile entry exists in `profiles.py`. There is no vision requirement — the model receives only token IDs, never images or screenshots. There is no structured output or tool-calling schema for the model itself to emit: the model never generates more than one token, and the answer is derived from logprobs of pre-selected label tokens, not from its generated response text. The server calls SGLang's `/generate` endpoint with raw `input_ids` and requests `token_ids_logprob` for candidate answer tokens. "Supported models" are effectively single-checkpoint deployments; the model is baked into the container at build time, not hot-swappable per request.

> **Editor's note.** Correction: weights are not baked into the container image. They are downloaded on first start into a persistent Modal Volume (`openjev-huggingface`) and reused by later containers; the model is still fixed per deployment, not per request.

Citations: [src/openjev/profiles.py:20-37](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/profiles.py#L20-L37) · [src/openjev/backend.py:39-60](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L39-L60) · [src/openjev/prompts.py:76-95](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/prompts.py#L76-L95)

### How are browser sessions, profiles, auth and anti-bot handled? (not applicable)

OpenJev has no browser sessions, profiles, or auth state. It is a pure inference API with no persistent user context between requests. Each HTTP request is stateless: the backend stores nothing between evaluations. The only "authentication" is an optional shared Bearer token checked against a static `OPENJEV_API_KEY` environment variable, validated in the ASGI middleware (`api.py:36-53`). There are no cookies, no stealth measures, no proxy configuration, no CAPTCHA handling, and no persistent browser profiles. SGLang itself has a radix attention cache for KV reuse across requests, but this is an opportunistic performance optimization (not a session) and is explicitly documented as not being a "pinned per-request KV session" (README). The Modal deployment has no auth on the public endpoint by default.


Citations: [src/openjev/api.py:34-78](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/api.py#L34-L78) · [src/openjev/config.py:12-50](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/config.py#L12-L50) · [src/openjev/backend.py:27-37](https://github.com/ekzhang/openjev-sglang/blob/bf6a53bbb1f75040b06de39abff1230524553ab5/src/openjev/backend.py#L27-L37)
