# tbphp/gpt-load

> Self-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/tbphp/gpt-load (reviewed at commit `a5c691bb0559f856ebb587879d82e341ff480202`, 2026-10-06)
- Stars: 7045 · Language: Go · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/gpt-load/

## Overview

GPT-Load is a single-binary, self-hosted gateway for people and small teams who hold many upstream credentials: API keys, relay accounts and subscription logins such as Codex, Claude, Antigravity or Grok. Clients use one base URL and one GPT-Load AccessKey, and keep their native protocol (OpenAI Chat Completions or Responses, Anthropic Messages, Gemini, embeddings, images, rerank). The gateway decides which credential in which group serves each request, retries on another credential when one fails, puts failing credentials on cooldown, and records usage and an estimated cost per request.

The 2.x code base at this commit is large and strongly typed. Its center is the credential scheduler, not the protocol translator. Each group binds one upstream "channel" (about 40 modules, from OpenAI and Bedrock to ZhipuAI, SiliconFlow and other gateways such as new-api and sub2api). A request can go out **native**, with the client's wire format kept as is, or **converted** to another provider's format. GPT-Load does not write those conversions itself. It imports Bifrost's Go provider packages (`github.com/maximhq/bifrost/core`) for converted API-key routes. Subscription channels go through an embedded build of CLIProxyAPI (the `cpa` executor).

Configuration lives in a database (SQLite, MySQL or PostgreSQL via gorm). It is compiled into an immutable `ConfigSnapshot` that the data plane reads through an atomic pointer, so admin changes take effect without restarts.

## Architecture

```mermaid
flowchart LR
  C["Client (OpenAI / Anthropic / Gemini SDK)"] --> G["gin data-plane module"]
  G --> A["authenticate: AccessKey hash, expiry, CIDR"]
  A --> H["Handler.Handle"]
  H --> L["Concurrency, cost quota, RPM"]
  L --> D["Dialect.InspectRequest"]
  D --> AF["Affinity: prompt prefix / cache key"]
  AF --> S["scheduler.Iterator.Next"]
  S --> X["Executor: native, Bifrost-converted, CPA"]
  X --> U["Upstream provider or account"]
  U --> J["health.JudgeExecution"]
  J -->|"retry next / refresh"| S
  J --> R["Request log, usage, price receipt"]
  SN["ConfigSnapshot (atomic)"] -.-> A
  SN -.-> S
```

| Component | Path | Role |
|---|---|---|
| Gateway | `internal/gateway/` | Data-plane routes, auth, `Handle`, attempt loop, affinity, request recording |
| Dialects | `internal/dialect/` | Per-protocol request inspection: operation, model, stream flag, affinity prefix |
| Scheduler | `internal/scheduler/` | Candidate groups and credentials, priority tiers, weighted fair selection |
| Health | `internal/health/` | Classifies each attempt into retry directive and cooldown/blacklist effect |
| Executors | `internal/execution/` | `bifrost` (converted and native HTTP), `cpa` (subscription accounts), websocket and live paths |
| Channels | `internal/channel/` | About 40 provider modules: base URLs, connection types, route capabilities |
| State | `internal/state/` | Immutable `ConfigSnapshot`, `Manager` with atomic publish |
| Accounting | `internal/accessquota/`, `internal/pricing/`, `internal/requestlog/` | Spend caps, price tables, per-request receipts and logs |
| Control and UI | `internal/control/`, `internal/webui/`, `web/` | Admin API and the embedded Vite/TypeScript UI |

## How a request flows

For `POST /v1/chat/completions` with `Authorization: Bearer <AccessKey>`:

1. **Route and authenticate.** The data module registers each endpoint from a fixed catalog (`/v1/chat/completions`, `/v1/messages`, `/v1beta/models/:model_action`, `/v1/responses`, embeddings, images, `/v1/systemone`) ([router.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/router.go#L10-L22), [http_routes.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/http_routes.go#L25-L57)). `authenticate` hashes the presented key, looks it up in the current snapshot, and checks expiry and allowed peer CIDRs ([auth.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/auth.go#L25-L52)).
2. **Admission.** `Handle` sends usage, Codex Live, Mistral realtime and websocket requests to their own paths. For normal requests it acquires a concurrency slot for the key, checks the AccessKey cost quota, and applies the RPM limiter ([handler.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L419-L460)).
3. **Inspect.** The protocol's dialect parses the body into `RequestMetadata` (operation, model, stream, route requirement, affinity prefix). If the model is an "auto" model, `prepareAutoModel` asks a classifier which preset to use first.
4. **Build the query.** The handler collects candidate groups for the key's filters, captures active credential refs, and resolves affinity. A `previous_response_id` pins the credential that served the earlier response. Otherwise a hashed prompt prefix or `prompt_cache_key` gives a preferred credential ([handler.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L640-L740), [affinity.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/affinity.go#L19-L60)).
5. **Attempt loop.** `executeAttempts` sets up per-request redaction, prepares the request once per group (parameter overrides), and calls `iterator.Next()` for each attempt ([handler.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L898-L960)). `Next` walks route-mode tiers (native first, then converted, or both together under `weighted_mix`) and marks each chosen credential as tried ([scheduler.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/scheduler/scheduler.go#L298-L338)).
6. **Execute and judge.** The provider adapter for the group's channel sends the request. `health.JudgeExecution` turns the outcome into a retry directive (none, refresh the credential, or next candidate) and an effect (cooldown, model cooldown, blacklist). A cancelled client or a failed downstream write never triggers a retry ([execution_judge.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/health/execution_judge.go#L24-L60)).
7. **Record.** On success the response is streamed or written back. Usage is captured from the body or stream events, priced into a frozen receipt, charged to the quota, and written to the request log.

## Key components

### Fair scheduler

Within a priority tier, each eligible credential has a progress watermark in a scheduling ledger. The lowest progress wins, with ties broken by last-selected time and then ID. New members are admitted at the current baseline, so they do not inherit a backlog ([fair.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/scheduler/fair.go#L60-L90)). Weights scale how fast progress grows. This is deterministic weighted round-robin rather than random choice, so load stays even even at low request rates.

### Executors and Bifrost reuse

`internal/execution/bifrost` maps channel kinds to Bifrost provider constructors (`openai.NewOpenAIProvider`, `anthropic.NewAnthropicProvider`, `gemini`, `azure`, `bedrock`, and more) and uses Bifrost's schemas and provider utilities for converted routes ([sdk_defaults.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/execution/bifrost/sdk_defaults.go#L35-L70)). GPT-Load keeps selection, retries and health for itself and uses Bifrost only for wire conversion. The `cpa` package wraps CLIProxyAPI for OAuth subscription accounts. It is pulled in through a `replace` directive to `third_party/cpaembedded`.

### Config snapshots

`state.Manager` holds `atomic.Pointer[ConfigSnapshot]` plus a reconciler that prepares infrastructure, such as proxies and provider targets, before a new snapshot becomes visible ([manager.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/manager.go#L11-L25)). Reads on the hot path are lock-free.

### Quotas, pricing and logs

AccessKeys carry an RPM limit, a concurrency limit, a price multiplier and cost-limit rules (lifetime or periodic, in nano-USD). Price tables cover input, output, cache read and write, and context tiers. Each request gets a frozen pricing receipt, so later price changes do not rewrite history.

### Redaction and audit

Optional regex redaction rewrites or encrypts matching text in outbound bodies and restores encrypted tokens in responses. `requestaudit` is an experimental guardrail: it sends up to 24 KB of request text to a "Jev" decision model with admin-written rules, and blocks or warns.

## Extending it

- **Channels:** add a module under `internal/channel/modules/` and bind it to an executor in the provider-adapter registry.
- **Compatible relays:** the `openai_compatible` channel and relay modules (new-api, sub2api, cliproxyapi) cover most OpenAI-shaped upstreams without code changes.
- **Parameter overrides, model aliases and auto models:** these are configured per group in the UI and compiled into the snapshot.
- There is no plugin API. Behaviour changes need code.

## Running it

- `docker compose up -d` with the published image `ghcr.io/tbphp/gpt-load:2`. The default port is 3001 ([config.go](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/platform/config/config.go#L20-L30)). Extra localhost ports are exposed for OAuth callbacks when you log in subscription accounts ([docker-compose.yml](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/docker-compose.yml#L1-L30)).
- Data goes to SQLite under `DATA_DIR` by default. MySQL and PostgreSQL are supported. Credentials are encrypted locally.
- Native binaries run as a service on Windows and other platforms (`service_*.go`). Go 1.27 is needed to build.
- A 1.x install cannot be upgraded in place. 2.0 starts from fresh data.

## Strengths and caveats

- **Strength: credential-pool depth.** Fair weighted scheduling, failure classification, cooldowns, blacklisting, refresh-and-replay for subscription tokens, and response-ID pinning go well beyond a simple key-rotation proxy.
- **Strength: native-first routing.** Requests keep their native protocol when an upstream supports it, and are converted only when needed, which avoids lossy translation.
- **Strength: good accounting.** Per-request price receipts, cost caps per AccessKey, and logs kept in your own database.
- **Caveat: single node.** RPM counters, concurrency, affinity and health live in process. Running several replicas against one database does not share them.
- **Caveat: no tenant hierarchy.** AccessKeys are the only principal, with filters on groups, protocols and models. There are no users, teams or SSO.
- **Caveat: no OpenTelemetry or Prometheus export.** Observability means the built-in UI, logs and request IDs.
- **Caveat: subscription channels depend on third-party code.** Using consumer subscriptions through a gateway may also conflict with those providers' terms of service. Check them before you rely on it.

*Sources: code at a5c691b, verified Q&A.*

## How tbphp/gpt-load answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

**Routing** in GPT-Load is a multi-stage pipeline of dialect inspection, affinity resolution, and weighted-fair credential selection. The `handler.Handle()` method in `internal/gateway/handler.go:419` is the main orchestrator. It first resolves the client protocol via `dialect.Dialect.InspectRequest()`, obtaining a `RequestMetadata` with `Operation`, `RouteRequirement` and `Model`. A `scheduler.Query` is built and handed to `scheduler.Iterator.Next()` (`internal/scheduler/scheduler.go:298`), which selects one credential from a pool of candidate targets.

**Route modes** — two wire strategies: `RouteNative` (preserve the client protocol upstream) and `RouteConverted` (translate to a provider-neutral format). `RouteRequirement` (`execution.RouteRequirement`, `internal/execution/contracts.go:171`) declares which modes are acceptable. `execution.RouteMode` (`internal/execution/contracts.go:129-133`) records the selected mode.

**Route strategies** — two global policies in `internal/state/runtime_settings.go:51-55`: `RouteStrategyNativeFirst` (try native routes first, fall back to converted) and `RouteStrategyWeightedMix` (treat both modes equally). The scheduler (`scheduler.go:153-154`) sets `routeModeTiers` accordingly — `[[native], [converted]]` for native-first, `[[native, converted]]` for weighted-mix.

**Weighted-fair scheduling** — within a group, credentials are selected via a fairness scheduler (`scheduler/fair.go`). Each credential carries a `WeightManual`, and the `SchedulingLedger` tracks progress watermarks. The lowest-progress eligible credential wins (`scheduler/fair.go:71-79`), ensuring even load across weighted candidates.

**Fallbacks and retries** — after an upstream attempt, `health.JudgeExecution()` (`internal/health/execution_judge.go:27`) decides the retry directive: `RetryNone`, `RetryRefreshCredential`, or `RetryNextCandidate`. Retries increment an attempt counter and call `Next()` again, skipping already-tried credentials (`scheduler.go:330`). Cooldowns and blacklisting come from `health.Decision.Effect` (`internal/health/decision.go:69`).

**Health checks** — `health.StatsStore` (`internal/health/stats.go:1`) tracks per-credential success/failure buckets in 1-minute windows (5 buckets). Problems trigger credential cooldown (`handler.go:341-361`), model-level cooldown, or blacklisting after exceeding `BlacklistThreshold` consecutive failures.

**Affinity** — `affinity.go` (`internal/gateway/affinity.go:19`) derives a prompt-affinity key from the first user message. If a prior request used the same prefix, the same credential is preferred. Prompt-cache-key affinity is also supported. The affinity cache (`affinity.Cache`) sits in `handler.affinityCache`.

**Model aliases** — `state.ModelConfig` includes an `Alias` field (`internal/state/snapshot.go:75-78`). `externalModelName()` returns the alias if set, otherwise the model ID. The scheduler resolves the external model name against upstream model IDs via `SelectModel()`.


Citations: [internal/gateway/handler.go:419-425](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L419-L425) · [internal/scheduler/scheduler.go:298-338](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/scheduler/scheduler.go#L298-L338) · [internal/state/runtime_settings.go:51-56](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/runtime_settings.go#L51-L56) · [internal/health/execution_judge.go:27-80](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/health/execution_judge.go#L27-L80)

### How are different provider APIs unified? (answered)

**Protocol translation** is handled by the `dialect` package plus the `execution/bifrost` executor. The system defines 11 `protocol.Protocol` values (`internal/protocol/protocol.go:6-19`): `OpenAICompletions`, `OpenAIResponses`, `OpenAIImages`, `OpenAIEmbeddings`, `Anthropic`, `Gemini`, `GeminiEmbeddings`, `Mistral`, `CodexLive`, `Rerank`, and `Decisions`.

**Dialect pattern** — each protocol implements `dialect.Dialect` (`internal/dialect/dialect.go:36`), which has two methods: `Protocol()` returns its identity, and `InspectRequest()` parses a raw HTTP request into `RequestMetadata`. Implementations: `dialect.OpenAI` (`internal/dialect/openai.go`), `dialect.Anthropic` (`internal/dialect/anthropic.go`), `dialect.Gemini` (`internal/dialect/gemini.go`), `dialect.Mistral` (`internal/dialect/mistral.go`), and others for images, embeddings, and rerank. All JSON-based protocols share the `inspectJSONRequestFields()` helper.

**Two execution modes** — `RouteNative` sends the request using the client's original protocol wire format to an upstream that supports it. `RouteConverted` translates to a provider-neutral format via `execution/bifrost` (the Bifrost executor, `internal/execution/bifrost/`). The conversion layer includes `chat_conversion.go`, `compatible.go`, `compatible_stream.go`, `responses_passthrough.go`, etc. For example, an OpenAI-format chat request can be converted to Anthropic's format and vice versa.

**Provider adapter registry** — `provideradapter.Registry` (`internal/provideradapter/registry.go:41-45`) compiles `ProviderKind`-to-`execution.Executor` bindings. Each adapter implements `RouteCapabilityValidator` and declares which `execution.Operation` values it can handle natively vs converted. The `execution/bifrost/executor.go` is the primary multi-protocol executor.

**Streaming** — abstracted via `StreamEvent` (`internal/dialect/stream_event.go:5`), a provider-neutral SSE event representation. `StreamEventClassifier` (`stream_event.go:38`) classifies events as continue/completed/failed. `UsageStreamEventObserver` captures token usage from stream events. Multiple SSE transformers exist for native event gates (`bifrost/model_alias.go`).

**Tool calls and multimodal** — handled in the Bifrost conversion layer: `tool_compatibility.go`, `compatible_conversion.go`, and `images_conversion_test.go` cover function-calling conversion across wire formats. Multimodal (images in chat) goes through `openai_images.go`, `openai_images_multipart.go` for the dialect layer, and `gemini_images_test.go` for Gemini-specific image handling.

**Provider count** — the `channel.ID` type lists about 40 upstream provider channels (`internal/channel/spec/definition.go:15-51`), including OpenAI, Anthropic, Claude, Gemini, Grok, DeepSeek, Cohere, HuggingFace, Azure, AWS Bedrock, Google Vertex, Mistral, and various Chinese providers (ZhipuAI, Alibaba, SiliconFlow, MoonshotAI).

> **Editor's note.** Correction: converted routes use the external Bifrost Go library (github.com/maximhq/bifrost/core) — GPT-Load instantiates Bifrost's provider packages in internal/execution/bifrost/sdk_defaults.go — and subscription channels (Codex, Claude, Antigravity) run through an embedded CLIProxyAPI in internal/execution/cpa.

Citations: [internal/protocol/protocol.go:6-20](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/protocol/protocol.go#L6-L20) · [internal/dialect/dialect.go:36-39](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/dialect/dialect.go#L36-L39) · [internal/provideradapter/registry.go:41-45](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/provideradapter/registry.go#L41-L45) · [internal/dialect/stream_event.go:1-50](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/dialect/stream_event.go#L1-L50)

### How are API keys, users and tenants managed? (answered)

**Authentication** works through virtual API keys (`AccessKey`). The `authenticate()` function in `internal/gateway/auth.go:25-52` extracts a key from the `Authorization` header (Bearer) or `key` query parameter. It hashes the plaintext with a `keyHasher` and looks it up in `snapshot.AccessKeysByHash`. The key is validated against expiration (`ExpiresAtMS`) and allowed peer CIDRs (`AllowedPeerCIDRs`), returning an `AccessKeyView` (`internal/state/snapshot.go:190-203`).

**AccessKey model** — each key has: `ID`, `Name`, `KeyHash`, `KeyPrefix`/`KeySuffix` (for display), `Status` (active/disabled), `ExpiresAtMS`, `AllowedPeerCIDRs`, `RPMLimit`, `ConcurrencyLimit`, `PriceMultiplier`, `Filters` (on Groups/Protocols/Models), and `CostLimitRules`.

**Credential management** — upstream API keys are stored encrypted via `encryption.Service`. The `runtimeCredentialRegistry` interface (`handler.go:80-92`) manages credential lifecycle: `ActiveEncryptedCredentialDataIfMatch()`, `SetCooldownWithChange()`, `IncrFailure()`, `SetBlacklistedWithChange()`, `ClearFailure()`. Credentials belong to `Groups` — each group connects to one channel (provider) with its own params, models, and proxy config. Secrets are identified via `connection.Type` — either `"api_key"` or a subscription driver.

**No user/tenant model** — the system does not implement user/team/tenant hierarchies. AccessKeys are the sole principal. Key-level `Filters` control which groups, protocols, and models a key can access (`state.FilterSet`, `snapshot.go:110-114`), providing capability-based access control without multi-tenant abstractions.

**Credential encryption** — `execution_boundary.go:24-65` normalizes channel credentials: it validates against the channel's connection type, optionally delegates to a subscription driver (`subscriptionruntime.Runtime.Driver()`), and extracts `api_key` plus secret values (`client_id`, `client_secret`, `tenant_id`, etc.). Proxy configs can be per-credential, encrypted and integrity-checked via `proxy.go:12-53`.

**Admin UI** — the web frontend (`web/src/main.ts`) supports both a "classic" and "modern" frontend. The `internal/webui/` package serves the admin dashboard with `http_routes.go` for API endpoints and `page_routes.json`/`modern_page_routes.json` for route definitions. The `control` package provides the configuration API surface for CRUD operations on keys, groups, and credentials.


Citations: [internal/gateway/auth.go:25-52](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/auth.go#L25-L52) · [internal/state/snapshot.go:87-101](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/snapshot.go#L87-L101) · [internal/gateway/handler.go:80-92](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L80-L92)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Per-key rate limits** — two levels: RPM (requests per minute) and concurrency. RPM is enforced by `AccessKeyRPMLimiter` (`handler.go:68-71`), implemented via `ratelimit.AccessKeyRPM` backed by `rpm.Store`. The `limiter.Allow(accessKey.ID, accessKey.RPMLimit)` call at `handler.go:543` returns a `LimitDecision`. Concurrency is enforced by `handler.acquireRequestConcurrency()` (`internal/gateway/concurrency.go:15`) against both per-key and global limits (`GlobalConcurrencyLimit`), tracked in `ratelimit.Concurrency`.

**Token and spend budgets** — the `accessquota` package (`internal/accessquota/runtime.go`) implements AccessKey-level cost limits. Two rule kinds exist: `KindTotal` (lifetime spend cap) and `KindPeriodic` (time-windowed cap like daily/monthly). Each `Rule` stores `LimitNanoUSD`, `PeriodSeconds`, window tracking, and `UsedNanoUSD`. The `accessquota.Runtime` provides `Check()` (pre-flight) and `Complete()` (post-request deduction) methods. Budgets are checked in `handler.go:522-541` before forwarding.

**Price tables** — `pricing.Table` (`internal/pricing/types.go`) maps channel+model identities to nano-USD-per-million-token prices. `pricing.Price` tracks per-input, per-output, per-cache-read, and per-cache-write rates. `ContextTier` allows different rates once input exceeds a threshold. `pricing.Quote()` (`internal/pricing/quote.go`) prices a finalized `usage.Result`. Group-level and AccessKey-level `PriceMultipliers` are applied via `QuoteForModeWithMultipliers()` for surcharging/discounting.

**Usage accounting** — `usage.Result` (`internal/usage`) holds token counts (uncached input, cache read, cache write, output). Usage is captured from upstream responses and stream events via `usageCaptureBoundary` (`internal/gateway/usage_capture.go:20`). Extracted usage flows to `pricing.Receipt` (`internal/pricing/receipt.go`), which is a frozen record of pricing method, rates, and total cost at request time (schema v6).

**Usage storage** — the `requestlog` package persists detailed request logs with cost. `accessquota` keeps in-memory rolling state for budget enforcement. The `state.Manager` uses gorm-backed snapshots for durable configuration, and `state/loader` loads config from the database. RPM state persists through `rpm.Store`.

**Cost estimation & budget enforcement** — `handler.freezeAttemptPricing()` (`handler.go:132-155`) captures the price table and multipliers at request time. After execution, `recorder.estimatedCostNanoUSD()` feeds into `accessQuota.Complete()`. If quota is exhausted, `Check()` returns `Allowed: false` and the request is blocked before reaching a provider.


Citations: [internal/gateway/concurrency.go:15-38](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/concurrency.go#L15-L38) · [internal/accessquota/runtime.go:1-79](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/accessquota/runtime.go#L1-L79) · [internal/pricing/quote.go:13-56](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/pricing/quote.go#L13-L56) · [internal/gateway/handler.go:132-155](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L132-L155)

### How are caching and guardrails implemented? (answered)

**Prompt caching** — the system does not implement a general-purpose response cache. Instead, it provides **prompt affinity routing**: `dialect.inspectPromptAffinityPrefix()` (`internal/dialect/prompt_affinity.go:43`) extracts the first user message content as a cache key. The `affinity.Cache` (`gateway/affinity.go:19-60`) stores a mapping from prompt prefix to the last-used credential ID, enabling session stickiness for providers like Anthropic that support prompt caching natively. Prompt cache keys can also be explicitly set by clients via the `prompt_cache_key` field (`internal/dialect/prompt_cache.go:12`).

**Semantic caching** — not implemented. No vector stores, embedding-based cache lookups, or cache hit/miss logic for responses exist in the codebase.

**PII redaction** — the `requestredact` package (`internal/requestredact/redact.go:1`) applies configurable regex-based redaction to outbound request content before it reaches the upstream provider. Up to 64 rules are supported, with two modes: `ModeReplace` (substitute matching text) and `ModeEncrypt` (encrypt in place with a per-credential cipher). Rules are configured globally or per-access-key. Redaction operates on raw HTTP body bytes with a 128 MB limit (`maxTextBytes`). The `redact.TokenCipher` interface (`redact.go:50-54`) supports token-aware encryption so API keys embedded in text can be redacted without breaking tokenization.

**Moderation and prompt injection** — handled by the `requestaudit` package (`internal/requestaudit/audit.go:1`), described as "experimental Jev guardrails without rewriting content." It defines `Rule` objects with `Instructions` (a prompt for the Jev LLM judge), `Threshold` (0-1 confidence), and `Action` (`block` or `warn`). Request content (up to 24 KB, `MaxRequestBytes`) is sent to the Jev classification engine (`jev` package). The `requestaudit.Cache` (`handler.go:94`) caches audit results. The Jev module (`internal/jev/`) is a mini LLM service specifically for classification decisions.

**Jev Decisions endpoint** — the system exposes a native `/v1/systemone` endpoint (`internal/dialect/decisions.go:16`) for making fast classification decisions. It accepts a `state` (string/object/array) and `questions` object, returning structured decisions. This is used internally for guardrail evaluation.

**Hook points** — the `requestredact` and `requestaudit` configurations are compiled into the immutable `ConfigSnapshot` (`internal/state/snapshot.go:206-207`) and applied per-request. Redaction happens inside the Bifrost executor before upstream dispatch; audit happens during handler processing. No plugin system exists for custom guardrails — both are configuration-driven.


Citations: [internal/requestredact/redact.go:1-80](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/requestredact/redact.go#L1-L80) · [internal/requestaudit/audit.go:1-70](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/requestaudit/audit.go#L1-L70) · [internal/dialect/prompt_affinity.go:43-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/dialect/prompt_affinity.go#L43-L60) · [internal/gateway/affinity.go:19-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/affinity.go#L19-L60)

### How is it observed, deployed and scaled? (answered)

**Logging** — the system uses `logrus` with structured fields. Request logs are written via `telemetry.RequestLogSink` (`handler.go:485-510`) — every request produces a `requestRecorder` that collects operation, model, cost, status, and timing. The `requestlog` package (`internal/requestlog/`) persists logs with configurable retention (default 7 days, `maxRequestLogRetentionDays` 365, `internal/state/runtime_settings.go:59-61`). Multiple log stores are supported: access-key usage, credential window usage, group usage, RPM history, and quota history.

**Metrics** — `health.StatsStore` (`internal/health/stats.go:14`) tracks per-credential success/failure/problem counts in sliding 5-minute windows (1-minute buckets). RPM counters are stored in `rpm.Store` and queried via `ratelimit.AccessKeyRPM`. No OpenTelemetry export is implemented — metrics are consumed in-process for health decisions and exposed via the control API.

**Traces** — no distributed tracing (OpenTelemetry or similar). The system uses request IDs (`handler.newRequestID()`, `handler.go:464`) that flow in `X-GPTLoad-Request-ID` response headers and appear in all log entries, enabling request correlation.

**Deployment** — packaged as a single Go binary via Docker (`Dockerfile`). The build uses multi-stage: Node.js for the web frontend, then Go compilation. The final image is Alpine-based. A `docker-compose.yml` and `docker-compose.voice.yml` are provided for container orchestration. The binary supports Windows as a first-class target with service management (`service_windows.go`, `service_cli.go`).

**Concurrency model** — Go language with goroutines (green threads). The `gin` HTTP framework provides non-blocking request handling. The `concurrency.go` package tracks active requests per access-key and globally. `channel.RouteMode` and the scheduler are pure in-memory operations with no per-request blocking IO beyond the upstream HTTP call.

**Configuration management** — `state.Manager` (`internal/state/manager.go:11`) holds an atomic pointer to the current `ConfigSnapshot`. Snapshots are immutable and replaced atomically on publish via `sync/atomic.Pointer`. The `SnapshotReconciler` interface synchronizes infrastructure resources before a snapshot becomes visible. Configuration is loaded from a database via `state/loader` and refreshed in-place without restarts.

**High availability** — the system is single-process; there is no built-in clustering, leader election, or shared state between instances. HA would rely on running multiple instances behind a load balancer with a shared database for configuration persistence. RPM and concurrency limits are in-process and not shared across nodes.

**Health endpoints** — the `health` package exposes credential health stats via the control API, and the auth layer checks credential cooldown/blacklist state inline during scheduling.


Citations: [internal/health/stats.go:1-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/health/stats.go#L1-L60) · [internal/gateway/handler.go:94-112](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L94-L112) · [internal/state/manager.go:1-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/manager.go#L1-L60) · [Dockerfile:1-30](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/Dockerfile#L1-L30)
