LLMs Technical Reviews
Home / LLM gateways / gpt-load

tbphp/gpt-load

Self-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas.

GitHub ↗★ 7.0kGoMITcommit a5c691b · 2026-10-06homepage ↗

Overview

GPT-Load is a single-binary, self-hosted gateway for people and small teams who hold many upstream credentials: API keys, relay accounts and subscription logins such as Codex, Claude, Antigravity or Grok. Clients use one base URL and one GPT-Load AccessKey, and keep their native protocol (OpenAI Chat Completions or Responses, Anthropic Messages, Gemini, embeddings, images, rerank). The gateway decides which credential in which group serves each request, retries on another credential when one fails, puts failing credentials on cooldown, and records usage and an estimated cost per request.

The 2.x code base at this commit is large and strongly typed. Its center is the credential scheduler, not the protocol translator. Each group binds one upstream “channel” (about 40 modules, from OpenAI and Bedrock to ZhipuAI, SiliconFlow and other gateways such as new-api and sub2api). A request can go out native, with the client’s wire format kept as is, or converted to another provider’s format. GPT-Load does not write those conversions itself. It imports Bifrost’s Go provider packages (github.com/maximhq/bifrost/core) for converted API-key routes. Subscription channels go through an embedded build of CLIProxyAPI (the cpa executor).

Configuration lives in a database (SQLite, MySQL or PostgreSQL via gorm). It is compiled into an immutable ConfigSnapshot that the data plane reads through an atomic pointer, so admin changes take effect without restarts.

Architecture

flowchart LR
  C["Client (OpenAI / Anthropic / Gemini SDK)"] --> G["gin data-plane module"]
  G --> A["authenticate: AccessKey hash, expiry, CIDR"]
  A --> H["Handler.Handle"]
  H --> L["Concurrency, cost quota, RPM"]
  L --> D["Dialect.InspectRequest"]
  D --> AF["Affinity: prompt prefix / cache key"]
  AF --> S["scheduler.Iterator.Next"]
  S --> X["Executor: native, Bifrost-converted, CPA"]
  X --> U["Upstream provider or account"]
  U --> J["health.JudgeExecution"]
  J -->|"retry next / refresh"| S
  J --> R["Request log, usage, price receipt"]
  SN["ConfigSnapshot (atomic)"] -.-> A
  SN -.-> S
Component Path Role
Gateway internal/gateway/ Data-plane routes, auth, Handle, attempt loop, affinity, request recording
Dialects internal/dialect/ Per-protocol request inspection: operation, model, stream flag, affinity prefix
Scheduler internal/scheduler/ Candidate groups and credentials, priority tiers, weighted fair selection
Health internal/health/ Classifies each attempt into retry directive and cooldown/blacklist effect
Executors internal/execution/ bifrost (converted and native HTTP), cpa (subscription accounts), websocket and live paths
Channels internal/channel/ About 40 provider modules: base URLs, connection types, route capabilities
State internal/state/ Immutable ConfigSnapshot, Manager with atomic publish
Accounting internal/accessquota/, internal/pricing/, internal/requestlog/ Spend caps, price tables, per-request receipts and logs
Control and UI internal/control/, internal/webui/, web/ Admin API and the embedded Vite/TypeScript UI

How a request flows

For POST /v1/chat/completions with Authorization: Bearer <AccessKey>:

  1. Route and authenticate. The data module registers each endpoint from a fixed catalog (/v1/chat/completions, /v1/messages, /v1beta/models/:model_action, /v1/responses, embeddings, images, /v1/systemone) (router.go, http_routes.go). authenticate hashes the presented key, looks it up in the current snapshot, and checks expiry and allowed peer CIDRs (auth.go).
  2. Admission. Handle sends usage, Codex Live, Mistral realtime and websocket requests to their own paths. For normal requests it acquires a concurrency slot for the key, checks the AccessKey cost quota, and applies the RPM limiter (handler.go).
  3. Inspect. The protocol’s dialect parses the body into RequestMetadata (operation, model, stream, route requirement, affinity prefix). If the model is an “auto” model, prepareAutoModel asks a classifier which preset to use first.
  4. Build the query. The handler collects candidate groups for the key’s filters, captures active credential refs, and resolves affinity. A previous_response_id pins the credential that served the earlier response. Otherwise a hashed prompt prefix or prompt_cache_key gives a preferred credential (handler.go, affinity.go).
  5. Attempt loop. executeAttempts sets up per-request redaction, prepares the request once per group (parameter overrides), and calls iterator.Next() for each attempt (handler.go). Next walks route-mode tiers (native first, then converted, or both together under weighted_mix) and marks each chosen credential as tried (scheduler.go).
  6. Execute and judge. The provider adapter for the group’s channel sends the request. health.JudgeExecution turns the outcome into a retry directive (none, refresh the credential, or next candidate) and an effect (cooldown, model cooldown, blacklist). A cancelled client or a failed downstream write never triggers a retry (execution_judge.go).
  7. Record. On success the response is streamed or written back. Usage is captured from the body or stream events, priced into a frozen receipt, charged to the quota, and written to the request log.

Key components

Fair scheduler

Within a priority tier, each eligible credential has a progress watermark in a scheduling ledger. The lowest progress wins, with ties broken by last-selected time and then ID. New members are admitted at the current baseline, so they do not inherit a backlog (fair.go). Weights scale how fast progress grows. This is deterministic weighted round-robin rather than random choice, so load stays even even at low request rates.

Executors and Bifrost reuse

internal/execution/bifrost maps channel kinds to Bifrost provider constructors (openai.NewOpenAIProvider, anthropic.NewAnthropicProvider, gemini, azure, bedrock, and more) and uses Bifrost’s schemas and provider utilities for converted routes (sdk_defaults.go). GPT-Load keeps selection, retries and health for itself and uses Bifrost only for wire conversion. The cpa package wraps CLIProxyAPI for OAuth subscription accounts. It is pulled in through a replace directive to third_party/cpaembedded.

Config snapshots

state.Manager holds atomic.Pointer[ConfigSnapshot] plus a reconciler that prepares infrastructure, such as proxies and provider targets, before a new snapshot becomes visible (manager.go). Reads on the hot path are lock-free.

Quotas, pricing and logs

AccessKeys carry an RPM limit, a concurrency limit, a price multiplier and cost-limit rules (lifetime or periodic, in nano-USD). Price tables cover input, output, cache read and write, and context tiers. Each request gets a frozen pricing receipt, so later price changes do not rewrite history.

Redaction and audit

Optional regex redaction rewrites or encrypts matching text in outbound bodies and restores encrypted tokens in responses. requestaudit is an experimental guardrail: it sends up to 24 KB of request text to a “Jev” decision model with admin-written rules, and blocks or warns.

Extending it

  • Channels: add a module under internal/channel/modules/ and bind it to an executor in the provider-adapter registry.
  • Compatible relays: the openai_compatible channel and relay modules (new-api, sub2api, cliproxyapi) cover most OpenAI-shaped upstreams without code changes.
  • Parameter overrides, model aliases and auto models: these are configured per group in the UI and compiled into the snapshot.
  • There is no plugin API. Behaviour changes need code.

Running it

  • docker compose up -d with the published image ghcr.io/tbphp/gpt-load:2. The default port is 3001 (config.go). Extra localhost ports are exposed for OAuth callbacks when you log in subscription accounts (docker-compose.yml).
  • Data goes to SQLite under DATA_DIR by default. MySQL and PostgreSQL are supported. Credentials are encrypted locally.
  • Native binaries run as a service on Windows and other platforms (service_*.go). Go 1.27 is needed to build.
  • A 1.x install cannot be upgraded in place. 2.0 starts from fresh data.

Strengths and caveats

  • Strength: credential-pool depth. Fair weighted scheduling, failure classification, cooldowns, blacklisting, refresh-and-replay for subscription tokens, and response-ID pinning go well beyond a simple key-rotation proxy.
  • Strength: native-first routing. Requests keep their native protocol when an upstream supports it, and are converted only when needed, which avoids lossy translation.
  • Strength: good accounting. Per-request price receipts, cost caps per AccessKey, and logs kept in your own database.
  • Caveat: single node. RPM counters, concurrency, affinity and health live in process. Running several replicas against one database does not share them.
  • Caveat: no tenant hierarchy. AccessKeys are the only principal, with filters on groups, protocols and models. There are no users, teams or SSO.
  • Caveat: no OpenTelemetry or Prometheus export. Observability means the built-in UI, logs and request IDs.
  • Caveat: subscription channels depend on third-party code. Using consumer subscriptions through a gateway may also conflict with those providers’ terms of service. Check them before you rely on it.

Sources: code at a5c691b, verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit a5c691b. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Routing in GPT-Load is a multi-stage pipeline of dialect inspection, affinity resolution, and weighted-fair credential selection. The handler.Handle() method in internal/gateway/handler.go:419 is the main orchestrator. It first resolves the client protocol via dialect.Dialect.InspectRequest(), obtaining a RequestMetadata with Operation, RouteRequirement and Model. A scheduler.Query is built and handed to scheduler.Iterator.Next() (internal/scheduler/scheduler.go:298), which selects one credential from a pool of candidate targets.

Route modes — two wire strategies: RouteNative (preserve the client protocol upstream) and RouteConverted (translate to a provider-neutral format). RouteRequirement (execution.RouteRequirement, internal/execution/contracts.go:171) declares which modes are acceptable. execution.RouteMode (internal/execution/contracts.go:129-133) records the selected mode.

Route strategies — two global policies in internal/state/runtime_settings.go:51-55: RouteStrategyNativeFirst (try native routes first, fall back to converted) and RouteStrategyWeightedMix (treat both modes equally). The scheduler (scheduler.go:153-154) sets routeModeTiers accordingly — [[native], [converted]] for native-first, [[native, converted]] for weighted-mix.

Weighted-fair scheduling — within a group, credentials are selected via a fairness scheduler (scheduler/fair.go). Each credential carries a WeightManual, and the SchedulingLedger tracks progress watermarks. The lowest-progress eligible credential wins (scheduler/fair.go:71-79), ensuring even load across weighted candidates.

Fallbacks and retries — after an upstream attempt, health.JudgeExecution() (internal/health/execution_judge.go:27) decides the retry directive: RetryNone, RetryRefreshCredential, or RetryNextCandidate. Retries increment an attempt counter and call Next() again, skipping already-tried credentials (scheduler.go:330). Cooldowns and blacklisting come from health.Decision.Effect (internal/health/decision.go:69).

Health checks — health.StatsStore (internal/health/stats.go:1) tracks per-credential success/failure buckets in 1-minute windows (5 buckets). Problems trigger credential cooldown (handler.go:341-361), model-level cooldown, or blacklisting after exceeding BlacklistThreshold consecutive failures.

Affinity — affinity.go (internal/gateway/affinity.go:19) derives a prompt-affinity key from the first user message. If a prior request used the same prefix, the same credential is preferred. Prompt-cache-key affinity is also supported. The affinity cache (affinity.Cache) sits in handler.affinityCache.

Model aliases — state.ModelConfig includes an Alias field (internal/state/snapshot.go:75-78). externalModelName() returns the alias if set, otherwise the model ID. The scheduler resolves the external model name against upstream model IDs via SelectModel().

How are different provider APIs unified?

answered

Protocol translation is handled by the dialect package plus the execution/bifrost executor. The system defines 11 protocol.Protocol values (internal/protocol/protocol.go:6-19): OpenAICompletions, OpenAIResponses, OpenAIImages, OpenAIEmbeddings, Anthropic, Gemini, GeminiEmbeddings, Mistral, CodexLive, Rerank, and Decisions.

Dialect pattern — each protocol implements dialect.Dialect (internal/dialect/dialect.go:36), which has two methods: Protocol() returns its identity, and InspectRequest() parses a raw HTTP request into RequestMetadata. Implementations: dialect.OpenAI (internal/dialect/openai.go), dialect.Anthropic (internal/dialect/anthropic.go), dialect.Gemini (internal/dialect/gemini.go), dialect.Mistral (internal/dialect/mistral.go), and others for images, embeddings, and rerank. All JSON-based protocols share the inspectJSONRequestFields() helper.

Two execution modes — RouteNative sends the request using the client's original protocol wire format to an upstream that supports it. RouteConverted translates to a provider-neutral format via execution/bifrost (the Bifrost executor, internal/execution/bifrost/). The conversion layer includes chat_conversion.go, compatible.go, compatible_stream.go, responses_passthrough.go, etc. For example, an OpenAI-format chat request can be converted to Anthropic's format and vice versa.

Provider adapter registry — provideradapter.Registry (internal/provideradapter/registry.go:41-45) compiles ProviderKind-to-execution.Executor bindings. Each adapter implements RouteCapabilityValidator and declares which execution.Operation values it can handle natively vs converted. The execution/bifrost/executor.go is the primary multi-protocol executor.

Streaming — abstracted via StreamEvent (internal/dialect/stream_event.go:5), a provider-neutral SSE event representation. StreamEventClassifier (stream_event.go:38) classifies events as continue/completed/failed. UsageStreamEventObserver captures token usage from stream events. Multiple SSE transformers exist for native event gates (bifrost/model_alias.go).

Tool calls and multimodal — handled in the Bifrost conversion layer: tool_compatibility.go, compatible_conversion.go, and images_conversion_test.go cover function-calling conversion across wire formats. Multimodal (images in chat) goes through openai_images.go, openai_images_multipart.go for the dialect layer, and gemini_images_test.go for Gemini-specific image handling.

Provider count — the channel.ID type lists about 40 upstream provider channels (internal/channel/spec/definition.go:15-51), including OpenAI, Anthropic, Claude, Gemini, Grok, DeepSeek, Cohere, HuggingFace, Azure, AWS Bedrock, Google Vertex, Mistral, and various Chinese providers (ZhipuAI, Alibaba, SiliconFlow, MoonshotAI).

Editor's note. Correction: converted routes use the external Bifrost Go library (github.com/maximhq/bifrost/core) — GPT-Load instantiates Bifrost's provider packages in internal/execution/bifrost/sdk_defaults.go — and subscription channels (Codex, Claude, Antigravity) run through an embedded CLIProxyAPI in internal/execution/cpa.

How are API keys, users and tenants managed?

answered

Authentication works through virtual API keys (AccessKey). The authenticate() function in internal/gateway/auth.go:25-52 extracts a key from the Authorization header (Bearer) or key query parameter. It hashes the plaintext with a keyHasher and looks it up in snapshot.AccessKeysByHash. The key is validated against expiration (ExpiresAtMS) and allowed peer CIDRs (AllowedPeerCIDRs), returning an AccessKeyView (internal/state/snapshot.go:190-203).

AccessKey model — each key has: ID, Name, KeyHash, KeyPrefix/KeySuffix (for display), Status (active/disabled), ExpiresAtMS, AllowedPeerCIDRs, RPMLimit, ConcurrencyLimit, PriceMultiplier, Filters (on Groups/Protocols/Models), and CostLimitRules.

Credential management — upstream API keys are stored encrypted via encryption.Service. The runtimeCredentialRegistry interface (handler.go:80-92) manages credential lifecycle: ActiveEncryptedCredentialDataIfMatch(), SetCooldownWithChange(), IncrFailure(), SetBlacklistedWithChange(), ClearFailure(). Credentials belong to Groups — each group connects to one channel (provider) with its own params, models, and proxy config. Secrets are identified via connection.Type — either "api_key" or a subscription driver.

No user/tenant model — the system does not implement user/team/tenant hierarchies. AccessKeys are the sole principal. Key-level Filters control which groups, protocols, and models a key can access (state.FilterSet, snapshot.go:110-114), providing capability-based access control without multi-tenant abstractions.

Credential encryption — execution_boundary.go:24-65 normalizes channel credentials: it validates against the channel's connection type, optionally delegates to a subscription driver (subscriptionruntime.Runtime.Driver()), and extracts api_key plus secret values (client_id, client_secret, tenant_id, etc.). Proxy configs can be per-credential, encrypted and integrity-checked via proxy.go:12-53.

Admin UI — the web frontend (web/src/main.ts) supports both a "classic" and "modern" frontend. The internal/webui/ package serves the admin dashboard with http_routes.go for API endpoints and page_routes.json/modern_page_routes.json for route definitions. The control package provides the configuration API surface for CRUD operations on keys, groups, and credentials.

How are rate limits, budgets and cost tracking implemented?

answered

Per-key rate limits — two levels: RPM (requests per minute) and concurrency. RPM is enforced by AccessKeyRPMLimiter (handler.go:68-71), implemented via ratelimit.AccessKeyRPM backed by rpm.Store. The limiter.Allow(accessKey.ID, accessKey.RPMLimit) call at handler.go:543 returns a LimitDecision. Concurrency is enforced by handler.acquireRequestConcurrency() (internal/gateway/concurrency.go:15) against both per-key and global limits (GlobalConcurrencyLimit), tracked in ratelimit.Concurrency.

Token and spend budgets — the accessquota package (internal/accessquota/runtime.go) implements AccessKey-level cost limits. Two rule kinds exist: KindTotal (lifetime spend cap) and KindPeriodic (time-windowed cap like daily/monthly). Each Rule stores LimitNanoUSD, PeriodSeconds, window tracking, and UsedNanoUSD. The accessquota.Runtime provides Check() (pre-flight) and Complete() (post-request deduction) methods. Budgets are checked in handler.go:522-541 before forwarding.

Price tables — pricing.Table (internal/pricing/types.go) maps channel+model identities to nano-USD-per-million-token prices. pricing.Price tracks per-input, per-output, per-cache-read, and per-cache-write rates. ContextTier allows different rates once input exceeds a threshold. pricing.Quote() (internal/pricing/quote.go) prices a finalized usage.Result. Group-level and AccessKey-level PriceMultipliers are applied via QuoteForModeWithMultipliers() for surcharging/discounting.

Usage accounting — usage.Result (internal/usage) holds token counts (uncached input, cache read, cache write, output). Usage is captured from upstream responses and stream events via usageCaptureBoundary (internal/gateway/usage_capture.go:20). Extracted usage flows to pricing.Receipt (internal/pricing/receipt.go), which is a frozen record of pricing method, rates, and total cost at request time (schema v6).

Usage storage — the requestlog package persists detailed request logs with cost. accessquota keeps in-memory rolling state for budget enforcement. The state.Manager uses gorm-backed snapshots for durable configuration, and state/loader loads config from the database. RPM state persists through rpm.Store.

Cost estimation & budget enforcement — handler.freezeAttemptPricing() (handler.go:132-155) captures the price table and multipliers at request time. After execution, recorder.estimatedCostNanoUSD() feeds into accessQuota.Complete(). If quota is exhausted, Check() returns Allowed: false and the request is blocked before reaching a provider.

How are caching and guardrails implemented?

answered

Prompt caching — the system does not implement a general-purpose response cache. Instead, it provides prompt affinity routing: dialect.inspectPromptAffinityPrefix() (internal/dialect/prompt_affinity.go:43) extracts the first user message content as a cache key. The affinity.Cache (gateway/affinity.go:19-60) stores a mapping from prompt prefix to the last-used credential ID, enabling session stickiness for providers like Anthropic that support prompt caching natively. Prompt cache keys can also be explicitly set by clients via the prompt_cache_key field (internal/dialect/prompt_cache.go:12).

Semantic caching — not implemented. No vector stores, embedding-based cache lookups, or cache hit/miss logic for responses exist in the codebase.

PII redaction — the requestredact package (internal/requestredact/redact.go:1) applies configurable regex-based redaction to outbound request content before it reaches the upstream provider. Up to 64 rules are supported, with two modes: ModeReplace (substitute matching text) and ModeEncrypt (encrypt in place with a per-credential cipher). Rules are configured globally or per-access-key. Redaction operates on raw HTTP body bytes with a 128 MB limit (maxTextBytes). The redact.TokenCipher interface (redact.go:50-54) supports token-aware encryption so API keys embedded in text can be redacted without breaking tokenization.

Moderation and prompt injection — handled by the requestaudit package (internal/requestaudit/audit.go:1), described as "experimental Jev guardrails without rewriting content." It defines Rule objects with Instructions (a prompt for the Jev LLM judge), Threshold (0-1 confidence), and Action (block or warn). Request content (up to 24 KB, MaxRequestBytes) is sent to the Jev classification engine (jev package). The requestaudit.Cache (handler.go:94) caches audit results. The Jev module (internal/jev/) is a mini LLM service specifically for classification decisions.

Jev Decisions endpoint — the system exposes a native /v1/systemone endpoint (internal/dialect/decisions.go:16) for making fast classification decisions. It accepts a state (string/object/array) and questions object, returning structured decisions. This is used internally for guardrail evaluation.

Hook points — the requestredact and requestaudit configurations are compiled into the immutable ConfigSnapshot (internal/state/snapshot.go:206-207) and applied per-request. Redaction happens inside the Bifrost executor before upstream dispatch; audit happens during handler processing. No plugin system exists for custom guardrails — both are configuration-driven.

How is it observed, deployed and scaled?

answered

Logging — the system uses logrus with structured fields. Request logs are written via telemetry.RequestLogSink (handler.go:485-510) — every request produces a requestRecorder that collects operation, model, cost, status, and timing. The requestlog package (internal/requestlog/) persists logs with configurable retention (default 7 days, maxRequestLogRetentionDays 365, internal/state/runtime_settings.go:59-61). Multiple log stores are supported: access-key usage, credential window usage, group usage, RPM history, and quota history.

Metrics — health.StatsStore (internal/health/stats.go:14) tracks per-credential success/failure/problem counts in sliding 5-minute windows (1-minute buckets). RPM counters are stored in rpm.Store and queried via ratelimit.AccessKeyRPM. No OpenTelemetry export is implemented — metrics are consumed in-process for health decisions and exposed via the control API.

Traces — no distributed tracing (OpenTelemetry or similar). The system uses request IDs (handler.newRequestID(), handler.go:464) that flow in X-GPTLoad-Request-ID response headers and appear in all log entries, enabling request correlation.

Deployment — packaged as a single Go binary via Docker (Dockerfile). The build uses multi-stage: Node.js for the web frontend, then Go compilation. The final image is Alpine-based. A docker-compose.yml and docker-compose.voice.yml are provided for container orchestration. The binary supports Windows as a first-class target with service management (service_windows.go, service_cli.go).

Concurrency model — Go language with goroutines (green threads). The gin HTTP framework provides non-blocking request handling. The concurrency.go package tracks active requests per access-key and globally. channel.RouteMode and the scheduler are pure in-memory operations with no per-request blocking IO beyond the upstream HTTP call.

Configuration management — state.Manager (internal/state/manager.go:11) holds an atomic pointer to the current ConfigSnapshot. Snapshots are immutable and replaced atomically on publish via sync/atomic.Pointer. The SnapshotReconciler interface synchronizes infrastructure resources before a snapshot becomes visible. Configuration is loaded from a database via state/loader and refreshed in-place without restarts.

High availability — the system is single-process; there is no built-in clustering, leader election, or shared state between instances. HA would rely on running multiple instances behind a load balancer with a shared database for configuration persistence. RPM and concurrency limits are in-process and not shared across nodes.

Health endpoints — the health package exposes credential health stats via the control API, and the auth layer checks credential cooldown/blacklist state inline during scheduling.