LLMs Technical Reviews
Home / LLM gateways / one-api

songquanpeng/one-api

Go/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI.

GitHub ↗★ 37kGoMITcommit 8df4a26 · 2025-02-21homepage ↗

Overview

One API is a self-hosted LLM gateway written in Go. It accepts OpenAI-format requests on /v1/*, picks an upstream “channel” (a provider endpoint plus its API key), converts the request to that provider’s format if needed, and streams the answer back in OpenAI format. Around that relay it adds the things a small team or a key reseller needs: users, per-user API tokens, quota billing in internal units, redemption codes, and a React admin console.

The design is deliberately small. There is one binary with the web UI embedded, one database (SQLite by default, MySQL or PostgreSQL in production), and optional Redis. Channel selection is a random pick inside the highest priority tier. Billing is a model-ratio table compiled into the code. There is no response cache, no content guardrail and no metrics exporter.

One API is the upstream of New API, which forked it and has grown much faster since. The pinned commit is from February 2025, and the code base at that SHA shows a project that is mostly in maintenance. If you need Claude- or Gemini-native inbound endpoints, weighted load balancing or newer billing models, those are in the fork, not here.

Architecture

flowchart LR
  C["Client (OpenAI SDK)"] --> R["Gin router /v1"]
  R --> A["TokenAuth"]
  A --> D["Distribute: pick channel"]
  D --> CC["In-memory channel cache"]
  D --> RL["controller.Relay + retry"]
  RL --> H["Relay helper: text / image / audio / proxy"]
  H --> Q["Quota pre-consume / post-consume"]
  H --> AD["Provider adaptor"]
  AD --> UP["Upstream provider API"]
  Q --> DB["SQL DB + optional Redis"]
  RL --> M["monitor: auto-disable channel"]
  UI["React admin themes"] --> API["/api management routes"]
  API --> DB
Component Path Role
Entry point main.go Init DB, Redis, option and channel caches, background sync, Gin server
Relay routes router/relay.go /v1/* OpenAI-shaped endpoints behind TokenAuth and Distribute
Auth middleware middleware/auth.go Session auth for the console, TokenAuth for relay calls
Distributor middleware/distributor.go Chooses a channel for (group, model) and stores it in the Gin context
Channel cache model/cache.go group2model2channels index, refreshed every SYNC_FREQUENCY seconds
Relay controller controller/relay.go Calls the relay helper, retries on other channels, reports errors to the monitor
Relay helpers relay/controller/ Text, image, audio and proxy flows, quota accounting
Adaptors relay/adaptor/, relay/adaptor.go One package per provider family; GetAdaptor maps API type to adaptor
Billing ratios relay/billing/ratio/ Per-model and per-group price multipliers
Monitor monitor/ Auto-disables channels on fatal errors or low success rate
Web UI web/default, web/berry, web/air Three React themes, built and embedded into the binary

How a request flows

Take POST /v1/chat/completions with Authorization: Bearer sk-...:

  1. Route. The /v1 group runs RelayPanicRecover, TokenAuth and Distribute before controller.Relay (router/relay.go).
  2. Authenticate. TokenAuth strips Bearer and sk-, validates the token, checks the optional subnet restriction, rejects banned or disabled users, and enforces the token’s model allow-list. A key of the form sk-<key>-<channelId> lets an admin pin a specific channel (middleware/auth.go).
  3. Pick a channel. Distribute reads the user’s group and calls CacheGetRandomSatisfiedChannel(group, model, false). SetupContextForSelectedChannel writes the channel’s key into the Authorization header, plus its base URL, model mapping, system prompt and config (middleware/distributor.go).
  4. Pre-charge. RelayTextHelper parses the body into GeneralOpenAIRequest, applies the model mapping, counts prompt tokens with the OpenAI tokenizer and pre-consumes quota as (prompt + max_tokens) × modelRatio × groupRatio (text.go, helper.go).
  5. Convert and send. For a plain OpenAI channel with no mapping or forced system prompt, the original body is forwarded as is. Otherwise the adaptor’s ConvertRequest builds the provider body (text.go). DoRequest sends it and DoResponse translates the reply or stream back to OpenAI format and returns usage.
  6. Settle. postConsumeQuota runs in a goroutine. It computes ceil((prompt + completion × completionRatio) × ratio), adjusts the token by the difference from the pre-charge, writes a consume log and updates user and channel counters (helper.go).
  7. Retry on failure. If the helper returns an error, Relay reports it to the monitor and retries up to RetryTimes times on another channel for 429, 5xx and most non-400 errors. The first retry stays in the top priority tier; later retries move to lower tiers (controller/relay.go).

Key components

Channels and selection

A channel is one row in the channels table: provider type, key, base URL, comma-separated models and groups, model mapping, priority and an optional forced system prompt (model/channel.go). InitChannelCache builds a group → model → []*Channel index sorted by priority. CacheGetRandomSatisfiedChannel picks uniformly at random among channels that share the highest priority (model/cache.go). The struct has a Weight field, but the selection code never reads it, so “weighted” balancing is not real here.

Adaptors

Every provider family implements a nine-method Adaptor interface: Init, GetRequestURL, SetupRequestHeader, ConvertRequest, ConvertImageRequest, DoRequest, DoResponse, GetModelList and GetChannelName (interface.go). GetAdaptor maps 19 API types to concrete adaptors (relay/adaptor.go). Most of the roughly 50 channel types (DeepSeek, Groq, Moonshot, SiliconFlow and others) are OpenAI-compatible and fall through to the OpenAI adaptor. The Anthropic adaptor, for example, targets /v1/messages, sets x-api-key and anthropic-version, and converts messages, tools and images (anthropic/adaptor.go).

Quota and billing

Quota is an integer balance on both the user and the token. Prices are multipliers: a model ratio from a large built-in map in relay/billing/ratio/model.go, a completion ratio for output tokens, and a group ratio. Admins edit them as JSON options in the console. A trusted-user shortcut skips the token pre-charge when the user balance is more than 100 times the estimate. With BATCH_UPDATE_ENABLED, balance changes are buffered and flushed periodically instead of one SQL UPDATE per request.

Health monitor

ShouldDisableChannel disables a channel on 401, on error types such as insufficient_quota or authentication_error, and on message substrings such as “credit” or “balance” (monitor/manage.go). With ENABLE_METRIC, an in-memory window of recent results per channel disables a channel whose success rate falls below the threshold (monitor/metric.go). CHANNEL_TEST_FREQUENCY adds periodic test calls.

Management API and UI

router/api.go exposes CRUD for channels, tokens, users, redemption codes, logs and options, gated by user, admin and root roles. Three themes ship in web/: default (Semantic UI), berry (MUI) and air (Semi UI). All three are built in the Dockerfile and embedded with go:embed. Login supports passwords, GitHub, WeChat, Lark and OIDC, with optional Turnstile.

Extending it

  • New provider. Add a package under relay/adaptor/, a channel type in relay/channeltype, an API type in relay/apitype, a case in GetAdaptor, and the channel option in each theme’s constants. If the provider speaks OpenAI’s wire format, a channel type that maps to the OpenAI adaptor is enough.
  • Prices. Edit model, completion and group ratios in the console. They are stored as options and override the compiled defaults.
  • Raw pass-through. /v1/oneapi/proxy/:channelid/*target forwards an arbitrary path to a specific channel through the Proxy adaptor.
  • No plugin system. There are no request hooks or middleware extension points. Behaviour changes mean editing Go code.

Running it

  • Docker. docker-compose.yml runs justsong/one-api with MySQL and Redis on port 3000 (docker-compose.yml). Without SQL_DSN it uses SQLite under /data.
  • Binary. The multi-stage Dockerfile builds the three React themes with Node 16 and then a static CGO binary on Alpine (Dockerfile).
  • First login. On an empty database the server creates root with password 123456 and a large quota (model/main.go). Change it immediately, or preset INITIAL_ROOT_TOKEN / INITIAL_ROOT_ACCESS_TOKEN.
  • Multiple nodes. Set NODE_TYPE=slave on replicas so only the master migrates the schema, share the database, Redis and SESSION_SECRET, and set SYNC_FREQUENCY so replicas refresh their channel cache. Load balancing across nodes is external.

Strengths and caveats

  • Strength: small and readable. The relay path is a few hundred lines across the distributor, Relay and RelayTextHelper. It is easy to audit and to patch.
  • Strength: cheap to run. One Go binary with SQLite works for a single box, and the same binary scales to MySQL plus Redis.
  • Strength: built for redistribution. Groups, ratios, quotas, redemption codes and per-token model and subnet limits fit the “share or resell keys” use case well.
  • Caveat: OpenAI format only on the way in. There is no /v1/messages, no Gemini endpoint and no Responses API. Assistants, files and fine-tuning routes return “not implemented”.
  • Caveat: simple routing. Random within a priority tier, the Weight field unused, and no latency- or cost-based selection.
  • Caveat: no caching, guardrails or metrics. Redis caches tokens, users and channels, not model output. There is no PII or moderation layer and no Prometheus or OpenTelemetry export.
  • Caveat: slow upstream. Development has largely moved to the New API fork. New providers and protocol changes land there first.

Sources: code at 8df4a26, deepwiki-open wiki (11 pages), verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit 8df4a26. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Load balancing and priority. Each channel (upstream provider endpoint) is assigned a priority and weight. The routing system (middleware/distributor.go) calls model.CacheGetRandomSatisfiedChannel which queries an in-memory index group2model2channels built on startup and refreshed every SYNC_FREQUENCY seconds (model/cache.go). Channels are sorted by priority within each (group, model) bucket; the system picks randomly from the top-priority tier. If ignoreFirstPriority is set (during retries), it picks from lower-priority tiers instead. This is a weighted-random-within-priority-tier scheme.

Model aliases. Each channel carries an optional ModelMapping JSON field that maps user-visible model names to provider-specific model names (model/channel.go:114-125). The mapping is applied right before outbound request conversion (controller/text.go:37-38, controller/image.go:117-118).

Retries and fallbacks. controller/relay.go:45-101 implements retry logic: after a request fails, it calls CacheGetRandomSatisfiedChannel again (with ignoreFirstPriority=true), skipping the last-failed channel ID. The number of retries is config.RetryTimes (default 0). Channels that return 401/403 or certain insufficient_quota/invalid_api_key errors are automatically disabled by the monitor (monitor/manage.go:11-44).

Health checks. The monitor/metric.go module tracks a sliding window of success/failure per channel. If a channel's success rate for the last METRIC_QUEUE_SIZE calls falls below METRIC_SUCCESS_RATE_THRESHOLD (default 80%), the channel is auto-disabled (monitor/channel.go:47-61). Additionally, CHANNEL_TEST_FREQUENCY can be set to automatically test channels periodically (main.go:79-84). There is no explicit latency-based or cost-based routing — routing is purely random-within-priority-tier.

Editor's note. Correction: selection inside the top priority tier is uniform (rand.Intn), not weighted; the channel Weight field exists but is never used. The first retry stays in the top tier and only later retries move to lower-priority channels.

How are different provider APIs unified?

answered

Architecture. Every provider has an adaptor implementing the adaptor.Adaptor interface (5 methods: Init, GetRequestURL, SetupRequestHeader, ConvertRequest/ConvertImageRequest, DoRequest, DoResponse, GetModelList, GetChannelName). The dispatcher relay/adaptor.go:27-68 maps APIType → adaptor instance (27 provider adaptors registered). channeltype/helper.go:5-47 maps channel type → API type.

OpenAI-compatible schema. All incoming requests are parsed into relay/model/general.go's GeneralOpenAIRequest struct — a superset of the OpenAI chat/completions, embeddings, images, and audio schemas. This is the internal canonical form. The relaymode/helper.go dispatches by URL path to decide relay mode (chat, completions, embeddings, images, audio, proxy).

Per-provider translation. Adaptors like Anthropic's (relay/adaptor/anthropic/main.go:39-146) convert the OpenAI-shaped request to the provider's native format: messages become Anthropic Messages, system prompts are extracted, tool definitions are reshaped to Claude's input_schema, and multi-part content (text + image) is converted to Claude's content blocks. Responses flow back through reverse converters: ResponseClaude2OpenAI (anthropic/main.go:210-247) maps Claude's Response → OpenAI TextResponse, including tool_use → function_call translation. Streaming uses StreamResponseClaude2OpenAI (anthropic/main.go:149-208) which translates each Claude SSE event (message_start, content_block_start, content_block_delta, message_delta) into OpenAI SSE chunks.

Streaming and tools. The OpenAI adaptor (openai/adaptor.go) passes through for pure OpenAI-compatible providers. For others, streaming is re-encoded event-by-event. Tool calls are converted bidirectionally: e.g., Anthropic tool_use content blocks → OpenAI tool_calls with function type. The anthropic/main.go:21-37 stopReasonClaude2OpenAI maps Claude stop reasons (end_turn, tool_use) to OpenAI equivalents (stop, tool_calls).

Multimodal. Image content is handled via Message.ParseContent() (relay/model/message.go:41-80) which extracts text and image_url parts from the OpenAI content array. The Anthropic adaptor fetches and base64-encodes images (anthropic/main.go:132-139). The relay/controller/image.go handles image generation relay with size/quality validation and cost ratio computation.

Provider count. The channel type enum (relay/channeltype/define.go) defines 56 channel types. The adaptor registry (relay/adaptor.go) explicitly maps 27 API types; the remaining channel types (Azure, OpenAI-compatible, etc.) fall through to the OpenAI adaptor or Custom handling.

Editor's note. Correction: GetAdaptor maps 19 API types (not 27) and the Adaptor interface has nine methods; the other channel types fall through to the OpenAI adaptor.

How are API keys, users and tenants managed?

answered

User model. Users have roles (model/user.go:19-24): RoleGuestUser (0), RoleCommonUser (1), RoleAdminUser (10), RoleRootUser (100). Fields include username, password (bcrypt-hashed), quota, used_quota, request_count, group (for routing), and OAuth IDs for GitHub, WeChat, Lark, OIDC (model/user.go:34-53). The root user is auto-created on first run with password 123456 (model/main.go:24-64).

Auth methods. The system supports password-based login (session cookie via gin-contrib/sessions), Bearer access tokens (for API/management), and OAuth (GitHub, WeChat, Lark, OIDC). The middleware/auth.go provides three auth middleware for the admin API: UserAuth(), AdminAuth(), RootAuth() checking minRole (auth.go:15-71). For the relay API, TokenAuth() (auth.go:91-151) validates user tokens sent as Authorization: Bearer sk-... headers.

Virtual keys (Tokens). Each token (model/token.go:23-37) is a 48-char Key linked to a UserId. Tokens have statuses (enabled/disabled/expired/exhausted), RemainQuota, UnlimitedQuota, Models (comma-separated whitelist), Subnet (CIDR restriction), and ExpiredTime. The token prefix sk- is stripped during parsing, and the remaining key is validated via Redis cache or DB (model/cache.go:28-56).

Token scoping. On validation, middleware/auth.go:125-131 checks if the requested model is in the token's allowed model list. The Subnet field is verified against c.ClientIP() (auth.go:104-109). An advanced sk-{key}-{channelId} format allows admins to pin a request to a specific channel (auth.go:135-142).

Admin UI/API. The REST API (router/api.go) exposes CRUD for channels (/api/channel), tokens (/api/token), users (/api/user), redemptions (/api/redemption), and system options (/api/option). Role-based access control gates each route. Three React frontends (default, berry, air) provide the web UI.

Upstream credential storage. Each Channel stores its provider API key in the Key field (plaintext in DB, model/channel.go:21-41). Additional config (region, SK, AK for AWS/Vertex) goes in the Config JSON field. The channel's key is injected as Authorization: Bearer {key} on outbound requests (middleware/distributor.go:73).

How are rate limits, budgets and cost tracking implemented?

answered

Rate limits. The middleware/rate-limit.go provides per-IP rate limiting with configurable windows: GlobalAPIRateLimit (480 requests/3 min), GlobalWebRateLimit (240/3 min), CriticalRateLimit (20/20 min), and upload/download limits. Supports both Redis-based (redisRateLimiter via LPUSH/EXPIRE/TTL checks) and in-memory (memoryRateLimiter) backends. The common/rate-limit.go InMemoryRateLimiter uses a sliding-window approach with per-key timestamp queues cleaned by a background goroutine.

Quota system. Users have a quota field (integer) and used_quota (model/user.go:48-49). Tokens have their own remain_quota and unlimited_quota (model/token.go:31-33). The quota is denominated in internal units where 1 unit = $0.002/1K tokens (configurable via QuotaPerUnit in common/config/config.go:21).

Pre-consumption. Before sending a request, the system pre-consumes quota: preConsumeQuota in controller/helper.go:68-95 estimates prompt tokens via the OpenAI tokenizer, adds max_tokens, multiplies by modelRatio × groupRatio, and checks the user + token balance. It then decrements the user's Redis-cached quota (CacheDecreaseUserQuota) and records a pre-consumption on the token. If user quota > 100× the estimate, pre-consumption is skipped (trusted user).

Post-consumption. After the response, postConsumeQuota (controller/helper.go:97-141) computes actual usage from usage.prompt_tokens and usage.completion_tokens, applies completionRatio for output tokens, calculates quota = ceil((promptTokens + completionTokens × completionRatio) × modelRatio × groupRatio), adjusts the token quota by the delta, updates the user quota cache, writes a Log record, and increments channel used_quota.

Price tables. relay/billing/ratio/model.go:27-622 contains an exhaustive ModelRatio map covering ~600+ models across OpenAI, Anthropic, Google, Baidu, Ali, Zhipu, Xunfei, Moonshot, DeepSeek, Mistral, Groq, Cohere, etc. The CompletionRatio map (model.go:624-633) adjusts for providers that price output tokens differently. GetModelRatio (model.go:686-710) looks up by exact name then falls back to defaults. Group-level multipliers are in relay/billing/ratio/group.go:10-13.

Batch updates. When BATCH_UPDATE_ENABLED is set, quota updates are queued into in-memory records and flushed periodically (BatchUpdateInterval, default 5s) to reduce DB write pressure (model/main.go:86-89). Without batching, each quota change issues an immediate SQL UPDATE ... SET quota = quota +/- N.

Storage. Usage logs are stored in the logs table (or a separate LOG_SQL_DSN database) with fields for user, token, model, prompt/completion tokens, quota consumed, elapsed time, and stream flag (model/log.go:15-32).

How are caching and guardrails implemented?

answered

Caching. There is no semantic caching of LLM responses (no exact-match or semantic cache for chat completions). The middleware/cache.go sets only HTTP-level Cache-Control headers (max-age=604800 for static assets, no-cache for the root route) — this is purely for the web frontend, not for API responses. The system caches database entities in Redis: tokens are cached by key with a TTL of SYNC_FREQUENCY seconds (model/cache.go:28-56), user groups are cached (cache.go:58-74), user quotas and enabled status are cached (cache.go:88-149), and group→models mappings are cached (cache.go:151-168). The in-memory channel index (group2model2channels) is refreshed on a timer (cache.go:170-225). None of these caches store model outputs.

PII redaction / prompt injection / moderation. The system has no built-in PII redaction, prompt-injection detection, or output moderation guardrails. It is a transparent proxy that does not inspect or sanitize request/response content. The /v1/moderations endpoint is simply relayed like any other endpoint (router/relay.go:46), passing through to the upstream provider's moderation API — not a local check. The only content-level control is the SystemPrompt feature on channels (model/channel.go:40), which lets an admin override or inject a system prompt into requests, but this is additive, not a guardrail.

Blacklist. A simple in-memory user ban list exists (common/blacklist/main.go) that blocks banned user IDs from auth. This is operational, not content-based.

Turnstile. CAPTCHA (Turnstile) support is available for registration, login, and password reset endpoints via middleware/turnstile-check.go, configured through TurnstileSiteKey and TurnstileSecretKey (common/config/config.go:89-90).

Plugin/hook points. There are no formal plugin or webhook systems. The adaptor pattern (relay/adaptor.go) is the extension mechanism for adding new providers — implementing the Adaptor interface gives you full control over request/response conversion. The message-pusher (common/message/message-pusher.go) can send notifications to external systems for channel disable events, but this is one-directional alerting, not an interceptor.

How is it observed, deployed and scaled?

answered

Language and concurrency. One API is written in Go using the Gin web framework (Go's analog of Express/Flask). Go's goroutine-based concurrency model and the Gin framework's non-blocking I/O provide good per-instance throughput. The server starts in main.go:103-120: a single Gin engine with middleware stack (session store, request ID, language, logger, rate limits) and route registration. No explicit worker count or multi-process configuration is visible — Gin runs as a single process with goroutine-based concurrency.

Logging. The middleware/logger.go sets up standard Gin request logging (method, path, status code, latency, client IP). Application-level logging uses common/logger/logger.go with log levels (Info, Error, Debug, SysLog). config.DebugEnabled and config.DebugSQLEnabled toggle verbose logging and SQL query logging. ONLY_ONE_LOG_FILE controls log file rotation (config.go:159).

Metrics and observability. No OpenTelemetry, Prometheus, or Datadog integration is present. The only built-in metric is the channel success-rate monitor (monitor/metric.go), which uses in-memory ring buffers per channel. There are no external metrics exporters.

Database. The system supports SQLite (default, file-based), MySQL, and PostgreSQL (model/main.go:67-109). Schema is auto-migrated via GORM's AutoMigrate on startup. A separate LOG_SQL_DSN can route usage logs to a secondary database. Connection pooling is configured via SQL_MAX_IDLE_CONNS, SQL_MAX_OPEN_CONNS, and SQL_MAX_LIFETIME environment variables (model/main.go:214-217).

Deployment. A multi-stage Dockerfile builds static Go binary + React frontends. The docker-compose.yml provides a full stack with Redis, MySQL, and the One API service. Environment variables control all configuration: SQL_DSN, REDIS_CONN_STRING, SESSION_SECRET, NODE_TYPE, SYNC_FREQUENCY, etc.

High-availability / multi-node. The NODE_TYPE environment variable (config.go:105) supports "master" vs "slave" nodes. Only the master node runs database migrations. Slave nodes can serve API traffic but delegate configuration reads to the database and rely on SYNC_FREQUENCY to refresh their in-memory channel cache. The FRONTEND_BASE_URL env var allows slave nodes to offload the React frontend to a separate server. There is no load balancer or clustering logic in the application itself — HA relies on an external LB + shared DB/Redis.

Health checks. The docker-compose defines a healthcheck at http://localhost:3000/api/status. The /api/status endpoint (controller/misc.go) provides a simple status probe.