# songquanpeng/one-api

> Go/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/songquanpeng/one-api (reviewed at commit `8df4a2670b98266bd287c698243fff327d9748cf`, 2025-02-21)
- Stars: 37077 · Language: Go · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/one-api/

## Overview

One API is a self-hosted LLM gateway written in Go. It accepts OpenAI-format requests on `/v1/*`, picks an upstream "channel" (a provider endpoint plus its API key), converts the request to that provider's format if needed, and streams the answer back in OpenAI format. Around that relay it adds the things a small team or a key reseller needs: users, per-user API tokens, quota billing in internal units, redemption codes, and a React admin console.

The design is deliberately small. There is one binary with the web UI embedded, one database (SQLite by default, MySQL or PostgreSQL in production), and optional Redis. Channel selection is a random pick inside the highest priority tier. Billing is a model-ratio table compiled into the code. There is no response cache, no content guardrail and no metrics exporter.

One API is the upstream of New API, which forked it and has grown much faster since. The pinned commit is from February 2025, and the code base at that SHA shows a project that is mostly in maintenance. If you need Claude- or Gemini-native inbound endpoints, weighted load balancing or newer billing models, those are in the fork, not here.

## Architecture

```mermaid
flowchart LR
  C["Client (OpenAI SDK)"] --> R["Gin router /v1"]
  R --> A["TokenAuth"]
  A --> D["Distribute: pick channel"]
  D --> CC["In-memory channel cache"]
  D --> RL["controller.Relay + retry"]
  RL --> H["Relay helper: text / image / audio / proxy"]
  H --> Q["Quota pre-consume / post-consume"]
  H --> AD["Provider adaptor"]
  AD --> UP["Upstream provider API"]
  Q --> DB["SQL DB + optional Redis"]
  RL --> M["monitor: auto-disable channel"]
  UI["React admin themes"] --> API["/api management routes"]
  API --> DB
```

| Component | Path | Role |
|---|---|---|
| Entry point | `main.go` | Init DB, Redis, option and channel caches, background sync, Gin server |
| Relay routes | `router/relay.go` | `/v1/*` OpenAI-shaped endpoints behind `TokenAuth` and `Distribute` |
| Auth middleware | `middleware/auth.go` | Session auth for the console, `TokenAuth` for relay calls |
| Distributor | `middleware/distributor.go` | Chooses a channel for (group, model) and stores it in the Gin context |
| Channel cache | `model/cache.go` | `group2model2channels` index, refreshed every `SYNC_FREQUENCY` seconds |
| Relay controller | `controller/relay.go` | Calls the relay helper, retries on other channels, reports errors to the monitor |
| Relay helpers | `relay/controller/` | Text, image, audio and proxy flows, quota accounting |
| Adaptors | `relay/adaptor/`, `relay/adaptor.go` | One package per provider family; `GetAdaptor` maps API type to adaptor |
| Billing ratios | `relay/billing/ratio/` | Per-model and per-group price multipliers |
| Monitor | `monitor/` | Auto-disables channels on fatal errors or low success rate |
| Web UI | `web/default`, `web/berry`, `web/air` | Three React themes, built and embedded into the binary |

## How a request flows

Take `POST /v1/chat/completions` with `Authorization: Bearer sk-...`:

1. **Route.** The `/v1` group runs `RelayPanicRecover`, `TokenAuth` and `Distribute` before `controller.Relay` ([router/relay.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/router/relay.go#L10-L74)).
2. **Authenticate.** `TokenAuth` strips `Bearer ` and `sk-`, validates the token, checks the optional subnet restriction, rejects banned or disabled users, and enforces the token's model allow-list. A key of the form `sk-<key>-<channelId>` lets an admin pin a specific channel ([middleware/auth.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/auth.go#L91-L151)).
3. **Pick a channel.** `Distribute` reads the user's group and calls `CacheGetRandomSatisfiedChannel(group, model, false)`. `SetupContextForSelectedChannel` writes the channel's key into the `Authorization` header, plus its base URL, model mapping, system prompt and config ([middleware/distributor.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/distributor.go#L20-L102)).
4. **Pre-charge.** `RelayTextHelper` parses the body into `GeneralOpenAIRequest`, applies the model mapping, counts prompt tokens with the OpenAI tokenizer and pre-consumes quota as `(prompt + max_tokens) × modelRatio × groupRatio` ([text.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/controller/text.go#L25-L88), [helper.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/controller/helper.go#L68-L95)).
5. **Convert and send.** For a plain OpenAI channel with no mapping or forced system prompt, the original body is forwarded as is. Otherwise the adaptor's `ConvertRequest` builds the provider body ([text.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/controller/text.go#L90-L115)). `DoRequest` sends it and `DoResponse` translates the reply or stream back to OpenAI format and returns `usage`.
6. **Settle.** `postConsumeQuota` runs in a goroutine. It computes `ceil((prompt + completion × completionRatio) × ratio)`, adjusts the token by the difference from the pre-charge, writes a consume log and updates user and channel counters ([helper.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/controller/helper.go#L97-L141)).
7. **Retry on failure.** If the helper returns an error, `Relay` reports it to the monitor and retries up to `RetryTimes` times on another channel for 429, 5xx and most non-400 errors. The first retry stays in the top priority tier; later retries move to lower tiers ([controller/relay.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/controller/relay.go#L45-L122)).

## Key components

### Channels and selection

A channel is one row in the `channels` table: provider type, key, base URL, comma-separated models and groups, model mapping, priority and an optional forced system prompt ([model/channel.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/channel.go#L20-L41)). `InitChannelCache` builds a `group → model → []*Channel` index sorted by priority. `CacheGetRandomSatisfiedChannel` picks uniformly at random among channels that share the highest priority ([model/cache.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/cache.go#L170-L255)). The struct has a `Weight` field, but the selection code never reads it, so "weighted" balancing is not real here.

### Adaptors

Every provider family implements a nine-method `Adaptor` interface: `Init`, `GetRequestURL`, `SetupRequestHeader`, `ConvertRequest`, `ConvertImageRequest`, `DoRequest`, `DoResponse`, `GetModelList` and `GetChannelName` ([interface.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor/interface.go#L11-L21)). `GetAdaptor` maps 19 API types to concrete adaptors ([relay/adaptor.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor.go#L27-L69)). Most of the roughly 50 channel types (DeepSeek, Groq, Moonshot, SiliconFlow and others) are OpenAI-compatible and fall through to the OpenAI adaptor. The Anthropic adaptor, for example, targets `/v1/messages`, sets `x-api-key` and `anthropic-version`, and converts messages, tools and images ([anthropic/adaptor.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor/anthropic/adaptor.go#L23-L44)).

### Quota and billing

Quota is an integer balance on both the user and the token. Prices are multipliers: a model ratio from a large built-in map in `relay/billing/ratio/model.go`, a completion ratio for output tokens, and a group ratio. Admins edit them as JSON options in the console. A trusted-user shortcut skips the token pre-charge when the user balance is more than 100 times the estimate. With `BATCH_UPDATE_ENABLED`, balance changes are buffered and flushed periodically instead of one SQL `UPDATE` per request.

### Health monitor

`ShouldDisableChannel` disables a channel on 401, on error types such as `insufficient_quota` or `authentication_error`, and on message substrings such as "credit" or "balance" ([monitor/manage.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/manage.go#L11-L44)). With `ENABLE_METRIC`, an in-memory window of recent results per channel disables a channel whose success rate falls below the threshold ([monitor/metric.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/metric.go#L18-L38)). `CHANNEL_TEST_FREQUENCY` adds periodic test calls.

### Management API and UI

`router/api.go` exposes CRUD for channels, tokens, users, redemption codes, logs and options, gated by user, admin and root roles. Three themes ship in `web/`: `default` (Semantic UI), `berry` (MUI) and `air` (Semi UI). All three are built in the Dockerfile and embedded with `go:embed`. Login supports passwords, GitHub, WeChat, Lark and OIDC, with optional Turnstile.

## Extending it

- **New provider.** Add a package under `relay/adaptor/`, a channel type in `relay/channeltype`, an API type in `relay/apitype`, a case in `GetAdaptor`, and the channel option in each theme's constants. If the provider speaks OpenAI's wire format, a channel type that maps to the OpenAI adaptor is enough.
- **Prices.** Edit model, completion and group ratios in the console. They are stored as options and override the compiled defaults.
- **Raw pass-through.** `/v1/oneapi/proxy/:channelid/*target` forwards an arbitrary path to a specific channel through the `Proxy` adaptor.
- **No plugin system.** There are no request hooks or middleware extension points. Behaviour changes mean editing Go code.

## Running it

- **Docker.** `docker-compose.yml` runs `justsong/one-api` with MySQL and Redis on port 3000 ([docker-compose.yml](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/docker-compose.yml#L3-L29)). Without `SQL_DSN` it uses SQLite under `/data`.
- **Binary.** The multi-stage Dockerfile builds the three React themes with Node 16 and then a static CGO binary on Alpine ([Dockerfile](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/Dockerfile#L1-L47)).
- **First login.** On an empty database the server creates `root` with password `123456` and a large quota ([model/main.go](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/main.go#L24-L65)). Change it immediately, or preset `INITIAL_ROOT_TOKEN` / `INITIAL_ROOT_ACCESS_TOKEN`.
- **Multiple nodes.** Set `NODE_TYPE=slave` on replicas so only the master migrates the schema, share the database, Redis and `SESSION_SECRET`, and set `SYNC_FREQUENCY` so replicas refresh their channel cache. Load balancing across nodes is external.

## Strengths and caveats

- **Strength: small and readable.** The relay path is a few hundred lines across the distributor, `Relay` and `RelayTextHelper`. It is easy to audit and to patch.
- **Strength: cheap to run.** One Go binary with SQLite works for a single box, and the same binary scales to MySQL plus Redis.
- **Strength: built for redistribution.** Groups, ratios, quotas, redemption codes and per-token model and subnet limits fit the "share or resell keys" use case well.
- **Caveat: OpenAI format only on the way in.** There is no `/v1/messages`, no Gemini endpoint and no Responses API. Assistants, files and fine-tuning routes return "not implemented".
- **Caveat: simple routing.** Random within a priority tier, the `Weight` field unused, and no latency- or cost-based selection.
- **Caveat: no caching, guardrails or metrics.** Redis caches tokens, users and channels, not model output. There is no PII or moderation layer and no Prometheus or OpenTelemetry export.
- **Caveat: slow upstream.** Development has largely moved to the New API fork. New providers and protocol changes land there first.

*Sources: code at 8df4a26, deepwiki-open wiki (11 pages), verified Q&A.*

## How songquanpeng/one-api answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

**Load balancing and priority.** Each channel (upstream provider endpoint) is assigned a priority and weight. The routing system (`middleware/distributor.go`) calls `model.CacheGetRandomSatisfiedChannel` which queries an in-memory index `group2model2channels` built on startup and refreshed every `SYNC_FREQUENCY` seconds (`model/cache.go`). Channels are sorted by priority within each (group, model) bucket; the system picks randomly from the top-priority tier. If `ignoreFirstPriority` is set (during retries), it picks from lower-priority tiers instead. This is a weighted-random-within-priority-tier scheme.

**Model aliases.** Each channel carries an optional `ModelMapping` JSON field that maps user-visible model names to provider-specific model names (`model/channel.go:114-125`). The mapping is applied right before outbound request conversion (`controller/text.go:37-38`, `controller/image.go:117-118`).

**Retries and fallbacks.** `controller/relay.go:45-101` implements retry logic: after a request fails, it calls `CacheGetRandomSatisfiedChannel` again (with `ignoreFirstPriority=true`), skipping the last-failed channel ID. The number of retries is `config.RetryTimes` (default 0). Channels that return 401/403 or certain `insufficient_quota`/`invalid_api_key` errors are automatically disabled by the monitor (`monitor/manage.go:11-44`).

**Health checks.** The `monitor/metric.go` module tracks a sliding window of success/failure per channel. If a channel's success rate for the last `METRIC_QUEUE_SIZE` calls falls below `METRIC_SUCCESS_RATE_THRESHOLD` (default 80%), the channel is auto-disabled (`monitor/channel.go:47-61`). Additionally, `CHANNEL_TEST_FREQUENCY` can be set to automatically test channels periodically (`main.go:79-84`). There is no explicit latency-based or cost-based routing — routing is purely random-within-priority-tier.

> **Editor's note.** Correction: selection inside the top priority tier is uniform (`rand.Intn`), not weighted; the channel `Weight` field exists but is never used. The first retry stays in the top tier and only later retries move to lower-priority channels.

Citations: [middleware/distributor.go:20-61](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/distributor.go#L20-L61) · [model/cache.go:227-255](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/cache.go#L227-L255) · [controller/relay.go:45-101](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/controller/relay.go#L45-L101) · [monitor/metric.go:11-39](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/metric.go#L11-L39) · [model/channel.go:100-125](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/channel.go#L100-L125)

### How are different provider APIs unified? (answered)

**Architecture.** Every provider has an adaptor implementing the `adaptor.Adaptor` interface (5 methods: `Init`, `GetRequestURL`, `SetupRequestHeader`, `ConvertRequest`/`ConvertImageRequest`, `DoRequest`, `DoResponse`, `GetModelList`, `GetChannelName`). The dispatcher `relay/adaptor.go:27-68` maps `APIType` → adaptor instance (27 provider adaptors registered). `channeltype/helper.go:5-47` maps channel type → API type.

**OpenAI-compatible schema.** All incoming requests are parsed into `relay/model/general.go`'s `GeneralOpenAIRequest` struct — a superset of the OpenAI chat/completions, embeddings, images, and audio schemas. This is the internal canonical form. The `relaymode/helper.go` dispatches by URL path to decide relay mode (chat, completions, embeddings, images, audio, proxy).

**Per-provider translation.** Adaptors like Anthropic's (`relay/adaptor/anthropic/main.go:39-146`) convert the OpenAI-shaped request to the provider's native format: messages become Anthropic Messages, system prompts are extracted, tool definitions are reshaped to Claude's `input_schema`, and multi-part content (text + image) is converted to Claude's content blocks. Responses flow back through reverse converters: `ResponseClaude2OpenAI` (`anthropic/main.go:210-247`) maps Claude's Response → OpenAI `TextResponse`, including tool_use → function_call translation. Streaming uses `StreamResponseClaude2OpenAI` (`anthropic/main.go:149-208`) which translates each Claude SSE event (message_start, content_block_start, content_block_delta, message_delta) into OpenAI SSE chunks.

**Streaming and tools.** The OpenAI adaptor (`openai/adaptor.go`) passes through for pure OpenAI-compatible providers. For others, streaming is re-encoded event-by-event. Tool calls are converted bidirectionally: e.g., Anthropic `tool_use` content blocks → OpenAI `tool_calls` with `function` type. The `anthropic/main.go:21-37` `stopReasonClaude2OpenAI` maps Claude stop reasons (`end_turn`, `tool_use`) to OpenAI equivalents (`stop`, `tool_calls`).

**Multimodal.** Image content is handled via `Message.ParseContent()` (`relay/model/message.go:41-80`) which extracts text and `image_url` parts from the OpenAI content array. The Anthropic adaptor fetches and base64-encodes images (`anthropic/main.go:132-139`). The `relay/controller/image.go` handles image generation relay with size/quality validation and cost ratio computation.

**Provider count.** The channel type enum (`relay/channeltype/define.go`) defines 56 channel types. The adaptor registry (`relay/adaptor.go`) explicitly maps 27 API types; the remaining channel types (Azure, OpenAI-compatible, etc.) fall through to the `OpenAI` adaptor or `Custom` handling.

> **Editor's note.** Correction: `GetAdaptor` maps 19 API types (not 27) and the `Adaptor` interface has nine methods; the other channel types fall through to the OpenAI adaptor.

Citations: [relay/adaptor.go:27-69](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor.go#L27-L69) · [relay/adaptor/anthropic/main.go:39-146](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor/anthropic/main.go#L39-L146) · [relay/adaptor/anthropic/main.go:210-247](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor/anthropic/main.go#L210-L247) · [relay/adaptor/openai/adaptor.go:84-96](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/adaptor/openai/adaptor.go#L84-L96) · [relay/model/general.go:24-69](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/model/general.go#L24-L69)

### How are API keys, users and tenants managed? (answered)

**User model.** Users have roles (`model/user.go:19-24`): `RoleGuestUser` (0), `RoleCommonUser` (1), `RoleAdminUser` (10), `RoleRootUser` (100). Fields include `username`, `password` (bcrypt-hashed), `quota`, `used_quota`, `request_count`, `group` (for routing), and OAuth IDs for GitHub, WeChat, Lark, OIDC (`model/user.go:34-53`). The root user is auto-created on first run with password `123456` (`model/main.go:24-64`).

**Auth methods.** The system supports password-based login (session cookie via `gin-contrib/sessions`), Bearer access tokens (for API/management), and OAuth (GitHub, WeChat, Lark, OIDC). The `middleware/auth.go` provides three auth middleware for the admin API: `UserAuth()`, `AdminAuth()`, `RootAuth()` checking `minRole` (`auth.go:15-71`). For the relay API, `TokenAuth()` (`auth.go:91-151`) validates user tokens sent as `Authorization: Bearer sk-...` headers.

**Virtual keys (Tokens).** Each token (`model/token.go:23-37`) is a 48-char `Key` linked to a `UserId`. Tokens have statuses (enabled/disabled/expired/exhausted), `RemainQuota`, `UnlimitedQuota`, `Models` (comma-separated whitelist), `Subnet` (CIDR restriction), and `ExpiredTime`. The token prefix `sk-` is stripped during parsing, and the remaining key is validated via Redis cache or DB (`model/cache.go:28-56`).

**Token scoping.** On validation, `middleware/auth.go:125-131` checks if the requested model is in the token's allowed model list. The `Subnet` field is verified against `c.ClientIP()` (`auth.go:104-109`). An advanced `sk-{key}-{channelId}` format allows admins to pin a request to a specific channel (`auth.go:135-142`).

**Admin UI/API.** The REST API (`router/api.go`) exposes CRUD for channels (`/api/channel`), tokens (`/api/token`), users (`/api/user`), redemptions (`/api/redemption`), and system options (`/api/option`). Role-based access control gates each route. Three React frontends (default, berry, air) provide the web UI.

**Upstream credential storage.** Each Channel stores its provider API key in the `Key` field (plaintext in DB, `model/channel.go:21-41`). Additional config (region, SK, AK for AWS/Vertex) goes in the `Config` JSON field. The channel's key is injected as `Authorization: Bearer {key}` on outbound requests (`middleware/distributor.go:73`).


Citations: [middleware/auth.go:91-151](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/auth.go#L91-L151) · [model/token.go:23-37](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/token.go#L23-L37) · [model/token.go:62-104](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/token.go#L62-L104) · [model/user.go:19-53](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/user.go#L19-L53) · [router/api.go:12-121](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/router/api.go#L12-L121)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limits.** The `middleware/rate-limit.go` provides per-IP rate limiting with configurable windows: `GlobalAPIRateLimit` (480 requests/3 min), `GlobalWebRateLimit` (240/3 min), `CriticalRateLimit` (20/20 min), and upload/download limits. Supports both Redis-based (`redisRateLimiter` via LPUSH/EXPIRE/TTL checks) and in-memory (`memoryRateLimiter`) backends. The `common/rate-limit.go` `InMemoryRateLimiter` uses a sliding-window approach with per-key timestamp queues cleaned by a background goroutine.

**Quota system.** Users have a `quota` field (integer) and `used_quota` (`model/user.go:48-49`). Tokens have their own `remain_quota` and `unlimited_quota` (`model/token.go:31-33`). The quota is denominated in internal units where 1 unit = $0.002/1K tokens (configurable via `QuotaPerUnit` in `common/config/config.go:21`).

**Pre-consumption.** Before sending a request, the system pre-consumes quota: `preConsumeQuota` in `controller/helper.go:68-95` estimates prompt tokens via the OpenAI tokenizer, adds max_tokens, multiplies by `modelRatio × groupRatio`, and checks the user + token balance. It then decrements the user's Redis-cached quota (`CacheDecreaseUserQuota`) and records a pre-consumption on the token. If user quota > 100× the estimate, pre-consumption is skipped (trusted user).

**Post-consumption.** After the response, `postConsumeQuota` (`controller/helper.go:97-141`) computes actual usage from `usage.prompt_tokens` and `usage.completion_tokens`, applies `completionRatio` for output tokens, calculates `quota = ceil((promptTokens + completionTokens × completionRatio) × modelRatio × groupRatio)`, adjusts the token quota by the delta, updates the user quota cache, writes a `Log` record, and increments channel used_quota.

**Price tables.** `relay/billing/ratio/model.go:27-622` contains an exhaustive `ModelRatio` map covering ~600+ models across OpenAI, Anthropic, Google, Baidu, Ali, Zhipu, Xunfei, Moonshot, DeepSeek, Mistral, Groq, Cohere, etc. The `CompletionRatio` map (`model.go:624-633`) adjusts for providers that price output tokens differently. `GetModelRatio` (`model.go:686-710`) looks up by exact name then falls back to defaults. Group-level multipliers are in `relay/billing/ratio/group.go:10-13`.

**Batch updates.** When `BATCH_UPDATE_ENABLED` is set, quota updates are queued into in-memory records and flushed periodically (`BatchUpdateInterval`, default 5s) to reduce DB write pressure (`model/main.go:86-89`). Without batching, each quota change issues an immediate SQL `UPDATE ... SET quota = quota +/- N`.

**Storage.** Usage logs are stored in the `logs` table (or a separate `LOG_SQL_DSN` database) with fields for user, token, model, prompt/completion tokens, quota consumed, elapsed time, and stream flag (`model/log.go:15-32`).


Citations: [middleware/rate-limit.go:19-91](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/rate-limit.go#L19-L91) · [relay/billing/ratio/model.go:27-90](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/billing/ratio/model.go#L27-L90) · [relay/billing/ratio/model.go:686-710](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/billing/ratio/model.go#L686-L710) · [common/config/config.go:86-99](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L86-L99)

### How are caching and guardrails implemented? (answered)

**Caching.** There is no semantic caching of LLM responses (no exact-match or semantic cache for chat completions). The `middleware/cache.go` sets only HTTP-level Cache-Control headers (`max-age=604800` for static assets, `no-cache` for the root route) — this is purely for the web frontend, not for API responses. The system caches database entities in Redis: tokens are cached by key with a TTL of `SYNC_FREQUENCY` seconds (`model/cache.go:28-56`), user groups are cached (`cache.go:58-74`), user quotas and enabled status are cached (`cache.go:88-149`), and group→models mappings are cached (`cache.go:151-168`). The in-memory channel index (`group2model2channels`) is refreshed on a timer (`cache.go:170-225`). None of these caches store model outputs.

**PII redaction / prompt injection / moderation.** The system has no built-in PII redaction, prompt-injection detection, or output moderation guardrails. It is a transparent proxy that does not inspect or sanitize request/response content. The `/v1/moderations` endpoint is simply relayed like any other endpoint (`router/relay.go:46`), passing through to the upstream provider's moderation API — not a local check. The only content-level control is the `SystemPrompt` feature on channels (`model/channel.go:40`), which lets an admin override or inject a system prompt into requests, but this is additive, not a guardrail.

**Blacklist.** A simple in-memory user ban list exists (`common/blacklist/main.go`) that blocks banned user IDs from auth. This is operational, not content-based.

**Turnstile.** CAPTCHA (Turnstile) support is available for registration, login, and password reset endpoints via `middleware/turnstile-check.go`, configured through `TurnstileSiteKey` and `TurnstileSecretKey` (`common/config/config.go:89-90`).

**Plugin/hook points.** There are no formal plugin or webhook systems. The adaptor pattern (`relay/adaptor.go`) is the extension mechanism for adding new providers — implementing the Adaptor interface gives you full control over request/response conversion. The message-pusher (`common/message/message-pusher.go`) can send notifications to external systems for channel disable events, but this is one-directional alerting, not an interceptor.


Citations: [common/config/config.go:86-91](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L86-L91) · [model/cache.go:28-56](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/cache.go#L28-L56)

### How is it observed, deployed and scaled? (answered)

**Language and concurrency.** One API is written in Go using the Gin web framework (Go's analog of Express/Flask). Go's goroutine-based concurrency model and the Gin framework's non-blocking I/O provide good per-instance throughput. The server starts in `main.go:103-120`: a single Gin engine with middleware stack (session store, request ID, language, logger, rate limits) and route registration. No explicit worker count or multi-process configuration is visible — Gin runs as a single process with goroutine-based concurrency.

**Logging.** The `middleware/logger.go` sets up standard Gin request logging (method, path, status code, latency, client IP). Application-level logging uses `common/logger/logger.go` with log levels (Info, Error, Debug, SysLog). `config.DebugEnabled` and `config.DebugSQLEnabled` toggle verbose logging and SQL query logging. `ONLY_ONE_LOG_FILE` controls log file rotation (`config.go:159`).

**Metrics and observability.** No OpenTelemetry, Prometheus, or Datadog integration is present. The only built-in metric is the channel success-rate monitor (`monitor/metric.go`), which uses in-memory ring buffers per channel. There are no external metrics exporters.

**Database.** The system supports SQLite (default, file-based), MySQL, and PostgreSQL (`model/main.go:67-109`). Schema is auto-migrated via GORM's `AutoMigrate` on startup. A separate `LOG_SQL_DSN` can route usage logs to a secondary database. Connection pooling is configured via `SQL_MAX_IDLE_CONNS`, `SQL_MAX_OPEN_CONNS`, and `SQL_MAX_LIFETIME` environment variables (`model/main.go:214-217`).

**Deployment.** A multi-stage Dockerfile builds static Go binary + React frontends. The `docker-compose.yml` provides a full stack with Redis, MySQL, and the One API service. Environment variables control all configuration: `SQL_DSN`, `REDIS_CONN_STRING`, `SESSION_SECRET`, `NODE_TYPE`, `SYNC_FREQUENCY`, etc.

**High-availability / multi-node.** The `NODE_TYPE` environment variable (`config.go:105`) supports "master" vs "slave" nodes. Only the master node runs database migrations. Slave nodes can serve API traffic but delegate configuration reads to the database and rely on `SYNC_FREQUENCY` to refresh their in-memory channel cache. The `FRONTEND_BASE_URL` env var allows slave nodes to offload the React frontend to a separate server. There is no load balancer or clustering logic in the application itself — HA relies on an external LB + shared DB/Redis.

**Health checks.** The docker-compose defines a healthcheck at `http://localhost:3000/api/status`. The `/api/status` endpoint (`controller/misc.go`) provides a simple status probe.


Citations: [main.go:29-124](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/main.go#L29-L124) · [model/main.go:67-109](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/main.go#L67-L109) · [common/config/config.go:105-111](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L105-L111) · [monitor/metric.go:1-79](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/metric.go#L1-L79)
