# QuantumNous/new-api

> AGPL fork of One API that converts between OpenAI, Claude and Gemini formats, with weighted routing, tiered billing and JS task plugins.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/QuantumNous/new-api (reviewed at commit `973cf8ef4600947a4270e95ada7916740fa8264c`, 2026-10-06)
- Stars: 49331 · Language: Go · License: AGPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/new-api/

## Overview

New API is a self-hosted LLM gateway in Go. It is a fork of One API: the README credits One API (MIT) as its base, and the core still shows it. The `Channel` table, the `TokenAuth` and `Distribute` middleware, the `group → model → channels` cache, the `sk-<key>-<channelId>` pin trick and the ratio-based quota system all come from One API. On top of that base, the fork has added a lot, and the result is a much larger and more capable product. It is now maintained under the QuantumNous organisation and licensed AGPL-3.0, with extra attribution terms in `NOTICE`.

The biggest functional differences from One API are on the protocol side. New API accepts OpenAI Chat Completions, the OpenAI Responses API (HTTP and WebSocket), Anthropic `/v1/messages`, Gemini `/v1beta/models/*`, Realtime, rerank, Midjourney and video endpoints. It converts between those formats through `relaykit`, a separate Go module. A Claude Code client can therefore call a GPT or Gemini channel, and an OpenAI SDK can call a Claude channel. One API only speaks OpenAI format inbound.

The other additions are operational. They include weighted and tiered channel selection with session affinity and auto-group failover, a billing engine with expression-based tiered prices, payment providers and subscriptions, JavaScript task plugins for async video and image providers, passkeys, TOTP and Casbin authorisation, ClickHouse for logs, and a new React 19 console. The audience is the same as One API's (teams and resellers sharing upstream keys), but the scope is much wider.

## Architecture

```mermaid
flowchart LR
  C["Client: OpenAI / Claude / Gemini SDK"] --> R["Gin relay router"]
  R --> A["TokenAuth (Bearer, x-api-key, x-goog-api-key)"]
  A --> RL["ModelRequestRateLimit"]
  RL --> D["Distribute -> SelectChannelForRequest"]
  D --> CC["Channel cache: priority + weight"]
  D --> CTL["controller.Relay: billing + retry loop"]
  CTL --> H["Format handlers: Text / Claude / Gemini / Responses"]
  H --> AD["Channel adaptor"]
  AD --> RK["relaykit converters"]
  AD --> UP["Upstream provider"]
  CTL --> B["Billing session: pre-consume / settle / refund"]
  B --> DB["SQL DB + Redis; logs to SQL or ClickHouse"]
  TP["JS task plugins (moejs)"] --> CTL
```

| Component | Path | Role |
|---|---|---|
| Relay router | `router/relay-router.go`, `router/task-plugin-protocol-router.go` | OpenAI, Claude, Gemini, Realtime and Midjourney routes; plugin-claimable Responses, image and video routes |
| Auth | `middleware/auth.go`, `service/authz/` | Token auth for relay calls; sessions, PATs, passkeys and Casbin for the console |
| Distributor | `middleware/distributor.go`, `service/channel_select.go` | Model extraction, token model limits, pin, affinity, auto-group channel choice |
| Channel cache | `model/channel_cache.go` | Priority-tiered, weight-based random pick |
| Relay controller | `controller/relay.go` | Billing reservation, retry loop, error policy |
| Format handlers | `relay/*_handler.go` | `TextHelper`, `ClaudeHelper`, `GeminiHelper`, `ResponsesHelper`, image, audio, rerank |
| Adaptors | `relay/channel/*`, `relay/relay_adaptor.go` | One package per provider; converts from OpenAI, Claude, Gemini and Responses inputs |
| relaykit | `relaykit/` | Independent module: DTOs and converters between Chat, Responses, Claude Messages and Gemini |
| Billing | `relay/request_billing.go`, `service/`, `pkg/billingexpr/` | Pre-consume, settle, refund, tiered expressions, violation fees |
| Task plugins | `plugins/tasks/*`, `pkg/jsplugin/` | JS plugins for async providers (Kling, Sora, Veo, Vidu, Jimeng and others) |
| Web console | `web/` | React 19 + TypeScript, Rsbuild, TanStack Router/Query, Tailwind |

## How a request flows

Take a Claude Code client calling `POST /v1/messages` with `x-api-key: sk-...`:

1. **Route and authenticate.** `/v1/messages` sits in the `/v1` group behind `TokenAuth`, `ModelRequestRateLimit` and `Distribute`, and calls `controller.Relay(c, RelayFormatClaude)` ([relay-router.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/router/relay-router.go#L71-L99)). `TokenAuth` copies `x-api-key` into `Authorization` for `/v1/messages`, and `x-goog-api-key` or `?key=` for Gemini paths, then parses the key the same way One API does ([auth.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/auth.go#L524-L564)).
2. **Pick a channel.** `SelectChannelForRequest` uses a pinned channel if one is set (admin pin, task plugin or origin task). On the first attempt it uses a session-affinity channel. Otherwise it makes a random eligible pick ([channel_select.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L283-L377)).
3. **Reserve billing.** `Relay` validates the request, builds `RelayInfo` and calls `PrepareRequestBilling`. That runs the optional sensitive-word check, estimates tokens, resolves the price and pre-consumes quota ([relay.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L68-L146), [request_billing.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/request_billing.go#L24-L67)).
4. **Convert.** `ClaudeHelper` applies model mapping, injects or overrides the channel's system prompt, and then either passes the body through or calls the adaptor's `ConvertClaudeRequest`. It also strips disabled fields and applies parameter overrides ([claude_handler.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/claude_handler.go#L37-L128)). On an OpenAI channel, `ConvertClaudeRequest` calls `service.ConvertRequest(..., RelayFormatOpenAI, ...)`, which goes through relaykit. On a Claude channel the Claude body stays native.
5. **Send and settle.** `DoRequest` sends the body and `DoResponse` streams the reply back in the client's format. Then `PostTextConsumeQuota` settles the real cost ([claude_handler.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/claude_handler.go#L130-L157)).
6. **Retry.** On failure, `DecideRelayRetry` classifies the error (channel errors retry; pinned or strict-session requests stop; status-code rules apply). The loop picks again with a higher retry index, which moves to the next priority tier. When all attempts fail, `RefundFailedRequestBilling` returns the reservation and may charge a configured violation fee ([relay.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L148-L221), [request_billing.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/request_billing.go#L69-L81)).

## Key components

### What changed from One API

| Area | One API | New API |
|---|---|---|
| Inbound formats | OpenAI only | OpenAI Chat, Responses (HTTP + WS), Claude Messages, Gemini, Realtime, rerank, Midjourney, video |
| Conversion | Per-adaptor OpenAI → provider | Adaptors accept OpenAI, Claude, Gemini and Responses; relaykit converts between all four |
| Channel pick | Uniform random in top priority; `Weight` unused | Weighted random per priority tier, affinity, pins, auto-group failover, multi-key channels |
| Billing | Model × group ratio, pre/post consume | Billing sessions, per-call prices, tiered expressions, violation fees, payments, subscriptions |
| Async tasks | None | JS task plugins with polling and settlement |
| Auth | Password, OAuth, sessions | Adds JWT sessions with refresh, passkeys, TOTP, PATs, Casbin, step-up verification |
| Logs | SQL (optional separate DSN) | SQL or ClickHouse |
| UI | Three themes (Semantic UI, MUI, Semi) | One React 19 / TypeScript console |
| License | MIT | AGPL-3.0 with attribution terms |

The data model shows the lineage directly. New API's `Channel` keeps One API's fields and adds `StatusCodeMapping`, `AutoBan`, `Tag`, `ParamOverride`, `HeaderOverride` and multi-key `ChannelInfo` ([model/channel.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L23-L60)).

### Channel selection

`GetRandomSatisfiedChannel` filters candidates for the group and model, falls back to a normalised model name, and sorts the distinct priorities. It then uses the retry index to choose the tier and draws by `Weight` inside it. When all weights are zero, the draw is uniform ([channel_cache.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel_cache.go#L117-L217)). With the `auto` token group, `CacheGetRandomSatisfiedChannel` walks the configured groups in order and resets the tier counter when it moves to the next group ([channel_select.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L114-L170)). A multi-key channel rotates its keys randomly or by polling.

### Adaptors and relaykit

The `Adaptor` interface grew from One API's single `ConvertRequest` to separate entry points: `ConvertOpenAIRequest`, `ConvertClaudeRequest`, `ConvertGeminiRequest`, `ConvertOpenAIResponsesRequest`, plus embedding, rerank, audio and image ([adapter.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/channel/adapter.go#L17-L34)). `GetAdaptor` maps 37 API types to adaptors, including Codex, Dify, Jina and a `newapi` adaptor for chaining gateways ([relay_adaptor.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/relay_adaptor.go#L50-L128)). Many adaptors call relaykit instead of hand-writing conversions. For example, the Claude adaptor's `ConvertOpenAIRequest` is one `service.ConvertRequest(..., RelayFormatClaude, ...)` call ([claude/adaptor.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/channel/claude/adaptor.go#L118-L127)). relaykit has no dependency on Gin, the database or billing. It records each conversion's steps and a quality grade (Good, Fair or Discouraged), and multi-hop routes such as Claude to Gemini go through an intermediate format.

### Billing

`PrepareRequestBilling` reserves a charge once per request, and the reservation survives channel retries. Prices come from model ratios or fixed per-call prices, group ratios and, optionally, `pkg/billingexpr` expressions evaluated against actual usage for tiered pricing. Top-ups and subscriptions are paid through Stripe, Epay, Creem or Waffo handlers in `controller/`. Usage logs record channel, group, ratios and a JSON `Other` field, in the main database or ClickHouse.

### Task plugins

Long-running providers (video, some image models) are JavaScript plugins in `plugins/tasks/`, run by the moejs engine in `pkg/jsplugin/`. A plugin declares its models, channel types, routes, accepted protocols (`openai_responses`, `openai_video`) and a usage schema for billing ([kling/plugin.js](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/plugins/tasks/kling/plugin.js#L1-L39)). The host registers the Responses, image and video routes from a protocol registry, so a plugin can claim a model and fall back to the normal relay otherwise ([task-plugin-protocol-router.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/router/task-plugin-protocol-router.go#L13-L60)).

### Content controls

There is no semantic cache and no PII or moderation guardrail. There is an admin-managed sensitive-word list: an Aho-Corasick match over the prompt text that rejects the request before billing ([sensitive.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/sensitive.go#L35-L49)). `/v1/moderations` is relayed upstream like any other call.

## Extending it

- **OpenAI-compatible provider.** Add a channel type that maps to the OpenAI adaptor, or use the advanced custom channel type and parameter or header overrides.
- **New native provider.** Add a package under `relay/channel/`, an API type and a `GetAdaptor` case. Reuse `service.ConvertRequest` so relaykit does the format work.
- **Async provider.** Write a JS task plugin with `meta`, routes, request builders, polling and usage extraction. No Go changes are needed.
- **Reuse the converters.** `relaykit` is published as its own Go module, so another gateway can import it without new-api.

## Running it

- **Docker Compose.** The bundled file runs `calciumion/new-api` with PostgreSQL and Redis on port 3000. MySQL and a ClickHouse `LOG_SQL_DSN` are commented options ([docker-compose.yml](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/docker-compose.yml#L17-L59)). SQLite is the default without `SQL_DSN`.
- **Build.** The Dockerfile builds the web console with Bun, compiles the Go binary and runs it on Debian slim ([Dockerfile](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/Dockerfile#L1-L39)). A systemd unit and an Electron wrapper are included.
- **Several instances.** As in One API, set `NODE_TYPE=slave` on replicas so only the master runs migrations and master-only jobs. Share the database, Redis and `SESSION_SECRET`, and set `NODE_NAME` for audit logs. There is no built-in clustering; put a load balancer in front.
- **Background jobs.** Startup launches channel and option sync, task plugin sync, policy sync, quota data aggregation, optional balance updates and Codex credential refresh ([main.go](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/main.go#L120-L135)).

## Strengths and caveats

- **Strength: multi-protocol in and out.** Claude, Gemini and Responses clients can use any channel, and the conversion layer is isolated, graded and tested.
- **Strength: real load balancing.** Weights, priority-tier failover, session affinity and auto-group failover cover what One API only stubbed.
- **Strength: built for running a paid service.** Tiered prices, top-ups, subscriptions, redemption codes and detailed usage logs work out of the box.
- **Caveat: size.** The controller, service and relay layers are large and interconnected. Billing and retries have many paths, and the repository's own contributor rules mark billing as a mandatory-read area.
- **Caveat: AGPL-3.0.** Running a modified version as a network service obliges you to publish the changes and keep the attribution required in `NOTICE`. This is stricter than One API's MIT licence.
- **Caveat: no response cache or content guardrails** beyond the keyword filter, and no OpenTelemetry. Observability is the request log, performance metrics and the console.
- **Caveat: fast-moving.** Frequent changes to protocols and settings mean upgrades need testing, especially if you depend on pass-through or parameter overrides.

*Sources: code at 973cf8e, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (22 pages), verified Q&A.*

## How QuantumNous/new-api answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

Requests are routed through a **middleware pipeline** that selects a channel (provider endpoint) for each request. The `Distribute()` middleware in `middleware/distributor.go` is the core dispatcher: it extracts the model name from the request body, applies **token model limits** (optional per-key model whitelist), resolves the user's group, and calls `SelectChannelForRequest` in `service/channel_select.go`.

**Selection priority (service/channel_select.go:283-376):** (1) A **pinned channel** (set by task plugins, origin-task affinity, or admin override) wins unconditionally. (2) On the first attempt (retry==0), **session affinity** is checked: if the same group+model was routed to a channel recently, that channel is reused (sticky sessions). (3) Otherwise `CacheGetRandomSatisfiedChannel` picks a random eligible channel by group+model+priority.

**Load balancing** is weighted: each `Channel` has a `Weight` field (`model/channel.go:31`), and multi-key channels support `MultiKeyModeRandom` (random selection among enabled keys) or `MultiKeyModePolling` (round-robin via an atomic polling index) (`model/channel.go:249-289`). Channels also have a `Priority` field (`model/channel.go:45`) for tiered fallback — when a priority level is exhausted, the system moves to the next priority group.

**Retry and fallback** work together: `controller/relay.go:158` loops while `retryParam.GetRetry() ≤ common.RetryTimes`. Each iteration calls `SelectChannelForRequest` again, which may select a different channel. The `DecideRelayRetry` function in `service/relay_error.go:21` decides whether to retry based on the error: channel-level errors trigger retries, 4xx status codes by default do not, and operation_setting can configure retryable status codes. Cross-group retry ("auto" group mode) cascades through configured groups: when one group has no eligible channels, the next group is tried (`service/channel_select.go:121-167`).

**Model aliases** work via `model_mapping` field on each channel (`model/channel.go:42`), parsed by `ModelMappedHelper` in `relay/helper/model_mapped.go:14`. It supports chain redirection with cycle detection. The `ResolveTaskModelAlias` in `model/task_model_alias.go:38` resolves task-plugin aliases through ASCII-folded name lookups. **Health checks** run periodically via the channel test endpoint (`controller/channel-test.go`) and Uptime Kuma integration (`controller/uptime_kuma.go:18`). Channels with `auto_ban` enabled (`model/channel.go:46`) are automatically disabled when their keys are all disabled.

> **Editor's note.** Correction: `controller/uptime_kuma.go` only reads an external Uptime Kuma status page for the console; it does not health-check channels. Channel health comes from scheduled channel tests (`CHANNEL_TEST_FREQUENCY` / monitor settings) plus auto-ban on errors.

Citations: [middleware/distributor.go:34-130](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/distributor.go#L34-L130) · [service/channel_select.go:80-204](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L80-L204) · [service/channel_select.go:283-377](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L283-L377) · [model/channel.go:23-60](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L23-L60) · [model/channel.go:206-290](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L206-L290) · [controller/relay.go:148-180](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L148-L180)

### How are different provider APIs unified? (answered)

The project unifies **40+ provider APIs** (listed in `constant/channel.go:4-63`) through an **OpenAI-compatible schema** as the lingua franca. Inbound requests arrive in OpenAI format (chat/completions, embeddings, images, audio, rerank, etc.) and are converted to each provider's native protocol via channel-specific **Adaptor** implementations.

Each channel type implements the `channel.Adaptor` interface defined in `relay/channel/adapter.go:17-34`. The key methods are `ConvertOpenAIRequest`, `ConvertClaudeRequest`, `ConvertGeminiRequest`, `DoRequest`, and `DoResponse`. The factory function `GetAdaptor` in `relay/relay_adaptor.go:50-128` maps an `APIType` constant to a concrete adaptor (e.g., APITypeAnthropic → `&claude.Adaptor{}`).

**Request translation** lives in `relaykit/relayconvert/internal/`, organized by target protocol:
- `oai_chat/to_claude_messages_req.go` — OpenAI chat → Claude Messages
- `oai_chat/to_gemini_chat_req.go` — OpenAI chat → Gemini
- `claude_messages/to_oai_chat_req.go` — Claude → OpenAI chat
- `gemini_chat/to_oai_chat_req.go` — Gemini → OpenAI chat

**Response translation** uses the reverse direction (`*_resp.go` files in the same directories). The `service/convert.go` wrapper exposes `ResponseOpenAI2Claude`, `StreamResponseOpenAI2Gemini`, etc. which call into `relaykit` conversion functions.

**Streaming** is handled by translating each SSE chunk from the upstream format back into OpenAI-compatible `ChatCompletionsStreamResponse` chunks. The `DoResponse` method on each adaptor receives the raw HTTP response and parses/extracts usage. The helper functions in `relay/helper/common.go` handle streaming headers (`SetEventStreamHeaders`) and format-specific chunk serialization (`ClaudeData`, etc.).

**Tool calls and multimodal** are fully converted: videos, images, audio, and tool_use blocks are translated between schemas in the internal converters.

**Relay format selection** happens in the router (`router/relay-router.go:97-153`): the controller receives a `RelayFormat` enum (OpenAI, Claude, Gemini, Embedding, etc.) based on the URL path structure. `GetAndValidateRequest` in `relay/helper/valid_request.go:22` dispatches to the appropriate request parser per format. The `relay_handler` in `controller/relay.go:33` then dispatches to a mode-specific relay helper (TextHelper, ImageHelper, AudioHelper, ResponsesHelper, GeminiHelper, etc.).


Citations: [relaykit/relayconvert/internal/oai_chat/to_claude_messages_req.go:1-5](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relaykit/relayconvert/internal/oai_chat/to_claude_messages_req.go#L1-L5) · [relaykit/relayconvert/internal/oai_chat/to_gemini_chat_req.go:1-5](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relaykit/relayconvert/internal/oai_chat/to_gemini_chat_req.go#L1-L5) · [relay/channel/adapter.go:17-34](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/channel/adapter.go#L17-L34) · [relay/relay_adaptor.go:50-128](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/relay/relay_adaptor.go#L50-L128) · [constant/channel.go:4-65](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/constant/channel.go#L4-L65)

### How are API keys, users and tenants managed? (answered)

The system has a **three-tier credential model**: Users → Tokens (virtual API keys with quotas) → Upstream channel keys.

**Virtual API keys** are represented by the `Token` model (`model/token.go:14-33`). Each Token belongs to a User, has its own `Key` string (the `sk-` prefix bearer token clients use), `RemainQuota` / `UnlimitedQuota`, `ModelLimits` (per-model allowlist), `AllowIps` (IP restriction), `ExpiredTime`, `Group` assignment, and `CrossGroupRetry` / `AutoGroups` for multi-group routing. Tokens are sent to clients; the gateway looks up the token by its key, identifies the user, and then maps it to an upstream channel.

**Upstream credential storage** is in the `Channel` model (`model/channel.go:23-60`). Each Channel stores its `Key` (the upstream API key, possibly multi-key via JSON array or newline-separated), `BaseURL`, API `Type`, and `ModelMapping`. Multi-key Channels support `MultiKeyModeRandom` or `MultiKeyModePolling` for key rotation (`model/channel.go:249-289`).

**Users and teams** use the `User` model (`model/user.go:79-115`) with fields for `Username`, `Password` (hashed), `Role` (common/admin/root), `Status`, `Email`, `Quota`/`UsedQuota`, `Group` assignment, and OAuth bindings (GitHub, Discord, WeChat, OIDC, Telegram, LinuxDo). User sessions use JWT access tokens with refresh token rotation (`model/user_session.go:42+`, `service/auth_token.go:19+`).

**Authentication methods:** (1) **Relay API auth** (`middleware/auth.go:524-564`) — `TokenAuth()` middleware extracts the bearer token from `Authorization: Bearer sk-...`, or from `x-api-key` (Anthropic) or `x-goog-api-key` (Gemini) headers for protocol compatibility. (2) **Dashboard auth** — browser sessions (cookies + JWTs) via `UserAuth()/AdminAuth()/RootAuth()` (`middleware/auth.go:132-148`), plus **Personal Access Tokens (PAT)** (`model/user_access_token.go:38-52`) with scoped permissions. (3) **Step-up verification** (`service/security_verification.go:18-48`) for sensitive actions (password change, 2FA setup, admin user management) using password, passkey, TOTP, or OAuth challenge. (4) **WebAuthn/Passkeys** and **TOTP** are fully supported.

**Admin UI and API** are served by the React frontend at the `/` path, with management endpoints under `/api/*` (routed in `router/main.go`). The admin dashboard manages channels, tokens, users, groups, pricing, and log viewing.


Citations: [model/token.go:14-33](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/token.go#L14-L33) · [model/user.go:79-115](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/user.go#L79-L115) · [model/channel.go:23-60](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L23-L60) · [middleware/auth.go:451-564](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/auth.go#L451-L564) · [model/user_access_token.go:38-52](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/user_access_token.go#L38-L52) · [service/security_verification.go:18-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/security_verification.go#L18-L80)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limiting** operates at multiple levels. (1) **Global rate limits** (`middleware/rate-limit.go:160-178`) — `GlobalWebRateLimit` and `GlobalAPIRateLimit` apply fixed-window Redis Lua scripts (or in-memory fallback) keyed by IP address. (2) **Model-level rate limits** (`middleware/model-rate-limit.go`) — `ModelRequestRateLimit()` tracks total requests and success counts per user using a Redis list (or in-memory bucket) with configurable max count and duration. Token bucket limiting (via `common/limiter`) is available as an alternative. (3) **Critical route rate limits** throttle brute-force attempts on auth endpoints. All rate limiters return `Retry-After` headers on 429 responses.

**Quota and budget tracking** is a two-phase system: **pre-consume** then **settle**. On request start, `PrepareRequestBilling` in `relay/request_billing.go` pre-computes the estimated quota cost based on model pricing and deducts it from the token's `RemainQuota` and the user's `Quota`. On completion, `service/task_billing.go` and `service/text_quota.go` calculate the actual usage (token counts from upstream response) and settle the difference (refund or additional charge).

**Price tables** are configured per-model in `model/model_pricing_config.go:24-57`. The `PricingValues` map supports key-based pricing (per-token, per-image, per-second, etc.). Group-level ratio multipliers and model-level ratios are applied through `ratio_setting`. The `price.go` helper in `relay/helper/price.go:44` resolves group ratios and tiered billing expressions. **Tiered billing** (`pkg/billingexpr`) supports complex expression-based pricing (e.g., "first 1000 tokens free, then $0.01/1K"), evaluated against actual usage facts captured in `TieredBillingSnapshot`.

**Usage accounting** is stored in the `Log` model (`model/log.go:59-80`), which records every request: user ID, token name, model, `PromptTokens`/`CompletionTokens`, quota consumed, channel, group, IP, streaming flag, and an `Other` JSON field for extended billing metadata (task details, model price, ratios, tiered billing snapshot). The log database supports both the main DB and a **separate ClickHouse** instance for high-volume analytics (`docker-compose.yml:32`).

**Wallet top-ups** and **subscriptions** are supported through multiple payment providers (Stripe, Epay, Waffle/Pancake, Creem) with redemption codes and affiliate tracking.


Citations: [middleware/rate-limit.go:22-178](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/rate-limit.go#L22-L178) · [middleware/model-rate-limit.go:28-160](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/model-rate-limit.go#L28-L160) · [model/log.go:59-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/log.go#L59-L80) · [model/model_pricing_config.go:24-57](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/model_pricing_config.go#L24-L57) · [service/task_billing.go:1-79](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/task_billing.go#L1-L79) · [common/quota_math.go:1-52](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/common/quota_math.go#L1-L52)

### How are caching and guardrails implemented? (answered)

**Caching** is minimal in this gateway. The only explicit HTTP cache is the `Cache()` middleware in `middleware/cache.go:7-17`, which sets `Cache-Control: max-age=604800` (one week) on static frontend assets and `no-cache` on the root path. There is **no semantic/response caching** for LLM completions — each relay request is forwarded to the upstream provider without an intermediary cache layer.

**Redis** is used as a shared data store for rate limit counters (`middleware/rate-limit.go`), channel polling state (`model/channel.go:206-290`), and in-memory caches for user/token/channel lookups via the memory cache system (`common` package). But this is operational caching (avoiding repeated DB queries), not semantic response caching.

**Content guardrails** are similarly absent as a core system feature. There is no built-in PII redaction, moderation filter, content safety check, or prompt-injection detection middleware in the request pipeline. The router (`router/relay-router.go:156`) has a single `/v1/moderations` endpoint that relays to OpenAI's moderation API, but this is a passthrough — the gateway does not inject its own moderation.

**What the project provides instead:** (1) **Request policy logging** (`service/request_policy.go`) records every routing decision (channel selection, retries, errors) per request for audit visibility. (2) **Admin security verification** (`service/security_verification.go`) protects sensitive dashboard actions with step-up authentication (2FA, passkey, password re-entry) — this is an admin guardrail, not a content guardrail. (3) **Token-level access controls** — `ModelLimits` (`model/token.go:25`) restricts which models a token can use, `AllowIps` restricts by source IP, and `ExpiredTime` enforces token expiry. (4) **Channel auto-ban** (`model/channel.go:345-350`) — when a channel's upstream keys are all disabled, the channel is automatically marked unavailable.

**Plugin hooks** provide extensibility points: JavaScript task plugins (in `plugins/tasks/`, run via `pkg/jsplugin/`) can implement custom pre-processing and billing logic. The `task_plugin` middleware (`middleware/task_plugin.go`) and plugin protocol routes enable custom endpoint protocols to override the standard relay path. However, these are not content-moderation-specific hooks — they are general-purpose plugin slots for custom request handling.

**In summary**, this gateway focuses on routing, billing, and access control rather than content safety; guardrails like PII filtering, moderation, or injection detection would need to be added at the upstream provider level or through custom plugin development.

> **Editor's note.** Correction: there is one built-in content control: an admin-managed sensitive-word list checked (Aho-Corasick, case-insensitive) in `PrepareRequestBilling` before any upstream call, which rejects matching prompts (service/sensitive.go, relay/request_billing.go).

Citations: [middleware/cache.go:1-17](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/cache.go#L1-L17) · [middleware/rate-limit.go:22-107](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/rate-limit.go#L22-L107) · [model/token.go:14-33](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/token.go#L14-L33) · [service/security_verification.go:18-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/security_verification.go#L18-L80) · [router/relay-router.go:156-158](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/router/relay-router.go#L156-L158)

### How is it observed, deployed and scaled? (answered)

**Observability.** The project has **no OpenTelemetry integration** — there are no trace exporters, metric collectors, or OTel SDK references in the code. Observability is instead built on: (1) **File-based logging** (`logger/logger.go`) — `SetupLogger()` writes to a timestamped log file in the configured `--log-dir`, rotating at startup. The `LogInfo`/`LogWarn`/`LogError`/`LogDebug` functions write structured messages with context from the Gin request context. (2) **Request logging** (`model/log.go`) — every API request is recorded in the `Log` table (or ClickHouse log DB), storing tokens, quota, channel, latency (`UseTime`), IP, model, and extended billing metadata. This serves as the audit trail and usage analytics source. (3) **Performance metrics** (`pkg/perf_metrics/`) — relay latency and business rejection rates are captured at the controller level (`controller/relay.go:134`). (4) **Active connection stats** (`middleware/stats.go`) — an atomic counter tracks concurrent requests. (5) **Admin dashboard** — the React UI provides real-time usage dashboards, log exploration, and performance charts.

**Performance architecture.** The project is written in **Go** (compiled, with a green-thread concurrency model via `gopool`). It uses the **Gin** web framework, which is an async Go framework. The thread pool handles concurrent HTTP relay connections efficiently — each relay request is a lightweight goroutine. **Redis** offloads rate limiting state, channel polling indices, and cache lookups, reducing database pressure. The database layer supports **ClickHouse** as a separate log database (`docker-compose.yml:32`) to isolate analytics writes from transactional data.

**Deployment modes.** (1) **Docker** — the primary deployment path (`Dockerfile`) uses a multi-stage build: Bun builds the React frontend, Go compiles the binary, and a Debian slim image runs it on port 3000. The `docker-compose.yml` bundles the app with PostgreSQL (or MySQL) and Redis, with config via environment variables. (2) **Standalone binary** — `go build` produces a single static binary with embedded web frontend. (3) **Electron desktop wrapper** — `electron/` provides a desktop app wrapping the same web UI. (4) **systemd service** (`new-api.service`) is included for Linux native deployment.

**High availability.** The architecture is **single-node** — there is no built-in clustering or HA mode. However, the system is **stateless** (session state in Redis, config in database), so multiple instances can share the same Postgres+Redis backend. The `NODE_NAME` env var (`docker-compose.yml:38`) identifies nodes in audit logs. The `SESSION_SECRET` must be shared across nodes (`docker-compose.yml:41`). Graceful connection handling is configurable via `STREAMING_TIMEOUT` and `RELAY_IDLE_CONN_TIMEOUT`.

**Deployment configuration** is entirely environment-variable driven: `SQL_DSN`, `LOG_SQL_DSN`, `REDIS_CONN_STRING`, `TZ`, `BATCH_UPDATE_ENABLED`, `ERROR_LOG_ENABLED`, plus session security settings (`SESSION_SECRET`, `SESSION_COOKIE_SECURE`, `SESSION_COOKIE_TRUSTED_URL`).


Citations: [logger/logger.go:1-74](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/logger/logger.go#L1-L74) · [Dockerfile:1-40](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/Dockerfile#L1-L40) · [docker-compose.yml:1-50](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/docker-compose.yml#L1-L50) · [controller/relay.go:127-139](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L127-L139) · [model/log.go:59-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/log.go#L59-L80)
