# How are rate limits, budgets and cost tracking implemented?

> LLM gateways — a good answer covers: Per-key/user limits; token and spend budgets; price tables; usage accounting and where it is stored.

Canonical page: https://llms-technical-reviews.com/llm-gateways/q/limits-cost/

## Verdict

[LiteLLM](/p/litellm/) has the most complete cost attribution, with a maintained price table and budgets at every level. [New API](/p/new-api/) is the most complete for running a paid service. [Bifrost](/p/bifrost/) and [GPT-Load](/p/gpt-load/) offer dollar caps in a smaller package.

**Dollar budgets from built-in prices.** LiteLLM ships a price file with about 4,500 model entries. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is written to PostgreSQL in batches. Bifrost attaches budgets with reset windows, plus token and request limits, to virtual keys, teams and customers. Its counters live in each process, so the open-source build does not share them across nodes. GPT-Load gives each AccessKey lifetime or periodic caps in nano-USD, an RPM limit and a concurrency limit. Every request gets a frozen price receipt. [OmniRoute](/p/omniroute/) enforces daily and weekly USD caps per key. Prices come from a short built-in map (about 20 models), synced or operator prices in SQLite, and an estimate for unknown models.

**Quota units for resale.** [One API](/p/one-api/) pre-charges `(prompt + max_tokens) × modelRatio × groupRatio` in internal units, then settles on real usage. Its rate limits are per IP (480 API requests per 3 minutes by default), not per key. New API keeps a billing reservation across channel retries. It adds tiered price expressions, top-ups through Stripe, Epay or Creem, and logs in SQL or ClickHouse.

**Token counters in the data plane.** [Higress](/p/higress/) keeps a Redis token quota per consumer and decrements it after the response, so one large request can overshoot. It has no price table. [Agent Router](/p/agent-router/) computes token or CEL cost expressions and enforces `QuotaPolicy` in Envoy's rate-limit service. Neither has a dollar budget. [Plano](/p/plano/) applies `governor` token buckets keyed by a header value. Its dollar tracking only caps how much switching models may add within a session (`max_switch_spend_pct`).

**Hook points only.** [Portkey Gateway](/p/portkey-gateway/) has no price tables and stores no logs. Budgets depend on a `preRequestValidator` that nothing in the repository sets, and its Redis token bucket is not wired into the request path.

Pick: LiteLLM for per-team spend tracking and budgets.
Pick: New API or One API to sell or ration access with prepaid quota.
Pick: Higress or Agent Router for token-based rate limits on an Envoy fleet.

## Per-project answers

### diegosouzapw/OmniRoute (answered)

Rate limiting is implemented via `src/shared/utils/rateLimiter.ts` using a Redis-backed (or in-memory fallback) leaky bucket strategy. API keys carry `maxRequestsPerDay`, `maxRequestsPerMinute`, and rich `rateLimits` rules (parsed by `parseRateLimits()`). Cost tracking centers on `open-sse/services/providerCostData.ts` which provides a `KNOWN_MODEL_PRICING` map covering ~25 key models (e.g., gpt-4o at $2.50/$10 per 1M tokens, claude-sonnet-4.6 at $3/$15) and a `getModelPricing()` resolver that falls back through: provider-specific KNOWN match → default pricing DB table → generic model name match → free-model check → estimated fallback ($5/$15). The tier resolver (`tierResolver.ts`) layers operator-configured pricing over the defaults using a DB snapshot from `pricing`/`pricing_synced`/`models_dev_pricing` key-value namespaces and caches the merged result. Usage accounting is stored in two SQLite tables: `usage_history` (per-request rows with `tokens_input`/`tokens_output` and ISO timestamps) and `daily_usage_summary` (rolled-up aggregates), queried by `src/lib/db/usageSummary.ts` and `src/lib/db/usageLogs.ts`. The `GET /v1/usage` endpoint exposes aggregated usage. Spend limits are enforced by `apiKeyPolicyService.ts` (`enforceApiKeyLimits()`) which checks daily/weekly USD caps before allowing requests. Circuit breakers, connection cooldown (exponential backoff on per-credential failures), and model lockout all act as runtime cost-control mechanisms by stopping traffic to failing-expensive or exhausted endpoints. The `connectionBillingCatalog.ts` classifies connections as `subscription` (plan-included) vs `metered` (per-token) to budget routing.


Citations: [open-sse/services/tierResolver.ts:1-80](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/tierResolver.ts#L1-L80) · [src/lib/db/apiKeys.ts:100-145](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/src/lib/db/apiKeys.ts#L100-L145) · [open-sse/config/connectionBillingCatalog.ts:31-60](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/config/connectionBillingCatalog.ts#L31-L60) · [src/shared/utils/rateLimiter.ts:1-25](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/src/shared/utils/rateLimiter.ts#L1-L25)

### BerriAI/litellm (answered)

**Rate limits.** The `LiteLLM_BudgetTable` in `schema.prisma` defines `max_parallel_requests`, `tpm_limit` (tokens per minute), `rpm_limit` (requests per minute), and `tpd_limit` (tokens per day) at every level of the hierarchy (org, team, user, key, end-user, tag). These are enforced via the budget throttle middleware (`litellm/proxy/auth/budget_throttle.py`) and pre-call RPM/TPM checks in the routing layer (`router.py:8984-9033`).

**Spend and budget tracking.** `litellm/proxy/spend_tracking/` implements a batched spend counter (`spend_counter_batch.py`) that increments in-memory counters and flushes to PostgreSQL asynchronously on a cadence. The `LiteLLM_SpendLog*` tables in the Prisma schema record every request's token usage, cost, and metadata. Budget enforcement checks (`_virtual_key_max_budget_check`, `_virtual_key_soft_budget_check` in `litellm/proxy/auth/auth_checks.py`) compare accumulated spend against `max_budget`/`soft_budget` thresholds.

**Price table.** `model_prices_and_context_window.json` (80,644 lines) is the bundled pricing catalog. Every model entry specifies `input_cost_per_token`, `output_cost_per_token`, `output_cost_per_reasoning_token`, regional uplift multipliers, prompt-caching discounts, batch rate reductions, and modality-specific costs (audio, image, video, web-search, computer-use). The cost calculator (`litellm/cost_calculator.py`) uses a chain-of-responsibility pattern: each provider module registers a `cost_per_token()` function (e.g. `openai_cost_per_token`, `anthropic_cost_per_token`, `bedrock_cost_per_token` in `litellm/llms/{provider}/cost_calculation.py`), and the base resolver selects the right one.

**Usage accounting.** The final `response_cost`, `total_tokens`, and breakdown are stored in the `LiteLLM_SpendLog*` tables. Budget carry-forward across reset periods is handled by `carried_budget_state.py`. Enterprise users get PTU (pay-per-use) flat-cost pricing via `ptu_flat_cost_rollup.py`.


Citations: [schema.prisma:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/schema.prisma#L1-L60) · [litellm/router.py:8984-9033](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L8984-L9033) · [litellm/proxy/spend_tracking/spend_counter_batch.py:1-40](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/spend_tracking/spend_counter_batch.py#L1-L40) · [litellm/cost_calculator.py:1-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/cost_calculator.py#L1-L80) · [litellm/proxy/auth/auth_checks.py:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/auth_checks.py#L1-L60)

### QuantumNous/new-api (answered)

**Rate limiting** operates at multiple levels. (1) **Global rate limits** (`middleware/rate-limit.go:160-178`) — `GlobalWebRateLimit` and `GlobalAPIRateLimit` apply fixed-window Redis Lua scripts (or in-memory fallback) keyed by IP address. (2) **Model-level rate limits** (`middleware/model-rate-limit.go`) — `ModelRequestRateLimit()` tracks total requests and success counts per user using a Redis list (or in-memory bucket) with configurable max count and duration. Token bucket limiting (via `common/limiter`) is available as an alternative. (3) **Critical route rate limits** throttle brute-force attempts on auth endpoints. All rate limiters return `Retry-After` headers on 429 responses.

**Quota and budget tracking** is a two-phase system: **pre-consume** then **settle**. On request start, `PrepareRequestBilling` in `relay/request_billing.go` pre-computes the estimated quota cost based on model pricing and deducts it from the token's `RemainQuota` and the user's `Quota`. On completion, `service/task_billing.go` and `service/text_quota.go` calculate the actual usage (token counts from upstream response) and settle the difference (refund or additional charge).

**Price tables** are configured per-model in `model/model_pricing_config.go:24-57`. The `PricingValues` map supports key-based pricing (per-token, per-image, per-second, etc.). Group-level ratio multipliers and model-level ratios are applied through `ratio_setting`. The `price.go` helper in `relay/helper/price.go:44` resolves group ratios and tiered billing expressions. **Tiered billing** (`pkg/billingexpr`) supports complex expression-based pricing (e.g., "first 1000 tokens free, then $0.01/1K"), evaluated against actual usage facts captured in `TieredBillingSnapshot`.

**Usage accounting** is stored in the `Log` model (`model/log.go:59-80`), which records every request: user ID, token name, model, `PromptTokens`/`CompletionTokens`, quota consumed, channel, group, IP, streaming flag, and an `Other` JSON field for extended billing metadata (task details, model price, ratios, tiered billing snapshot). The log database supports both the main DB and a **separate ClickHouse** instance for high-volume analytics (`docker-compose.yml:32`).

**Wallet top-ups** and **subscriptions** are supported through multiple payment providers (Stripe, Epay, Waffle/Pancake, Creem) with redemption codes and affiliate tracking.


Citations: [middleware/rate-limit.go:22-178](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/rate-limit.go#L22-L178) · [middleware/model-rate-limit.go:28-160](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/model-rate-limit.go#L28-L160) · [model/log.go:59-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/log.go#L59-L80) · [model/model_pricing_config.go:24-57](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/model_pricing_config.go#L24-L57) · [service/task_billing.go:1-79](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/task_billing.go#L1-L79) · [common/quota_math.go:1-52](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/common/quota_math.go#L1-L52)

### songquanpeng/one-api (answered)

**Rate limits.** The `middleware/rate-limit.go` provides per-IP rate limiting with configurable windows: `GlobalAPIRateLimit` (480 requests/3 min), `GlobalWebRateLimit` (240/3 min), `CriticalRateLimit` (20/20 min), and upload/download limits. Supports both Redis-based (`redisRateLimiter` via LPUSH/EXPIRE/TTL checks) and in-memory (`memoryRateLimiter`) backends. The `common/rate-limit.go` `InMemoryRateLimiter` uses a sliding-window approach with per-key timestamp queues cleaned by a background goroutine.

**Quota system.** Users have a `quota` field (integer) and `used_quota` (`model/user.go:48-49`). Tokens have their own `remain_quota` and `unlimited_quota` (`model/token.go:31-33`). The quota is denominated in internal units where 1 unit = $0.002/1K tokens (configurable via `QuotaPerUnit` in `common/config/config.go:21`).

**Pre-consumption.** Before sending a request, the system pre-consumes quota: `preConsumeQuota` in `controller/helper.go:68-95` estimates prompt tokens via the OpenAI tokenizer, adds max_tokens, multiplies by `modelRatio × groupRatio`, and checks the user + token balance. It then decrements the user's Redis-cached quota (`CacheDecreaseUserQuota`) and records a pre-consumption on the token. If user quota > 100× the estimate, pre-consumption is skipped (trusted user).

**Post-consumption.** After the response, `postConsumeQuota` (`controller/helper.go:97-141`) computes actual usage from `usage.prompt_tokens` and `usage.completion_tokens`, applies `completionRatio` for output tokens, calculates `quota = ceil((promptTokens + completionTokens × completionRatio) × modelRatio × groupRatio)`, adjusts the token quota by the delta, updates the user quota cache, writes a `Log` record, and increments channel used_quota.

**Price tables.** `relay/billing/ratio/model.go:27-622` contains an exhaustive `ModelRatio` map covering ~600+ models across OpenAI, Anthropic, Google, Baidu, Ali, Zhipu, Xunfei, Moonshot, DeepSeek, Mistral, Groq, Cohere, etc. The `CompletionRatio` map (`model.go:624-633`) adjusts for providers that price output tokens differently. `GetModelRatio` (`model.go:686-710`) looks up by exact name then falls back to defaults. Group-level multipliers are in `relay/billing/ratio/group.go:10-13`.

**Batch updates.** When `BATCH_UPDATE_ENABLED` is set, quota updates are queued into in-memory records and flushed periodically (`BatchUpdateInterval`, default 5s) to reduce DB write pressure (`model/main.go:86-89`). Without batching, each quota change issues an immediate SQL `UPDATE ... SET quota = quota +/- N`.

**Storage.** Usage logs are stored in the `logs` table (or a separate `LOG_SQL_DSN` database) with fields for user, token, model, prompt/completion tokens, quota consumed, elapsed time, and stream flag (`model/log.go:15-32`).


Citations: [middleware/rate-limit.go:19-91](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/rate-limit.go#L19-L91) · [relay/billing/ratio/model.go:27-90](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/billing/ratio/model.go#L27-L90) · [relay/billing/ratio/model.go:686-710](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/relay/billing/ratio/model.go#L686-L710) · [common/config/config.go:86-99](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L86-L99)

### Portkey-AI/gateway (answered)

**Rate limiting** is implemented server-side only when Redis is available, using a token-bucket algorithm in Lua via `src/shared/services/cache/utils/rateLimiter.ts:84`. The `RedisRateLimiter` class takes a `capacity`, `windowSize`, and `key`, and runs a Lua script (`RATE_LIMIT_LUA` at line 5) atomically to decide if a request is allowed. It supports both consume and check-only modes. However, this rate limiter is not wired into the request path by default — it's a utility available for use.

**Per-key/user limits** are delegated to the `preRequestValidator` hook. The `PreRequestValidatorService` at `src/handlers/services/preRequestValidatorService.ts:6` calls `honoContext.get('preRequestValidator')`, which is an externally injected function (from Portkey SaaS, not defined in the open-source core). It returns an optional `response` (to deny the request) and `modelPricingConfig` (to track cost).

**Price tables and cost tracking** are stored per-provider as `modelPricingConfig` (an optional `Record<string, any>` on the provider option, see `src/types/requestBody.ts:179`). This is set by the pre-request validator and attached to the log object (`src/handlers/services/logsService.ts:288-293`). There are no built-in price tables in the gateway code.

**Usage accounting** happens via the logging pipeline. Every request produces a `LogObject` through `LogObjectBuilder` in `src/handlers/services/logsService.ts:215`, capturing `providerOptions`, `transformedRequest`, `requestParams`, `originalResponse` (including usage tokens from provider), `executionTime`, and `modelPricingConfig`. These logs are stored in-memory on the Hono context (`c.get('requestOptions')`) and broadcast via SSE to connected clients at `/log/stream` by `src/middlewares/log/index.ts:80-110`. The logs are not persisted to a database by the gateway — they are ephemeral in-memory structures.

In summary: cost infrastructure exists (modelPricingConfig, usage capture in log objects) but is passive — recording rather than enforcing budgets. Enforcement requires an external pre-request validator.


Citations: [src/shared/services/cache/utils/rateLimiter.ts:5-188](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/shared/services/cache/utils/rateLimiter.ts#L5-L188) · [src/handlers/services/preRequestValidatorService.ts:6-34](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/preRequestValidatorService.ts#L6-L34) · [src/handlers/services/logsService.ts:215-360](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/logsService.ts#L215-L360) · [src/middlewares/log/index.ts:80-166](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/log/index.ts#L80-L166)

### higress-group/higress (answered)

**Rate limits, budgets, and cost tracking are handled by three separate plugins: `ai-quota`, `ai-token-ratelimit`, and `ai-statistics`.**

**Per-key/user quotas (`ai-quota`)**: The plugin checks a Redis-stored integer quota for each consumer (identified by `x-mse-consumer` header) before allowing a request. The quota is checked as `GET chat_quota:{consumer}` and if ≤0, the request is denied with 403 — `ai-quota/main.go:194-213`. After response completion, the plugin extracts token usage from the response and decrements the quota via `DECRBY chat_quota:{consumer} {totalTokens}` — `ai-quota/main.go:267-303`. The `GetTokenUsage` function parses token counts from both OpenAI and Anthropic response formats.

**Token-based rate limiting (`ai-token-ratelimit`)**: This plugin enforces sliding-window rate limits using **server-side Lua scripts in Redis**. It supports global thresholds and per-key limits with configurable windows per rule — `ai-token-ratelimit/main.go:47-79`. The Lua script `MultiKeyRequestPhaseScript` atomically checks multiple rate limit counters in one Redis call, returning thresholds, current counts, and TTLs — `ai-token-ratelimit/main.go:58-79`. The response-phase script increments counters only after a successful LLM response.

**Usage statistics and cost tracking (`ai-statistics`)**: This plugin records per-request token usage (input, output, total, reasoning tokens, cached tokens, token details) from both streaming and non-streaming responses — `ai-statistics/main.go:990-1045`. Token usage is extracted via the `tokenusage.GetTokenUsage` library which handles multiple response formats: OpenAI chat completions, Anthropic Messages, Gemini, Doubao — `ai-statistics/main.go:960-962`. Metrics are recorded as **Envoy distributed counters** at the route/cluster/model/consumer granularity — `ai-statistics/main.go:481-524`. The metric name follows the pattern `route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName}` — `ai-statistics/main.go:481-483`.

**Built-in metrics** recorded include: `llm_first_token_duration`, `llm_service_duration`, `llm_failure_count`, and token counts — `ai-statistics/main.go:93-97`. These are accumulated into Envoy-native counter metrics and also written to the ARMS (Application Resource Monitoring System) tracing spans with `gen_ai.*` prefix — `ai-statistics/main.go:102-108`.

**Price tables**: There is no built-in price table in the codebase — cost tracking is purely token-based, leaving monetary cost calculation to external systems.


Citations: [plugins/wasm-go/extensions/ai-quota/main.go:194-213](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L194-L213) · [plugins/wasm-go/extensions/ai-quota/main.go:267-303](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L267-L303) · [plugins/wasm-go/extensions/ai-token-ratelimit/main.go:47-79](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-token-ratelimit/main.go#L47-L79) · [plugins/wasm-go/extensions/ai-statistics/main.go:481-524](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L481-L524) · [plugins/wasm-go/extensions/ai-statistics/main.go:990-1045](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L990-L1045) · [plugins/wasm-go/extensions/ai-statistics/main.go:93-108](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L93-L108)

### maximhq/bifrost (answered)

Rate limits and budgets are enforced per entity via the governance plugin's `LocalGovernanceStore`. **Budgets** (`TableBudget` in `configstore/tables/budget.go`) carry a `MaxLimit` (dollars), `ResetDuration` (e.g. "1d", "1M"), `CurrentUsage`, and `LastReset`, plus override support. Budgets are attached at the virtual-key provider-config level, and also at team/customer levels for hierarchical cost control. **Rate limits** (`TableRateLimit`) track both token-based (`TokenMaxLimit`) and request-based (`RequestMaxLimit`) quotas with independent reset durations. Both budgets and rate limits use a sliding-window reset mechanism computed from `LastReset + ResetDuration`. The `UsageTracker` (`plugins/governance/tracker.go`) processes `UsageUpdate` events asynchronously through a background goroutine that batches writes to the in-memory store and periodically dumps to the DB. **Cost calculation:** The `datasheet.Store.CalculateCost` (`framework/modelcatalog/datasheet/cost.go:19`) computes response cost from model pricing. The `TableModelPricing` table (`configstore/tables/modelpricing.go`) stores per-model per-provider per-mode pricing fields: text token rates (input/output, with batch/priority/fast/ultrafast/flex tiers), image/video/audio pricing, and tiered rates for contexts above 128k/200k/272k tokens. Custom pricing overrides can be scoped globally or per-provider. **Usage accounting** stores spend in the `LocalGovernanceStore` in-memory maps and periodically persists to PostgreSQL. The data is exposed via the admin API and used for real-time governance decisions (budget-exceeded/rate-limited denials).


Citations: [framework/configstore/tables/budget.go:39-52](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/framework/configstore/tables/budget.go#L39-L52) · [framework/configstore/tables/ratelimit.go:10-42](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/framework/configstore/tables/ratelimit.go#L10-L42) · [framework/configstore/tables/modelpricing.go:10-50](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/framework/configstore/tables/modelpricing.go#L10-L50) · [framework/modelcatalog/datasheet/cost.go:19-50](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/framework/modelcatalog/datasheet/cost.go#L19-L50) · [plugins/governance/tracker.go:16-56](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/governance/tracker.go#L16-L56)

### katanemo/plano (answered)

**Rate limits and cost tracking are implemented via the `governor` crate for rate limiting and a per-model pricing catalog for cost tracking, with a session-level switch-cost budget.**

**Rate limits** are configured under `ratelimits` in the YAML config, each specifying a `model`, a `selector` (HTTP header key+value), and a `Limit` (tokens + time unit: second/minute/hour/day). The `RatelimitMap` in `crates/common/src/ratelimit.rs` structures them as `Provider → {Header → KeyedRateLimiter}`, using the `governor` crate's `DefaultKeyedRateLimiter` for token-bucket enforcement. The selector header value (or empty string for wildcard) acts as the key. Failed checks return `Error::ExceededLimit` with which provider/selector/tokens were exceeded.

Per-model rate limits (`LlmRatelimit`) can also be attached directly to `LlmProvider` config blocks via the `rate_limits` field (`crates/common/src/configuration.rs:771`), using an HTTP header selector.

**Cost tracking** operates at the model level via `ModelRates` (`crates/brightstaff/src/router/model_metrics.rs:23`), which stores `input_per_million`, `output_per_million`, and optional `cache_read_per_million` USD rates. The `ModelMetricsService` (line 84) fetches live pricing catalogs from DigitalOcean (`api.digitalocean.com/v2/gen-ai/models/catalog`) or models.dev, with configurable refresh intervals, and supports `model_aliases` for catalog-to-Plano model name mapping. `request_cost_usd()` (line 52) computes the actual dollar cost of one request from token usage, distinguishing cached vs uncached input tokens with three pricing tiers: plain input rate, cached read rate (defaults to 10% of input rate via `cache_read_discount`), and cache creation at plain input rate.

**Usage accounting** is extracted from provider responses in `streaming.rs`'s `ExtractedUsage` struct (lines 36-100), which parses both OpenAI-shape (`prompt_tokens` includes cached) and Anthropic-shape (separate `input_tokens`, `cache_read_input_tokens`) usage from JSON. This feeds into the `SessionBinding` which tracks `session_cost_usd`, `baseline_usd`, and `switch_spend_usd` across turns. Session data is stored in either an in-memory (`MemorySessionCache`) or Redis backend (`SessionBinding` in `crates/brightstaff/src/session_cache/mod.rs`).

The **routing budget** (`EffectiveRoutingBudget` in `configuration.rs:336`) caps cumulative model-switch overhead at `max_switch_spend_pct`% of the session's never-switch baseline. This is a cost gate, not a hard spend limit — it ensures quality-driven switches don't inflate the bill beyond a configurable overhead.


Citations: [crates/common/src/ratelimit.rs:27-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/ratelimit.rs#L27-L80) · [crates/brightstaff/src/router/model_metrics.rs:1-74](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/router/model_metrics.rs#L1-L74) · [crates/brightstaff/src/session_cache/mod.rs:19-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/session_cache/mod.rs#L19-L80) · [crates/common/src/configuration.rs:287-350](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L287-L350)

### tbphp/gpt-load (answered)

**Per-key rate limits** — two levels: RPM (requests per minute) and concurrency. RPM is enforced by `AccessKeyRPMLimiter` (`handler.go:68-71`), implemented via `ratelimit.AccessKeyRPM` backed by `rpm.Store`. The `limiter.Allow(accessKey.ID, accessKey.RPMLimit)` call at `handler.go:543` returns a `LimitDecision`. Concurrency is enforced by `handler.acquireRequestConcurrency()` (`internal/gateway/concurrency.go:15`) against both per-key and global limits (`GlobalConcurrencyLimit`), tracked in `ratelimit.Concurrency`.

**Token and spend budgets** — the `accessquota` package (`internal/accessquota/runtime.go`) implements AccessKey-level cost limits. Two rule kinds exist: `KindTotal` (lifetime spend cap) and `KindPeriodic` (time-windowed cap like daily/monthly). Each `Rule` stores `LimitNanoUSD`, `PeriodSeconds`, window tracking, and `UsedNanoUSD`. The `accessquota.Runtime` provides `Check()` (pre-flight) and `Complete()` (post-request deduction) methods. Budgets are checked in `handler.go:522-541` before forwarding.

**Price tables** — `pricing.Table` (`internal/pricing/types.go`) maps channel+model identities to nano-USD-per-million-token prices. `pricing.Price` tracks per-input, per-output, per-cache-read, and per-cache-write rates. `ContextTier` allows different rates once input exceeds a threshold. `pricing.Quote()` (`internal/pricing/quote.go`) prices a finalized `usage.Result`. Group-level and AccessKey-level `PriceMultipliers` are applied via `QuoteForModeWithMultipliers()` for surcharging/discounting.

**Usage accounting** — `usage.Result` (`internal/usage`) holds token counts (uncached input, cache read, cache write, output). Usage is captured from upstream responses and stream events via `usageCaptureBoundary` (`internal/gateway/usage_capture.go:20`). Extracted usage flows to `pricing.Receipt` (`internal/pricing/receipt.go`), which is a frozen record of pricing method, rates, and total cost at request time (schema v6).

**Usage storage** — the `requestlog` package persists detailed request logs with cost. `accessquota` keeps in-memory rolling state for budget enforcement. The `state.Manager` uses gorm-backed snapshots for durable configuration, and `state/loader` loads config from the database. RPM state persists through `rpm.Store`.

**Cost estimation & budget enforcement** — `handler.freezeAttemptPricing()` (`handler.go:132-155`) captures the price table and multipliers at request time. After execution, `recorder.estimatedCostNanoUSD()` feeds into `accessQuota.Complete()`. If quota is exhausted, `Check()` returns `Allowed: false` and the request is blocked before reaching a provider.


Citations: [internal/gateway/concurrency.go:15-38](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/concurrency.go#L15-L38) · [internal/accessquota/runtime.go:1-79](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/accessquota/runtime.go#L1-L79) · [internal/pricing/quote.go:13-56](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/pricing/quote.go#L13-L56) · [internal/gateway/handler.go:132-155](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L132-L155)

### theagentrouter/agent-router (answered)

Rate limits are implemented via the `QuotaPolicy` Kubernetes CRD, which translates into Envoy rate limit configuration. Each QuotaPolicy specifies `PerModelQuotas` (per-model rate limits) and `ServiceQuota` (service-wide catch-all), with nested `BucketRules` for client-based differentiation using header selectors (`internal/ratelimit/translator/translator.go:92-177`). Rate limits are expressed as requests per unit (1s/1m/1h/1d) converted to Envoy's `RateLimitUnit` (`internal/ratelimit/translator/translator.go:392-417`). All backends share a single rate limit domain (`ai-gateway-quota`), distinguished by `backend_name` and `model_name_override` descriptors (`internal/ratelimit/translator/translator.go:21-33`). The `QuotaPolicyController` translates policy definitions into `RateLimitConfig` objects consumed by Envoy's rate limit service and caches them for incremental updates (`internal/controller/quota_policy.go:30-59`). Cost tracking is configured through `LLMRequestCost` entries on AIGatewayRoutes (and `GlobalLLMRequestCost` on GatewayConfig). Cost types include InputToken, OutputToken, TotalToken, CachedInputToken, CacheCreationInputToken, ReasoningToken, and custom CEL expressions (`internal/filterapi/filterconfig.go:112-130`). CEL expressions support variables like `input_tokens`, `output_tokens`, `total_tokens`, `model`, `backend`, `route_name` for custom cost formulas (`internal/llmcostcel/cel.go:18-48`). Costs are computed by the ext_proc response handler and written into Envoy's dynamic metadata under the namespace `io.envoy.ai_gateway` for consumption by Envoy's rate limit `HitsAddend` filter (`internal/extproc/processor_impl.go:651-661`). The `TokenUsage` struct tracks input, output, total, cached input, cache creation input, and reasoning tokens separately (`internal/metrics/metrics.go:143-158`). QuotaPolicy cost expressions are injected as `LLMRequestCost` entries during reconciliation, filtered by backend and model name (`internal/controller/gateway.go:1053-1130`).


Citations: [internal/ratelimit/translator/translator.go:92-177](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/ratelimit/translator/translator.go#L92-L177) · [internal/ratelimit/translator/translator.go:21-33](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/ratelimit/translator/translator.go#L21-L33) · [internal/filterapi/filterconfig.go:112-130](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/filterapi/filterconfig.go#L112-L130) · [internal/llmcostcel/cel.go:18-48](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/llmcostcel/cel.go#L18-L48) · [internal/extproc/processor_impl.go:651-661](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/extproc/processor_impl.go#L651-L661) · [internal/metrics/metrics.go:143-158](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/metrics/metrics.go#L143-L158)
