LLMs Technical Reviews

How are rate limits, budgets and cost tracking implemented?

Per-key/user limits; token and spend budgets; price tables; usage accounting and where it is stored.

Verdict

LiteLLM has the most complete cost attribution, with a maintained price table and budgets at every level. New API is the most complete for running a paid service. Bifrost and GPT-Load offer dollar caps in a smaller package.

Dollar budgets from built-in prices. LiteLLM ships a price file with about 4,500 model entries. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is written to PostgreSQL in batches. Bifrost attaches budgets with reset windows, plus token and request limits, to virtual keys, teams and customers. Its counters live in each process, so the open-source build does not share them across nodes. GPT-Load gives each AccessKey lifetime or periodic caps in nano-USD, an RPM limit and a concurrency limit. Every request gets a frozen price receipt. OmniRoute enforces daily and weekly USD caps per key. Prices come from a short built-in map (about 20 models), synced or operator prices in SQLite, and an estimate for unknown models.

Quota units for resale. One API pre-charges (prompt + max_tokens) × modelRatio × groupRatio in internal units, then settles on real usage. Its rate limits are per IP (480 API requests per 3 minutes by default), not per key. New API keeps a billing reservation across channel retries. It adds tiered price expressions, top-ups through Stripe, Epay or Creem, and logs in SQL or ClickHouse.

Token counters in the data plane. Higress keeps a Redis token quota per consumer and decrements it after the response, so one large request can overshoot. It has no price table. Agent Router computes token or CEL cost expressions and enforces QuotaPolicy in Envoy’s rate-limit service. Neither has a dollar budget. Plano applies governor token buckets keyed by a header value. Its dollar tracking only caps how much switching models may add within a session (max_switch_spend_pct).

Hook points only. Portkey Gateway has no price tables and stores no logs. Budgets depend on a preRequestValidator that nothing in the repository sets, and its Redis token bucket is not wired into the request path.

Pick: LiteLLM for per-team spend tracking and budgets. Pick: New API or One API to sell or ration access with prepaid quota. Pick: Higress or Agent Router for token-based rate limits on an Envoy fleet.

Per-project answers

diegosouzapw/OmniRoute

answered

Rate limiting is implemented via src/shared/utils/rateLimiter.ts using a Redis-backed (or in-memory fallback) leaky bucket strategy. API keys carry maxRequestsPerDay, maxRequestsPerMinute, and rich rateLimits rules (parsed by parseRateLimits()). Cost tracking centers on open-sse/services/providerCostData.ts which provides a KNOWN_MODEL_PRICING map covering ~25 key models (e.g., gpt-4o at $2.50/$10 per 1M tokens, claude-sonnet-4.6 at $3/$15) and a getModelPricing() resolver that falls back through: provider-specific KNOWN match → default pricing DB table → generic model name match → free-model check → estimated fallback ($5/$15). The tier resolver (tierResolver.ts) layers operator-configured pricing over the defaults using a DB snapshot from pricing/pricing_synced/models_dev_pricing key-value namespaces and caches the merged result. Usage accounting is stored in two SQLite tables: usage_history (per-request rows with tokens_input/tokens_output and ISO timestamps) and daily_usage_summary (rolled-up aggregates), queried by src/lib/db/usageSummary.ts and src/lib/db/usageLogs.ts. The GET /v1/usage endpoint exposes aggregated usage. Spend limits are enforced by apiKeyPolicyService.ts (enforceApiKeyLimits()) which checks daily/weekly USD caps before allowing requests. Circuit breakers, connection cooldown (exponential backoff on per-credential failures), and model lockout all act as runtime cost-control mechanisms by stopping traffic to failing-expensive or exhausted endpoints. The connectionBillingCatalog.ts classifies connections as subscription (plan-included) vs metered (per-token) to budget routing.

BerriAI/litellm

answered

Rate limits. The LiteLLM_BudgetTable in schema.prisma defines max_parallel_requests, tpm_limit (tokens per minute), rpm_limit (requests per minute), and tpd_limit (tokens per day) at every level of the hierarchy (org, team, user, key, end-user, tag). These are enforced via the budget throttle middleware (litellm/proxy/auth/budget_throttle.py) and pre-call RPM/TPM checks in the routing layer (router.py:8984-9033).

Spend and budget tracking. litellm/proxy/spend_tracking/ implements a batched spend counter (spend_counter_batch.py) that increments in-memory counters and flushes to PostgreSQL asynchronously on a cadence. The LiteLLM_SpendLog* tables in the Prisma schema record every request's token usage, cost, and metadata. Budget enforcement checks (_virtual_key_max_budget_check, _virtual_key_soft_budget_check in litellm/proxy/auth/auth_checks.py) compare accumulated spend against max_budget/soft_budget thresholds.

Price table. model_prices_and_context_window.json (80,644 lines) is the bundled pricing catalog. Every model entry specifies input_cost_per_token, output_cost_per_token, output_cost_per_reasoning_token, regional uplift multipliers, prompt-caching discounts, batch rate reductions, and modality-specific costs (audio, image, video, web-search, computer-use). The cost calculator (litellm/cost_calculator.py) uses a chain-of-responsibility pattern: each provider module registers a cost_per_token() function (e.g. openai_cost_per_token, anthropic_cost_per_token, bedrock_cost_per_token in litellm/llms/{provider}/cost_calculation.py), and the base resolver selects the right one.

Usage accounting. The final response_cost, total_tokens, and breakdown are stored in the LiteLLM_SpendLog* tables. Budget carry-forward across reset periods is handled by carried_budget_state.py. Enterprise users get PTU (pay-per-use) flat-cost pricing via ptu_flat_cost_rollup.py.

QuantumNous/new-api

answered

Rate limiting operates at multiple levels. (1) Global rate limits (middleware/rate-limit.go:160-178) — GlobalWebRateLimit and GlobalAPIRateLimit apply fixed-window Redis Lua scripts (or in-memory fallback) keyed by IP address. (2) Model-level rate limits (middleware/model-rate-limit.go) — ModelRequestRateLimit() tracks total requests and success counts per user using a Redis list (or in-memory bucket) with configurable max count and duration. Token bucket limiting (via common/limiter) is available as an alternative. (3) Critical route rate limits throttle brute-force attempts on auth endpoints. All rate limiters return Retry-After headers on 429 responses.

Quota and budget tracking is a two-phase system: pre-consume then settle. On request start, PrepareRequestBilling in relay/request_billing.go pre-computes the estimated quota cost based on model pricing and deducts it from the token's RemainQuota and the user's Quota. On completion, service/task_billing.go and service/text_quota.go calculate the actual usage (token counts from upstream response) and settle the difference (refund or additional charge).

Price tables are configured per-model in model/model_pricing_config.go:24-57. The PricingValues map supports key-based pricing (per-token, per-image, per-second, etc.). Group-level ratio multipliers and model-level ratios are applied through ratio_setting. The price.go helper in relay/helper/price.go:44 resolves group ratios and tiered billing expressions. Tiered billing (pkg/billingexpr) supports complex expression-based pricing (e.g., "first 1000 tokens free, then $0.01/1K"), evaluated against actual usage facts captured in TieredBillingSnapshot.

Usage accounting is stored in the Log model (model/log.go:59-80), which records every request: user ID, token name, model, PromptTokens/CompletionTokens, quota consumed, channel, group, IP, streaming flag, and an Other JSON field for extended billing metadata (task details, model price, ratios, tiered billing snapshot). The log database supports both the main DB and a separate ClickHouse instance for high-volume analytics (docker-compose.yml:32).

Wallet top-ups and subscriptions are supported through multiple payment providers (Stripe, Epay, Waffle/Pancake, Creem) with redemption codes and affiliate tracking.

songquanpeng/one-api

answered

Rate limits. The middleware/rate-limit.go provides per-IP rate limiting with configurable windows: GlobalAPIRateLimit (480 requests/3 min), GlobalWebRateLimit (240/3 min), CriticalRateLimit (20/20 min), and upload/download limits. Supports both Redis-based (redisRateLimiter via LPUSH/EXPIRE/TTL checks) and in-memory (memoryRateLimiter) backends. The common/rate-limit.go InMemoryRateLimiter uses a sliding-window approach with per-key timestamp queues cleaned by a background goroutine.

Quota system. Users have a quota field (integer) and used_quota (model/user.go:48-49). Tokens have their own remain_quota and unlimited_quota (model/token.go:31-33). The quota is denominated in internal units where 1 unit = $0.002/1K tokens (configurable via QuotaPerUnit in common/config/config.go:21).

Pre-consumption. Before sending a request, the system pre-consumes quota: preConsumeQuota in controller/helper.go:68-95 estimates prompt tokens via the OpenAI tokenizer, adds max_tokens, multiplies by modelRatio × groupRatio, and checks the user + token balance. It then decrements the user's Redis-cached quota (CacheDecreaseUserQuota) and records a pre-consumption on the token. If user quota > 100× the estimate, pre-consumption is skipped (trusted user).

Post-consumption. After the response, postConsumeQuota (controller/helper.go:97-141) computes actual usage from usage.prompt_tokens and usage.completion_tokens, applies completionRatio for output tokens, calculates quota = ceil((promptTokens + completionTokens × completionRatio) × modelRatio × groupRatio), adjusts the token quota by the delta, updates the user quota cache, writes a Log record, and increments channel used_quota.

Price tables. relay/billing/ratio/model.go:27-622 contains an exhaustive ModelRatio map covering ~600+ models across OpenAI, Anthropic, Google, Baidu, Ali, Zhipu, Xunfei, Moonshot, DeepSeek, Mistral, Groq, Cohere, etc. The CompletionRatio map (model.go:624-633) adjusts for providers that price output tokens differently. GetModelRatio (model.go:686-710) looks up by exact name then falls back to defaults. Group-level multipliers are in relay/billing/ratio/group.go:10-13.

Batch updates. When BATCH_UPDATE_ENABLED is set, quota updates are queued into in-memory records and flushed periodically (BatchUpdateInterval, default 5s) to reduce DB write pressure (model/main.go:86-89). Without batching, each quota change issues an immediate SQL UPDATE ... SET quota = quota +/- N.

Storage. Usage logs are stored in the logs table (or a separate LOG_SQL_DSN database) with fields for user, token, model, prompt/completion tokens, quota consumed, elapsed time, and stream flag (model/log.go:15-32).

Portkey-AI/gateway

answered

Rate limiting is implemented server-side only when Redis is available, using a token-bucket algorithm in Lua via src/shared/services/cache/utils/rateLimiter.ts:84. The RedisRateLimiter class takes a capacity, windowSize, and key, and runs a Lua script (RATE_LIMIT_LUA at line 5) atomically to decide if a request is allowed. It supports both consume and check-only modes. However, this rate limiter is not wired into the request path by default — it's a utility available for use.

Per-key/user limits are delegated to the preRequestValidator hook. The PreRequestValidatorService at src/handlers/services/preRequestValidatorService.ts:6 calls honoContext.get('preRequestValidator'), which is an externally injected function (from Portkey SaaS, not defined in the open-source core). It returns an optional response (to deny the request) and modelPricingConfig (to track cost).

Price tables and cost tracking are stored per-provider as modelPricingConfig (an optional Record<string, any> on the provider option, see src/types/requestBody.ts:179). This is set by the pre-request validator and attached to the log object (src/handlers/services/logsService.ts:288-293). There are no built-in price tables in the gateway code.

Usage accounting happens via the logging pipeline. Every request produces a LogObject through LogObjectBuilder in src/handlers/services/logsService.ts:215, capturing providerOptions, transformedRequest, requestParams, originalResponse (including usage tokens from provider), executionTime, and modelPricingConfig. These logs are stored in-memory on the Hono context (c.get('requestOptions')) and broadcast via SSE to connected clients at /log/stream by src/middlewares/log/index.ts:80-110. The logs are not persisted to a database by the gateway — they are ephemeral in-memory structures.

In summary: cost infrastructure exists (modelPricingConfig, usage capture in log objects) but is passive — recording rather than enforcing budgets. Enforcement requires an external pre-request validator.

higress-group/higress

answered

Rate limits, budgets, and cost tracking are handled by three separate plugins: ai-quota, ai-token-ratelimit, and ai-statistics.

Per-key/user quotas (ai-quota): The plugin checks a Redis-stored integer quota for each consumer (identified by x-mse-consumer header) before allowing a request. The quota is checked as GET chat_quota:{consumer} and if ≤0, the request is denied with 403 — ai-quota/main.go:194-213. After response completion, the plugin extracts token usage from the response and decrements the quota via DECRBY chat_quota:{consumer} {totalTokens} — ai-quota/main.go:267-303. The GetTokenUsage function parses token counts from both OpenAI and Anthropic response formats.

Token-based rate limiting (ai-token-ratelimit): This plugin enforces sliding-window rate limits using server-side Lua scripts in Redis. It supports global thresholds and per-key limits with configurable windows per rule — ai-token-ratelimit/main.go:47-79. The Lua script MultiKeyRequestPhaseScript atomically checks multiple rate limit counters in one Redis call, returning thresholds, current counts, and TTLs — ai-token-ratelimit/main.go:58-79. The response-phase script increments counters only after a successful LLM response.

Usage statistics and cost tracking (ai-statistics): This plugin records per-request token usage (input, output, total, reasoning tokens, cached tokens, token details) from both streaming and non-streaming responses — ai-statistics/main.go:990-1045. Token usage is extracted via the tokenusage.GetTokenUsage library which handles multiple response formats: OpenAI chat completions, Anthropic Messages, Gemini, Doubao — ai-statistics/main.go:960-962. Metrics are recorded as Envoy distributed counters at the route/cluster/model/consumer granularity — ai-statistics/main.go:481-524. The metric name follows the pattern route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName} — ai-statistics/main.go:481-483.

Built-in metrics recorded include: llm_first_token_duration, llm_service_duration, llm_failure_count, and token counts — ai-statistics/main.go:93-97. These are accumulated into Envoy-native counter metrics and also written to the ARMS (Application Resource Monitoring System) tracing spans with gen_ai.* prefix — ai-statistics/main.go:102-108.

Price tables: There is no built-in price table in the codebase — cost tracking is purely token-based, leaving monetary cost calculation to external systems.

maximhq/bifrost

answered

Rate limits and budgets are enforced per entity via the governance plugin's LocalGovernanceStore. Budgets (TableBudget in configstore/tables/budget.go) carry a MaxLimit (dollars), ResetDuration (e.g. "1d", "1M"), CurrentUsage, and LastReset, plus override support. Budgets are attached at the virtual-key provider-config level, and also at team/customer levels for hierarchical cost control. Rate limits (TableRateLimit) track both token-based (TokenMaxLimit) and request-based (RequestMaxLimit) quotas with independent reset durations. Both budgets and rate limits use a sliding-window reset mechanism computed from LastReset + ResetDuration. The UsageTracker (plugins/governance/tracker.go) processes UsageUpdate events asynchronously through a background goroutine that batches writes to the in-memory store and periodically dumps to the DB. Cost calculation: The datasheet.Store.CalculateCost (framework/modelcatalog/datasheet/cost.go:19) computes response cost from model pricing. The TableModelPricing table (configstore/tables/modelpricing.go) stores per-model per-provider per-mode pricing fields: text token rates (input/output, with batch/priority/fast/ultrafast/flex tiers), image/video/audio pricing, and tiered rates for contexts above 128k/200k/272k tokens. Custom pricing overrides can be scoped globally or per-provider. Usage accounting stores spend in the LocalGovernanceStore in-memory maps and periodically persists to PostgreSQL. The data is exposed via the admin API and used for real-time governance decisions (budget-exceeded/rate-limited denials).

katanemo/plano

answered

Rate limits and cost tracking are implemented via the governor crate for rate limiting and a per-model pricing catalog for cost tracking, with a session-level switch-cost budget.

Rate limits are configured under ratelimits in the YAML config, each specifying a model, a selector (HTTP header key+value), and a Limit (tokens + time unit: second/minute/hour/day). The RatelimitMap in crates/common/src/ratelimit.rs structures them as Provider → {Header → KeyedRateLimiter}, using the governor crate's DefaultKeyedRateLimiter for token-bucket enforcement. The selector header value (or empty string for wildcard) acts as the key. Failed checks return Error::ExceededLimit with which provider/selector/tokens were exceeded.

Per-model rate limits (LlmRatelimit) can also be attached directly to LlmProvider config blocks via the rate_limits field (crates/common/src/configuration.rs:771), using an HTTP header selector.

Cost tracking operates at the model level via ModelRates (crates/brightstaff/src/router/model_metrics.rs:23), which stores input_per_million, output_per_million, and optional cache_read_per_million USD rates. The ModelMetricsService (line 84) fetches live pricing catalogs from DigitalOcean (api.digitalocean.com/v2/gen-ai/models/catalog) or models.dev, with configurable refresh intervals, and supports model_aliases for catalog-to-Plano model name mapping. request_cost_usd() (line 52) computes the actual dollar cost of one request from token usage, distinguishing cached vs uncached input tokens with three pricing tiers: plain input rate, cached read rate (defaults to 10% of input rate via cache_read_discount), and cache creation at plain input rate.

Usage accounting is extracted from provider responses in streaming.rs's ExtractedUsage struct (lines 36-100), which parses both OpenAI-shape (prompt_tokens includes cached) and Anthropic-shape (separate input_tokens, cache_read_input_tokens) usage from JSON. This feeds into the SessionBinding which tracks session_cost_usd, baseline_usd, and switch_spend_usd across turns. Session data is stored in either an in-memory (MemorySessionCache) or Redis backend (SessionBinding in crates/brightstaff/src/session_cache/mod.rs).

The routing budget (EffectiveRoutingBudget in configuration.rs:336) caps cumulative model-switch overhead at max_switch_spend_pct% of the session's never-switch baseline. This is a cost gate, not a hard spend limit — it ensures quality-driven switches don't inflate the bill beyond a configurable overhead.

tbphp/gpt-load

answered

Per-key rate limits — two levels: RPM (requests per minute) and concurrency. RPM is enforced by AccessKeyRPMLimiter (handler.go:68-71), implemented via ratelimit.AccessKeyRPM backed by rpm.Store. The limiter.Allow(accessKey.ID, accessKey.RPMLimit) call at handler.go:543 returns a LimitDecision. Concurrency is enforced by handler.acquireRequestConcurrency() (internal/gateway/concurrency.go:15) against both per-key and global limits (GlobalConcurrencyLimit), tracked in ratelimit.Concurrency.

Token and spend budgets — the accessquota package (internal/accessquota/runtime.go) implements AccessKey-level cost limits. Two rule kinds exist: KindTotal (lifetime spend cap) and KindPeriodic (time-windowed cap like daily/monthly). Each Rule stores LimitNanoUSD, PeriodSeconds, window tracking, and UsedNanoUSD. The accessquota.Runtime provides Check() (pre-flight) and Complete() (post-request deduction) methods. Budgets are checked in handler.go:522-541 before forwarding.

Price tables — pricing.Table (internal/pricing/types.go) maps channel+model identities to nano-USD-per-million-token prices. pricing.Price tracks per-input, per-output, per-cache-read, and per-cache-write rates. ContextTier allows different rates once input exceeds a threshold. pricing.Quote() (internal/pricing/quote.go) prices a finalized usage.Result. Group-level and AccessKey-level PriceMultipliers are applied via QuoteForModeWithMultipliers() for surcharging/discounting.

Usage accounting — usage.Result (internal/usage) holds token counts (uncached input, cache read, cache write, output). Usage is captured from upstream responses and stream events via usageCaptureBoundary (internal/gateway/usage_capture.go:20). Extracted usage flows to pricing.Receipt (internal/pricing/receipt.go), which is a frozen record of pricing method, rates, and total cost at request time (schema v6).

Usage storage — the requestlog package persists detailed request logs with cost. accessquota keeps in-memory rolling state for budget enforcement. The state.Manager uses gorm-backed snapshots for durable configuration, and state/loader loads config from the database. RPM state persists through rpm.Store.

Cost estimation & budget enforcement — handler.freezeAttemptPricing() (handler.go:132-155) captures the price table and multipliers at request time. After execution, recorder.estimatedCostNanoUSD() feeds into accessQuota.Complete(). If quota is exhausted, Check() returns Allowed: false and the request is blocked before reaching a provider.

theagentrouter/agent-router

answered

Rate limits are implemented via the QuotaPolicy Kubernetes CRD, which translates into Envoy rate limit configuration. Each QuotaPolicy specifies PerModelQuotas (per-model rate limits) and ServiceQuota (service-wide catch-all), with nested BucketRules for client-based differentiation using header selectors (internal/ratelimit/translator/translator.go:92-177). Rate limits are expressed as requests per unit (1s/1m/1h/1d) converted to Envoy's RateLimitUnit (internal/ratelimit/translator/translator.go:392-417). All backends share a single rate limit domain (ai-gateway-quota), distinguished by backend_name and model_name_override descriptors (internal/ratelimit/translator/translator.go:21-33). The QuotaPolicyController translates policy definitions into RateLimitConfig objects consumed by Envoy's rate limit service and caches them for incremental updates (internal/controller/quota_policy.go:30-59). Cost tracking is configured through LLMRequestCost entries on AIGatewayRoutes (and GlobalLLMRequestCost on GatewayConfig). Cost types include InputToken, OutputToken, TotalToken, CachedInputToken, CacheCreationInputToken, ReasoningToken, and custom CEL expressions (internal/filterapi/filterconfig.go:112-130). CEL expressions support variables like input_tokens, output_tokens, total_tokens, model, backend, route_name for custom cost formulas (internal/llmcostcel/cel.go:18-48). Costs are computed by the ext_proc response handler and written into Envoy's dynamic metadata under the namespace io.envoy.ai_gateway for consumption by Envoy's rate limit HitsAddend filter (internal/extproc/processor_impl.go:651-661). The TokenUsage struct tracks input, output, total, cached input, cache creation input, and reasoning tokens separately (internal/metrics/metrics.go:143-158). QuotaPolicy cost expressions are injected as LLMRequestCost entries during reconciliation, filtered by backend and model name (internal/controller/gateway.go:1053-1130).

← How are API keys, users and tenants managed? · How are caching and guardrails implemented? →