# LLM gateways: comparison

> Proxies that put many model providers behind one API: routing and fallbacks, keys and budgets, caching, guardrails and usage tracking.

Canonical page: https://llms-technical-reviews.com/compare/llm-gateways/

## How are requests routed across providers and models?

[LiteLLM](/p/litellm/) and [Bifrost](/p/bifrost/) have the most complete built-in routing. [GPT-Load](/p/gpt-load/) and [OmniRoute](/p/omniroute/) go deepest on failover across many keys or accounts. [Plano](/p/plano/) is the only one that routes by intent.

**Strategy engines inside the gateway.** LiteLLM's `Router` defaults to `simple-shuffle`, a weighted draw by `weight`, `rpm` or `tpm`. It adds least-busy, latency, cost and usage strategies, cooldowns, and fallback in mid-stream. Bifrost first evaluates CEL rules with the scope order VirtualKey > User > Team > Customer > Global. Its governance plugin then makes a weighted provider pick and builds fallbacks from the remaining providers. OmniRoute "combos" offer about 19 strategies, including cost-optimised and quota-reset-aware. They sit on per-account cooldowns and a circuit breaker stored in SQLite. Many pooled accounts are subscription or web-session logins. OmniRoute's own catalogue marks 17 providers `avoid` under their terms and warns of account bans, but by default that flag does not stop routing. [Portkey Gateway](/p/portkey-gateway/) reads a routing tree from each request's `x-portkey-config` header, with fallback, weighted and conditional (`$eq`, `$regex`) nodes. Its circuit breaker is inert: nothing in the repository sets `isOpen`.

**Channel and credential pools.** [One API](/p/one-api/) picks uniformly at random among the channels in the top priority tier. Its `Weight` field is never read. [New API](/p/new-api/) adds weighted draws per tier, session affinity, pins and failover across groups. Channel health comes from scheduled tests and auto-ban, not its Uptime Kuma page. GPT-Load uses a deterministic fair scheduler and prefers the client's native protocol before converted routes.

**Routing left to Envoy.** [Agent Router](/p/agent-router/) only sets the `x-ai-eg-model` header. Envoy route weights, priorities and `BackendTrafficPolicy` retries do the rest. [Higress](/p/higress/) fixes one `activeProviderId` per route rule. Per-request provider choice needs `model-router` to copy the model into a header that Envoy routes on. Failover happens between the API tokens of one provider.

**Intent routing.** In Plano, an orchestrator LLM picks a declared preference and ranks its models as `cheapest` or `fastest`. A session gate then keeps a model whose prompt cache is still warm. The main path sends only the chosen model and does not walk the ranked list.

Pick: LiteLLM or Bifrost for policy-driven load balancing and fallbacks.
Pick: GPT-Load, New API or OmniRoute to spread traffic over many keys or accounts.
Pick: Agent Router or Higress if Envoy should own traffic management.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/routing/index.md

## How are different provider APIs unified?

[LiteLLM](/p/litellm/) covers the most providers behind an OpenAI-style API. [New API](/p/new-api/) and [Plano](/p/plano/) are the strongest when clients speak several formats, such as Claude Code next to an OpenAI SDK. [Agent Router](/p/agent-router/) is the most careful about cross-provider retries.

**OpenAI format in, adapters out.** LiteLLM has about 150 provider folders. Each provider is a `Config` class with `transform_request` and `transform_response`, driven by one shared HTTP handler. It also serves `/v1/messages` and raw pass-through routes. [Portkey Gateway](/p/portkey-gateway/) registers 72 providers. Each one maps gateway parameters to provider parameters, with renames, defaults and min/max clamps, and streams come back as OpenAI SSE. [One API](/p/one-api/) accepts OpenAI format only. `GetAdaptor` maps 19 API types to a nine-method `Adaptor` interface, and most other channel types fall through to the OpenAI adaptor.

**Many formats in, many out.** New API accepts OpenAI Chat, Responses (HTTP and WebSocket), Claude Messages and Gemini. Its `relaykit` converters are a separate Go module, and they grade each conversion path. Plano's `hermesllm` crate maps three client APIs to five upstream shapes, including Bedrock Converse, in one `TryFrom` match. [OmniRoute](/p/omniroute/) converts through an OpenAI middle format unless a direct translator exists. It replays cached `reasoning_content` to thinking models to avoid upstream 400 errors.

**A shared schema library.** [Bifrost](/p/bifrost/) converts every request to a canonical `BifrostRequest` and has about 30 provider packages. Its HTTP server adds routes for the OpenAI, Anthropic, Bedrock, GenAI, Cohere and LiteLLM wire formats. [GPT-Load](/p/gpt-load/) writes no converters of its own. It forwards the client's native format when the upstream supports it. Converted routes use Bifrost's Go provider packages, and subscription accounts go through an embedded CLIProxyAPI.

**Translation inside Envoy.** [Higress](/p/higress/)'s `ai-proxy` WASM plugin supports 38 provider types. `protocol: original` skips conversion, and Claude-format requests are converted for providers that lack native support. Agent Router's ext_proc sidecar translates once per upstream attempt, from the stored original body. A retry from OpenAI to Bedrock therefore gets a fresh, correctly signed request. Its inputs include Anthropic `/v1/messages` and Cohere rerank, not only OpenAI.

Pick: LiteLLM or Portkey Gateway for the widest OpenAI-compatible provider coverage.
Pick: New API, Plano or OmniRoute when Claude, Gemini or Responses clients must reach any backend.
Pick: Bifrost to embed the translation layer as a Go library.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/translation/index.md

## How are API keys, users and tenants managed?

[LiteLLM](/p/litellm/) has the deepest tenant model, with organisations, teams, users, end users and keys. [Bifrost](/p/bifrost/) is a lighter alternative with virtual keys and OIDC. [One API](/p/one-api/) and [New API](/p/new-api/) suit groups that share or resell upstream keys.

**Tenant hierarchies with virtual keys.** LiteLLM stores hashed keys in `LiteLLM_VerificationToken`. It accepts the key in `Authorization`, `x-api-key` or Azure's `api-key`, and also supports JWT and SSO. Upstream keys can come from AWS Secrets Manager, Vault and other secret managers. Bifrost's `sk-bf-…` virtual keys map to provider configs and upstream keys, and can belong to teams and customers. OIDC login can sync teams and roles from a directory. [OmniRoute](/p/omniroute/) keeps client keys in SQLite with model, combo and connection allowlists, access schedules and IP allowlists. It encrypts upstream credentials with AES-256-GCM.

**User and token consoles.** One API has guest, common, admin and root roles and 48-character `sk-` tokens with a model list and a subnet limit. On first start it creates `root` with password `123456`, and channel keys are stored in plain text. New API keeps that model. It adds passkeys, TOTP, personal access tokens, step-up checks for sensitive actions, and `x-api-key` and `x-goog-api-key` headers. [GPT-Load](/p/gpt-load/) is narrower. Its hashed AccessKey is the only principal, with expiry, CIDR limits and filters on groups, protocols and models. There are no users or teams.

**Identity from the platform.** [Agent Router](/p/agent-router/) stores upstream credentials in Kubernetes Secrets referenced by a `BackendSecurityPolicy`. It signs AWS requests with SigV4 and rotates AWS, Azure and GCP tokens. It has no virtual keys, users or admin UI. [Higress](/p/higress/) reads the caller from an `x-mse-consumer` header set by an auth plugin, and holds an `apiTokens` list per provider.

**No key management.** [Plano](/p/plano/) has no users, teams or virtual keys. It has an `access_key` per provider, or `passthrough_auth` to forward the client's own header. [Portkey Gateway](/p/portkey-gateway/) receives provider keys with every request. It passes the virtual-key header through without resolving it, and `/v1/*` does not authenticate callers. `admin_token` guards only the UI and the log stream.

Pick: LiteLLM for multi-team cost attribution with SSO.
Pick: New API or One API to hand out quota-limited keys from one shared pool.
Pick: Agent Router or Higress when identity already lives in your gateway or cluster.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/keys-auth/index.md

## How are rate limits, budgets and cost tracking implemented?

[LiteLLM](/p/litellm/) has the most complete cost attribution, with a maintained price table and budgets at every level. [New API](/p/new-api/) is the most complete for running a paid service. [Bifrost](/p/bifrost/) and [GPT-Load](/p/gpt-load/) offer dollar caps in a smaller package.

**Dollar budgets from built-in prices.** LiteLLM ships a price file with about 4,500 model entries. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is written to PostgreSQL in batches. Bifrost attaches budgets with reset windows, plus token and request limits, to virtual keys, teams and customers. Its counters live in each process, so the open-source build does not share them across nodes. GPT-Load gives each AccessKey lifetime or periodic caps in nano-USD, an RPM limit and a concurrency limit. Every request gets a frozen price receipt. [OmniRoute](/p/omniroute/) enforces daily and weekly USD caps per key. Prices come from a short built-in map (about 20 models), synced or operator prices in SQLite, and an estimate for unknown models.

**Quota units for resale.** [One API](/p/one-api/) pre-charges `(prompt + max_tokens) × modelRatio × groupRatio` in internal units, then settles on real usage. Its rate limits are per IP (480 API requests per 3 minutes by default), not per key. New API keeps a billing reservation across channel retries. It adds tiered price expressions, top-ups through Stripe, Epay or Creem, and logs in SQL or ClickHouse.

**Token counters in the data plane.** [Higress](/p/higress/) keeps a Redis token quota per consumer and decrements it after the response, so one large request can overshoot. It has no price table. [Agent Router](/p/agent-router/) computes token or CEL cost expressions and enforces `QuotaPolicy` in Envoy's rate-limit service. Neither has a dollar budget. [Plano](/p/plano/) applies `governor` token buckets keyed by a header value. Its dollar tracking only caps how much switching models may add within a session (`max_switch_spend_pct`).

**Hook points only.** [Portkey Gateway](/p/portkey-gateway/) has no price tables and stores no logs. Budgets depend on a `preRequestValidator` that nothing in the repository sets, and its Redis token bucket is not wired into the request path.

Pick: LiteLLM for per-team spend tracking and budgets.
Pick: New API or One API to sell or ration access with prepaid quota.
Pick: Higress or Agent Router for token-based rate limits on an Envoy fleet.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/limits-cost/index.md

## How are caching and guardrails implemented?

[LiteLLM](/p/litellm/) offers the most here: exact and semantic caches on many backends, and about 60 guardrail integrations. [Portkey Gateway](/p/portkey-gateway/) has the most flexible guardrail hooks but only a process-local cache. [Bifrost](/p/bifrost/) has a strong semantic cache and no open-source guardrails.

**Caches and guardrails built in.** LiteLLM caches to memory, Redis, disk, S3, GCS or Azure Blob, with semantic caches on Redis, Qdrant or Valkey. Guardrails run before, during or after the call, including on stream chunks. Integrations include Presidio, Lakera and Bedrock Guardrails. [OmniRoute](/p/omniroute/) checks a SHA-256 exact match first, then embedding similarity on an in-memory or Redis vector store. Its guardrail registry holds a PII masker (off by default), a credential masker and a regex prompt-injection guard. [Higress](/p/higress/)'s `ai-cache` checks Redis, then a vector store. Its default key is the last message's content, without the model or system prompt. `ai-security-guard` sends text to Alibaba Cloud's moderation service.

**One strong half.** Bifrost tries a hashed exact lookup, then cosine similarity (default 0.8) on chromem, Redis, Qdrant, Weaviate or Pinecone. It skips conversations over three messages by default. Guardrails exist only as enterprise schema stubs. Portkey runs guardrails as before- and after-request hooks from about 20 plugin packs, and a failed `deny` check returns HTTP 446. Its cache is exact-match and in memory only. A request with an explicit `stream: false` is never stored, and `SEMANTIC_HIT` is just a label.

**Provider prompt caches, no response cache.** [Plano](/p/plano/) injects provider cache-control markers and keeps a session on the model whose cache is warm. Its `prompt_guards` setting is parsed but never applied. Guardrails work through HTTP agents in input and output filter chains instead. [GPT-Load](/p/gpt-load/) pins requests that share a prompt prefix to one credential. It adds regex redaction that replaces or encrypts text, and an experimental LLM-judge audit.

**Little or nothing.** [One API](/p/one-api/) has no response cache and no content checks. Redis caches only tokens, users and channels. [New API](/p/new-api/) adds one control: an admin word list (Aho-Corasick) that rejects matching prompts before billing. [Agent Router](/p/agent-router/) redacts content only in debug logs.

Pick: LiteLLM for semantic caching plus PII and moderation guardrails in one place.
Pick: Portkey Gateway for configurable per-request guardrails from many vendors.
Pick: Bifrost or Higress when a semantic cache matters more than guardrails.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/caching-guardrails/index.md

## How is it observed, deployed and scaled?

[Agent Router](/p/agent-router/) and [Higress](/p/higress/) scale most like ordinary infrastructure, because Envoy carries the traffic. [Bifrost](/p/bifrost/) is the most complete single binary for metrics and tracing. This site runs no benchmarks. The speed figures below are the projects' own claims, and only the architecture is compared.

**Envoy data planes.** Agent Router injects a Go ext_proc into each Envoy pod, as a regular container or as a native sidecar. It exports OpenTelemetry metrics (Prometheus always on) and traces with GenAI conventions. Each request makes gRPC round trips to that sidecar. Higress runs Go plugins compiled to WASM inside Envoy, with an Istio-based controller. `ai-statistics` emits Envoy counters per route, model and consumer. [Plano](/p/plano/) pairs Envoy and Rust WASM filters with a separate async Rust service, so a request crosses Envoy twice. It exports OTLP traces, and Redis lets replicas share session state.

**Go services.** Bifrost gives each provider a bounded queue with a fixed worker pool and uses fasthttp and object pools. It has Prometheus and OpenTelemetry plugins. Its README reports 11 µs of added overhead at 5,000 RPS. Counters and session affinity are per node in the open-source build. [One API](/p/one-api/) and [New API](/p/new-api/) are Gin binaries on SQLite, MySQL or PostgreSQL. Replicas set `NODE_TYPE=slave` and share the database and Redis. Neither exports OpenTelemetry, and New API can send logs to ClickHouse. [GPT-Load](/p/gpt-load/) swaps its config snapshot atomically. Its RPM limits, concurrency limits and health state stay in one process.

**Python and Node services.** [LiteLLM](/p/litellm/) runs FastAPI under uvicorn workers. Its chat path is still Python, and only a few routes use its Rust extension. Shared limits need PostgreSQL and Redis. It has Prometheus, OpenTelemetry and many logging callbacks, and its README claims 8 ms P95 latency at 1,000 RPS. [Portkey Gateway](/p/portkey-gateway/) is a stateless Hono app for Node or Cloudflare Workers. Its README claims under 1 ms of latency. Its logs go to `console` and an SSE stream, with no exporter. [OmniRoute](/p/omniroute/) is a Next.js app on SQLite. Its OTLP exporter is off unless an endpoint is set, and its in-memory rotation stores are not shared between replicas.

Pick: Agent Router or Higress on Kubernetes with an Envoy fleet.
Pick: Bifrost for a single binary with built-in Prometheus and OpenTelemetry.
Pick: LiteLLM when logging integrations matter more than per-instance overhead.

Per-project answers: https://llms-technical-reviews.com/llm-gateways/q/ops/index.md

## Projects

- [diegosouzapw/OmniRoute](https://llms-technical-reviews.com/p/omniroute/index.md) — Self-hosted TypeScript AI gateway that pools provider free tiers, subscription OAuth logins and web sessions behind one OpenAI-style API.
- [BerriAI/litellm](https://llms-technical-reviews.com/p/litellm/index.md) — Python SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails.
- [QuantumNous/new-api](https://llms-technical-reviews.com/p/new-api/index.md) — AGPL fork of One API that converts between OpenAI, Claude and Gemini formats, with weighted routing, tiered billing and JS task plugins.
- [songquanpeng/one-api](https://llms-technical-reviews.com/p/one-api/index.md) — Go/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI.
- [Portkey-AI/gateway](https://llms-technical-reviews.com/p/portkey-gateway/index.md) — Stateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs.
- [higress-group/higress](https://llms-technical-reviews.com/p/higress/index.md) — Envoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types.
- [maximhq/bifrost](https://llms-technical-reviews.com/p/bifrost/index.md) — Go LLM gateway and embeddable SDK with per-provider worker queues, a hook-based plugin pipeline, virtual-key governance and fallbacks.
- [katanemo/plano](https://llms-technical-reviews.com/p/plano/index.md) — Envoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs.
- [tbphp/gpt-load](https://llms-technical-reviews.com/p/gpt-load/index.md) — Self-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas.
- [theagentrouter/agent-router](https://llms-technical-reviews.com/p/agent-router/index.md) — Kubernetes control plane plus Envoy ext_proc sidecar that routes, translates, authenticates and meters LLM and MCP traffic.