LLM gateways: OmniRoute vs litellm vs new-api vs one-api vs gateway vs higress vs bifrost vs plano vs gpt-load vs agent-router
Proxies that put many model providers behind one API: routing and fallbacks, keys and budgets, caching, guardrails and usage tracking. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.
At a glance
● answered from code · — not applicable (the project does not do this) · ? insufficient evidence
How are requests routed across providers and models?
LiteLLM and Bifrost have the most complete built-in routing. GPT-Load and OmniRoute go deepest on failover across many keys or accounts. Plano is the only one that routes by intent.
Strategy engines inside the gateway. LiteLLM's Router defaults to simple-shuffle, a weighted draw by weight, rpm or tpm. It adds least-busy, latency, cost and usage strategies, cooldowns, and fallback in mid-stream. Bifrost first evaluates CEL rules with the scope order VirtualKey > User > Team > Customer > Global. Its governance plugin then makes a weighted provider pick and builds fallbacks from the remaining providers. OmniRoute "combos" offer about 19 strategies, including cost-optimised and quota-reset-aware. They sit on per-account cooldowns and a circuit breaker stored in SQLite. Many pooled accounts are subscription or web-session logins. OmniRoute's own catalogue marks 17 providers avoid under their terms and warns of account bans, but by default that flag does not stop routing. Portkey Gateway reads a routing tree from each request's x-portkey-config header, with fallback, weighted and conditional ($eq, $regex) nodes. Its circuit breaker is inert: nothing in the repository sets isOpen.
Channel and credential pools. One API picks uniformly at random among the channels in the top priority tier. Its Weight field is never read. New API adds weighted draws per tier, session affinity, pins and failover across groups. Channel health comes from scheduled tests and auto-ban, not its Uptime Kuma page. GPT-Load uses a deterministic fair scheduler and prefers the client's native protocol before converted routes.
Routing left to Envoy. Agent Router only sets the x-ai-eg-model header. Envoy route weights, priorities and BackendTrafficPolicy retries do the rest. Higress fixes one activeProviderId per route rule. Per-request provider choice needs model-router to copy the model into a header that Envoy routes on. Failover happens between the API tokens of one provider.
Intent routing. In Plano, an orchestrator LLM picks a declared preference and ranks its models as cheapest or fastest. A session gate then keeps a model whose prompt cache is still warm. The main path sends only the chosen model and does not walk the ranked list.
Pick: LiteLLM or Bifrost for policy-driven load balancing and fallbacks. Pick: GPT-Load, New API or OmniRoute to spread traffic over many keys or accounts. Pick: Agent Router or Higress if Envoy should own traffic management.
How are different provider APIs unified?
LiteLLM covers the most providers behind an OpenAI-style API. New API and Plano are the strongest when clients speak several formats, such as Claude Code next to an OpenAI SDK. Agent Router is the most careful about cross-provider retries.
OpenAI format in, adapters out. LiteLLM has about 150 provider folders. Each provider is a Config class with transform_request and transform_response, driven by one shared HTTP handler. It also serves /v1/messages and raw pass-through routes. Portkey Gateway registers 72 providers. Each one maps gateway parameters to provider parameters, with renames, defaults and min/max clamps, and streams come back as OpenAI SSE. One API accepts OpenAI format only. GetAdaptor maps 19 API types to a nine-method Adaptor interface, and most other channel types fall through to the OpenAI adaptor.
Many formats in, many out. New API accepts OpenAI Chat, Responses (HTTP and WebSocket), Claude Messages and Gemini. Its relaykit converters are a separate Go module, and they grade each conversion path. Plano's hermesllm crate maps three client APIs to five upstream shapes, including Bedrock Converse, in one TryFrom match. OmniRoute converts through an OpenAI middle format unless a direct translator exists. It replays cached reasoning_content to thinking models to avoid upstream 400 errors.
A shared schema library. Bifrost converts every request to a canonical BifrostRequest and has about 30 provider packages. Its HTTP server adds routes for the OpenAI, Anthropic, Bedrock, GenAI, Cohere and LiteLLM wire formats. GPT-Load writes no converters of its own. It forwards the client's native format when the upstream supports it. Converted routes use Bifrost's Go provider packages, and subscription accounts go through an embedded CLIProxyAPI.
Translation inside Envoy. Higress's ai-proxy WASM plugin supports 38 provider types. protocol: original skips conversion, and Claude-format requests are converted for providers that lack native support. Agent Router's ext_proc sidecar translates once per upstream attempt, from the stored original body. A retry from OpenAI to Bedrock therefore gets a fresh, correctly signed request. Its inputs include Anthropic /v1/messages and Cohere rerank, not only OpenAI.
Pick: LiteLLM or Portkey Gateway for the widest OpenAI-compatible provider coverage. Pick: New API, Plano or OmniRoute when Claude, Gemini or Responses clients must reach any backend. Pick: Bifrost to embed the translation layer as a Go library.
How are API keys, users and tenants managed?
LiteLLM has the deepest tenant model, with organisations, teams, users, end users and keys. Bifrost is a lighter alternative with virtual keys and OIDC. One API and New API suit groups that share or resell upstream keys.
Tenant hierarchies with virtual keys. LiteLLM stores hashed keys in LiteLLM_VerificationToken. It accepts the key in Authorization, x-api-key or Azure's api-key, and also supports JWT and SSO. Upstream keys can come from AWS Secrets Manager, Vault and other secret managers. Bifrost's sk-bf-… virtual keys map to provider configs and upstream keys, and can belong to teams and customers. OIDC login can sync teams and roles from a directory. OmniRoute keeps client keys in SQLite with model, combo and connection allowlists, access schedules and IP allowlists. It encrypts upstream credentials with AES-256-GCM.
User and token consoles. One API has guest, common, admin and root roles and 48-character sk- tokens with a model list and a subnet limit. On first start it creates root with password 123456, and channel keys are stored in plain text. New API keeps that model. It adds passkeys, TOTP, personal access tokens, step-up checks for sensitive actions, and x-api-key and x-goog-api-key headers. GPT-Load is narrower. Its hashed AccessKey is the only principal, with expiry, CIDR limits and filters on groups, protocols and models. There are no users or teams.
Identity from the platform. Agent Router stores upstream credentials in Kubernetes Secrets referenced by a BackendSecurityPolicy. It signs AWS requests with SigV4 and rotates AWS, Azure and GCP tokens. It has no virtual keys, users or admin UI. Higress reads the caller from an x-mse-consumer header set by an auth plugin, and holds an apiTokens list per provider.
No key management. Plano has no users, teams or virtual keys. It has an access_key per provider, or passthrough_auth to forward the client's own header. Portkey Gateway receives provider keys with every request. It passes the virtual-key header through without resolving it, and /v1/* does not authenticate callers. admin_token guards only the UI and the log stream.
Pick: LiteLLM for multi-team cost attribution with SSO. Pick: New API or One API to hand out quota-limited keys from one shared pool. Pick: Agent Router or Higress when identity already lives in your gateway or cluster.
How are rate limits, budgets and cost tracking implemented?
LiteLLM has the most complete cost attribution, with a maintained price table and budgets at every level. New API is the most complete for running a paid service. Bifrost and GPT-Load offer dollar caps in a smaller package.
Dollar budgets from built-in prices. LiteLLM ships a price file with about 4,500 model entries. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is written to PostgreSQL in batches. Bifrost attaches budgets with reset windows, plus token and request limits, to virtual keys, teams and customers. Its counters live in each process, so the open-source build does not share them across nodes. GPT-Load gives each AccessKey lifetime or periodic caps in nano-USD, an RPM limit and a concurrency limit. Every request gets a frozen price receipt. OmniRoute enforces daily and weekly USD caps per key. Prices come from a short built-in map (about 20 models), synced or operator prices in SQLite, and an estimate for unknown models.
Quota units for resale. One API pre-charges (prompt + max_tokens) × modelRatio × groupRatio in internal units, then settles on real usage. Its rate limits are per IP (480 API requests per 3 minutes by default), not per key. New API keeps a billing reservation across channel retries. It adds tiered price expressions, top-ups through Stripe, Epay or Creem, and logs in SQL or ClickHouse.
Token counters in the data plane. Higress keeps a Redis token quota per consumer and decrements it after the response, so one large request can overshoot. It has no price table. Agent Router computes token or CEL cost expressions and enforces QuotaPolicy in Envoy's rate-limit service. Neither has a dollar budget. Plano applies governor token buckets keyed by a header value. Its dollar tracking only caps how much switching models may add within a session (max_switch_spend_pct).
Hook points only. Portkey Gateway has no price tables and stores no logs. Budgets depend on a preRequestValidator that nothing in the repository sets, and its Redis token bucket is not wired into the request path.
Pick: LiteLLM for per-team spend tracking and budgets. Pick: New API or One API to sell or ration access with prepaid quota. Pick: Higress or Agent Router for token-based rate limits on an Envoy fleet.
How are caching and guardrails implemented?
LiteLLM offers the most here: exact and semantic caches on many backends, and about 60 guardrail integrations. Portkey Gateway has the most flexible guardrail hooks but only a process-local cache. Bifrost has a strong semantic cache and no open-source guardrails.
Caches and guardrails built in. LiteLLM caches to memory, Redis, disk, S3, GCS or Azure Blob, with semantic caches on Redis, Qdrant or Valkey. Guardrails run before, during or after the call, including on stream chunks. Integrations include Presidio, Lakera and Bedrock Guardrails. OmniRoute checks a SHA-256 exact match first, then embedding similarity on an in-memory or Redis vector store. Its guardrail registry holds a PII masker (off by default), a credential masker and a regex prompt-injection guard. Higress's ai-cache checks Redis, then a vector store. Its default key is the last message's content, without the model or system prompt. ai-security-guard sends text to Alibaba Cloud's moderation service.
One strong half. Bifrost tries a hashed exact lookup, then cosine similarity (default 0.8) on chromem, Redis, Qdrant, Weaviate or Pinecone. It skips conversations over three messages by default. Guardrails exist only as enterprise schema stubs. Portkey runs guardrails as before- and after-request hooks from about 20 plugin packs, and a failed deny check returns HTTP 446. Its cache is exact-match and in memory only. A request with an explicit stream: false is never stored, and SEMANTIC_HIT is just a label.
Provider prompt caches, no response cache. Plano injects provider cache-control markers and keeps a session on the model whose cache is warm. Its prompt_guards setting is parsed but never applied. Guardrails work through HTTP agents in input and output filter chains instead. GPT-Load pins requests that share a prompt prefix to one credential. It adds regex redaction that replaces or encrypts text, and an experimental LLM-judge audit.
Little or nothing. One API has no response cache and no content checks. Redis caches only tokens, users and channels. New API adds one control: an admin word list (Aho-Corasick) that rejects matching prompts before billing. Agent Router redacts content only in debug logs.
Pick: LiteLLM for semantic caching plus PII and moderation guardrails in one place. Pick: Portkey Gateway for configurable per-request guardrails from many vendors. Pick: Bifrost or Higress when a semantic cache matters more than guardrails.
How is it observed, deployed and scaled?
Agent Router and Higress scale most like ordinary infrastructure, because Envoy carries the traffic. Bifrost is the most complete single binary for metrics and tracing. This site runs no benchmarks. The speed figures below are the projects' own claims, and only the architecture is compared.
Envoy data planes. Agent Router injects a Go ext_proc into each Envoy pod, as a regular container or as a native sidecar. It exports OpenTelemetry metrics (Prometheus always on) and traces with GenAI conventions. Each request makes gRPC round trips to that sidecar. Higress runs Go plugins compiled to WASM inside Envoy, with an Istio-based controller. ai-statistics emits Envoy counters per route, model and consumer. Plano pairs Envoy and Rust WASM filters with a separate async Rust service, so a request crosses Envoy twice. It exports OTLP traces, and Redis lets replicas share session state.
Go services. Bifrost gives each provider a bounded queue with a fixed worker pool and uses fasthttp and object pools. It has Prometheus and OpenTelemetry plugins. Its README reports 11 µs of added overhead at 5,000 RPS. Counters and session affinity are per node in the open-source build. One API and New API are Gin binaries on SQLite, MySQL or PostgreSQL. Replicas set NODE_TYPE=slave and share the database and Redis. Neither exports OpenTelemetry, and New API can send logs to ClickHouse. GPT-Load swaps its config snapshot atomically. Its RPM limits, concurrency limits and health state stay in one process.
Python and Node services. LiteLLM runs FastAPI under uvicorn workers. Its chat path is still Python, and only a few routes use its Rust extension. Shared limits need PostgreSQL and Redis. It has Prometheus, OpenTelemetry and many logging callbacks, and its README claims 8 ms P95 latency at 1,000 RPS. Portkey Gateway is a stateless Hono app for Node or Cloudflare Workers. Its README claims under 1 ms of latency. Its logs go to console and an SSE stream, with no exporter. OmniRoute is a Next.js app on SQLite. Its OTLP exporter is off unless an endpoint is set, and its in-memory rotation stores are not shared between replicas.
Pick: Agent Router or Higress on Kubernetes with an Envoy fleet. Pick: Bifrost for a single binary with built-in Prometheus and OpenTelemetry. Pick: LiteLLM when logging integrations matter more than per-instance overhead.
The projects
- diegosouzapw/OmniRoute: Self-hosted TypeScript AI gateway that pools provider free tiers, subscription OAuth logins and web sessions behind one OpenAI-style API.
- BerriAI/litellm: Python SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails.
- QuantumNous/new-api: AGPL fork of One API that converts between OpenAI, Claude and Gemini formats, with weighted routing, tiered billing and JS task plugins.
- songquanpeng/one-api: Go/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI.
- Portkey-AI/gateway: Stateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs.
- higress-group/higress: Envoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types.
- maximhq/bifrost: Go LLM gateway and embeddable SDK with per-provider worker queues, a hook-based plugin pipeline, virtual-key governance and fallbacks.
- katanemo/plano: Envoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs.
- tbphp/gpt-load: Self-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas.
- theagentrouter/agent-router: Kubernetes control plane plus Envoy ext_proc sidecar that routes, translates, authenticates and meters LLM and MCP traffic.