# LLM gateways

> Proxies that put many model providers behind one API: routing and fallbacks, keys and budgets, caching, guardrails and usage tracking.

LLM gateways put many model providers behind one API. A client sends one request format, usually OpenAI's, and the gateway picks a provider, translates the call, retries or falls back on errors, and records usage. The ten projects here split along four axes. The first is form: an SDK plus proxy, a self-hosted console for sharing keys with billing, a lightweight proxy, or plugins inside an Envoy-based gateway. The second is routing: weighted or rule-based load balancing, fair scheduling over pools of keys or accounts, or intent-based model choice. The third is governance: how deep the user, team and budget model goes, and whether money is counted from a price table or only tokens. The fourth is where state lives: in one process, or shared through a database and Redis. When choosing, check which inbound formats you need. Check whether limits hold across replicas, and which features are only hooks or enterprise extras. Also check the providers' terms for pooled or subscription credentials.


## Projects

- [diegosouzapw/OmniRoute](https://llms-technical-reviews.com/p/omniroute/) — Self-hosted TypeScript AI gateway that pools provider free tiers, subscription OAuth logins and web sessions behind one OpenAI-style API. (★73707, TypeScript)
- [BerriAI/litellm](https://llms-technical-reviews.com/p/litellm/) — Python SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails. (★60246, Python)
- [QuantumNous/new-api](https://llms-technical-reviews.com/p/new-api/) — AGPL fork of One API that converts between OpenAI, Claude and Gemini formats, with weighted routing, tiered billing and JS task plugins. (★49331, Go)
- [songquanpeng/one-api](https://llms-technical-reviews.com/p/one-api/) — Go/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI. (★37077, Go)
- [Portkey-AI/gateway](https://llms-technical-reviews.com/p/portkey-gateway/) — Stateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs. (★13136, TypeScript)
- [higress-group/higress](https://llms-technical-reviews.com/p/higress/) — Envoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types. (★9496, Go)
- [maximhq/bifrost](https://llms-technical-reviews.com/p/bifrost/) — Go LLM gateway and embeddable SDK with per-provider worker queues, a hook-based plugin pipeline, virtual-key governance and fallbacks. (★8591, Go)
- [katanemo/plano](https://llms-technical-reviews.com/p/plano/) — Envoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs. (★7080, Rust)
- [tbphp/gpt-load](https://llms-technical-reviews.com/p/gpt-load/) — Self-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas. (★7045, Go)
- [theagentrouter/agent-router](https://llms-technical-reviews.com/p/agent-router/) — Kubernetes control plane plus Envoy ext_proc sidecar that routes, translates, authenticates and meters LLM and MCP traffic. (★2192, Go)


## Comparison questions

- [How are requests routed across providers and models?](https://llms-technical-reviews.com/llm-gateways/q/routing/) — Load balancing strategies; fallbacks and retries; latency- or cost-based routing; model aliases; health checks.
- [How are different provider APIs unified?](https://llms-technical-reviews.com/llm-gateways/q/translation/) — OpenAI-compatible (or other) schema; request/response translation; streaming; tool calls and multimodal; provider count.
- [How are API keys, users and tenants managed?](https://llms-technical-reviews.com/llm-gateways/q/keys-auth/) — Virtual keys; upstream credential storage; user/team/tenant model; auth methods; admin UI or API.
- [How are rate limits, budgets and cost tracking implemented?](https://llms-technical-reviews.com/llm-gateways/q/limits-cost/) — Per-key/user limits; token and spend budgets; price tables; usage accounting and where it is stored.
- [How are caching and guardrails implemented?](https://llms-technical-reviews.com/llm-gateways/q/caching-guardrails/) — Exact and semantic caching; PII redaction, moderation and prompt-injection checks; plugin or hook points.
- [How is it observed, deployed and scaled?](https://llms-technical-reviews.com/llm-gateways/q/ops/) — Logs, metrics and traces (OpenTelemetry); performance architecture (language, concurrency); deployment modes; HA.

Full comparison: https://llms-technical-reviews.com/compare/llm-gateways/