LLMs Technical Reviews
Home / LLM gateways

LLM gateways

Proxies that put many model providers behind one API: routing and fallbacks, keys and budgets, caching, guardrails and usage tracking.

LLM gateways put many model providers behind one API. A client sends one request format, usually OpenAI's, and the gateway picks a provider, translates the call, retries or falls back on errors, and records usage. The ten projects here split along four axes. The first is form: an SDK plus proxy, a self-hosted console for sharing keys with billing, a lightweight proxy, or plugins inside an Envoy-based gateway. The second is routing: weighted or rule-based load balancing, fair scheduling over pools of keys or accounts, or intent-based model choice. The third is governance: how deep the user, team and budget model goes, and whether money is counted from a price table or only tokens. The fourth is where state lives: in one process, or shared through a database and Redis. When choosing, check which inbound formats you need. Check whether limits hold across replicas, and which features are only hooks or enterprise extras. Also check the providers' terms for pooled or subscription credentials.

Projects (10)

ProjectStarsLanguageLicense
diegosouzapw/OmniRouteSelf-hosted TypeScript AI gateway that pools provider free tiers, subscription OAuth logins and web sessions behind one OpenAI-style API.★ 74kTypeScriptMIT
BerriAI/litellmPython SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails.★ 60kPython—
QuantumNous/new-apiAGPL fork of One API that converts between OpenAI, Claude and Gemini formats, with weighted routing, tiered billing and JS task plugins.★ 49kGoAGPL-3.0
songquanpeng/one-apiGo/Gin LLM gateway that puts many providers behind one OpenAI-format API, with per-user tokens, quota billing and a React admin UI.★ 37kGoMIT
Portkey-AI/gatewayStateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs.★ 13kTypeScriptMIT
higress-group/higressEnvoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types.★ 9.5kGoApache-2.0
maximhq/bifrostGo LLM gateway and embeddable SDK with per-provider worker queues, a hook-based plugin pipeline, virtual-key governance and fallbacks.★ 8.6kGoApache-2.0
katanemo/planoEnvoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs.★ 7.1kRustApache-2.0
tbphp/gpt-loadSelf-hosted Go gateway that pools API keys and subscription accounts, with fair scheduling, health-based failover and cost quotas.★ 7.0kGoMIT
theagentrouter/agent-routerKubernetes control plane plus Envoy ext_proc sidecar that routes, translates, authenticates and meters LLM and MCP traffic.★ 2.2kGoApache-2.0

Comparison questions

Each question is answered separately for every project in this category, from that project's source code.

All verdicts on one page →

  1. How are requests routed across providers and models?Load balancing strategies; fallbacks and retries; latency- or cost-based routing; model aliases; health checks.
  2. How are different provider APIs unified?OpenAI-compatible (or other) schema; request/response translation; streaming; tool calls and multimodal; provider count.
  3. How are API keys, users and tenants managed?Virtual keys; upstream credential storage; user/team/tenant model; auth methods; admin UI or API.
  4. How are rate limits, budgets and cost tracking implemented?Per-key/user limits; token and spend budgets; price tables; usage accounting and where it is stored.
  5. How are caching and guardrails implemented?Exact and semantic caching; PII redaction, moderation and prompt-injection checks; plugin or hook points.
  6. How is it observed, deployed and scaled?Logs, metrics and traces (OpenTelemetry); performance architecture (language, concurrency); deployment modes; HA.