LLMs Technical Reviews

How are requests routed across providers and models?

Load balancing strategies; fallbacks and retries; latency- or cost-based routing; model aliases; health checks.

Verdict

LiteLLM and Bifrost have the most complete built-in routing. GPT-Load and OmniRoute go deepest on failover across many keys or accounts. Plano is the only one that routes by intent.

Strategy engines inside the gateway. LiteLLM’s Router defaults to simple-shuffle, a weighted draw by weight, rpm or tpm. It adds least-busy, latency, cost and usage strategies, cooldowns, and fallback in mid-stream. Bifrost first evaluates CEL rules with the scope order VirtualKey > User > Team > Customer > Global. Its governance plugin then makes a weighted provider pick and builds fallbacks from the remaining providers. OmniRoute “combos” offer about 19 strategies, including cost-optimised and quota-reset-aware. They sit on per-account cooldowns and a circuit breaker stored in SQLite. Many pooled accounts are subscription or web-session logins. OmniRoute’s own catalogue marks 17 providers avoid under their terms and warns of account bans, but by default that flag does not stop routing. Portkey Gateway reads a routing tree from each request’s x-portkey-config header, with fallback, weighted and conditional ($eq, $regex) nodes. Its circuit breaker is inert: nothing in the repository sets isOpen.

Channel and credential pools. One API picks uniformly at random among the channels in the top priority tier. Its Weight field is never read. New API adds weighted draws per tier, session affinity, pins and failover across groups. Channel health comes from scheduled tests and auto-ban, not its Uptime Kuma page. GPT-Load uses a deterministic fair scheduler and prefers the client’s native protocol before converted routes.

Routing left to Envoy. Agent Router only sets the x-ai-eg-model header. Envoy route weights, priorities and BackendTrafficPolicy retries do the rest. Higress fixes one activeProviderId per route rule. Per-request provider choice needs model-router to copy the model into a header that Envoy routes on. Failover happens between the API tokens of one provider.

Intent routing. In Plano, an orchestrator LLM picks a declared preference and ranks its models as cheapest or fastest. A session gate then keeps a model whose prompt cache is still warm. The main path sends only the chosen model and does not walk the ranked list.

Pick: LiteLLM or Bifrost for policy-driven load balancing and fallbacks. Pick: GPT-Load, New API or OmniRoute to spread traffic over many keys or accounts. Pick: Agent Router or Higress if Envoy should own traffic management.

Per-project answers

diegosouzapw/OmniRoute

answered

Routing is orchestrated by handleComboChat() in open-sse/services/combo.ts, which implements 19+ strategies (priority, weighted, round-robin, random, least-used, cost-optimized, reset-aware, reset-window, strict-random, auto, fill-first, p2c, lkgp, context-optimized, context-relay, fusion, pipeline, quota-share, chaos). The dispatch passes through a prelude sequence: pinned model → fusion (parallel judge-synthesis) → chaos (multi-model parallel) → pipeline → runtime-unit → round-robin → target iteration with cooldown-aware retry. Target resolution in targetResolution.ts performs provider-wildcard expansion, weighted step-group resolution, prompt-cache affinity, session stickiness, eval scores, and request-compatibility filtering. The auto strategy generates scored candidates via buildAutoCandidates() using 16 factors: cost per token, historical p95 latency, error rate, circuit-breaker state, quota remaining, session affinity, reset-window affinity, OAuth session availability, connection pool size, quality scores, and speed telemetry. Fallbacks operate at three resilience layers (detailed in AGENTS.md): provider circuit breakers (src/shared/utils/circuitBreaker.ts — 4 states: CLOSED/DEGRADED/OPEN/HALF_OPEN, failure-kind-aware thresholds, DB-persisted), connection cooldown (accountFallback.ts — per-credential exponential backoff), and model lockout (per provider+connection+model). Retries use full-jitter backoff with status-code-aware classifiers (408/500/502/503/504 trip the provider breaker; 401/403/429 route to connection cooldown). Model aliases resolve through open-sse/config/providerRegistry.ts where user-supplied names map to provider-specific model IDs. Health checks run via circuit-breaker lazy recovery: expired OPEN states auto-transition to HALF_OPEN on the next getStatus(), and canExecute() gates target selection.

BerriAI/litellm

answered

Load balancing. The Router class (litellm/router.py) supports multiple strategies: simple-shuffle (weighted random: deployments with higher weight, rpm, or tpm config values get proportionally more traffic — litellm/router_strategy/simple_shuffle.py:43-67), least-busy (picks the deployment with fewest in-flight requests, tracked via a Redis-backed counter incremented pre-call and decremented on success/failure — litellm/router_strategy/least_busy.py:114-225), latency-based-routing (selects the deployment with the lowest average or percentile time-to-first-token over a configurable window — litellm/router_strategy/lowest_latency.py:53-80), cost-based-routing (tracks token usage per deployment in minute buckets via cost_map:{model_group} cache keys to pick the cheapest — litellm/router_strategy/lowest_cost.py:15-80), usage-based-routing-v2 (T/RPM-aware), and LAR-1 (semantic routing by agent confidence thresholds — litellm/router_strategy/lar1_routing.py:1-60). Additional strategies include complexity-based (complexity_router/), quality-based (quality_router/), and adaptive (adaptive_router/) routers.

Fallbacks and retries. The key entry point Router.async_function_with_fallbacks() (router.py:7865) wraps every provider call. On failure it checks two levels: order-based fallback (tries higher target_order deployments within the same model group) then configured external fallbacks (explicit fallbacks=[...] or per-error-type context_window_fallbacks, content_policy_fallbacks). A RetryPolicy per deployment controls retry counts per error category. Failed deployments enter cooldown (configurable allowed_fails and cooldown_time). Mid-stream fallbacks are supported via MidStreamFallbackError + FallbackAwareStreamWrapper (router.py:828-867,3061-3167).

Model aliases. model_group_alias (router.py:1119) maps alias names to deployment groups, and Router.get_model_from_alias() resolves them at routing time (router.py:5382).

Health checks. async_get_available_deployment (router.py:13711-13831) runs pre-routing hooks and checks deployment health via configurable staleness thresholds before selecting a target. The health-check subsystem (litellm/proxy/health_check.py) pings endpoints periodically.

QuantumNous/new-api

answered

Requests are routed through a middleware pipeline that selects a channel (provider endpoint) for each request. The Distribute() middleware in middleware/distributor.go is the core dispatcher: it extracts the model name from the request body, applies token model limits (optional per-key model whitelist), resolves the user's group, and calls SelectChannelForRequest in service/channel_select.go.

Selection priority (service/channel_select.go:283-376): (1) A pinned channel (set by task plugins, origin-task affinity, or admin override) wins unconditionally. (2) On the first attempt (retry==0), session affinity is checked: if the same group+model was routed to a channel recently, that channel is reused (sticky sessions). (3) Otherwise CacheGetRandomSatisfiedChannel picks a random eligible channel by group+model+priority.

Load balancing is weighted: each Channel has a Weight field (model/channel.go:31), and multi-key channels support MultiKeyModeRandom (random selection among enabled keys) or MultiKeyModePolling (round-robin via an atomic polling index) (model/channel.go:249-289). Channels also have a Priority field (model/channel.go:45) for tiered fallback — when a priority level is exhausted, the system moves to the next priority group.

Retry and fallback work together: controller/relay.go:158 loops while retryParam.GetRetry() ≤ common.RetryTimes. Each iteration calls SelectChannelForRequest again, which may select a different channel. The DecideRelayRetry function in service/relay_error.go:21 decides whether to retry based on the error: channel-level errors trigger retries, 4xx status codes by default do not, and operation_setting can configure retryable status codes. Cross-group retry ("auto" group mode) cascades through configured groups: when one group has no eligible channels, the next group is tried (service/channel_select.go:121-167).

Model aliases work via model_mapping field on each channel (model/channel.go:42), parsed by ModelMappedHelper in relay/helper/model_mapped.go:14. It supports chain redirection with cycle detection. The ResolveTaskModelAlias in model/task_model_alias.go:38 resolves task-plugin aliases through ASCII-folded name lookups. Health checks run periodically via the channel test endpoint (controller/channel-test.go) and Uptime Kuma integration (controller/uptime_kuma.go:18). Channels with auto_ban enabled (model/channel.go:46) are automatically disabled when their keys are all disabled.

Editor's note. Correction: controller/uptime_kuma.go only reads an external Uptime Kuma status page for the console; it does not health-check channels. Channel health comes from scheduled channel tests (CHANNEL_TEST_FREQUENCY / monitor settings) plus auto-ban on errors.

songquanpeng/one-api

answered

Load balancing and priority. Each channel (upstream provider endpoint) is assigned a priority and weight. The routing system (middleware/distributor.go) calls model.CacheGetRandomSatisfiedChannel which queries an in-memory index group2model2channels built on startup and refreshed every SYNC_FREQUENCY seconds (model/cache.go). Channels are sorted by priority within each (group, model) bucket; the system picks randomly from the top-priority tier. If ignoreFirstPriority is set (during retries), it picks from lower-priority tiers instead. This is a weighted-random-within-priority-tier scheme.

Model aliases. Each channel carries an optional ModelMapping JSON field that maps user-visible model names to provider-specific model names (model/channel.go:114-125). The mapping is applied right before outbound request conversion (controller/text.go:37-38, controller/image.go:117-118).

Retries and fallbacks. controller/relay.go:45-101 implements retry logic: after a request fails, it calls CacheGetRandomSatisfiedChannel again (with ignoreFirstPriority=true), skipping the last-failed channel ID. The number of retries is config.RetryTimes (default 0). Channels that return 401/403 or certain insufficient_quota/invalid_api_key errors are automatically disabled by the monitor (monitor/manage.go:11-44).

Health checks. The monitor/metric.go module tracks a sliding window of success/failure per channel. If a channel's success rate for the last METRIC_QUEUE_SIZE calls falls below METRIC_SUCCESS_RATE_THRESHOLD (default 80%), the channel is auto-disabled (monitor/channel.go:47-61). Additionally, CHANNEL_TEST_FREQUENCY can be set to automatically test channels periodically (main.go:79-84). There is no explicit latency-based or cost-based routing — routing is purely random-within-priority-tier.

Editor's note. Correction: selection inside the top priority tier is uniform (rand.Intn), not weighted; the channel Weight field exists but is never used. The first retry stays in the top tier and only later retries move to lower-priority channels.

Portkey-AI/gateway

answered

Requests are routed through the tryTargetsRecursively function in src/handlers/handlerUtils.ts:476 which evaluates a JSON config (passed via x-portkey-config header or the request body). This config supports five strategy modes: Single — pick the first target; Fallback (strategy.mode: "fallback") — iterate targets sequentially, stopping on the first success (or matching onStatusCodes); Loadbalance (mode: "loadbalance") — weighted random selection via selectProviderByWeight at handlerUtils.ts:204, where each target has a numeric weight property (defaulting to 1); Conditional (mode: "conditional") — a ConditionalRouter class at src/services/conditionalRouter.ts:32 matches conditions against request metadata, URL path, and params using operators like $eq, $regex, $in, $and, $or, routing to a named target; Circuit Breaker — targets with isOpen flag are filtered out before selection (handlerUtils.ts:648-658).

Retries are handled by retryRequest in src/handlers/retryHandler.ts:65, which uses the async-retry library to retry on configurable onStatusCodes (default: 429, 500, 502, 503, 504 per src/globals.ts:38). It respects provider retry-after headers when followProviderRetry is set, with a 60-second max retry budget (MAX_RETRY_LIMIT_MS, globals.ts:5). Request timeout is configurable per-target via requestTimeout.

The gateway does not implement latency-based or cost-based dynamic routing — weights are static. Model aliasing is not built into the router; each provider option specifies its model directly in overrideParams. Health checks are passive via the circuit-breaker mechanism (currentTarget.targets.filter(t => !t.isOpen) at handlerUtils.ts:648-653), which tracks per-target open/close state.

Editor's note. Correction: the circuit-breaker path is a hook point, not a working feature in this repository. tryTargetsRecursively filters targets on isOpen and calls c.get('handleCircuitBreakerResponse'), but nothing in the open-source code sets either, so no target is ever marked open unless an embedding host supplies that callback.

higress-group/higress

answered

Routing is primarily handled by the ai-proxy plugin's provider config system and the ai-load-balancer/ai-endpoint-picker plugins.

Provider selection is defined per-request by the ai-proxy config's activeProviderId field in the PluginConfig, which selects one of the configured providers[] entries (each with id, type, apiTokens, modelMapping, etc.) — config/config.go:27-31. The model-router plugin further intercepts the body and rewrites the model field; it parses provider/model (e.g. "openai/gpt-4") and splits it into a provider header and model, enabling provider-routing — model-router/main.go:256-275. It also supports auto-routing with regex rules matched against the last user message: when the model is higress/auto, rules match on message content to select the destination model — model-router/main.go:194-248.

Load balancing is provided by ai-load-balancer, supporting cluster-level (cluster_metrics, cluster_hash) and endpoint-level (endpoint_metrics, global_least_request, prefix_cache) policies — ai-load-balancer/main.go:48-86. The ai-endpoint-picker plugin does fine-grained endpoint selection based on signals including queue depth, KV cache, prefix cache, LoRA affinity, inflight requests, and failure counts, all scored by a weighted pipeline — ai-endpoint-picker/scheduling/types.go:6-12.

Retries and failover are configured per-provider. The retry mechanism (retry.go) re-sends failed requests using different API tokens with configurable maxRetries and status code matching — ai-proxy/provider/retry.go:21-29. The failover mechanism (failover.go) tracks per-token failure counts and automatically removes failing tokens, then performs periodic health checks to restore them — ai-proxy/provider/failover.go:370-403. Token selection uses random choice or FNV-1a consistent hashing for stateful APIs — ai-proxy/provider/provider.go:834-919.

Model aliases are done via modelMapping with exact match, wildcard (*), prefix (prefix*), and regex (~pattern) rules — ai-proxy/provider/provider.go:1053-1100.

Health checks for failover send real chat completions to unavailable tokens and count successes before restoring — ai-proxy/provider/failover.go:192-218.

Editor's note. Correction: activeProviderId is applied when the plugin config (global or per route/domain rule) is parsed, not per request; each rule has exactly one active provider. Choosing a provider per request means routing to different Envoy routes, typically by having model-router copy the model or provider prefix into a header that the route matches on.

maximhq/bifrost

answered

Bifrost uses a three-layer routing pipeline. Layer 1 — CEL rule engine: The routing plugin (plugins/routing/main.go) evaluates Google CEL (Common Expression Language) rules with scope precedence (VirtualKey > User > Team > Customer > Global). Each TableRoutingRule has a CelExpression (e.g. headers["x-region"] == "eu") that matches on model, provider, request_type, headers, query params, complexity_tier, budget_used, tokens_used, and request count. Rules can be terminal or chainable (recursive re-evaluation after a matched rule rewrites provider/model). Layer 2 — Weighted provider load balancing: After rules settle, GovernancePlugin.LoadBalanceProvider (plugins/governance/main.go:430) selects among eligible weighted provider configs using weighted random selection. The eligible set is filtered by model allowance, budget limits, rate limits, and key grants. Fallbacks are automatically derived from remaining weighted providers sorted by weight. Layer 3 — Ordered fallback chain in core: Bifrost.handleRequest (core/bifrost.go:5689) iterates the fallback list sequentially. Each fallback calls tryRequest on the next provider/model, with full tracing spans per attempt. TTFT (time-to-first-token) deadlines can cut short slow streaming attempts via AttemptAbort (core/providers/utils/stream.go:29). Session affinity (core/sessionaffinity.go) remembers which provider/key served a session, routing subsequent requests to the same provider. Health checks are implicit — failures trigger fallback traversal; a provider that fails in tryRequest is skipped for subsequent fallbacks. The model catalog supports aliases via keyconfig.Store.ResolveAlias, allowing one model name to map to different wire names per key or provider.

katanemo/plano

answered

Routing in Plano is a two-stage process: a quality-based router picks a candidate model, then a session-cache-aware gate decides whether to honor it or stick to the warm anchor.

The quality router lives in crates/brightstaff/src/handlers/llm/model_selection.rs, which calls orchestrator_service.determine_route() (crates/brightstaff/src/router/orchestrator.rs) and returns a ranked list of models for fallback. The router converts all request shapes to a common ChatCompletionsRequest via ProviderRequestType::try_from(), then invokes a downstream orchestration model (an LLM used for routing decisions). Route preferences are configured as TopLevelRoutingPreference with optional SelectionPolicy (prefer: cheapest | fastest | none) and can be backed by live cost/latency metrics from ModelMetricsService (crates/brightstaff/src/router/model_metrics.rs), which fetches pricing catalogs from DigitalOcean or models.dev.

Fallbacks and retries use the ranked models: Vec<String> returned by router_chat_get_upstream_model. If the primary model returns 429/5xx, the ranked list provides ordered alternatives. There is no explicit load-balancing across providers — the router always picks the single best model for the request.

Session-cache-aware switch gating (crates/brightstaff/src/handlers/llm/session_router.rs): The route() function checks whether the session's provider cache is still warm (comparing last_used against the provider's cache TTL from provider_cache_capability() in crates/hermesllm/src/providers/id.rs). When warm, it either sticks to the anchor model or allows a switch within a configurable cost envelope (RoutingBudget — cumulative switch spend capped at max_switch_spend_pct% of the never-switch baseline, priced from the model rates feed). Switch cost is computed in USD: switch_cost_in_usd() in orchestrator.rs compares the candidate's uncached read cost against the anchor's cached read cost. A switch that is outright cheaper (negative cost) is always allowed.

Model aliases come from the model_aliases config map (type HashMap<String, ModelAlias>) in common/configuration.rs:48. Health checks are not implemented as active pings — the system relies on upstream HTTP error codes.

Editor's note. Correction: the main proxy path (llm_chat) sends only the single decided model and does not walk the ranked list on 429/5xx. The ranked list is returned by the /routing/* decision endpoints so a client can fall back itself; Envoy only retries the same provider cluster when llm_gateway_listener.max_retries is set.

tbphp/gpt-load

answered

Routing in GPT-Load is a multi-stage pipeline of dialect inspection, affinity resolution, and weighted-fair credential selection. The handler.Handle() method in internal/gateway/handler.go:419 is the main orchestrator. It first resolves the client protocol via dialect.Dialect.InspectRequest(), obtaining a RequestMetadata with Operation, RouteRequirement and Model. A scheduler.Query is built and handed to scheduler.Iterator.Next() (internal/scheduler/scheduler.go:298), which selects one credential from a pool of candidate targets.

Route modes — two wire strategies: RouteNative (preserve the client protocol upstream) and RouteConverted (translate to a provider-neutral format). RouteRequirement (execution.RouteRequirement, internal/execution/contracts.go:171) declares which modes are acceptable. execution.RouteMode (internal/execution/contracts.go:129-133) records the selected mode.

Route strategies — two global policies in internal/state/runtime_settings.go:51-55: RouteStrategyNativeFirst (try native routes first, fall back to converted) and RouteStrategyWeightedMix (treat both modes equally). The scheduler (scheduler.go:153-154) sets routeModeTiers accordingly — [[native], [converted]] for native-first, [[native, converted]] for weighted-mix.

Weighted-fair scheduling — within a group, credentials are selected via a fairness scheduler (scheduler/fair.go). Each credential carries a WeightManual, and the SchedulingLedger tracks progress watermarks. The lowest-progress eligible credential wins (scheduler/fair.go:71-79), ensuring even load across weighted candidates.

Fallbacks and retries — after an upstream attempt, health.JudgeExecution() (internal/health/execution_judge.go:27) decides the retry directive: RetryNone, RetryRefreshCredential, or RetryNextCandidate. Retries increment an attempt counter and call Next() again, skipping already-tried credentials (scheduler.go:330). Cooldowns and blacklisting come from health.Decision.Effect (internal/health/decision.go:69).

Health checks — health.StatsStore (internal/health/stats.go:1) tracks per-credential success/failure buckets in 1-minute windows (5 buckets). Problems trigger credential cooldown (handler.go:341-361), model-level cooldown, or blacklisting after exceeding BlacklistThreshold consecutive failures.

Affinity — affinity.go (internal/gateway/affinity.go:19) derives a prompt-affinity key from the first user message. If a prior request used the same prefix, the same credential is preferred. Prompt-cache-key affinity is also supported. The affinity cache (affinity.Cache) sits in handler.affinityCache.

Model aliases — state.ModelConfig includes an Alias field (internal/state/snapshot.go:75-78). externalModelName() returns the alias if set, otherwise the model ID. The scheduler resolves the external model name against upstream model IDs via SelectModel().

theagentrouter/agent-router

answered

Routing is defined declaratively via the AIGatewayRoute Kubernetes CRD. Each AIGatewayRoute contains Rules, and each rule has Matches (header-based conditions) and BackendRefs (references to AIServiceBackend or InferencePool resources). The x-ai-eg-model header is injected by the ext_proc filter after parsing the request body (e.g., the model field from a chat completion request), making model-based routing possible via header matching in Envoy Gateway's HTTPRoute system (api/v1beta1/ai_gateway_route.go:57-98). Multiple backends per rule with Weight fields (default 1) support weighted traffic distribution, and Priority fields enable priority-based load balancing in Envoy (api/v1beta1/ai_gateway_route.go:387-407). Fallback and retries are not implemented in this codebase directly — they are delegated to Envoy Gateway's BackendTrafficPolicy resource, which provides retry and failover at the data plane level (api/v1beta1/ai_gateway_route.go:226-241). There is no latency- or cost-based routing in this codebase; routing is purely rule/header-based via Envoy's existing routing engine. Model aliases are implemented via ModelNameOverride on each backend ref, which the translator applies by rewriting the model field in the request body (internal/translator/openai_openai.go:69-95). Health checks are provided via the aigw healthcheck CLI command (cmd/aigw/main.go:39-40,60). The actual routing decision is made by Envoy's external processor filter (ext_proc), whose Go sidecar receives the request headers/body and can return CONTINUE/CONTINUE_AND_REPLACE to influence where Envoy sends the request (internal/extproc/processor.go:21-40). InferencePool support provides endpoint picker integration for dynamic backend selection (internal/controller/inference_pool.go:28-41).

How are different provider APIs unified? →