# How are requests routed across providers and models?

> LLM gateways — a good answer covers: Load balancing strategies; fallbacks and retries; latency- or cost-based routing; model aliases; health checks.

Canonical page: https://llms-technical-reviews.com/llm-gateways/q/routing/

## Verdict

[LiteLLM](/p/litellm/) and [Bifrost](/p/bifrost/) have the most complete built-in routing. [GPT-Load](/p/gpt-load/) and [OmniRoute](/p/omniroute/) go deepest on failover across many keys or accounts. [Plano](/p/plano/) is the only one that routes by intent.

**Strategy engines inside the gateway.** LiteLLM's `Router` defaults to `simple-shuffle`, a weighted draw by `weight`, `rpm` or `tpm`. It adds least-busy, latency, cost and usage strategies, cooldowns, and fallback in mid-stream. Bifrost first evaluates CEL rules with the scope order VirtualKey > User > Team > Customer > Global. Its governance plugin then makes a weighted provider pick and builds fallbacks from the remaining providers. OmniRoute "combos" offer about 19 strategies, including cost-optimised and quota-reset-aware. They sit on per-account cooldowns and a circuit breaker stored in SQLite. Many pooled accounts are subscription or web-session logins. OmniRoute's own catalogue marks 17 providers `avoid` under their terms and warns of account bans, but by default that flag does not stop routing. [Portkey Gateway](/p/portkey-gateway/) reads a routing tree from each request's `x-portkey-config` header, with fallback, weighted and conditional (`$eq`, `$regex`) nodes. Its circuit breaker is inert: nothing in the repository sets `isOpen`.

**Channel and credential pools.** [One API](/p/one-api/) picks uniformly at random among the channels in the top priority tier. Its `Weight` field is never read. [New API](/p/new-api/) adds weighted draws per tier, session affinity, pins and failover across groups. Channel health comes from scheduled tests and auto-ban, not its Uptime Kuma page. GPT-Load uses a deterministic fair scheduler and prefers the client's native protocol before converted routes.

**Routing left to Envoy.** [Agent Router](/p/agent-router/) only sets the `x-ai-eg-model` header. Envoy route weights, priorities and `BackendTrafficPolicy` retries do the rest. [Higress](/p/higress/) fixes one `activeProviderId` per route rule. Per-request provider choice needs `model-router` to copy the model into a header that Envoy routes on. Failover happens between the API tokens of one provider.

**Intent routing.** In Plano, an orchestrator LLM picks a declared preference and ranks its models as `cheapest` or `fastest`. A session gate then keeps a model whose prompt cache is still warm. The main path sends only the chosen model and does not walk the ranked list.

Pick: LiteLLM or Bifrost for policy-driven load balancing and fallbacks.
Pick: GPT-Load, New API or OmniRoute to spread traffic over many keys or accounts.
Pick: Agent Router or Higress if Envoy should own traffic management.

## Per-project answers

### diegosouzapw/OmniRoute (answered)

Routing is orchestrated by `handleComboChat()` in `open-sse/services/combo.ts`, which implements 19+ strategies (priority, weighted, round-robin, random, least-used, cost-optimized, reset-aware, reset-window, strict-random, auto, fill-first, p2c, lkgp, context-optimized, context-relay, fusion, pipeline, quota-share, chaos). The dispatch passes through a prelude sequence: pinned model → fusion (parallel judge-synthesis) → chaos (multi-model parallel) → pipeline → runtime-unit → round-robin → target iteration with cooldown-aware retry. Target resolution in `targetResolution.ts` performs provider-wildcard expansion, weighted step-group resolution, prompt-cache affinity, session stickiness, eval scores, and request-compatibility filtering. The auto strategy generates scored candidates via `buildAutoCandidates()` using 16 factors: cost per token, historical p95 latency, error rate, circuit-breaker state, quota remaining, session affinity, reset-window affinity, OAuth session availability, connection pool size, quality scores, and speed telemetry. Fallbacks operate at three resilience layers (detailed in `AGENTS.md`): provider circuit breakers (`src/shared/utils/circuitBreaker.ts` — 4 states: CLOSED/DEGRADED/OPEN/HALF_OPEN, failure-kind-aware thresholds, DB-persisted), connection cooldown (`accountFallback.ts` — per-credential exponential backoff), and model lockout (per provider+connection+model). Retries use full-jitter backoff with status-code-aware classifiers (408/500/502/503/504 trip the provider breaker; 401/403/429 route to connection cooldown). Model aliases resolve through `open-sse/config/providerRegistry.ts` where user-supplied names map to provider-specific model IDs. Health checks run via circuit-breaker lazy recovery: expired OPEN states auto-transition to HALF_OPEN on the next `getStatus()`, and `canExecute()` gates target selection.


Citations: [open-sse/services/combo.ts:669-710](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/combo.ts#L669-L710) · [open-sse/services/combo.ts:317-395](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/combo.ts#L317-L395) · [src/shared/utils/circuitBreaker.ts:129-175](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/src/shared/utils/circuitBreaker.ts#L129-L175) · [open-sse/services/combo/targetResolution.ts:1-30](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/combo/targetResolution.ts#L1-L30) · [open-sse/services/accountFallback.ts:1-80](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/accountFallback.ts#L1-L80) · [open-sse/config/providerRegistry.ts:1-15](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/config/providerRegistry.ts#L1-L15)

### BerriAI/litellm (answered)

**Load balancing.** The `Router` class (`litellm/router.py`) supports multiple strategies: **simple-shuffle** (weighted random: deployments with higher `weight`, `rpm`, or `tpm` config values get proportionally more traffic — `litellm/router_strategy/simple_shuffle.py:43-67`), **least-busy** (picks the deployment with fewest in-flight requests, tracked via a Redis-backed counter incremented pre-call and decremented on success/failure — `litellm/router_strategy/least_busy.py:114-225`), **latency-based-routing** (selects the deployment with the lowest average or percentile time-to-first-token over a configurable window — `litellm/router_strategy/lowest_latency.py:53-80`), **cost-based-routing** (tracks token usage per deployment in minute buckets via `cost_map:{model_group}` cache keys to pick the cheapest — `litellm/router_strategy/lowest_cost.py:15-80`), **usage-based-routing-v2** (T/RPM-aware), and **LAR-1** (semantic routing by agent confidence thresholds — `litellm/router_strategy/lar1_routing.py:1-60`). Additional strategies include complexity-based (`complexity_router/`), quality-based (`quality_router/`), and adaptive (`adaptive_router/`) routers.

**Fallbacks and retries.** The key entry point `Router.async_function_with_fallbacks()` (`router.py:7865`) wraps every provider call. On failure it checks two levels: order-based fallback (tries higher `target_order` deployments within the same model group) then configured external fallbacks (explicit `fallbacks=[...]` or per-error-type `context_window_fallbacks`, `content_policy_fallbacks`). A `RetryPolicy` per deployment controls retry counts per error category. Failed deployments enter cooldown (configurable `allowed_fails` and `cooldown_time`). Mid-stream fallbacks are supported via `MidStreamFallbackError` + `FallbackAwareStreamWrapper` (`router.py:828-867,3061-3167`).

**Model aliases.** `model_group_alias` (`router.py:1119`) maps alias names to deployment groups, and `Router.get_model_from_alias()` resolves them at routing time (`router.py:5382`).

**Health checks.** `async_get_available_deployment` (`router.py:13711-13831`) runs pre-routing hooks and checks deployment health via configurable staleness thresholds before selecting a target. The health-check subsystem (`litellm/proxy/health_check.py`) pings endpoints periodically.


Citations: [litellm/router_strategy/simple_shuffle.py:43-67](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/simple_shuffle.py#L43-L67) · [litellm/router_strategy/least_busy.py:114-225](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/least_busy.py#L114-L225) · [litellm/router_strategy/lowest_latency.py:53-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/lowest_latency.py#L53-L80) · [litellm/router_strategy/lowest_cost.py:15-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/lowest_cost.py#L15-L80) · [litellm/router.py:7865-7870](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L7865-L7870) · [litellm/router.py:13711-13831](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L13711-L13831)

### QuantumNous/new-api (answered)

Requests are routed through a **middleware pipeline** that selects a channel (provider endpoint) for each request. The `Distribute()` middleware in `middleware/distributor.go` is the core dispatcher: it extracts the model name from the request body, applies **token model limits** (optional per-key model whitelist), resolves the user's group, and calls `SelectChannelForRequest` in `service/channel_select.go`.

**Selection priority (service/channel_select.go:283-376):** (1) A **pinned channel** (set by task plugins, origin-task affinity, or admin override) wins unconditionally. (2) On the first attempt (retry==0), **session affinity** is checked: if the same group+model was routed to a channel recently, that channel is reused (sticky sessions). (3) Otherwise `CacheGetRandomSatisfiedChannel` picks a random eligible channel by group+model+priority.

**Load balancing** is weighted: each `Channel` has a `Weight` field (`model/channel.go:31`), and multi-key channels support `MultiKeyModeRandom` (random selection among enabled keys) or `MultiKeyModePolling` (round-robin via an atomic polling index) (`model/channel.go:249-289`). Channels also have a `Priority` field (`model/channel.go:45`) for tiered fallback — when a priority level is exhausted, the system moves to the next priority group.

**Retry and fallback** work together: `controller/relay.go:158` loops while `retryParam.GetRetry() ≤ common.RetryTimes`. Each iteration calls `SelectChannelForRequest` again, which may select a different channel. The `DecideRelayRetry` function in `service/relay_error.go:21` decides whether to retry based on the error: channel-level errors trigger retries, 4xx status codes by default do not, and operation_setting can configure retryable status codes. Cross-group retry ("auto" group mode) cascades through configured groups: when one group has no eligible channels, the next group is tried (`service/channel_select.go:121-167`).

**Model aliases** work via `model_mapping` field on each channel (`model/channel.go:42`), parsed by `ModelMappedHelper` in `relay/helper/model_mapped.go:14`. It supports chain redirection with cycle detection. The `ResolveTaskModelAlias` in `model/task_model_alias.go:38` resolves task-plugin aliases through ASCII-folded name lookups. **Health checks** run periodically via the channel test endpoint (`controller/channel-test.go`) and Uptime Kuma integration (`controller/uptime_kuma.go:18`). Channels with `auto_ban` enabled (`model/channel.go:46`) are automatically disabled when their keys are all disabled.

> **Editor's note.** Correction: `controller/uptime_kuma.go` only reads an external Uptime Kuma status page for the console; it does not health-check channels. Channel health comes from scheduled channel tests (`CHANNEL_TEST_FREQUENCY` / monitor settings) plus auto-ban on errors.

Citations: [middleware/distributor.go:34-130](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/distributor.go#L34-L130) · [service/channel_select.go:80-204](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L80-L204) · [service/channel_select.go:283-377](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/channel_select.go#L283-L377) · [model/channel.go:23-60](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L23-L60) · [model/channel.go:206-290](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/channel.go#L206-L290) · [controller/relay.go:148-180](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L148-L180)

### songquanpeng/one-api (answered)

**Load balancing and priority.** Each channel (upstream provider endpoint) is assigned a priority and weight. The routing system (`middleware/distributor.go`) calls `model.CacheGetRandomSatisfiedChannel` which queries an in-memory index `group2model2channels` built on startup and refreshed every `SYNC_FREQUENCY` seconds (`model/cache.go`). Channels are sorted by priority within each (group, model) bucket; the system picks randomly from the top-priority tier. If `ignoreFirstPriority` is set (during retries), it picks from lower-priority tiers instead. This is a weighted-random-within-priority-tier scheme.

**Model aliases.** Each channel carries an optional `ModelMapping` JSON field that maps user-visible model names to provider-specific model names (`model/channel.go:114-125`). The mapping is applied right before outbound request conversion (`controller/text.go:37-38`, `controller/image.go:117-118`).

**Retries and fallbacks.** `controller/relay.go:45-101` implements retry logic: after a request fails, it calls `CacheGetRandomSatisfiedChannel` again (with `ignoreFirstPriority=true`), skipping the last-failed channel ID. The number of retries is `config.RetryTimes` (default 0). Channels that return 401/403 or certain `insufficient_quota`/`invalid_api_key` errors are automatically disabled by the monitor (`monitor/manage.go:11-44`).

**Health checks.** The `monitor/metric.go` module tracks a sliding window of success/failure per channel. If a channel's success rate for the last `METRIC_QUEUE_SIZE` calls falls below `METRIC_SUCCESS_RATE_THRESHOLD` (default 80%), the channel is auto-disabled (`monitor/channel.go:47-61`). Additionally, `CHANNEL_TEST_FREQUENCY` can be set to automatically test channels periodically (`main.go:79-84`). There is no explicit latency-based or cost-based routing — routing is purely random-within-priority-tier.

> **Editor's note.** Correction: selection inside the top priority tier is uniform (`rand.Intn`), not weighted; the channel `Weight` field exists but is never used. The first retry stays in the top tier and only later retries move to lower-priority channels.

Citations: [middleware/distributor.go:20-61](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/middleware/distributor.go#L20-L61) · [model/cache.go:227-255](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/cache.go#L227-L255) · [controller/relay.go:45-101](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/controller/relay.go#L45-L101) · [monitor/metric.go:11-39](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/metric.go#L11-L39) · [model/channel.go:100-125](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/channel.go#L100-L125)

### Portkey-AI/gateway (answered)

Requests are routed through the `tryTargetsRecursively` function in `src/handlers/handlerUtils.ts:476` which evaluates a JSON config (passed via `x-portkey-config` header or the request body). This config supports five strategy modes: **Single** — pick the first target; **Fallback** (`strategy.mode: "fallback"`) — iterate targets sequentially, stopping on the first success (or matching `onStatusCodes`); **Loadbalance** (`mode: "loadbalance"`) — weighted random selection via `selectProviderByWeight` at `handlerUtils.ts:204`, where each target has a numeric `weight` property (defaulting to 1); **Conditional** (`mode: "conditional"`) — a `ConditionalRouter` class at `src/services/conditionalRouter.ts:32` matches conditions against request metadata, URL path, and params using operators like `$eq`, `$regex`, `$in`, `$and`, `$or`, routing to a named target; **Circuit Breaker** — targets with `isOpen` flag are filtered out before selection (`handlerUtils.ts:648-658`).

Retries are handled by `retryRequest` in `src/handlers/retryHandler.ts:65`, which uses the `async-retry` library to retry on configurable `onStatusCodes` (default: 429, 500, 502, 503, 504 per `src/globals.ts:38`). It respects provider `retry-after` headers when `followProviderRetry` is set, with a 60-second max retry budget (`MAX_RETRY_LIMIT_MS`, `globals.ts:5`). Request timeout is configurable per-target via `requestTimeout`.

The gateway does not implement latency-based or cost-based dynamic routing — weights are static. Model aliasing is not built into the router; each provider option specifies its model directly in `overrideParams`. Health checks are passive via the circuit-breaker mechanism (`currentTarget.targets.filter(t => !t.isOpen)` at `handlerUtils.ts:648-653`), which tracks per-target open/close state.

> **Editor's note.** Correction: the circuit-breaker path is a hook point, not a working feature in this repository. `tryTargetsRecursively` filters targets on `isOpen` and calls `c.get('handleCircuitBreakerResponse')`, but nothing in the open-source code sets either, so no target is ever marked open unless an embedding host supplies that callback.

Citations: [src/handlers/handlerUtils.ts:476-834](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L476-L834) · [src/handlers/handlerUtils.ts:204-231](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L204-L231) · [src/services/conditionalRouter.ts:32-156](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/services/conditionalRouter.ts#L32-L156) · [src/handlers/retryHandler.ts:65-220](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/retryHandler.ts#L65-L220) · [src/globals.ts:5-39](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/globals.ts#L5-L39)

### higress-group/higress (answered)

**Routing is primarily handled by the `ai-proxy` plugin's provider config system and the `ai-load-balancer`/`ai-endpoint-picker` plugins.**

**Provider selection** is defined per-request by the `ai-proxy` config's `activeProviderId` field in the `PluginConfig`, which selects one of the configured `providers[]` entries (each with `id`, `type`, `apiTokens`, `modelMapping`, etc.) — `config/config.go:27-31`. The `model-router` plugin further intercepts the body and rewrites the `model` field; it parses `provider/model` (e.g. "openai/gpt-4") and splits it into a provider header and model, enabling provider-routing — `model-router/main.go:256-275`. It also supports **auto-routing** with regex rules matched against the last user message: when the model is `higress/auto`, rules match on message content to select the destination model — `model-router/main.go:194-248`.

**Load balancing** is provided by `ai-load-balancer`, supporting cluster-level (cluster_metrics, cluster_hash) and endpoint-level (endpoint_metrics, global_least_request, prefix_cache) policies — `ai-load-balancer/main.go:48-86`. The `ai-endpoint-picker` plugin does fine-grained endpoint selection based on **signals** including queue depth, KV cache, prefix cache, LoRA affinity, inflight requests, and failure counts, all scored by a weighted pipeline — `ai-endpoint-picker/scheduling/types.go:6-12`.

**Retries and failover** are configured per-provider. The retry mechanism (`retry.go`) re-sends failed requests using different API tokens with configurable `maxRetries` and status code matching — `ai-proxy/provider/retry.go:21-29`. The failover mechanism (`failover.go`) tracks per-token failure counts and automatically removes failing tokens, then performs periodic health checks to restore them — `ai-proxy/provider/failover.go:370-403`. Token selection uses random choice or FNV-1a consistent hashing for stateful APIs — `ai-proxy/provider/provider.go:834-919`.

**Model aliases** are done via `modelMapping` with exact match, wildcard (`*`), prefix (`prefix*`), and regex (`~pattern`) rules — `ai-proxy/provider/provider.go:1053-1100`.

**Health checks** for failover send real chat completions to unavailable tokens and count successes before restoring — `ai-proxy/provider/failover.go:192-218`.

> **Editor's note.** Correction: `activeProviderId` is applied when the plugin config (global or per route/domain rule) is parsed, not per request; each rule has exactly one active provider. Choosing a provider per request means routing to different Envoy routes, typically by having `model-router` copy the model or provider prefix into a header that the route matches on.

Citations: [plugins/wasm-go/extensions/ai-proxy/config/config.go:27-31](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/config/config.go#L27-L31) · [plugins/wasm-go/extensions/model-router/main.go:194-248](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/model-router/main.go#L194-L248) · [plugins/wasm-go/extensions/ai-load-balancer/main.go:48-86](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-load-balancer/main.go#L48-L86) · [plugins/wasm-go/extensions/ai-endpoint-picker/scheduling/types.go:6-12](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-endpoint-picker/scheduling/types.go#L6-L12) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:834-919](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L834-L919) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:370-403](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L370-L403)

### maximhq/bifrost (answered)

Bifrost uses a three-layer routing pipeline. **Layer 1 — CEL rule engine:** The `routing` plugin (`plugins/routing/main.go`) evaluates Google CEL (Common Expression Language) rules with scope precedence (VirtualKey > User > Team > Customer > Global). Each `TableRoutingRule` has a `CelExpression` (e.g. `headers["x-region"] == "eu"`) that matches on model, provider, request_type, headers, query params, complexity_tier, budget_used, tokens_used, and request count. Rules can be terminal or chainable (recursive re-evaluation after a matched rule rewrites provider/model). **Layer 2 — Weighted provider load balancing:** After rules settle, `GovernancePlugin.LoadBalanceProvider` (`plugins/governance/main.go:430`) selects among eligible weighted provider configs using weighted random selection. The eligible set is filtered by model allowance, budget limits, rate limits, and key grants. Fallbacks are automatically derived from remaining weighted providers sorted by weight. **Layer 3 — Ordered fallback chain in core:** `Bifrost.handleRequest` (`core/bifrost.go:5689`) iterates the fallback list sequentially. Each fallback calls `tryRequest` on the next provider/model, with full tracing spans per attempt. TTFT (time-to-first-token) deadlines can cut short slow streaming attempts via `AttemptAbort` (`core/providers/utils/stream.go:29`). Session affinity (`core/sessionaffinity.go`) remembers which provider/key served a session, routing subsequent requests to the same provider. Health checks are implicit — failures trigger fallback traversal; a provider that fails in `tryRequest` is skipped for subsequent fallbacks. The model catalog supports aliases via `keyconfig.Store.ResolveAlias`, allowing one model name to map to different wire names per key or provider.


Citations: [plugins/routing/main.go:1-8](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/routing/main.go#L1-L8) · [plugins/routing/rules/engine.go:125-145](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/routing/rules/engine.go#L125-L145) · [plugins/governance/main.go:423-440](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/governance/main.go#L423-L440) · [core/bifrost.go:5689-5720](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/core/bifrost.go#L5689-L5720) · [core/bifrost.go:5787-5810](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/core/bifrost.go#L5787-L5810) · [core/sessionaffinity.go:45-60](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/core/sessionaffinity.go#L45-L60)

### katanemo/plano (answered)

**Routing in Plano is a two-stage process: a quality-based router picks a candidate model, then a session-cache-aware gate decides whether to honor it or stick to the warm anchor.**

The quality router lives in `crates/brightstaff/src/handlers/llm/model_selection.rs`, which calls `orchestrator_service.determine_route()` (`crates/brightstaff/src/router/orchestrator.rs`) and returns a ranked list of models for fallback. The router converts all request shapes to a common `ChatCompletionsRequest` via `ProviderRequestType::try_from()`, then invokes a downstream orchestration model (an LLM used for routing decisions). Route preferences are configured as `TopLevelRoutingPreference` with optional `SelectionPolicy` (`prefer: cheapest | fastest | none`) and can be backed by live cost/latency metrics from `ModelMetricsService` (`crates/brightstaff/src/router/model_metrics.rs`), which fetches pricing catalogs from DigitalOcean or models.dev.

Fallbacks and retries use the ranked `models: Vec<String>` returned by `router_chat_get_upstream_model`. If the primary model returns 429/5xx, the ranked list provides ordered alternatives. There is no explicit load-balancing across providers — the router always picks the single best model for the request.

**Session-cache-aware switch gating** (`crates/brightstaff/src/handlers/llm/session_router.rs`): The `route()` function checks whether the session's provider cache is still warm (comparing `last_used` against the provider's cache TTL from `provider_cache_capability()` in `crates/hermesllm/src/providers/id.rs`). When warm, it either sticks to the anchor model or allows a switch within a configurable cost envelope (`RoutingBudget` — cumulative switch spend capped at `max_switch_spend_pct`% of the never-switch baseline, priced from the model rates feed). Switch cost is computed in USD: `switch_cost_in_usd()` in `orchestrator.rs` compares the candidate's uncached read cost against the anchor's cached read cost. A switch that is outright cheaper (negative cost) is always allowed. 

**Model aliases** come from the `model_aliases` config map (type `HashMap<String, ModelAlias>`) in `common/configuration.rs:48`. **Health checks** are not implemented as active pings — the system relies on upstream HTTP error codes.

> **Editor's note.** Correction: the main proxy path (llm_chat) sends only the single decided model and does not walk the ranked list on 429/5xx. The ranked list is returned by the /routing/* decision endpoints so a client can fall back itself; Envoy only retries the same provider cluster when llm_gateway_listener.max_retries is set.

Citations: [crates/brightstaff/src/handlers/llm/model_selection.rs:53-174](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/model_selection.rs#L53-L174) · [crates/brightstaff/src/handlers/llm/session_router.rs:333-539](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/session_router.rs#L333-L539) · [crates/brightstaff/src/router/orchestrator.rs:43-57](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/router/orchestrator.rs#L43-L57) · [crates/common/src/configuration.rs:170-210](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L170-L210)

### tbphp/gpt-load (answered)

**Routing** in GPT-Load is a multi-stage pipeline of dialect inspection, affinity resolution, and weighted-fair credential selection. The `handler.Handle()` method in `internal/gateway/handler.go:419` is the main orchestrator. It first resolves the client protocol via `dialect.Dialect.InspectRequest()`, obtaining a `RequestMetadata` with `Operation`, `RouteRequirement` and `Model`. A `scheduler.Query` is built and handed to `scheduler.Iterator.Next()` (`internal/scheduler/scheduler.go:298`), which selects one credential from a pool of candidate targets.

**Route modes** — two wire strategies: `RouteNative` (preserve the client protocol upstream) and `RouteConverted` (translate to a provider-neutral format). `RouteRequirement` (`execution.RouteRequirement`, `internal/execution/contracts.go:171`) declares which modes are acceptable. `execution.RouteMode` (`internal/execution/contracts.go:129-133`) records the selected mode.

**Route strategies** — two global policies in `internal/state/runtime_settings.go:51-55`: `RouteStrategyNativeFirst` (try native routes first, fall back to converted) and `RouteStrategyWeightedMix` (treat both modes equally). The scheduler (`scheduler.go:153-154`) sets `routeModeTiers` accordingly — `[[native], [converted]]` for native-first, `[[native, converted]]` for weighted-mix.

**Weighted-fair scheduling** — within a group, credentials are selected via a fairness scheduler (`scheduler/fair.go`). Each credential carries a `WeightManual`, and the `SchedulingLedger` tracks progress watermarks. The lowest-progress eligible credential wins (`scheduler/fair.go:71-79`), ensuring even load across weighted candidates.

**Fallbacks and retries** — after an upstream attempt, `health.JudgeExecution()` (`internal/health/execution_judge.go:27`) decides the retry directive: `RetryNone`, `RetryRefreshCredential`, or `RetryNextCandidate`. Retries increment an attempt counter and call `Next()` again, skipping already-tried credentials (`scheduler.go:330`). Cooldowns and blacklisting come from `health.Decision.Effect` (`internal/health/decision.go:69`).

**Health checks** — `health.StatsStore` (`internal/health/stats.go:1`) tracks per-credential success/failure buckets in 1-minute windows (5 buckets). Problems trigger credential cooldown (`handler.go:341-361`), model-level cooldown, or blacklisting after exceeding `BlacklistThreshold` consecutive failures.

**Affinity** — `affinity.go` (`internal/gateway/affinity.go:19`) derives a prompt-affinity key from the first user message. If a prior request used the same prefix, the same credential is preferred. Prompt-cache-key affinity is also supported. The affinity cache (`affinity.Cache`) sits in `handler.affinityCache`.

**Model aliases** — `state.ModelConfig` includes an `Alias` field (`internal/state/snapshot.go:75-78`). `externalModelName()` returns the alias if set, otherwise the model ID. The scheduler resolves the external model name against upstream model IDs via `SelectModel()`.


Citations: [internal/gateway/handler.go:419-425](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L419-L425) · [internal/scheduler/scheduler.go:298-338](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/scheduler/scheduler.go#L298-L338) · [internal/state/runtime_settings.go:51-56](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/runtime_settings.go#L51-L56) · [internal/health/execution_judge.go:27-80](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/health/execution_judge.go#L27-L80)

### theagentrouter/agent-router (answered)

Routing is defined declaratively via the `AIGatewayRoute` Kubernetes CRD. Each `AIGatewayRoute` contains `Rules`, and each rule has `Matches` (header-based conditions) and `BackendRefs` (references to `AIServiceBackend` or `InferencePool` resources). The `x-ai-eg-model` header is injected by the ext_proc filter after parsing the request body (e.g., the model field from a chat completion request), making model-based routing possible via header matching in Envoy Gateway's HTTPRoute system (`api/v1beta1/ai_gateway_route.go:57-98`). Multiple backends per rule with `Weight` fields (default 1) support weighted traffic distribution, and `Priority` fields enable priority-based load balancing in Envoy (`api/v1beta1/ai_gateway_route.go:387-407`). Fallback and retries are not implemented in this codebase directly — they are delegated to Envoy Gateway's `BackendTrafficPolicy` resource, which provides retry and failover at the data plane level (`api/v1beta1/ai_gateway_route.go:226-241`). There is no latency- or cost-based routing in this codebase; routing is purely rule/header-based via Envoy's existing routing engine. Model aliases are implemented via `ModelNameOverride` on each backend ref, which the translator applies by rewriting the model field in the request body (`internal/translator/openai_openai.go:69-95`). Health checks are provided via the `aigw healthcheck` CLI command (`cmd/aigw/main.go:39-40,60`). The actual routing decision is made by Envoy's external processor filter (ext_proc), whose Go sidecar receives the request headers/body and can return CONTINUE/CONTINUE_AND_REPLACE to influence where Envoy sends the request (`internal/extproc/processor.go:21-40`). InferencePool support provides endpoint picker integration for dynamic backend selection (`internal/controller/inference_pool.go:28-41`).


Citations: [api/v1beta1/ai_gateway_route.go:57-98](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/api/v1beta1/ai_gateway_route.go#L57-L98) · [api/v1beta1/ai_gateway_route.go:211-260](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/api/v1beta1/ai_gateway_route.go#L211-L260) · [api/v1beta1/ai_gateway_route.go:387-407](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/api/v1beta1/ai_gateway_route.go#L387-L407) · [internal/translator/openai_openai.go:69-95](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/translator/openai_openai.go#L69-L95) · [internal/extproc/processor.go:21-40](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/extproc/processor.go#L21-L40) · [internal/controller/inference_pool.go:28-41](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/controller/inference_pool.go#L28-L41)
