theagentrouter/agent-router
Kubernetes control plane plus Envoy ext_proc sidecar that routes, translates, authenticates and meters LLM and MCP traffic.
Overview
Agent Router is the project formerly called Envoy AI Gateway (envoyproxy/ai-gateway). The README says the code, maintainers, CRDs (AIGatewayRoute, AIServiceBackend, BackendSecurityPolicy, API group aigateway.envoyproxy.io), the aigw CLI, the images and the Go module path (github.com/envoyproxy/ai-gateway) are all unchanged. The repository and website moved, and it is now an Agentic AI Foundation project. The go.mod and the copyright headers at this commit still use the old name.
It is not a standalone proxy. It extends Envoy Gateway. A Kubernetes controller turns AI-specific CRDs into Gateway API HTTPRoutes plus Envoy Gateway resources. An Envoy Gateway extension server patches the generated xDS. A Go external processor (ext_proc) runs as a sidecar in every Envoy pod and does the AI-specific work on each request: it reads the model from the body, translates between API schemas, signs or authenticates upstream calls, counts tokens, and emits cost metadata. Envoy itself still does the routing, load balancing, retries, priority failover and rate limiting.
That split is the main design decision. Agent Router adds little logic of its own on the hot path. The cost is a gRPC round trip over a Unix socket to the sidecar for each filter phase, in exchange for Envoy’s mature traffic handling. The audience is platform teams who already run Kubernetes and Envoy Gateway. For local use, aigw run starts the same stack on a laptop.
Architecture
flowchart LR
CRD["CRDs: AIGatewayRoute, AIServiceBackend, BackendSecurityPolicy, QuotaPolicy, MCPRoute"] --> CTRL["Controller (controller-runtime)"]
CTRL --> EG["Envoy Gateway: HTTPRoute, policies"]
CTRL --> BUNDLE["Filter config bundle (Secrets)"]
EG --> EXT["Extension server: xDS patches"]
EXT --> ENVOY["Envoy proxy"]
CL["Client"] --> ENVOY
ENVOY -->|"router-level ext_proc"| EP["ext_proc sidecar"]
ENVOY -->|"upstream-level ext_proc per try"| EP
BUNDLE --> EP
ENVOY --> RL["Envoy rate limit service"]
ENVOY --> UP["Provider / InferencePool"]
| Component | Path | Role |
|---|---|---|
| CRD types | api/v1beta1/, api/v1alpha1/ |
AIGatewayRoute, AIServiceBackend, BackendSecurityPolicy, GatewayConfig, MCPRoute; QuotaPolicy in v1alpha1 |
| Controller | internal/controller/, cmd/controller/ |
Reconcilers, sidecar-injecting webhook, filter-config bundle, token rotators (AWS OIDC, Azure, GCP) |
| Extension server | internal/extensionserver/ |
Envoy Gateway hooks: cluster/route/listener patches, InferencePool EPP clusters, quota rate limits |
| ext_proc server | internal/extproc/, cmd/extproc/ |
gRPC ExternalProcessor; router and upstream processors per endpoint |
| Endpoint specs and translators | internal/endpointspec/, internal/translator/ |
Per-endpoint parsing and schema-to-schema translation |
| Backend auth | internal/backendauth/ |
API key, Azure, Anthropic, AWS SigV4, GCP handlers and per-request credential override |
| Cost and quotas | internal/llmcostcel/, internal/ratelimit/ |
CEL cost expressions and QuotaPolicy-to-Envoy rate limit translation |
| MCP proxy | internal/mcpproxy/ |
MCP gateway (sessions, authorization, SSE) |
| CLI | cmd/aigw/ |
aigw run, healthcheck, download-envoy |
How a request flows
For POST /v1/chat/completions through a Gateway with an AIGatewayRoute:
- Pick a processor. Envoy opens an ext_proc stream to the sidecar.
Server.Processreads:pathfrom the first headers message, picks the registered factory (chat completions, messages, embeddings, responses, audio, images, rerank and so on), and decides from the presence of attributes whether this is the router-level or the upstream-level filter (server.go, main.go). - Router phase.
routerProcessor.ProcessRequestBodyparses the body, sets thex-ai-eg-modelheader (overwriting any value the client sent), records the original path, starts a trace span, and returnsClearRouteCache: true(processor_impl.go). Envoy then re-matches the route. The controller generatedHTTPRouterules that match on that header. - Envoy chooses a backend. Normal Envoy load balancing applies the rule’s
backendRefsweights and priorities. Fallback across backends comes from priority levels plus Envoy Gateway’sBackendTrafficPolicyretries (ai_gateway_route.go, L380-L410). - Upstream phase, once per try. An upstream-level ext_proc stream starts for the chosen backend.
SetBackendpicks the translator for that backend’sAPISchemaand model override.ProcessRequestHeaderstranslates the stored body, applies header and body mutations, and runs the backend auth handler. It returnsCONTINUE_AND_REPLACE, so Envoy does not send the body to the sidecar a second time (processor_impl.go). On a retry to another backend this repeats against the original body, so each provider gets its own translation. - Response. The upstream processor converts headers and body (or SSE chunks) back to the client’s schema, extracts token usage, records TTFT and inter-token latency, and calls
buildDynamicMetadata. That writes CEL-computed costs under the filter’s metadata namespace for rate limiting and access logs (processor_impl.go).
Key components
Two-level ext_proc
The router-level filter runs once per request. The upstream-level filter runs once per upstream attempt. The server keeps router processors in a map keyed by internal request ID, behind an RWMutex, so the upstream filter can find the parsed body (server.go). This is what lets Envoy-native retries fail over from OpenAI to, for example, Bedrock with a correctly translated body.
Translators
Inputs are not only OpenAI. The sidecar registers OpenAI-style endpoints, Anthropic /v1/messages and count_tokens, Cohere /v2/rerank and TypeSafe /v1/systemone. Output schemas are OpenAI, Cohere, AWSBedrock, AzureOpenAI, GCPVertexAI, GCPAnthropic, Anthropic, AWSAnthropic, AWSOpenAI and TypeSafe (filterconfig.go). Each endpoint spec’s GetTranslator decides which pairs are allowed. For example, Messages input can go to Anthropic, GCP/AWS Anthropic, Bedrock, or OpenAI Chat Completions (endpointspec.go).
Backend auth
BackendSecurityPolicy selects a handler: Bearer API key, Azure key, Anthropic key, AWS SigV4, Azure token or GCP token. Each can be wrapped in a CredentialOverride that takes the credential from trusted request headers or dynamic metadata, with a fallback to the static secret (auth.go). Rotator controllers refresh short-lived cloud tokens into Secrets.
Cost and quotas
LLMRequestCost entries are token counters or CEL expressions over model, backend, route_name and the token counts (cel.go). QuotaPolicy compiles to Envoy rate limit configs in one domain, ai-gateway-quota, keyed by backend and model descriptors (translator.go). Enforcement happens in Envoy’s rate limit service, not in Go.
Sidecar injection
The controller’s admission webhook injects the ext_proc container into Envoy Gateway’s proxy pods. With sidecar mode on, it goes in as a native sidecar (an init container with restartPolicy: Always) so that it shuts down after Envoy. The filter-config bundle is mounted read-only (gateway_mutator.go, extproc_builder.go).
Extending it
- New backends: add an
APISchemaName, translators, and aGetTranslatorbranch in each endpoint spec that should support it. - Self-hosted inference: point a rule at a Gateway API Inference Extension
InferencePool. Its endpoint picker chooses pods, and the extension server adds the EPP cluster (post_translate_modify.go). - Policies: anything Envoy Gateway offers (JWT/OIDC auth on the Gateway, retries, timeouts, external auth) composes with the AI CRDs.
- MCP:
MCPRouteputs MCP servers behind the same gateway, with its own authorization and session handling.
Running it
- Kubernetes: install Envoy Gateway with the provided values, then the
ai-gateway-crds-helmandai-gateway-helmcharts frommanifests/charts, in theenvoy-gateway-systemandenvoy-ai-gateway-systemnamespaces. - Standalone:
OPENAI_API_KEY=… aigw run(or with a config YAML). It downloads Envoy, runs Envoy Gateway’s standalone mode with the ext_proc in process, and serves/metricsand/healthon admin port 1064 (main.go). - Observability: OpenTelemetry metrics (Prometheus always on) and traces with GenAI and OpenInference conventions, configured by
OTEL_*environment variables.
Strengths and caveats
- Strength: Envoy does the traffic work. Load balancing, priority failover, retries, rate limiting, TLS and access logs come from Envoy and Envoy Gateway rather than custom Go.
- Strength: per-try translation. Cross-provider failover gets a correct, freshly translated and authenticated request on every attempt.
- Strength: declarative and Kubernetes-native. CRDs, ReferenceGrants for cross-namespace secrets, and cloud-token rotators fit GitOps workflows.
- Caveat: heavy prerequisites. You need Kubernetes and Envoy Gateway, or the
aigwstandalone wrapper around them. There is no single binary to drop in. - Caveat: no tenant or key product. There are no virtual keys, users, budgets or admin UI. Identity comes from Gateway policies or request headers, and spend control is rate limiting on CEL-computed cost.
- Caveat: no caching or content guardrails. Redaction exists only for debug logs. Prompt caching is only reflected in token accounting.
- Caveat: an extra hop per request. Every request makes gRPC round trips to the sidecar over a Unix socket (router phase, then upstream phase), and the generated filter buffers the request body before it is processed (post_translate_modify.go).
Sources: code at daa9f89, verified Q&A.
How it answers the LLM gateways questions
Each answer was drafted by a code-reading agent at commit daa9f89. Its citations were checked mechanically. Compare with the other llm gateways →
How are requests routed across providers and models?
answeredRouting is defined declaratively via the AIGatewayRoute Kubernetes CRD. Each AIGatewayRoute contains Rules, and each rule has Matches (header-based conditions) and BackendRefs (references to AIServiceBackend or InferencePool resources). The x-ai-eg-model header is injected by the ext_proc filter after parsing the request body (e.g., the model field from a chat completion request), making model-based routing possible via header matching in Envoy Gateway's HTTPRoute system (api/v1beta1/ai_gateway_route.go:57-98). Multiple backends per rule with Weight fields (default 1) support weighted traffic distribution, and Priority fields enable priority-based load balancing in Envoy (api/v1beta1/ai_gateway_route.go:387-407). Fallback and retries are not implemented in this codebase directly — they are delegated to Envoy Gateway's BackendTrafficPolicy resource, which provides retry and failover at the data plane level (api/v1beta1/ai_gateway_route.go:226-241). There is no latency- or cost-based routing in this codebase; routing is purely rule/header-based via Envoy's existing routing engine. Model aliases are implemented via ModelNameOverride on each backend ref, which the translator applies by rewriting the model field in the request body (internal/translator/openai_openai.go:69-95). Health checks are provided via the aigw healthcheck CLI command (cmd/aigw/main.go:39-40,60). The actual routing decision is made by Envoy's external processor filter (ext_proc), whose Go sidecar receives the request headers/body and can return CONTINUE/CONTINUE_AND_REPLACE to influence where Envoy sends the request (internal/extproc/processor.go:21-40). InferencePool support provides endpoint picker integration for dynamic backend selection (internal/controller/inference_pool.go:28-41).
How are different provider APIs unified?
answeredProvider APIs are unified through a multi-layer translation architecture. The Translator interface (internal/translator/translator.go:48-83) defines RequestBody, ResponseBody, ResponseHeaders, and ResponseError methods, parameterized with request/response types (ReqT/SpanT). Ten endpoint-specific translator types are defined, covering chat completions, completions, embeddings, image generation, responses, speech, transcription, translation, rerank, tokenize, and count-tokens (internal/translator/translator.go:132-160). Nine API schemas are supported: OpenAI, Cohere, AWSBedrock, AzureOpenAI, GCPVertexAI, GCPAnthropic, Anthropic, AWSAnthropic, AWSOpenAI, and TypeSafe (internal/apischema/typesafe/systemone.go). The EndpointSpec.GetTranslator() method selects the correct translator based on the backend's API schema — for example, ChatCompletionsEndpointSpec.GetTranslator() has seven schema branches, translating OpenAI input to each backend's native format (internal/endpointspec/endpointspec.go:202-221). OpenAI-compatible schema is the canonical input format: all requests enter as OpenAI-shaped JSON and are translated to the backend's format. Streaming is handled per-endpoint: the upstream processor sets streaming mode in response headers when stream=true and processes streamed SSE chunks (internal/extproc/processor_impl.go:517-520,538-666). Tool calls and multimodal content are supported through the full OpenAI chat completions schema including tool_calls, ToolCalls, ContentPart, image URL, audio data, and file parts — all of which are preserved through translation (internal/endpointspec/endpointspec.go:689-859). Each translator pair (e.g., openai_awsbedrock.go, anthropic_gcpanthropic.go, openai_azureopenai.go) handles the specific request/response format conversion between OpenAI input and the target backend API.
How are API keys, users and tenants managed?
answeredAPI keys, users, and tenants are managed via the Kubernetes-native BackendSecurityPolicy CRD. Each AIServiceBackend has an associated BackendSecurityPolicy that defines the authentication type and stores credentials. Six auth types are supported: APIKey (Bearer token), AzureAPIKey, AnthropicAPIKey, AWSCredentials (SigV4 signing with region), AzureCredentials (access token), and GCPCredentials (access token with region/project) (internal/controller/gateway.go:913-1004). Static credentials are stored in Kubernetes Secrets referenced by the BackendSecurityPolicy's secretRef, resolved via getBSPSecretRefData() (internal/controller/gateway.go:1035-1047). The apiKeyInSecret constant defines the standard key used within secrets (internal/controller/ai_gateway_route.go:48). Cross-namespace credential references are validated through ReferenceGrant resources, enforced by the referenceGrantValidator (internal/controller/gateway.go:490-505). Per-request credential override is supported via CredentialOverride with two modes: FromRequestHeaders (injected by trusted ingress filters using x-aigw-* headers) and FromDynamicMetadata (from Envoy dynamic metadata). Default input header names per auth type include x-aigw-api-key, x-aigw-anthropic-api-key, etc. (internal/controller/gateway.go:827-853). The backendauth package implements handlers per auth type: apiKeyHandler sets the Authorization header, newAWSHandler performs SigV4 signing, newAzureHandler for Azure tokens, etc. (internal/backendauth/auth.go:16-67). Token rotation is handled by dedicated rotator controllers for AWS OIDC, Azure, and GCP tokens (internal/controller/rotators/). There is no built-in user/tenant model or admin UI — user identification is done via request headers (e.g., x-tenant-id) for downstream rate limiting.
How are rate limits, budgets and cost tracking implemented?
answeredRate limits are implemented via the QuotaPolicy Kubernetes CRD, which translates into Envoy rate limit configuration. Each QuotaPolicy specifies PerModelQuotas (per-model rate limits) and ServiceQuota (service-wide catch-all), with nested BucketRules for client-based differentiation using header selectors (internal/ratelimit/translator/translator.go:92-177). Rate limits are expressed as requests per unit (1s/1m/1h/1d) converted to Envoy's RateLimitUnit (internal/ratelimit/translator/translator.go:392-417). All backends share a single rate limit domain (ai-gateway-quota), distinguished by backend_name and model_name_override descriptors (internal/ratelimit/translator/translator.go:21-33). The QuotaPolicyController translates policy definitions into RateLimitConfig objects consumed by Envoy's rate limit service and caches them for incremental updates (internal/controller/quota_policy.go:30-59). Cost tracking is configured through LLMRequestCost entries on AIGatewayRoutes (and GlobalLLMRequestCost on GatewayConfig). Cost types include InputToken, OutputToken, TotalToken, CachedInputToken, CacheCreationInputToken, ReasoningToken, and custom CEL expressions (internal/filterapi/filterconfig.go:112-130). CEL expressions support variables like input_tokens, output_tokens, total_tokens, model, backend, route_name for custom cost formulas (internal/llmcostcel/cel.go:18-48). Costs are computed by the ext_proc response handler and written into Envoy's dynamic metadata under the namespace io.envoy.ai_gateway for consumption by Envoy's rate limit HitsAddend filter (internal/extproc/processor_impl.go:651-661). The TokenUsage struct tracks input, output, total, cached input, cache creation input, and reasoning tokens separately (internal/metrics/metrics.go:143-158). QuotaPolicy cost expressions are injected as LLMRequestCost entries during reconciliation, filtered by backend and model name (internal/controller/gateway.go:1053-1130).
How are caching and guardrails implemented?
answeredThis repository does not implement exact-response caching, semantic caching, PII redaction guardrails, content moderation, or prompt-injection detection. The codebase focuses on routing, translation, and rate limiting — caching and safety guardrails are not present in the data-plane processing path. What does exist: (1) Response redaction for debug logging: the ResponseRedactor interface allows translators to create redacted copies of responses for safe logging, replacing sensitive content (message text, tool call arguments, audio data) with [REDACTED LENGTH=n HASH=xxxx] placeholders while preserving structure for debugging (internal/translator/openai_openai.go:238-266). The RedactString function uses SHA256 hashing for content correlation in logs (internal/redaction/redaction.go:29-47). (2) Request redaction for debug logging: EndpointSpec.RedactSensitiveInfoFromRequest() creates redacted copies of parsed request bodies, stripping user content while keeping structural fields (tool definitions, response formats) (internal/endpointspec/endpointspec.go:223-245). (3) Token-level caching awareness: the TokenUsage struct tracks cachedInputTokens and cacheCreationInputTokens, and these values are extracted from provider responses and propagated into dynamic metadata for cost tracking (internal/metrics/metrics.go:151-152). The ExtractTokenUsageFromExplicitCaching function normalizes Anthropic/AWS Bedrock explicit caching token counts into the unified format (internal/metrics/metrics.go:298-313). (4) Header/body mutation: routes and backends can configure header and body field removal/modification (e.g., stripping sensitive headers before forwarding to upstream) (internal/controller/gateway.go:287-379). These mutations are not automated guardrails — they are declarative per-route configurations.
How is it observed, deployed and scaled?
answeredThe project is observed via OpenTelemetry for both metrics and traces. Metrics are backed by OpenTelemetry with configurable exporters: Prometheus (always enabled), console exporter, and OTLP exporter. Environment variables OTEL_METRICS_EXPORTER, OTEL_EXPORTER_OTLP_ENDPOINT, and OTEL_SDK_DISABLED control configuration (internal/metrics/metrics.go:35-94). The Metrics interface tracks request start/completion timing, token usage, backend selection, time-to-first-token, and inter-token latency for streaming (internal/metrics/metrics.go:97-127). Traces support the same exporters (console, OTLP) with OTEL_TRACES_EXPORTER and auto-propagation via autoprop. Span recorders exist for every endpoint type (chat completions, embeddings, responses, speech, transcription, MCP, etc.) with semantic conventions for both OpenInference and OTEL GenAI (internal/tracing/tracing.go:137-328). Performance architecture: The Go-based ext_proc sidecar handles request/response transformations per-stream via gRPC. Router processors are tracked per-request-ID with a read-write mutex for concurrent access (internal/extproc/server.go:54-64). The data plane itself runs on Envoy Proxy (C++), providing battle-tested HTTP/2, connection pooling, and async I/O. Deployment modes: (1) Standalone CLI: aigw run downloads and configures an Envoy binary with the ext_proc sidecar as a local process (cmd/aigw/main.go:38-40). (2) Kubernetes: runs as a controller-runtime operator that reconciles Gateway resources, injects the ext_proc as a sidecar container (DaemonSet or init container), and manages rolling updates via workload template annotations (internal/controller/gateway.go:1298-1391). (3) Admin server: AdminPort exposes /metrics and /health HTTP endpoints (cmd/aigw/main.go:49). HA: Kubernetes controller pattern provides self-healing via reconciliation; Envoy's data plane handles connection retries and failover. Config bundle checksums ensure data-plane integrity (internal/filterapi/config_bundle.go:21-91).