LLMs Technical Reviews
Home / LLM gateways / higress

higress-group/higress

Envoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types.

GitHub ↗★ 9.5kGoApache-2.0commit bda81f1 · 2026-10-05homepage ↗

Overview

Higress is an Envoy-based API gateway from Alibaba. Its AI features are a set of WebAssembly plugins that run inside Envoy’s request path. It is not a separate LLM proxy service. The control plane (Go, built on Istio) turns Kubernetes Ingress, Gateway API resources and Higress CRDs into Envoy configuration. The data plane is Envoy with Go plugins compiled by TinyGo to WASM, plus some Rust, C++ and native Go (“golang-filter”) extensions.

For LLM traffic the key plugin is ai-proxy. It takes OpenAI-format requests (and Anthropic /v1/messages), rewrites them for one of 38 provider types, swaps in an upstream API token, and converts responses and SSE streams back. Other features are separate plugins that you attach to routes as needed: model-router, ai-load-balancer, ai-cache, ai-quota, ai-token-ratelimit, ai-statistics, ai-security-guard and about a dozen more. MCP servers are hosted through mcp-server plugins and a golang-filter.

Higress suits teams that already run, or want, a Kubernetes ingress and would like LLM routing to live in the same Envoy fleet as their other APIs. For a single developer it is heavy. The all-in-one Docker image helps, but the configuration model is still Envoy routes plus per-route plugin config.

Architecture

flowchart LR
  K8S["Ingress / Gateway API / CRDs"] --> CP["Higress controller (Istio-based)"]
  CP -->|"xDS + WasmPlugin config"| ENV["Envoy data plane"]
  C["Client"] --> ENV
  ENV --> MR["model-router (WASM)"]
  MR --> RT["Envoy route match"]
  RT --> AP["ai-proxy (WASM)"]
  AP --> Q["ai-quota / ai-token-ratelimit"]
  AP --> CA["ai-cache"]
  AP --> UP["LLM provider"]
  AP --> ST["ai-statistics: metrics, ai_log, spans"]
Component Path Role
Controller cmd/higress, pkg/ingress/ Watches Ingress, Gateway API, McpBridge, WasmPlugin; translates to Istio/Envoy config
Registry bridge registry/ Service discovery from Nacos, Consul, Eureka, ZooKeeper (the McpBridge CRD)
ai-proxy plugins/wasm-go/extensions/ai-proxy/ Protocol translation, provider auth, token failover, retries
Providers ai-proxy/provider/*.go One file per provider type implementing the Provider handler interfaces
model-router plugins/wasm-go/extensions/model-router/ Moves model into a header for route matching; regex auto-routing
ai-load-balancer plugins/wasm-go/extensions/ai-load-balancer/ Cluster and endpoint policies (metrics, least-request, prefix cache)
ai-cache plugins/wasm-go/extensions/ai-cache/ Redis exact cache and vector similarity cache
ai-quota, ai-token-ratelimit plugins/wasm-go/extensions/ Per-consumer token quotas and Redis sliding-window token limits
ai-statistics plugins/wasm-go/extensions/ai-statistics/ Token usage, latency metrics, AI access log, tracing attributes
MCP hosting plugins/golang-filter/mcp-server/, plugins/wasm-go/mcp-servers/ MCP server framework and ready-made servers

How a request flows

Take POST /v1/chat/completions on a route that has ai-proxy attached:

  1. Config load. When Envoy pushes plugin config, parseGlobalConfig or parseOverrideRuleConfig reads providers[] and activeProviderId. Complete() builds the provider and starts API-token failover (main.go, config.go). The active provider is fixed per route or domain rule, not chosen per request.
  2. Model routing (optional). If model-router runs first, it copies model into a header or splits provider/model into a provider header. Envoy then matches a route whose ai-proxy config targets that provider. With higress/auto, regex rules over the last user message pick the model (model-router/main.go).
  3. Request headers. onHttpRequestHeader maps the path to an ApiName. If the request is Anthropic-format and the provider lacks native support, it rewrites the path to chat completions and marks the response for conversion. It then picks an API token and calls the provider’s OnRequestHeaders (main.go).
  4. Request body. onHttpRequestBody applies custom settings, adds stream_options for usage stats on OpenAI-protocol providers, and calls the provider’s OnRequestBody, which applies modelMapping and transforms the body (main.go).
  5. Upstream. Envoy sends the request to the provider cluster. Other plugins on the route (quota, rate limit, cache, security guard) run in their configured phase order.
  6. Response. onHttpResponseHeaders, onStreamingResponseBody and onHttpResponseBody let the provider transform the output to OpenAI format, or to Claude format if the request came in as Claude. On failure statuses, failover counts the token and the retry path may resend with another token.

Key components

Provider abstraction

Provider is a small interface. Optional handler interfaces (RequestHeadersHandler, RequestBodyHandler, TransformRequestBodyHandler, StreamingResponseBodyHandler, StreamingEventHandler and others) let each provider override only what it needs (provider.go). providerInitializers registers 38 types, including OpenAI, Azure, Claude, Gemini, Vertex, Bedrock, Qwen, DeepSeek, Ollama, vLLM, Triton and a generic pass-through (provider.go).

Token selection, failover and retry

Each provider config holds a list of apiTokens. selectApiToken picks one at random, or uses consumer affinity (via the x-mse-consumer header) for stateful APIs such as Responses or files (provider.go). With failover on, a token that reaches failureThreshold failures moves to an unavailable list in Envoy shared data. A health checker, elected through a CAS lease across WASM VMs, sends real requests to bring it back (failover.go). retryOnFailure resends the saved body with a different token, by default on any 4xx or 5xx status (retry.go). Failover is between keys of one provider. Fallback across providers comes from Envoy routes and x-higress-fallback-from internal redirects, and initContext takes care to keep the client’s original Authorization across those hops (main.go).

Model mapping

modelMapping supports exact keys, prefix*, ~regex with capture-group replacement, and a * default (provider.go). Prefix and regex keys are checked while ranging over a Go map, so when two of them match the same model, which one wins is not defined. Regex keys are compiled with MustCompile on each lookup.

Quotas, limits and stats

ai-quota reads the consumer from x-mse-consumer, denies with 403 when the Redis counter chat_quota:<consumer> is not positive, and decrements by total tokens once the response ends (ai-quota/main.go, L267-L303). The check is before the request and the charge is after, so one large request can overshoot. Quotas count tokens. No price table exists, so cost in money is left to downstream systems. ai-statistics writes Envoy counters named route.<r>.upstream.<c>.model.<m>.consumer.<u>.metric.<name> plus ai_log fields and tracing attributes (ai-statistics/main.go).

Cache and safety

ai-cache builds its key from the last message’s content by default ([email protected]), or from all user messages. It checks Redis, then falls back to embedding similarity search on a vector store (ai-cache/main.go, core.go). The key does not include the model or system prompt, so scope the plugin per route. ai-security-guard sends request and response text (and images) to Alibaba Cloud’s content-moderation service.

Extending it

  • New provider. Add a file under ai-proxy/provider/ implementing the handler interfaces you need, and register it in providerInitializers.
  • New plugin. Write a Go WASM plugin with wrapper.SetCtx and phase callbacks, build it with the plugins/wasm-go Makefile, and attach it with a WasmPlugin resource or through the console.
  • MCP servers. Register a server in the golang-filter mcp-server registry, or expose REST APIs as MCP tools with the mcp-server WASM plugin.
  • Service discovery. McpBridge brings in Nacos, Consul, Eureka or ZooKeeper services as upstreams. The name refers to multi-registry discovery, not to the Model Context Protocol.

Running it

  • Local. One docker run of the all-in-one image exposes the console on 8001 and the gateway on 8080/8443, with config written to a mounted directory (README.md).
  • Kubernetes. Helm charts in helm/ install the controller, gateway and console. Plugins are delivered as OCI images.
  • Required services. Redis for ai-quota, ai-token-ratelimit and ai-cache. An embedding provider and a vector DB for semantic caching.

Strengths and caveats

  • Strength: production data plane. LLM traffic gets Envoy’s connection handling, TLS, observability and route model, and sits next to ordinary API traffic in one gateway.
  • Strength: modular. Each AI feature is a separate plugin with its own config. You turn on only what you use, per route.
  • Strength: careful key failover. Shared-data token health, a leased health checker and auth-preserving internal redirects are well thought out.
  • Caveat: one provider per route. ai-proxy has no in-plugin multi-provider fallback or cost routing. Cross-provider logic is assembled from Envoy routes, model-router and redirects.
  • Caveat: Alibaba-leaning defaults. The security guard calls Alibaba Cloud moderation, consumer identity uses the x-mse-consumer header, and images default to an Alibaba Cloud registry.
  • Caveat: token-only accounting. No price tables. Quotas charge after the response.
  • Caveat: plugin edge cases. Non-deterministic prefix/regex model mapping and a model-agnostic default cache key need care in config.

Sources: code at bda81f1, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit bda81f1. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Routing is primarily handled by the ai-proxy plugin's provider config system and the ai-load-balancer/ai-endpoint-picker plugins.

Provider selection is defined per-request by the ai-proxy config's activeProviderId field in the PluginConfig, which selects one of the configured providers[] entries (each with id, type, apiTokens, modelMapping, etc.) — config/config.go:27-31. The model-router plugin further intercepts the body and rewrites the model field; it parses provider/model (e.g. "openai/gpt-4") and splits it into a provider header and model, enabling provider-routing — model-router/main.go:256-275. It also supports auto-routing with regex rules matched against the last user message: when the model is higress/auto, rules match on message content to select the destination model — model-router/main.go:194-248.

Load balancing is provided by ai-load-balancer, supporting cluster-level (cluster_metrics, cluster_hash) and endpoint-level (endpoint_metrics, global_least_request, prefix_cache) policies — ai-load-balancer/main.go:48-86. The ai-endpoint-picker plugin does fine-grained endpoint selection based on signals including queue depth, KV cache, prefix cache, LoRA affinity, inflight requests, and failure counts, all scored by a weighted pipeline — ai-endpoint-picker/scheduling/types.go:6-12.

Retries and failover are configured per-provider. The retry mechanism (retry.go) re-sends failed requests using different API tokens with configurable maxRetries and status code matching — ai-proxy/provider/retry.go:21-29. The failover mechanism (failover.go) tracks per-token failure counts and automatically removes failing tokens, then performs periodic health checks to restore them — ai-proxy/provider/failover.go:370-403. Token selection uses random choice or FNV-1a consistent hashing for stateful APIs — ai-proxy/provider/provider.go:834-919.

Model aliases are done via modelMapping with exact match, wildcard (*), prefix (prefix*), and regex (~pattern) rules — ai-proxy/provider/provider.go:1053-1100.

Health checks for failover send real chat completions to unavailable tokens and count successes before restoring — ai-proxy/provider/failover.go:192-218.

Editor's note. Correction: activeProviderId is applied when the plugin config (global or per route/domain rule) is parsed, not per request; each rule has exactly one active provider. Choosing a provider per request means routing to different Envoy routes, typically by having model-router copy the model or provider prefix into a header that the route matches on.

How are different provider APIs unified?

answered

Provider API unification is the central function of the ai-proxy plugin: all requests enter the gateway as OpenAI-format requests and are translated to each provider's native format.

API Name Resolution: Incoming request paths are mapped to ApiName constants (e.g. /v1/chat/completions → ApiNameChatCompletion) via suffix and regex matching — ai-proxy/main.go:56-104. Over 30 API types are defined across OpenAI, Anthropic, Cohere, Gemini, and Qwen formats — ai-proxy/provider/provider.go:37-82.

OpenAI-compatible schema is the internal standard: chatCompletionRequest, chatCompletionResponse, embeddingsRequest, imageGenerationRequest, etc. are defined in model.go with proper JSON field names — ai-proxy/provider/model.go:43-78. The protocol field can be set to "openai" (default) or "original" to bypass conversion — ai-proxy/provider/provider.go:411.

Request/response translation happens per-provider via the Provider interface with TransformRequestHeadersHandler, TransformRequestBodyHandler, and TransformResponseBodyHandler — ai-proxy/provider/provider.go:302-322. Each provider file (48 provider implementations) implements these interfaces. For example, openai.go passes through with path rewriting to api.openai.com, while claude.go translates to Anthropic's Messages API format (api.anthropic.com) — ai-proxy/provider/claude.go:18-20. The claude_to_openai.go converter handles bidirectional protocol translation for providers that don't natively support Claude's protocol — ai-proxy/provider/claude_to_openai.go:13-35.

Auto protocol detection converts Claude-format requests to OpenAI-format when the upstream provider doesn't natively support Claude — ai-proxy/main.go:254-266. The response is converted back to Claude format on streaming/non-streaming response callbacks — ai-proxy/main.go:688-771.

Streaming is handled by StreamingResponseBodyHandler or StreamingEventHandler interfaces, with SSE framing and event extraction — ai-proxy/provider/provider.go:290-296. The ExtractStreamingEvents function parses raw SSE chunks into structured StreamEvent objects, handling partial chunk buffering — ai-proxy/provider/provider.go:1130-1206.

Tool calls and multimodal are fully supported: the chatMessage struct includes tool_calls, ToolCallId, FunctionCall, and multimodal content types (text, image_url, input_audio, file) — ai-proxy/provider/model.go:207-224,348-370.

Provider count: 38 provider types are registered in the providerInitializers map — ai-proxy/provider/provider.go:236-275.

How are API keys, users and tenants managed?

answered

API keys, users, and tenant management uses a consumer-based model with the ai-quota and key-auth plugins, plus per-provider API token configuration.

Virtual keys/consumer identification: The ai-quota plugin reads the x-mse-consumer HTTP header to identify the consumer (user/tenant) and enforces quota against Redis — ai-quota/main.go:161-167. This consumer identity is propagated downstream and reused by ai-statistics for per-consumer metrics — ai-statistics/main.go:726-728. A working ai-quota key-auth example shows how API keys map to consumers: credentials are extracted from request headers and validated — plugins/wasm-go/examples/key-auth/main.go.

Upstream credential storage: Each provider config holds apiTokens []string — one or more API tokens for the upstream LLM service — ai-proxy/provider/provider.go:333. At request time, the SetApiTokenInUse method selects a token (via failover-global random or context-based selection) and stores it in the request context — ai-proxy/provider/provider.go:724-734. Token failover uses a CAS-based shared data store to track available vs. unavailable tokens per provider instance — ai-proxy/provider/failover.go:139-224.

Per-request authentication header handling: The plugin saves the original Authorization header at first hop and restores it on internal redirects (using X-HI-ORIGINAL-AUTH), distinguishing first-hop from re-entry requests via the x-higress-fallback-from header — ai-proxy/main.go:156-197. The upstream token replaces the Authorization header for the LLM call.

Admin UI/API: The ai-quota plugin exposes an admin API at /v1/chat/completions/quota with refresh, delta (increment/decrement), and query operations — ai-quota/main.go:316-331. These are authenticated by comparing the caller's consumer identity against a configured admin_consumer — ai-quota/main.go:336-339. Data is stored in Redis with key prefix chat_quota: — ai-quota/main.go:117-118.

User/team model: The system does not have a built-in user/team/tenant model in the Wasm plugins — consumer identity via the x-mse-consumer header is the primary mechanism, suitable for integration with an upstream identity provider. The Kubernetes CRD-based control plane in api/kubernetes/ manages plugin configuration through CustomResourceDefinitions — api/kubernetes/crd_contract.go:1-27.

How are rate limits, budgets and cost tracking implemented?

answered

Rate limits, budgets, and cost tracking are handled by three separate plugins: ai-quota, ai-token-ratelimit, and ai-statistics.

Per-key/user quotas (ai-quota): The plugin checks a Redis-stored integer quota for each consumer (identified by x-mse-consumer header) before allowing a request. The quota is checked as GET chat_quota:{consumer} and if ≤0, the request is denied with 403 — ai-quota/main.go:194-213. After response completion, the plugin extracts token usage from the response and decrements the quota via DECRBY chat_quota:{consumer} {totalTokens} — ai-quota/main.go:267-303. The GetTokenUsage function parses token counts from both OpenAI and Anthropic response formats.

Token-based rate limiting (ai-token-ratelimit): This plugin enforces sliding-window rate limits using server-side Lua scripts in Redis. It supports global thresholds and per-key limits with configurable windows per rule — ai-token-ratelimit/main.go:47-79. The Lua script MultiKeyRequestPhaseScript atomically checks multiple rate limit counters in one Redis call, returning thresholds, current counts, and TTLs — ai-token-ratelimit/main.go:58-79. The response-phase script increments counters only after a successful LLM response.

Usage statistics and cost tracking (ai-statistics): This plugin records per-request token usage (input, output, total, reasoning tokens, cached tokens, token details) from both streaming and non-streaming responses — ai-statistics/main.go:990-1045. Token usage is extracted via the tokenusage.GetTokenUsage library which handles multiple response formats: OpenAI chat completions, Anthropic Messages, Gemini, Doubao — ai-statistics/main.go:960-962. Metrics are recorded as Envoy distributed counters at the route/cluster/model/consumer granularity — ai-statistics/main.go:481-524. The metric name follows the pattern route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName} — ai-statistics/main.go:481-483.

Built-in metrics recorded include: llm_first_token_duration, llm_service_duration, llm_failure_count, and token counts — ai-statistics/main.go:93-97. These are accumulated into Envoy-native counter metrics and also written to the ARMS (Application Resource Monitoring System) tracing spans with gen_ai.* prefix — ai-statistics/main.go:102-108.

Price tables: There is no built-in price table in the codebase — cost tracking is purely token-based, leaving monetary cost calculation to external systems.

How are caching and guardrails implemented?

answered

Caching and guardrails are provided by the ai-cache and ai-security-guard plugins.

Exact caching (ai-cache): The plugin intercepts LLM chat requests and checks Redis for an exact cache hit using the last user message (or all user messages) as the key — ai-cache/main.go:95-116. On a hit, the cached response is returned directly as an SSE stream or JSON body — ai-cache/core.go:80-84. On a miss, the request is forwarded upstream and the response is cached for future use — ai-cache/core.go:207-219. Cache status (hit/miss/skip) is recorded in the AI log — ai-cache/main.go:142-144.

Semantic caching: When enabled, the ai-cache plugin generates text embeddings (using OpenAI, Cohere, DashScope, HuggingFace, Ollama, etc.) and queries a vector database for semantically similar requests — ai-cache/core.go:88-107. Vector providers implement EmbeddingQuerier or StringQuerier interfaces — ai-cache/core.go:100-107. An embedding-based similarity search is performed when no exact cache hit is found — ai-cache/core.go:121-148. Results are compared against a configurable similarity threshold — ai-cache/core.go:163-186. The embedding upload after caching enables building the semantic index progressively — ai-cache/core.go:222-258. Multiple embedding providers are supported: Azure, Cohere, DashScope, HuggingFace, Ollama, OpenAI, TextIn, XFYun — ai-cache/embedding/.

PII redaction and moderation (ai-security-guard): This plugin checks both request and response bodies for harmful content. Two guard modes are available:

  • MultiModalGuard: Handles multimodal content with image/text checking — ai-security-guard/main.go:46-53.
  • TextModerationPlus: Text-based content moderation — ai-security-guard/main.go:48-49. The guard runs on request bodies (user input), response headers, streaming bodies, and full response bodies — ai-security-guard/main.go:56-105. Response checking can be disabled independently — ai-security-guard/main.go:57-58.

Additional security extensions: The ai-prompt-decorator and ai-prompt-template plugins can inject system prompts for guardrails. The qwen3guard extension provides model-specific safety checks. The waf extension provides Web Application Firewall capabilities. The ip-restriction extension enables IP-based access control — extensions/ip-restriction/.

Prompt injection: There is no dedicated prompt-injection detection module; the security guard's text moderation plus and multimodal guard serve this purpose. The ai-security-guard delegates to the "Lvwang" moderation service for actual content checking — ai-security-guard/main.go:6-7.

How is it observed, deployed and scaled?

answered

Higress provides extensive observability, deploys as a Docker container or via Helm on Kubernetes, and runs as a high-performance Envoy-based proxy.

Logs and traces: The ai-statistics plugin writes structured AI logs (ai_log) and OpenTelemetry-style trace spans. At the request header phase, it sets the span kind gen_ai.span.kind = LLM and extracts the route, cluster, API name — ai-statistics/main.go:700-750. On the response, it records model name, input/output/total tokens, llm_first_token_duration, and llm_service_duration — ai-statistics/main.go:1069-1098. Span attributes use the gen_ai.* prefix for ARMS (Alibaba Resource Monitoring System) compatibility — ai-statistics/main.go:102-108. Built-in attributes like question, answer, reasoning, tool_calls, system, reasoning_tokens, cached_tokens, and full input_token_details/output_token_details maps are extracted from both OpenAI and Anthropic response formats — ai-statistics/main.go:115-144. Streaming responses are re-assembled using an SSE framer for accurate token extraction — ai-statistics/main.go:895-919.

Metrics: The plugin defines Envoy-native counter metrics at the granularity of route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName} — ai-statistics/main.go:481-483. Metrics include token counts, request duration, stream duration count, and failure count. Error detection works across both streaming and non-streaming responses, checking response-level error fields and HTTP status codes — ai-statistics/main.go:1519-1560.

Performance architecture: Higress is built on Envoy (C++), with extension logic running as WASM plugins compiled from Go using TinyGo — plugins/wasm-go/Makefile:1-8. The WASM sandbox provides process-level isolation. Request body buffering is configurable per-extension (default 100MB) — ai-proxy/main.go:30. The WASM VM memory limit triggers automatic VM rebuild when exceeded — ai-endpoint-picker/main.go:230-238. Lua scripts in Redis are used for atomic rate-limit operations — ai-token-ratelimit/main.go:58-79.

Deployment modes: The quickstart uses Docker with included configuration (docker run -d --rm --name higress-ai) — README.md:70-72. Kubernetes deployment uses Helm charts and CustomResourceDefinitions for plugin configuration management — helm/. The API proto definitions at api/networking/v1/ define the control plane data models — api/networking/v1/http_2_rpc.proto.

HA and scaling: The failover mechanism (failover.go) uses VM-level leader election (CAS-based lease) to coordinate health checks across Wasm VMs — ai-proxy/provider/failover.go:292-341. Token failover tracks per-token health with configurable failure/success thresholds and cooldown recovery — ai-proxy/provider/failover.go:349-403. The x-higress-fallback-from header enables safe internal redirects across the gateway chain — ai-proxy/main.go:189-196.