# higress-group/higress

> Envoy-based API gateway whose WASM plugins translate, route, cache, rate-limit and meter LLM traffic for 38 provider types.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/higress-group/higress (reviewed at commit `bda81f1067e1285775e650644a20a635f51b0a6d`, 2026-10-05)
- Stars: 9496 · Language: Go · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/higress/

## Overview

Higress is an Envoy-based API gateway from Alibaba. Its AI features are a set of WebAssembly plugins that run inside Envoy's request path. It is not a separate LLM proxy service. The control plane (Go, built on Istio) turns Kubernetes Ingress, Gateway API resources and Higress CRDs into Envoy configuration. The data plane is Envoy with Go plugins compiled by TinyGo to WASM, plus some Rust, C++ and native Go ("golang-filter") extensions.

For LLM traffic the key plugin is `ai-proxy`. It takes OpenAI-format requests (and Anthropic `/v1/messages`), rewrites them for one of 38 provider types, swaps in an upstream API token, and converts responses and SSE streams back. Other features are separate plugins that you attach to routes as needed: `model-router`, `ai-load-balancer`, `ai-cache`, `ai-quota`, `ai-token-ratelimit`, `ai-statistics`, `ai-security-guard` and about a dozen more. MCP servers are hosted through `mcp-server` plugins and a golang-filter.

Higress suits teams that already run, or want, a Kubernetes ingress and would like LLM routing to live in the same Envoy fleet as their other APIs. For a single developer it is heavy. The all-in-one Docker image helps, but the configuration model is still Envoy routes plus per-route plugin config.

## Architecture

```mermaid
flowchart LR
  K8S["Ingress / Gateway API / CRDs"] --> CP["Higress controller (Istio-based)"]
  CP -->|"xDS + WasmPlugin config"| ENV["Envoy data plane"]
  C["Client"] --> ENV
  ENV --> MR["model-router (WASM)"]
  MR --> RT["Envoy route match"]
  RT --> AP["ai-proxy (WASM)"]
  AP --> Q["ai-quota / ai-token-ratelimit"]
  AP --> CA["ai-cache"]
  AP --> UP["LLM provider"]
  AP --> ST["ai-statistics: metrics, ai_log, spans"]
```

| Component | Path | Role |
|---|---|---|
| Controller | `cmd/higress`, `pkg/ingress/` | Watches Ingress, Gateway API, `McpBridge`, `WasmPlugin`; translates to Istio/Envoy config |
| Registry bridge | `registry/` | Service discovery from Nacos, Consul, Eureka, ZooKeeper (the `McpBridge` CRD) |
| ai-proxy | `plugins/wasm-go/extensions/ai-proxy/` | Protocol translation, provider auth, token failover, retries |
| Providers | `ai-proxy/provider/*.go` | One file per provider type implementing the `Provider` handler interfaces |
| model-router | `plugins/wasm-go/extensions/model-router/` | Moves `model` into a header for route matching; regex auto-routing |
| ai-load-balancer | `plugins/wasm-go/extensions/ai-load-balancer/` | Cluster and endpoint policies (metrics, least-request, prefix cache) |
| ai-cache | `plugins/wasm-go/extensions/ai-cache/` | Redis exact cache and vector similarity cache |
| ai-quota, ai-token-ratelimit | `plugins/wasm-go/extensions/` | Per-consumer token quotas and Redis sliding-window token limits |
| ai-statistics | `plugins/wasm-go/extensions/ai-statistics/` | Token usage, latency metrics, AI access log, tracing attributes |
| MCP hosting | `plugins/golang-filter/mcp-server/`, `plugins/wasm-go/mcp-servers/` | MCP server framework and ready-made servers |

## How a request flows

Take `POST /v1/chat/completions` on a route that has `ai-proxy` attached:

1. **Config load.** When Envoy pushes plugin config, `parseGlobalConfig` or `parseOverrideRuleConfig` reads `providers[]` and `activeProviderId`. `Complete()` builds the provider and starts API-token failover ([main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L109-L154), [config.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/config/config.go#L40-L100)). The active provider is fixed per route or domain rule, not chosen per request.
2. **Model routing (optional).** If `model-router` runs first, it copies `model` into a header or splits `provider/model` into a provider header. Envoy then matches a route whose `ai-proxy` config targets that provider. With `higress/auto`, regex rules over the last user message pick the model ([model-router/main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/model-router/main.go#L193-L260)).
3. **Request headers.** `onHttpRequestHeader` maps the path to an `ApiName`. If the request is Anthropic-format and the provider lacks native support, it rewrites the path to chat completions and marks the response for conversion. It then picks an API token and calls the provider's `OnRequestHeaders` ([main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L220-L326)).
4. **Request body.** `onHttpRequestBody` applies custom settings, adds `stream_options` for usage stats on OpenAI-protocol providers, and calls the provider's `OnRequestBody`, which applies `modelMapping` and transforms the body ([main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L362-L401)).
5. **Upstream.** Envoy sends the request to the provider cluster. Other plugins on the route (quota, rate limit, cache, security guard) run in their configured phase order.
6. **Response.** `onHttpResponseHeaders`, `onStreamingResponseBody` and `onHttpResponseBody` let the provider transform the output to OpenAI format, or to Claude format if the request came in as Claude. On failure statuses, failover counts the token and the retry path may resend with another token.

## Key components

### Provider abstraction

`Provider` is a small interface. Optional handler interfaces (`RequestHeadersHandler`, `RequestBodyHandler`, `TransformRequestBodyHandler`, `StreamingResponseBodyHandler`, `StreamingEventHandler` and others) let each provider override only what it needs ([provider.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L286-L323)). `providerInitializers` registers 38 types, including OpenAI, Azure, Claude, Gemini, Vertex, Bedrock, Qwen, DeepSeek, Ollama, vLLM, Triton and a `generic` pass-through ([provider.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L236-L275)).

### Token selection, failover and retry

Each provider config holds a list of `apiTokens`. `selectApiToken` picks one at random, or uses consumer affinity (via the `x-mse-consumer` header) for stateful APIs such as Responses or files ([provider.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L834-L870)). With failover on, a token that reaches `failureThreshold` failures moves to an unavailable list in Envoy shared data. A health checker, elected through a CAS lease across WASM VMs, sends real requests to bring it back ([failover.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L370-L403)). `retryOnFailure` resends the saved body with a different token, by default on any 4xx or 5xx status ([retry.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/retry.go#L21-L49)). Failover is between keys of one provider. Fallback across providers comes from Envoy routes and `x-higress-fallback-from` internal redirects, and `initContext` takes care to keep the client's original `Authorization` across those hops ([main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L156-L197)).

### Model mapping

`modelMapping` supports exact keys, `prefix*`, `~regex` with capture-group replacement, and a `*` default ([provider.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L1053-L1100)). Prefix and regex keys are checked while ranging over a Go map, so when two of them match the same model, which one wins is not defined. Regex keys are compiled with `MustCompile` on each lookup.

### Quotas, limits and stats

`ai-quota` reads the consumer from `x-mse-consumer`, denies with 403 when the Redis counter `chat_quota:<consumer>` is not positive, and decrements by total tokens once the response ends ([ai-quota/main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L194-L213), [L267-L303](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L267-L303)). The check is before the request and the charge is after, so one large request can overshoot. Quotas count tokens. No price table exists, so cost in money is left to downstream systems. `ai-statistics` writes Envoy counters named `route.<r>.upstream.<c>.model.<m>.consumer.<u>.metric.<name>` plus `ai_log` fields and tracing attributes ([ai-statistics/main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L478-L490)).

### Cache and safety

`ai-cache` builds its key from the last message's content by default (`messages.@reverse.0.content`), or from all user messages. It checks Redis, then falls back to embedding similarity search on a vector store ([ai-cache/main.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/main.go#L95-L130), [core.go](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L80-L110)). The key does not include the model or system prompt, so scope the plugin per route. `ai-security-guard` sends request and response text (and images) to Alibaba Cloud's content-moderation service.

## Extending it

- **New provider.** Add a file under `ai-proxy/provider/` implementing the handler interfaces you need, and register it in `providerInitializers`.
- **New plugin.** Write a Go WASM plugin with `wrapper.SetCtx` and phase callbacks, build it with the `plugins/wasm-go` Makefile, and attach it with a `WasmPlugin` resource or through the console.
- **MCP servers.** Register a server in the golang-filter `mcp-server` registry, or expose REST APIs as MCP tools with the `mcp-server` WASM plugin.
- **Service discovery.** `McpBridge` brings in Nacos, Consul, Eureka or ZooKeeper services as upstreams. The name refers to multi-registry discovery, not to the Model Context Protocol.

## Running it

- **Local.** One `docker run` of the `all-in-one` image exposes the console on 8001 and the gateway on 8080/8443, with config written to a mounted directory ([README.md](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/README.md#L62-L80)).
- **Kubernetes.** Helm charts in `helm/` install the controller, gateway and console. Plugins are delivered as OCI images.
- **Required services.** Redis for `ai-quota`, `ai-token-ratelimit` and `ai-cache`. An embedding provider and a vector DB for semantic caching.

## Strengths and caveats

- **Strength: production data plane.** LLM traffic gets Envoy's connection handling, TLS, observability and route model, and sits next to ordinary API traffic in one gateway.
- **Strength: modular.** Each AI feature is a separate plugin with its own config. You turn on only what you use, per route.
- **Strength: careful key failover.** Shared-data token health, a leased health checker and auth-preserving internal redirects are well thought out.
- **Caveat: one provider per route.** `ai-proxy` has no in-plugin multi-provider fallback or cost routing. Cross-provider logic is assembled from Envoy routes, `model-router` and redirects.
- **Caveat: Alibaba-leaning defaults.** The security guard calls Alibaba Cloud moderation, consumer identity uses the `x-mse-consumer` header, and images default to an Alibaba Cloud registry.
- **Caveat: token-only accounting.** No price tables. Quotas charge after the response.
- **Caveat: plugin edge cases.** Non-deterministic prefix/regex model mapping and a model-agnostic default cache key need care in config.

*Sources: code at bda81f1, deepwiki-open wiki (12 pages), verified Q&A.*

## How higress-group/higress answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

**Routing is primarily handled by the `ai-proxy` plugin's provider config system and the `ai-load-balancer`/`ai-endpoint-picker` plugins.**

**Provider selection** is defined per-request by the `ai-proxy` config's `activeProviderId` field in the `PluginConfig`, which selects one of the configured `providers[]` entries (each with `id`, `type`, `apiTokens`, `modelMapping`, etc.) — `config/config.go:27-31`. The `model-router` plugin further intercepts the body and rewrites the `model` field; it parses `provider/model` (e.g. "openai/gpt-4") and splits it into a provider header and model, enabling provider-routing — `model-router/main.go:256-275`. It also supports **auto-routing** with regex rules matched against the last user message: when the model is `higress/auto`, rules match on message content to select the destination model — `model-router/main.go:194-248`.

**Load balancing** is provided by `ai-load-balancer`, supporting cluster-level (cluster_metrics, cluster_hash) and endpoint-level (endpoint_metrics, global_least_request, prefix_cache) policies — `ai-load-balancer/main.go:48-86`. The `ai-endpoint-picker` plugin does fine-grained endpoint selection based on **signals** including queue depth, KV cache, prefix cache, LoRA affinity, inflight requests, and failure counts, all scored by a weighted pipeline — `ai-endpoint-picker/scheduling/types.go:6-12`.

**Retries and failover** are configured per-provider. The retry mechanism (`retry.go`) re-sends failed requests using different API tokens with configurable `maxRetries` and status code matching — `ai-proxy/provider/retry.go:21-29`. The failover mechanism (`failover.go`) tracks per-token failure counts and automatically removes failing tokens, then performs periodic health checks to restore them — `ai-proxy/provider/failover.go:370-403`. Token selection uses random choice or FNV-1a consistent hashing for stateful APIs — `ai-proxy/provider/provider.go:834-919`.

**Model aliases** are done via `modelMapping` with exact match, wildcard (`*`), prefix (`prefix*`), and regex (`~pattern`) rules — `ai-proxy/provider/provider.go:1053-1100`.

**Health checks** for failover send real chat completions to unavailable tokens and count successes before restoring — `ai-proxy/provider/failover.go:192-218`.

> **Editor's note.** Correction: `activeProviderId` is applied when the plugin config (global or per route/domain rule) is parsed, not per request; each rule has exactly one active provider. Choosing a provider per request means routing to different Envoy routes, typically by having `model-router` copy the model or provider prefix into a header that the route matches on.

Citations: [plugins/wasm-go/extensions/ai-proxy/config/config.go:27-31](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/config/config.go#L27-L31) · [plugins/wasm-go/extensions/model-router/main.go:194-248](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/model-router/main.go#L194-L248) · [plugins/wasm-go/extensions/ai-load-balancer/main.go:48-86](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-load-balancer/main.go#L48-L86) · [plugins/wasm-go/extensions/ai-endpoint-picker/scheduling/types.go:6-12](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-endpoint-picker/scheduling/types.go#L6-L12) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:834-919](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L834-L919) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:370-403](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L370-L403)

### How are different provider APIs unified? (answered)

**Provider API unification is the central function of the `ai-proxy` plugin: all requests enter the gateway as OpenAI-format requests and are translated to each provider's native format.**

**API Name Resolution**: Incoming request paths are mapped to `ApiName` constants (e.g. `/v1/chat/completions` → `ApiNameChatCompletion`) via suffix and regex matching — `ai-proxy/main.go:56-104`. Over 30 API types are defined across OpenAI, Anthropic, Cohere, Gemini, and Qwen formats — `ai-proxy/provider/provider.go:37-82`.

**OpenAI-compatible schema** is the internal standard: `chatCompletionRequest`, `chatCompletionResponse`, `embeddingsRequest`, `imageGenerationRequest`, etc. are defined in `model.go` with proper JSON field names — `ai-proxy/provider/model.go:43-78`. The `protocol` field can be set to `"openai"` (default) or `"original"` to bypass conversion — `ai-proxy/provider/provider.go:411`.

**Request/response translation** happens per-provider via the `Provider` interface with `TransformRequestHeadersHandler`, `TransformRequestBodyHandler`, and `TransformResponseBodyHandler` — `ai-proxy/provider/provider.go:302-322`. Each provider file (48 provider implementations) implements these interfaces. For example, `openai.go` passes through with path rewriting to `api.openai.com`, while `claude.go` translates to Anthropic's Messages API format (`api.anthropic.com`) — `ai-proxy/provider/claude.go:18-20`. The `claude_to_openai.go` converter handles bidirectional protocol translation for providers that don't natively support Claude's protocol — `ai-proxy/provider/claude_to_openai.go:13-35`.

**Auto protocol detection** converts Claude-format requests to OpenAI-format when the upstream provider doesn't natively support Claude — `ai-proxy/main.go:254-266`. The response is converted back to Claude format on streaming/non-streaming response callbacks — `ai-proxy/main.go:688-771`.

**Streaming** is handled by `StreamingResponseBodyHandler` or `StreamingEventHandler` interfaces, with SSE framing and event extraction — `ai-proxy/provider/provider.go:290-296`. The `ExtractStreamingEvents` function parses raw SSE chunks into structured `StreamEvent` objects, handling partial chunk buffering — `ai-proxy/provider/provider.go:1130-1206`.

**Tool calls and multimodal** are fully supported: the `chatMessage` struct includes `tool_calls`, `ToolCallId`, `FunctionCall`, and multimodal content types (text, image_url, input_audio, file) — `ai-proxy/provider/model.go:207-224,348-370`.

**Provider count**: 38 provider types are registered in the `providerInitializers` map — `ai-proxy/provider/provider.go:236-275`.


Citations: [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:37-82](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L37-L82) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:302-322](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L302-L322) · [plugins/wasm-go/extensions/ai-proxy/main.go:254-266](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L254-L266) · [plugins/wasm-go/extensions/ai-proxy/provider/model.go:43-78](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/model.go#L43-L78) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:1130-1206](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L1130-L1206) · [plugins/wasm-go/extensions/ai-proxy/provider/claude_to_openai.go:13-35](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/claude_to_openai.go#L13-L35)

### How are API keys, users and tenants managed? (answered)

**API keys, users, and tenant management uses a consumer-based model with the `ai-quota` and `key-auth` plugins, plus per-provider API token configuration.**

**Virtual keys/consumer identification**: The `ai-quota` plugin reads the `x-mse-consumer` HTTP header to identify the consumer (user/tenant) and enforces quota against Redis — `ai-quota/main.go:161-167`. This consumer identity is propagated downstream and reused by `ai-statistics` for per-consumer metrics — `ai-statistics/main.go:726-728`. A working `ai-quota` key-auth example shows how API keys map to consumers: credentials are extracted from request headers and validated — `plugins/wasm-go/examples/key-auth/main.go`.

**Upstream credential storage**: Each provider config holds `apiTokens []string` — one or more API tokens for the upstream LLM service — `ai-proxy/provider/provider.go:333`. At request time, the `SetApiTokenInUse` method selects a token (via failover-global random or context-based selection) and stores it in the request context — `ai-proxy/provider/provider.go:724-734`. Token failover uses a CAS-based shared data store to track available vs. unavailable tokens per provider instance — `ai-proxy/provider/failover.go:139-224`.

**Per-request authentication header handling**: The plugin saves the original Authorization header at first hop and restores it on internal redirects (using `X-HI-ORIGINAL-AUTH`), distinguishing first-hop from re-entry requests via the `x-higress-fallback-from` header — `ai-proxy/main.go:156-197`. The upstream token replaces the Authorization header for the LLM call.

**Admin UI/API**: The `ai-quota` plugin exposes an admin API at `/v1/chat/completions/quota` with `refresh`, `delta` (increment/decrement), and `query` operations — `ai-quota/main.go:316-331`. These are authenticated by comparing the caller's consumer identity against a configured `admin_consumer` — `ai-quota/main.go:336-339`. Data is stored in Redis with key prefix `chat_quota:` — `ai-quota/main.go:117-118`.

**User/team model**: The system does not have a built-in user/team/tenant model in the Wasm plugins — consumer identity via the `x-mse-consumer` header is the primary mechanism, suitable for integration with an upstream identity provider. The Kubernetes CRD-based control plane in `api/kubernetes/` manages plugin configuration through CustomResourceDefinitions — `api/kubernetes/crd_contract.go:1-27`.


Citations: [plugins/wasm-go/extensions/ai-quota/main.go:161-167](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L161-L167) · [plugins/wasm-go/extensions/ai-quota/main.go:316-331](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L316-L331) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:333-333](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L333-L333) · [plugins/wasm-go/extensions/ai-proxy/provider/provider.go:724-734](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/provider.go#L724-L734) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:139-224](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L139-L224) · [plugins/wasm-go/extensions/ai-proxy/main.go:156-197](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L156-L197)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limits, budgets, and cost tracking are handled by three separate plugins: `ai-quota`, `ai-token-ratelimit`, and `ai-statistics`.**

**Per-key/user quotas (`ai-quota`)**: The plugin checks a Redis-stored integer quota for each consumer (identified by `x-mse-consumer` header) before allowing a request. The quota is checked as `GET chat_quota:{consumer}` and if ≤0, the request is denied with 403 — `ai-quota/main.go:194-213`. After response completion, the plugin extracts token usage from the response and decrements the quota via `DECRBY chat_quota:{consumer} {totalTokens}` — `ai-quota/main.go:267-303`. The `GetTokenUsage` function parses token counts from both OpenAI and Anthropic response formats.

**Token-based rate limiting (`ai-token-ratelimit`)**: This plugin enforces sliding-window rate limits using **server-side Lua scripts in Redis**. It supports global thresholds and per-key limits with configurable windows per rule — `ai-token-ratelimit/main.go:47-79`. The Lua script `MultiKeyRequestPhaseScript` atomically checks multiple rate limit counters in one Redis call, returning thresholds, current counts, and TTLs — `ai-token-ratelimit/main.go:58-79`. The response-phase script increments counters only after a successful LLM response.

**Usage statistics and cost tracking (`ai-statistics`)**: This plugin records per-request token usage (input, output, total, reasoning tokens, cached tokens, token details) from both streaming and non-streaming responses — `ai-statistics/main.go:990-1045`. Token usage is extracted via the `tokenusage.GetTokenUsage` library which handles multiple response formats: OpenAI chat completions, Anthropic Messages, Gemini, Doubao — `ai-statistics/main.go:960-962`. Metrics are recorded as **Envoy distributed counters** at the route/cluster/model/consumer granularity — `ai-statistics/main.go:481-524`. The metric name follows the pattern `route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName}` — `ai-statistics/main.go:481-483`.

**Built-in metrics** recorded include: `llm_first_token_duration`, `llm_service_duration`, `llm_failure_count`, and token counts — `ai-statistics/main.go:93-97`. These are accumulated into Envoy-native counter metrics and also written to the ARMS (Application Resource Monitoring System) tracing spans with `gen_ai.*` prefix — `ai-statistics/main.go:102-108`.

**Price tables**: There is no built-in price table in the codebase — cost tracking is purely token-based, leaving monetary cost calculation to external systems.


Citations: [plugins/wasm-go/extensions/ai-quota/main.go:194-213](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L194-L213) · [plugins/wasm-go/extensions/ai-quota/main.go:267-303](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-quota/main.go#L267-L303) · [plugins/wasm-go/extensions/ai-token-ratelimit/main.go:47-79](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-token-ratelimit/main.go#L47-L79) · [plugins/wasm-go/extensions/ai-statistics/main.go:481-524](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L481-L524) · [plugins/wasm-go/extensions/ai-statistics/main.go:990-1045](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L990-L1045) · [plugins/wasm-go/extensions/ai-statistics/main.go:93-108](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L93-L108)

### How are caching and guardrails implemented? (answered)

**Caching and guardrails are provided by the `ai-cache` and `ai-security-guard` plugins.**

**Exact caching (`ai-cache`)**: The plugin intercepts LLM chat requests and checks Redis for an exact cache hit using the last user message (or all user messages) as the key — `ai-cache/main.go:95-116`. On a hit, the cached response is returned directly as an SSE stream or JSON body — `ai-cache/core.go:80-84`. On a miss, the request is forwarded upstream and the response is cached for future use — `ai-cache/core.go:207-219`. Cache status (hit/miss/skip) is recorded in the AI log — `ai-cache/main.go:142-144`.

**Semantic caching**: When enabled, the `ai-cache` plugin generates text embeddings (using OpenAI, Cohere, DashScope, HuggingFace, Ollama, etc.) and queries a vector database for semantically similar requests — `ai-cache/core.go:88-107`. Vector providers implement `EmbeddingQuerier` or `StringQuerier` interfaces — `ai-cache/core.go:100-107`. An embedding-based similarity search is performed when no exact cache hit is found — `ai-cache/core.go:121-148`. Results are compared against a configurable similarity threshold — `ai-cache/core.go:163-186`. The embedding upload after caching enables building the semantic index progressively — `ai-cache/core.go:222-258`. Multiple embedding providers are supported: Azure, Cohere, DashScope, HuggingFace, Ollama, OpenAI, TextIn, XFYun — `ai-cache/embedding/`.

**PII redaction and moderation (`ai-security-guard`)**: This plugin checks both request and response bodies for harmful content. Two guard modes are available:
 - **MultiModalGuard**: Handles multimodal content with image/text checking — `ai-security-guard/main.go:46-53`.
 - **TextModerationPlus**: Text-based content moderation — `ai-security-guard/main.go:48-49`.
The guard runs on request bodies (user input), response headers, streaming bodies, and full response bodies — `ai-security-guard/main.go:56-105`. Response checking can be disabled independently — `ai-security-guard/main.go:57-58`.

**Additional security extensions**: The `ai-prompt-decorator` and `ai-prompt-template` plugins can inject system prompts for guardrails. The `qwen3guard` extension provides model-specific safety checks. The `waf` extension provides Web Application Firewall capabilities. The `ip-restriction` extension enables IP-based access control — `extensions/ip-restriction/`.

**Prompt injection**: There is no dedicated prompt-injection detection module; the security guard's text moderation plus and multimodal guard serve this purpose. The `ai-security-guard` delegates to the "Lvwang" moderation service for actual content checking — `ai-security-guard/main.go:6-7`.


Citations: [plugins/wasm-go/extensions/ai-cache/main.go:95-116](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/main.go#L95-L116) · [plugins/wasm-go/extensions/ai-cache/core.go:88-107](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L88-L107) · [plugins/wasm-go/extensions/ai-cache/core.go:121-148](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L121-L148) · [plugins/wasm-go/extensions/ai-cache/core.go:163-186](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L163-L186) · [plugins/wasm-go/extensions/ai-cache/core.go:222-258](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L222-L258) · [plugins/wasm-go/extensions/ai-security-guard/main.go:43-105](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-security-guard/main.go#L43-L105)

### How is it observed, deployed and scaled? (answered)

**Higress provides extensive observability, deploys as a Docker container or via Helm on Kubernetes, and runs as a high-performance Envoy-based proxy.**

**Logs and traces**: The `ai-statistics` plugin writes structured AI logs (`ai_log`) and OpenTelemetry-style trace spans. At the request header phase, it sets the span kind `gen_ai.span.kind = LLM` and extracts the route, cluster, API name — `ai-statistics/main.go:700-750`. On the response, it records model name, input/output/total tokens, `llm_first_token_duration`, and `llm_service_duration` — `ai-statistics/main.go:1069-1098`. Span attributes use the `gen_ai.*` prefix for ARMS (Alibaba Resource Monitoring System) compatibility — `ai-statistics/main.go:102-108`. Built-in attributes like `question`, `answer`, `reasoning`, `tool_calls`, `system`, `reasoning_tokens`, `cached_tokens`, and full `input_token_details`/`output_token_details` maps are extracted from both OpenAI and Anthropic response formats — `ai-statistics/main.go:115-144`. Streaming responses are re-assembled using an SSE framer for accurate token extraction — `ai-statistics/main.go:895-919`.

**Metrics**: The plugin defines Envoy-native counter metrics at the granularity of `route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName}` — `ai-statistics/main.go:481-483`. Metrics include token counts, request duration, stream duration count, and failure count. Error detection works across both streaming and non-streaming responses, checking response-level error fields and HTTP status codes — `ai-statistics/main.go:1519-1560`.

**Performance architecture**: Higress is built on Envoy (C++), with extension logic running as **WASM plugins** compiled from Go using TinyGo — `plugins/wasm-go/Makefile:1-8`. The WASM sandbox provides process-level isolation. Request body buffering is configurable per-extension (default 100MB) — `ai-proxy/main.go:30`. The WASM VM memory limit triggers automatic VM rebuild when exceeded — `ai-endpoint-picker/main.go:230-238`. Lua scripts in Redis are used for atomic rate-limit operations — `ai-token-ratelimit/main.go:58-79`.

**Deployment modes**: The quickstart uses Docker with included configuration (`docker run -d --rm --name higress-ai`) — `README.md:70-72`. Kubernetes deployment uses Helm charts and CustomResourceDefinitions for plugin configuration management — `helm/`. The API proto definitions at `api/networking/v1/` define the control plane data models — `api/networking/v1/http_2_rpc.proto`.

**HA and scaling**: The failover mechanism (`failover.go`) uses VM-level leader election (CAS-based lease) to coordinate health checks across Wasm VMs — `ai-proxy/provider/failover.go:292-341`. Token failover tracks per-token health with configurable failure/success thresholds and cooldown recovery — `ai-proxy/provider/failover.go:349-403`. The `x-higress-fallback-from` header enables safe internal redirects across the gateway chain — `ai-proxy/main.go:189-196`.


Citations: [plugins/wasm-go/extensions/ai-statistics/main.go:700-750](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L700-L750) · [plugins/wasm-go/extensions/ai-statistics/main.go:1069-1098](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L1069-L1098) · [plugins/wasm-go/extensions/ai-statistics/main.go:481-524](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L481-L524) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:292-341](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L292-L341) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:349-403](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L349-L403) · [plugins/wasm-go/extensions/ai-proxy/main.go:189-196](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L189-L196)
