LLMs Technical Reviews
Home / LLM gateways / plano

katanemo/plano

Envoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs.

GitHub ↗★ 7.1kRustApache-2.0commit 72002a6 · 2026-09-28homepage ↗

Overview

Plano, from Katanemo, is a data plane for LLM and agent traffic built on Envoy. Envoy takes the traffic. Two Rust crates compiled to wasm32-wasip1 run inside Envoy as Proxy-WASM filters: llm_gateway handles provider selection, API translation, auth and rate limits, and prompt_gateway handles prompt-target function calling. A native Rust service, brightstaff, runs next to Envoy and owns everything that needs async I/O or state: model routing, agent orchestration, filter chains, conversation state for the Responses API, “signals” quality analysis, and tracing. A Python CLI, planoai, renders the Envoy config from one plano_config.yaml and starts the stack.

Routing is the most distinctive part. Plano does not use weighted load balancing. It routes by intent. You declare routing preferences, each a name, a description and a list of models. An orchestration LLM (“Plano-Orchestrator”) reads the conversation and picks a preference. The models for that preference can then be re-ranked by live price or latency. On top of that, a session layer tries to keep a conversation on the model whose provider prompt cache is still warm. It allows a switch only if the switch stays within a cost budget.

Plano is not a multi-tenant key-management gateway. There are no virtual keys, users or teams. Upstream credentials are per provider in the config, or passed through from the client.

Architecture

flowchart LR
  C["Client"] --> L["Envoy model listener"]
  L --> BS["brightstaff (Rust, tokio/hyper)"]
  BS --> F["Input filter chain (agents)"]
  BS --> OR["Plano-Orchestrator route choice"]
  OR --> SR["Session router: warm-cache gate"]
  SR --> E2["Envoy listener 12001"]
  E2 --> WASM["llm_gateway.wasm"]
  WASM --> UP["Provider cluster (OpenAI, Anthropic, Bedrock...)"]
  BS --> ST["State: memory / Postgres / Redis"]
  BS --> OT["OTel traces, PostHog, signals"]
  PG["prompt_gateway.wasm"] --> FC["Arch-Function / prompt targets"]
Component Path Role
brightstaff crates/brightstaff/ Native HTTP server: /v1/chat/completions, /v1/messages, /v1/responses, /agents/*, /routing/*; routing, sessions, state, signals, tracing
llm_gateway crates/llm_gateway/ Proxy-WASM filter: provider pick by header, request/response translation, auth injection, token rate limits, TTFT metrics
prompt_gateway crates/prompt_gateway/ Proxy-WASM filter for prompt targets (function calling against configured endpoints)
hermesllm crates/hermesllm/ Provider IDs and typed OpenAI / Anthropic / Bedrock / Responses request and response conversions, including streaming
common crates/common/ Config types, rate limiter (governor), routing helpers, tokenizer, PII header masking
CLI and config cli/planoai/, config/ planoai up/down/build/logs/trace, JSON schema, envoy.template.yaml, supervisord

How a request flows

A model-listener request (POST /v1/chat/completions):

  1. Envoy to brightstaff. The egress_traffic listener has a single catch-all route to the bright_staff cluster, with /healthz answered directly (envoy.template.yaml). brightstaff’s dispatch sends chat, messages and responses paths to llm_chat, /agents/* to agent_chat, and /routing/* to a decision-only endpoint (main.rs).
  2. Parse and filter. llm_chat_inner reads the affinity, tenant and X-Plano-Cache headers, parses the body into a ProviderRequestType, resolves model aliases, then runs any configured input filter chain (mod.rs).
  3. Pick a route. router_chat_get_upstream_model converts the request to Chat Completions shape and calls determine_route (model_selection.rs). The orchestrator model picks a route by name. If it detects several intents it takes the first. Then rank_models orders that route’s models by the prefer: cheapest | fastest policy (orchestrator.rs).
  4. Session gate. session_router::route checks whether the session’s anchor model is still warm (inside the provider’s cache TTL and with an unchanged prefix hash). It then keeps the anchor or allows the switch if the switch fits the max_switch_spend_pct budget (session_router.rs).
  5. Back into Envoy. send_upstream sets x-arch-llm-provider to the chosen model and posts to LLM_PROVIDER_ENDPOINT, which defaults to http://localhost:12001 (mod.rs, main.rs). On that listener, routes match the header to a provider cluster. They can retry 429/5xx on the same cluster when max_retries is set (envoy.template.yaml).
  6. WASM translate. llm_gateway.wasm runs on that listener (envoy.template.yaml). select_llm_provider resolves the hint, or uses the default provider (stream_context.rs). The request body is checked against token rate limits, converted to the upstream API with ProviderRequestType::try_from, and normalised for the provider (stream_context.rs). The response, streaming or not, is converted back to the client’s API.
  7. Account. brightstaff extracts usage from the response, updates the session binding’s cost and switch spend, stores Responses-API state if configured, and closes the trace.

Key components

Two runtimes, one proxy

The WASM filters follow Proxy-WASM rules. They are synchronous and callback-driven, with no tokio and no sockets, and they reach out only through dispatch_http_call. Their config is parsed once in on_configure. Anything async or stateful was moved to brightstaff, a hyper server on 0.0.0.0:9091 that spawns a task per connection. As a result a model request crosses Envoy twice: once into brightstaff, and once back out through the llm_gateway listener.

hermesllm

This crate handles translation. Clients can send OpenAI Chat Completions, Anthropic Messages or OpenAI Responses. Upstreams can be Chat Completions, Messages, Bedrock Converse (plain and streaming) or Responses. Every client-to-upstream pair is one arm of a single TryFrom<(ProviderRequestType, &SupportedUpstreamAPIs)> match (request.rs). normalize_for_upstream patches provider quirks. Adding a provider means adding a ProviderId variant, its request and response types, and the match arms.

Routing, sessions and cost

ModelMetricsService pulls price catalogs from models.dev or DigitalOcean and, optionally, latency metrics. switch_cost_in_usd compares the candidate’s uncached read cost with the anchor’s cached read cost, and a switch that is cheaper is always allowed. Session bindings live in memory or Redis. The /routing/* endpoints return the decision plus a ranked model list, so a client can do its own fallback. The main llm_chat path sends only the decided model and does not walk that list.

Agents, filters and signals

/agents/* uses AgentSelector to pick agents for a listener, with the same orchestrator model, then runs a PipelineProcessor over JSON-RPC agents. Input and output filters are agents too, which is how guardrail-style checks are plugged in. signals/ is a Rust port of Katanemo’s signals reference implementation. It scores conversations for misalignment, stagnation, loops, failures and exhaustion, and marks flagged spans with an emoji marker.

Rate limits

common::ratelimit keys governor token buckets by provider and a header selector. The WASM filter applies them to the input token count before forwarding and returns 429 when a limit is exceeded.

Extending it

  • Providers: add to ProviderId, the hermesllm API types and provider_models.yaml.
  • Routes: declare routing_preferences (name, description, models, selection policy) in config, or send them inline per request.
  • Guardrails and transforms: write an HTTP agent and list it in a listener’s input_filters or output_filters.
  • Agents: register agents on an agent listener and let the orchestrator dispatch to them.
  • Observability: OTLP via the tracing config, plus an optional PostHog exporter keyed by a distinct_id header.

Running it

  • planoai up plano_config.yaml starts the stack either from native binaries or in Docker (--docker). The config declares listeners, model_providers (with access_key: $ENV), and optionally agents, filters, routing preferences and prompt caching.
  • In Docker, supervisord runs a one-shot config generator, then brightstaff and Envoy (supervisord.conf).
  • Optional services: Redis for shared session bindings, PostgreSQL for Responses-API state, and an OTLP collector.
  • Building from source needs a Rust toolchain with the wasm32-wasip1 target.

Strengths and caveats

  • Strength: intent routing with cost awareness. Routing by a described intent, then ranking by price or latency, then guarding prompt-cache warmth, is a combination none of the weighted-routing gateways here offer.
  • Strength: Envoy underneath. Connection pooling, retries, timeouts, compression and access logs come from Envoy instead of custom code.
  • Strength: agent features in the proxy. Agent dispatch, filter chains and behavioural signals sit in the same data plane as model routing.
  • Caveat: an extra LLM call per routed request. The orchestrator call adds latency and a dependency whenever routing preferences are configured.
  • Caveat: no tenant model. There are no virtual keys, budgets or per-user spend controls. Rate limits key on a header value.
  • Caveat: fallback is partial. Envoy retries the same provider cluster. Cross-model fallback from the ranked list is left to callers of /routing/*.
  • Caveat: prompt_guards config is not enforced. The prompt gateway parses it into its root context but never passes it to the request context (filter_context.rs). Use filter chains for guardrails.
  • Caveat: many moving parts. Envoy, two WASM modules, brightstaff, a Python config generator and supervisord all have to agree on generated config.

Sources: code at 72002a6, verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit 72002a6. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Routing in Plano is a two-stage process: a quality-based router picks a candidate model, then a session-cache-aware gate decides whether to honor it or stick to the warm anchor.

The quality router lives in crates/brightstaff/src/handlers/llm/model_selection.rs, which calls orchestrator_service.determine_route() (crates/brightstaff/src/router/orchestrator.rs) and returns a ranked list of models for fallback. The router converts all request shapes to a common ChatCompletionsRequest via ProviderRequestType::try_from(), then invokes a downstream orchestration model (an LLM used for routing decisions). Route preferences are configured as TopLevelRoutingPreference with optional SelectionPolicy (prefer: cheapest | fastest | none) and can be backed by live cost/latency metrics from ModelMetricsService (crates/brightstaff/src/router/model_metrics.rs), which fetches pricing catalogs from DigitalOcean or models.dev.

Fallbacks and retries use the ranked models: Vec<String> returned by router_chat_get_upstream_model. If the primary model returns 429/5xx, the ranked list provides ordered alternatives. There is no explicit load-balancing across providers — the router always picks the single best model for the request.

Session-cache-aware switch gating (crates/brightstaff/src/handlers/llm/session_router.rs): The route() function checks whether the session's provider cache is still warm (comparing last_used against the provider's cache TTL from provider_cache_capability() in crates/hermesllm/src/providers/id.rs). When warm, it either sticks to the anchor model or allows a switch within a configurable cost envelope (RoutingBudget — cumulative switch spend capped at max_switch_spend_pct% of the never-switch baseline, priced from the model rates feed). Switch cost is computed in USD: switch_cost_in_usd() in orchestrator.rs compares the candidate's uncached read cost against the anchor's cached read cost. A switch that is outright cheaper (negative cost) is always allowed.

Model aliases come from the model_aliases config map (type HashMap<String, ModelAlias>) in common/configuration.rs:48. Health checks are not implemented as active pings — the system relies on upstream HTTP error codes.

Editor's note. Correction: the main proxy path (llm_chat) sends only the single decided model and does not walk the ranked list on 429/5xx. The ranked list is returned by the /routing/* decision endpoints so a client can fall back itself; Envoy only retries the same provider cluster when llm_gateway_listener.max_retries is set.

How are different provider APIs unified?

answered

Provider API unification is done through a layered type system that normalizes all request shapes into one of 4 internal formats and translates responses back to the client's expected API.

The hermesllm crate owns this layer. Three client-facing APIs are accepted: OpenAI Chat Completions (/v1/chat/completions), Anthropic Messages (/v1/messages), and OpenAI Responses (/v1/responses), enumerated in SupportedAPIsFromClient (crates/hermesllm/src/clients/endpoints.rs:7). Upstream, the system translates to 5 possible shapes: OpenAI Chat Completions, Anthropic Messages API, Amazon Bedrock Converse (streaming and non-streaming), and OpenAI Responses API (SupportedUpstreamAPIs, same file lines 14-20).

Request translation is handled by ProviderRequestType (crates/hermesllm/src/providers/request.rs:17), an enum that wraps ChatCompletionsRequest, MessagesRequest, BedrockConverse, BedrockConverseStream, or ResponsesAPIRequest. Conversion between formats is chained via TryFrom<(ProviderRequestType, &SupportedUpstreamAPIs)> (line 319), which contains ~30 match arms covering every cross-format pair — for example ResponsesAPI → ChatCompletions → Anthropic Messages is done by chaining two conversions. The ProviderRequest trait provides a uniform interface (model(), set_messages(), get_messages(), is_streaming(), to_bytes()) across all types.

Provider-specific normalization is applied in normalize_for_upstream() (line 81) — the system strips xAI's deprecated web_search_options from chat completions, normalizes Moonshot's Kimi models (strips fixed sampling fields for first-party K3, only reasoning_effort kept), and adds mandatory instructions for ChatGPT Codex.

Response translation (crates/hermesllm/src/providers/response.rs:103) is symmetric: ProviderResponseType wraps ChatCompletionsResponse, MessagesResponse, or ResponsesAPIResponse, and conversion between formats goes through transform modules in crates/hermesllm/src/transforms/response/to_openai.rs and to_anthropic.rs. Bedrock responses are first converted to the canonical format, then to the client's target. Token usage is extracted via the TokenUsage trait with support for cached input tokens (cached_input_tokens()), cache creation tokens (cache_creation_tokens()), and reasoning tokens (reasoning_tokens()).

Streaming uses separate transform modules (crates/hermesllm/src/transforms/response_streaming/) that convert between provider streaming formats (SSE chunk processors). The system supports 27+ providers via the ProviderId enum, covering OpenAI, Anthropic, Deepseek, Gemini, Mistral, Groq, Amazon Bedrock, Moonshot AI, Zhipu, Qwen, DigitalOcean, OpenRouter, and more.

Tool calls and multimodal are supported through the OpenAI-style Message type which includes tool_calls, function fields, and content parts including image_url for images. These are preserved across format conversions.

How are API keys, users and tenants managed?

insufficient evidence

Plano does not implement a user/tenant/API-key management system with virtual keys, user/team/tenant models, or an admin UI. These features are not present in the codebase at the reviewed commit.

API keys for upstream LLM providers are stored in each LlmProvider config entry's access_key: Option<String> field (crates/common/src/configuration.rs:765), which holds the credential used when proxying to that provider. The passthrough_auth: Option<bool> flag (line 776) controls whether the client's original Authorization header is forwarded directly to the upstream, bypassing the configured access_key. The Authorization header is obfuscated in logs via obfuscate_auth_header() in crates/common/src/pii.rs.

The tenant_header optional field on SessionCacheConfig (crates/common/src/configuration.rs:26) allows scoping session-cache keys to a tenant namespace by reading a named HTTP header, but this is for cache isolation, not authentication — there is no tenant identity provisioning or validation. No user model, team model, organization hierarchy, or virtual-key abstraction was found. The project instead delegates authentication to the client's existing infrastructure (the caller authenticates, Plano proxies). A search of the entire crates/ tree for "virtual_key", "user_model", "team", "organization", and "tenant_model" with no results.

How are rate limits, budgets and cost tracking implemented?

answered

Rate limits and cost tracking are implemented via the governor crate for rate limiting and a per-model pricing catalog for cost tracking, with a session-level switch-cost budget.

Rate limits are configured under ratelimits in the YAML config, each specifying a model, a selector (HTTP header key+value), and a Limit (tokens + time unit: second/minute/hour/day). The RatelimitMap in crates/common/src/ratelimit.rs structures them as Provider → {Header → KeyedRateLimiter}, using the governor crate's DefaultKeyedRateLimiter for token-bucket enforcement. The selector header value (or empty string for wildcard) acts as the key. Failed checks return Error::ExceededLimit with which provider/selector/tokens were exceeded.

Per-model rate limits (LlmRatelimit) can also be attached directly to LlmProvider config blocks via the rate_limits field (crates/common/src/configuration.rs:771), using an HTTP header selector.

Cost tracking operates at the model level via ModelRates (crates/brightstaff/src/router/model_metrics.rs:23), which stores input_per_million, output_per_million, and optional cache_read_per_million USD rates. The ModelMetricsService (line 84) fetches live pricing catalogs from DigitalOcean (api.digitalocean.com/v2/gen-ai/models/catalog) or models.dev, with configurable refresh intervals, and supports model_aliases for catalog-to-Plano model name mapping. request_cost_usd() (line 52) computes the actual dollar cost of one request from token usage, distinguishing cached vs uncached input tokens with three pricing tiers: plain input rate, cached read rate (defaults to 10% of input rate via cache_read_discount), and cache creation at plain input rate.

Usage accounting is extracted from provider responses in streaming.rs's ExtractedUsage struct (lines 36-100), which parses both OpenAI-shape (prompt_tokens includes cached) and Anthropic-shape (separate input_tokens, cache_read_input_tokens) usage from JSON. This feeds into the SessionBinding which tracks session_cost_usd, baseline_usd, and switch_spend_usd across turns. Session data is stored in either an in-memory (MemorySessionCache) or Redis backend (SessionBinding in crates/brightstaff/src/session_cache/mod.rs).

The routing budget (EffectiveRoutingBudget in configuration.rs:336) caps cumulative model-switch overhead at max_switch_spend_pct% of the session's never-switch baseline. This is a cost gate, not a hard spend limit — it ensures quality-driven switches don't inflate the bill beyond a configurable overhead.

How are caching and guardrails implemented?

answered

Caching and guardrails are implemented as separate mechanisms: provider prompt-cache marker injection for caching, and filter chains with external guard services for safety checks.

Prompt caching is controlled by the prompt_caching config block (crates/common/src/configuration.rs:270). When enabled, Plano does two things: (a) auto-injects provider-specific cache-control markers into outbound requests via inject_cache_markers() in crates/brightstaff/src/handlers/llm/prompt_caching.rs:25, and (b) derives an implicit session key from the stable prompt prefix so follow-up turns reuse the same warm cache. The cache-marking strategy is resolved from the (gateway × model family × upstream API) combination in cache_marker_strategy() (crates/hermesllm/src/providers/id.rs:186). Three strategies exist: AnthropicMessagesBreakpoints (injects ephemeral breakpoints on native Anthropic API), OpenAiContentPartCacheControl (attaches cache_control to OpenAI content parts for Anthropic-family models behind gateways like DigitalOcean/OpenRouter), and Automatic (OpenAI-family models cache stably without markers). A min_prefix_tokens threshold avoids injecting markers below the provider's minimum cacheable prefix (~1024 tokens).

Implicit session affinity (crates/brightstaff/src/affinity.rs:66) derives a session key from hash(system + tools + first_user_message) using FNV-1a 64-bit with a deterministic salt, so an explicit X-Model-Affinity header is not required. The prefix hash (system + tools only) enables drift detection — if the stable prefix changes, the provider cache is already invalidated.

Exact vs semantic caching: Plano does not implement a local response cache (exact or semantic). Caching is entirely delegated to upstream provider prompt caches — Plano just optimizes the marker injection to maximize hit rate. (README claim: none found.)

Guardrails are divided between the WASM prompt_gateway and the brightstaff backend. PromptGuards (crates/common/src/configuration.rs:545) configures input_guards keyed by GuardType (currently only Jailbreak), with on_exception handling (forward to error target, custom error handler, or message). The PromptGuardRequest/PromptGuardResponse types (crates/common/src/api/prompt_guard.rs) define the API contract for jailbreak and toxicity detection endpoints. ZeroShotClassificationRequest (crates/common/src/api/zero_shot.rs) provides intent classification via external ML models.

PII redaction is limited to auth-header obfuscation in logs (crates/common/src/pii.rs). No content-level PII redaction was found.

Filter chains are configured per-listener with input_filters/output_filters (configuration.rs:123), resolved to FilterPipeline objects that reference named Agent entries for executing guard logic. The prompt_gateway WASM filter (crates/prompt_gateway/src/filter_context.rs:62) loads config including prompt_guards at plugin initialization time and uses dispatch_http_call() to invoke guard endpoints upstream.

Moderation: No built-in moderation is implemented — the system delegates to external services via the prompt guard and zero-shot classification endpoints.

Plugin/hook points: The WASM architecture itself is the extensibility mechanism — new guardrails can be added as Envoy filter chains or external agents referenced from the filter config.

Editor's note. Correction: prompt_guards is parsed by the prompt_gateway WASM filter but never passed to its per-request context, so no jailbreak guard call is made at this commit. Guardrails in practice are HTTP agents wired as input_filters/output_filters in brightstaff's filter chains.

How is it observed, deployed and scaled?

answered

Plano is observed through OpenTelemetry tracing and structured logging, deployed as a Docker container with supervisord managing three processes, and scales horizontally behind a load balancer.

Logs, metrics, and traces: OpenTelemetry is the primary observability pillar. The tracing module in crates/brightstaff/src/tracing/ initializes a tracer (init.rs) with configurable sampling rate and OpenTelemetry gRPC endpoint (configuration.rs:480-485). Custom span attributes are injected from HTTP headers via collect_custom_trace_attributes(). LLM-specific semantic conventions are defined in tracing/constants.rs:49 (model name, provider, request/response content length, temperature, token counts). A dedicated PostHogExporter (tracing/posthog_exporter.rs:43) translates LLM spans into $ai_generation events and POSTs them to PostHog's batch/ API, supporting distinct_id from a configurable request header. The otlp_exporter gRPC target is also natively supported via Envoy's OpenTelemetry tracing integration (envoy.template.yaml:47-57).

Metrics are instrumented at two levels: WASM-level counters/gauges via proxy_wasm::hostcalls (crates/common/src/stats.rs) and application-level metrics collected throughout brightstaff (crates/brightstaff/src/metrics/). The time_to_first_token histogram is configured in envoy.template.yaml:7-28. Access logs are written to /var/log/access_ingress.log in standard Envoy format.

Performance architecture: The system is written in Rust (edition 2021, safe Rust throughout). The two WASM plugins (prompt_gateway, llm_gateway) are cdylib crates compiled for wasm32-wasip1 and run inside Envoy's WASM sandbox — meaning synchronous, callback-driven execution with no tokio/async, no networking, and no filesystem access. The native brightstaff binary runs as a separate async Tokio process for routing, state, and tracing. Communication between Envoy and brightstaff happens over HTTP; brightstaff exposes an orchestrator API (orchestrator_url) for routing decisions.

Deployment: A single Docker image bundles all components. supervisord.conf orchestrates three programs: config_generator (Python CLI, runs once to render YAML configs), brightstaff (native binary), and envoy (the proxy). The config generator runs envsubst on rendered templates to inject environment variables. The planoai Python CLI (cli/planoai/main.py) provides up, down, build, logs, trace and init commands for local development and deployment management.

High availability: Session bindings can be stored in Redis (crates/brightstaff/src/session_cache/redis.rs) rather than in-memory, enabling multiple brightstaff instances to share session state behind a load balancer. The Redis ConnectionManager with automatic reconnection handles connection drops. Session-cache lookups use short timeouts (connect: 5s, response: 2s) so a slow backend never blocks routing decisions. The project does not implement active health checks — it relies on Envoy's upstream cluster health monitoring and the response code fallback chain. Key rotation is handled through env-substituted configs and passthrough_auth mode.