# katanemo/plano

> Envoy-based LLM and agent proxy where a Rust sidecar picks models by intent and Proxy-WASM filters translate APIs.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/katanemo/plano (reviewed at commit `72002a62d90ad13dd246cf3ef8d99d8c98b075ff`, 2026-09-28)
- Stars: 7080 · Language: Rust · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/plano/

## Overview

Plano, from Katanemo, is a data plane for LLM and agent traffic built on Envoy. Envoy takes the traffic. Two Rust crates compiled to `wasm32-wasip1` run inside Envoy as Proxy-WASM filters: `llm_gateway` handles provider selection, API translation, auth and rate limits, and `prompt_gateway` handles prompt-target function calling. A native Rust service, `brightstaff`, runs next to Envoy and owns everything that needs async I/O or state: model routing, agent orchestration, filter chains, conversation state for the Responses API, "signals" quality analysis, and tracing. A Python CLI, `planoai`, renders the Envoy config from one `plano_config.yaml` and starts the stack.

Routing is the most distinctive part. Plano does not use weighted load balancing. It routes by intent. You declare routing preferences, each a name, a description and a list of models. An orchestration LLM ("Plano-Orchestrator") reads the conversation and picks a preference. The models for that preference can then be re-ranked by live price or latency. On top of that, a session layer tries to keep a conversation on the model whose provider prompt cache is still warm. It allows a switch only if the switch stays within a cost budget.

Plano is not a multi-tenant key-management gateway. There are no virtual keys, users or teams. Upstream credentials are per provider in the config, or passed through from the client.

## Architecture

```mermaid
flowchart LR
  C["Client"] --> L["Envoy model listener"]
  L --> BS["brightstaff (Rust, tokio/hyper)"]
  BS --> F["Input filter chain (agents)"]
  BS --> OR["Plano-Orchestrator route choice"]
  OR --> SR["Session router: warm-cache gate"]
  SR --> E2["Envoy listener 12001"]
  E2 --> WASM["llm_gateway.wasm"]
  WASM --> UP["Provider cluster (OpenAI, Anthropic, Bedrock...)"]
  BS --> ST["State: memory / Postgres / Redis"]
  BS --> OT["OTel traces, PostHog, signals"]
  PG["prompt_gateway.wasm"] --> FC["Arch-Function / prompt targets"]
```

| Component | Path | Role |
|---|---|---|
| brightstaff | `crates/brightstaff/` | Native HTTP server: `/v1/chat/completions`, `/v1/messages`, `/v1/responses`, `/agents/*`, `/routing/*`; routing, sessions, state, signals, tracing |
| llm_gateway | `crates/llm_gateway/` | Proxy-WASM filter: provider pick by header, request/response translation, auth injection, token rate limits, TTFT metrics |
| prompt_gateway | `crates/prompt_gateway/` | Proxy-WASM filter for prompt targets (function calling against configured endpoints) |
| hermesllm | `crates/hermesllm/` | Provider IDs and typed OpenAI / Anthropic / Bedrock / Responses request and response conversions, including streaming |
| common | `crates/common/` | Config types, rate limiter (`governor`), routing helpers, tokenizer, PII header masking |
| CLI and config | `cli/planoai/`, `config/` | `planoai up/down/build/logs/trace`, JSON schema, `envoy.template.yaml`, supervisord |

## How a request flows

A model-listener request (`POST /v1/chat/completions`):

1. **Envoy to brightstaff.** The `egress_traffic` listener has a single catch-all route to the `bright_staff` cluster, with `/healthz` answered directly ([envoy.template.yaml](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/config/envoy.template.yaml#L397-L436)). brightstaff's `dispatch` sends chat, messages and responses paths to `llm_chat`, `/agents/*` to `agent_chat`, and `/routing/*` to a decision-only endpoint ([main.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/main.rs#L563-L627)).
2. **Parse and filter.** `llm_chat_inner` reads the affinity, tenant and `X-Plano-Cache` headers, parses the body into a `ProviderRequestType`, resolves model aliases, then runs any configured input filter chain ([mod.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/mod.rs#L87-L200)).
3. **Pick a route.** `router_chat_get_upstream_model` converts the request to Chat Completions shape and calls `determine_route` ([model_selection.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/model_selection.rs#L53-L120)). The orchestrator model picks a route by name. If it detects several intents it takes the first. Then `rank_models` orders that route's models by the `prefer: cheapest | fastest` policy ([orchestrator.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/router/orchestrator.rs#L320-L400)).
4. **Session gate.** `session_router::route` checks whether the session's anchor model is still warm (inside the provider's cache TTL and with an unchanged prefix hash). It then keeps the anchor or allows the switch if the switch fits the `max_switch_spend_pct` budget ([session_router.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/session_router.rs#L333-L380)).
5. **Back into Envoy.** `send_upstream` sets `x-arch-llm-provider` to the chosen model and posts to `LLM_PROVIDER_ENDPOINT`, which defaults to `http://localhost:12001` ([mod.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/mod.rs#L822-L860), [main.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/main.rs#L105-L109)). On that listener, routes match the header to a provider cluster. They can retry 429/5xx on the same cluster when `max_retries` is set ([envoy.template.yaml](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/config/envoy.template.yaml#L528-L557)).
6. **WASM translate.** `llm_gateway.wasm` runs on that listener ([envoy.template.yaml](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/config/envoy.template.yaml#L575-L592)). `select_llm_provider` resolves the hint, or uses the default provider ([stream_context.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/llm_gateway/src/stream_context.rs#L133-L178)). The request body is checked against token rate limits, converted to the upstream API with `ProviderRequestType::try_from`, and normalised for the provider ([stream_context.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/llm_gateway/src/stream_context.rs#L1040-L1100)). The response, streaming or not, is converted back to the client's API.
7. **Account.** brightstaff extracts usage from the response, updates the session binding's cost and switch spend, stores Responses-API state if configured, and closes the trace.

## Key components

### Two runtimes, one proxy

The WASM filters follow Proxy-WASM rules. They are synchronous and callback-driven, with no tokio and no sockets, and they reach out only through `dispatch_http_call`. Their config is parsed once in `on_configure`. Anything async or stateful was moved to brightstaff, a hyper server on `0.0.0.0:9091` that spawns a task per connection. As a result a model request crosses Envoy twice: once into brightstaff, and once back out through the `llm_gateway` listener.

### hermesllm

This crate handles translation. Clients can send OpenAI Chat Completions, Anthropic Messages or OpenAI Responses. Upstreams can be Chat Completions, Messages, Bedrock Converse (plain and streaming) or Responses. Every client-to-upstream pair is one arm of a single `TryFrom<(ProviderRequestType, &SupportedUpstreamAPIs)>` match ([request.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/request.rs#L318-L360)). `normalize_for_upstream` patches provider quirks. Adding a provider means adding a `ProviderId` variant, its request and response types, and the match arms.

### Routing, sessions and cost

`ModelMetricsService` pulls price catalogs from models.dev or DigitalOcean and, optionally, latency metrics. `switch_cost_in_usd` compares the candidate's uncached read cost with the anchor's cached read cost, and a switch that is cheaper is always allowed. Session bindings live in memory or Redis. The `/routing/*` endpoints return the decision plus a ranked model list, so a client can do its own fallback. The main `llm_chat` path sends only the decided model and does not walk that list.

### Agents, filters and signals

`/agents/*` uses `AgentSelector` to pick agents for a listener, with the same orchestrator model, then runs a `PipelineProcessor` over JSON-RPC agents. Input and output filters are agents too, which is how guardrail-style checks are plugged in. `signals/` is a Rust port of Katanemo's signals reference implementation. It scores conversations for misalignment, stagnation, loops, failures and exhaustion, and marks flagged spans with an emoji marker.

### Rate limits

`common::ratelimit` keys `governor` token buckets by provider and a header selector. The WASM filter applies them to the input token count before forwarding and returns 429 when a limit is exceeded.

## Extending it

- **Providers:** add to `ProviderId`, the hermesllm API types and `provider_models.yaml`.
- **Routes:** declare `routing_preferences` (name, description, models, selection policy) in config, or send them inline per request.
- **Guardrails and transforms:** write an HTTP agent and list it in a listener's `input_filters` or `output_filters`.
- **Agents:** register agents on an agent listener and let the orchestrator dispatch to them.
- **Observability:** OTLP via the tracing config, plus an optional PostHog exporter keyed by a `distinct_id` header.

## Running it

- `planoai up plano_config.yaml` starts the stack either from native binaries or in Docker (`--docker`). The config declares `listeners`, `model_providers` (with `access_key: $ENV`), and optionally agents, filters, routing preferences and prompt caching.
- In Docker, supervisord runs a one-shot config generator, then brightstaff and Envoy ([supervisord.conf](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/config/supervisord.conf#L5-L52)).
- Optional services: Redis for shared session bindings, PostgreSQL for Responses-API state, and an OTLP collector.
- Building from source needs a Rust toolchain with the `wasm32-wasip1` target.

## Strengths and caveats

- **Strength: intent routing with cost awareness.** Routing by a described intent, then ranking by price or latency, then guarding prompt-cache warmth, is a combination none of the weighted-routing gateways here offer.
- **Strength: Envoy underneath.** Connection pooling, retries, timeouts, compression and access logs come from Envoy instead of custom code.
- **Strength: agent features in the proxy.** Agent dispatch, filter chains and behavioural signals sit in the same data plane as model routing.
- **Caveat: an extra LLM call per routed request.** The orchestrator call adds latency and a dependency whenever routing preferences are configured.
- **Caveat: no tenant model.** There are no virtual keys, budgets or per-user spend controls. Rate limits key on a header value.
- **Caveat: fallback is partial.** Envoy retries the same provider cluster. Cross-model fallback from the ranked list is left to callers of `/routing/*`.
- **Caveat: `prompt_guards` config is not enforced.** The prompt gateway parses it into its root context but never passes it to the request context ([filter_context.rs](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/prompt_gateway/src/filter_context.rs#L60-L100)). Use filter chains for guardrails.
- **Caveat: many moving parts.** Envoy, two WASM modules, brightstaff, a Python config generator and supervisord all have to agree on generated config.

*Sources: code at 72002a6, verified Q&A.*

## How katanemo/plano answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

**Routing in Plano is a two-stage process: a quality-based router picks a candidate model, then a session-cache-aware gate decides whether to honor it or stick to the warm anchor.**

The quality router lives in `crates/brightstaff/src/handlers/llm/model_selection.rs`, which calls `orchestrator_service.determine_route()` (`crates/brightstaff/src/router/orchestrator.rs`) and returns a ranked list of models for fallback. The router converts all request shapes to a common `ChatCompletionsRequest` via `ProviderRequestType::try_from()`, then invokes a downstream orchestration model (an LLM used for routing decisions). Route preferences are configured as `TopLevelRoutingPreference` with optional `SelectionPolicy` (`prefer: cheapest | fastest | none`) and can be backed by live cost/latency metrics from `ModelMetricsService` (`crates/brightstaff/src/router/model_metrics.rs`), which fetches pricing catalogs from DigitalOcean or models.dev.

Fallbacks and retries use the ranked `models: Vec<String>` returned by `router_chat_get_upstream_model`. If the primary model returns 429/5xx, the ranked list provides ordered alternatives. There is no explicit load-balancing across providers — the router always picks the single best model for the request.

**Session-cache-aware switch gating** (`crates/brightstaff/src/handlers/llm/session_router.rs`): The `route()` function checks whether the session's provider cache is still warm (comparing `last_used` against the provider's cache TTL from `provider_cache_capability()` in `crates/hermesllm/src/providers/id.rs`). When warm, it either sticks to the anchor model or allows a switch within a configurable cost envelope (`RoutingBudget` — cumulative switch spend capped at `max_switch_spend_pct`% of the never-switch baseline, priced from the model rates feed). Switch cost is computed in USD: `switch_cost_in_usd()` in `orchestrator.rs` compares the candidate's uncached read cost against the anchor's cached read cost. A switch that is outright cheaper (negative cost) is always allowed. 

**Model aliases** come from the `model_aliases` config map (type `HashMap<String, ModelAlias>`) in `common/configuration.rs:48`. **Health checks** are not implemented as active pings — the system relies on upstream HTTP error codes.

> **Editor's note.** Correction: the main proxy path (llm_chat) sends only the single decided model and does not walk the ranked list on 429/5xx. The ranked list is returned by the /routing/* decision endpoints so a client can fall back itself; Envoy only retries the same provider cluster when llm_gateway_listener.max_retries is set.

Citations: [crates/brightstaff/src/handlers/llm/model_selection.rs:53-174](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/model_selection.rs#L53-L174) · [crates/brightstaff/src/handlers/llm/session_router.rs:333-539](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/session_router.rs#L333-L539) · [crates/brightstaff/src/router/orchestrator.rs:43-57](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/router/orchestrator.rs#L43-L57) · [crates/common/src/configuration.rs:170-210](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L170-L210)

### How are different provider APIs unified? (answered)

**Provider API unification is done through a layered type system that normalizes all request shapes into one of 4 internal formats and translates responses back to the client's expected API.**

The `hermesllm` crate owns this layer. Three client-facing APIs are accepted: OpenAI Chat Completions (`/v1/chat/completions`), Anthropic Messages (`/v1/messages`), and OpenAI Responses (`/v1/responses`), enumerated in `SupportedAPIsFromClient` (`crates/hermesllm/src/clients/endpoints.rs:7`). Upstream, the system translates to 5 possible shapes: OpenAI Chat Completions, Anthropic Messages API, Amazon Bedrock Converse (streaming and non-streaming), and OpenAI Responses API (`SupportedUpstreamAPIs`, same file lines 14-20).

Request translation is handled by `ProviderRequestType` (`crates/hermesllm/src/providers/request.rs:17`), an enum that wraps `ChatCompletionsRequest`, `MessagesRequest`, `BedrockConverse`, `BedrockConverseStream`, or `ResponsesAPIRequest`. Conversion between formats is chained via `TryFrom<(ProviderRequestType, &SupportedUpstreamAPIs)>` (line 319), which contains ~30 match arms covering every cross-format pair — for example ResponsesAPI → ChatCompletions → Anthropic Messages is done by chaining two conversions. The `ProviderRequest` trait provides a uniform interface (`model()`, `set_messages()`, `get_messages()`, `is_streaming()`, `to_bytes()`) across all types.

Provider-specific normalization is applied in `normalize_for_upstream()` (line 81) — the system strips xAI's deprecated `web_search_options` from chat completions, normalizes Moonshot's Kimi models (strips fixed sampling fields for first-party K3, only `reasoning_effort` kept), and adds mandatory instructions for ChatGPT Codex.

**Response translation** (`crates/hermesllm/src/providers/response.rs:103`) is symmetric: `ProviderResponseType` wraps `ChatCompletionsResponse`, `MessagesResponse`, or `ResponsesAPIResponse`, and conversion between formats goes through transform modules in `crates/hermesllm/src/transforms/response/to_openai.rs` and `to_anthropic.rs`. Bedrock responses are first converted to the canonical format, then to the client's target. Token usage is extracted via the `TokenUsage` trait with support for cached input tokens (`cached_input_tokens()`), cache creation tokens (`cache_creation_tokens()`), and reasoning tokens (`reasoning_tokens()`).

**Streaming** uses separate transform modules (`crates/hermesllm/src/transforms/response_streaming/`) that convert between provider streaming formats (SSE chunk processors). The system supports 27+ providers via the `ProviderId` enum, covering OpenAI, Anthropic, Deepseek, Gemini, Mistral, Groq, Amazon Bedrock, Moonshot AI, Zhipu, Qwen, DigitalOcean, OpenRouter, and more.

**Tool calls and multimodal** are supported through the OpenAI-style `Message` type which includes `tool_calls`, `function` fields, and content parts including `image_url` for images. These are preserved across format conversions.


Citations: [crates/hermesllm/src/providers/request.rs:17-24](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/request.rs#L17-L24) · [crates/hermesllm/src/providers/request.rs:318-400](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/request.rs#L318-L400) · [crates/hermesllm/src/providers/response.rs:13-20](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/response.rs#L13-L20) · [crates/hermesllm/src/providers/response.rs:103-160](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/response.rs#L103-L160) · [crates/hermesllm/src/providers/id.rs:30-57](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/id.rs#L30-L57) · [crates/hermesllm/src/clients/endpoints.rs:7-20](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/clients/endpoints.rs#L7-L20)

### How are API keys, users and tenants managed? (insufficient evidence)

**Plano does not implement a user/tenant/API-key management system with virtual keys, user/team/tenant models, or an admin UI.** These features are not present in the codebase at the reviewed commit.

API keys for *upstream* LLM providers are stored in each `LlmProvider` config entry's `access_key: Option<String>` field (`crates/common/src/configuration.rs:765`), which holds the credential used when proxying to that provider. The `passthrough_auth: Option<bool>` flag (line 776) controls whether the client's original `Authorization` header is forwarded directly to the upstream, bypassing the configured `access_key`. The `Authorization` header is obfuscated in logs via `obfuscate_auth_header()` in `crates/common/src/pii.rs`.

The `tenant_header` optional field on `SessionCacheConfig` (`crates/common/src/configuration.rs:26`) allows scoping session-cache keys to a tenant namespace by reading a named HTTP header, but this is for cache isolation, not authentication — there is no tenant identity provisioning or validation. No user model, team model, organization hierarchy, or virtual-key abstraction was found. The project instead delegates authentication to the client's existing infrastructure (the caller authenticates, Plano proxies). A search of the entire `crates/` tree for "virtual_key", "user_model", "team", "organization", and "tenant_model" with no results.


Citations: [crates/common/src/configuration.rs:762-778](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L762-L778) · [crates/common/src/pii.rs:1-14](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/pii.rs#L1-L14) · [crates/common/src/configuration.rs:18-27](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L18-L27)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limits and cost tracking are implemented via the `governor` crate for rate limiting and a per-model pricing catalog for cost tracking, with a session-level switch-cost budget.**

**Rate limits** are configured under `ratelimits` in the YAML config, each specifying a `model`, a `selector` (HTTP header key+value), and a `Limit` (tokens + time unit: second/minute/hour/day). The `RatelimitMap` in `crates/common/src/ratelimit.rs` structures them as `Provider → {Header → KeyedRateLimiter}`, using the `governor` crate's `DefaultKeyedRateLimiter` for token-bucket enforcement. The selector header value (or empty string for wildcard) acts as the key. Failed checks return `Error::ExceededLimit` with which provider/selector/tokens were exceeded.

Per-model rate limits (`LlmRatelimit`) can also be attached directly to `LlmProvider` config blocks via the `rate_limits` field (`crates/common/src/configuration.rs:771`), using an HTTP header selector.

**Cost tracking** operates at the model level via `ModelRates` (`crates/brightstaff/src/router/model_metrics.rs:23`), which stores `input_per_million`, `output_per_million`, and optional `cache_read_per_million` USD rates. The `ModelMetricsService` (line 84) fetches live pricing catalogs from DigitalOcean (`api.digitalocean.com/v2/gen-ai/models/catalog`) or models.dev, with configurable refresh intervals, and supports `model_aliases` for catalog-to-Plano model name mapping. `request_cost_usd()` (line 52) computes the actual dollar cost of one request from token usage, distinguishing cached vs uncached input tokens with three pricing tiers: plain input rate, cached read rate (defaults to 10% of input rate via `cache_read_discount`), and cache creation at plain input rate.

**Usage accounting** is extracted from provider responses in `streaming.rs`'s `ExtractedUsage` struct (lines 36-100), which parses both OpenAI-shape (`prompt_tokens` includes cached) and Anthropic-shape (separate `input_tokens`, `cache_read_input_tokens`) usage from JSON. This feeds into the `SessionBinding` which tracks `session_cost_usd`, `baseline_usd`, and `switch_spend_usd` across turns. Session data is stored in either an in-memory (`MemorySessionCache`) or Redis backend (`SessionBinding` in `crates/brightstaff/src/session_cache/mod.rs`).

The **routing budget** (`EffectiveRoutingBudget` in `configuration.rs:336`) caps cumulative model-switch overhead at `max_switch_spend_pct`% of the session's never-switch baseline. This is a cost gate, not a hard spend limit — it ensures quality-driven switches don't inflate the bill beyond a configurable overhead.


Citations: [crates/common/src/ratelimit.rs:27-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/ratelimit.rs#L27-L80) · [crates/brightstaff/src/router/model_metrics.rs:1-74](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/router/model_metrics.rs#L1-L74) · [crates/brightstaff/src/session_cache/mod.rs:19-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/session_cache/mod.rs#L19-L80) · [crates/common/src/configuration.rs:287-350](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L287-L350)

### How are caching and guardrails implemented? (answered)

**Caching and guardrails are implemented as separate mechanisms: provider prompt-cache marker injection for caching, and filter chains with external guard services for safety checks.**

**Prompt caching** is controlled by the `prompt_caching` config block (`crates/common/src/configuration.rs:270`). When enabled, Plano does two things: (a) auto-injects provider-specific cache-control markers into outbound requests via `inject_cache_markers()` in `crates/brightstaff/src/handlers/llm/prompt_caching.rs:25`, and (b) derives an implicit session key from the stable prompt prefix so follow-up turns reuse the same warm cache. The cache-marking strategy is resolved from the `(gateway × model family × upstream API)` combination in `cache_marker_strategy()` (`crates/hermesllm/src/providers/id.rs:186`). Three strategies exist: `AnthropicMessagesBreakpoints` (injects ephemeral breakpoints on native Anthropic API), `OpenAiContentPartCacheControl` (attaches `cache_control` to OpenAI content parts for Anthropic-family models behind gateways like DigitalOcean/OpenRouter), and `Automatic` (OpenAI-family models cache stably without markers). A `min_prefix_tokens` threshold avoids injecting markers below the provider's minimum cacheable prefix (~1024 tokens).

**Implicit session affinity** (`crates/brightstaff/src/affinity.rs:66`) derives a session key from `hash(system + tools + first_user_message)` using FNV-1a 64-bit with a deterministic salt, so an explicit `X-Model-Affinity` header is not required. The prefix hash (system + tools only) enables drift detection — if the stable prefix changes, the provider cache is already invalidated.

**Exact vs semantic caching**: Plano does not implement a local response cache (exact or semantic). Caching is entirely delegated to upstream provider prompt caches — Plano just optimizes the marker injection to maximize hit rate. (README claim: none found.)

**Guardrails** are divided between the WASM prompt_gateway and the brightstaff backend. `PromptGuards` (`crates/common/src/configuration.rs:545`) configures `input_guards` keyed by `GuardType` (currently only `Jailbreak`), with `on_exception` handling (forward to error target, custom error handler, or message). The `PromptGuardRequest`/`PromptGuardResponse` types (`crates/common/src/api/prompt_guard.rs`) define the API contract for jailbreak and toxicity detection endpoints. `ZeroShotClassificationRequest` (`crates/common/src/api/zero_shot.rs`) provides intent classification via external ML models.

**PII redaction** is limited to auth-header obfuscation in logs (`crates/common/src/pii.rs`). No content-level PII redaction was found.

**Filter chains** are configured per-listener with `input_filters`/`output_filters` (`configuration.rs:123`), resolved to `FilterPipeline` objects that reference named `Agent` entries for executing guard logic. The prompt_gateway WASM filter (`crates/prompt_gateway/src/filter_context.rs:62`) loads config including `prompt_guards` at plugin initialization time and uses `dispatch_http_call()` to invoke guard endpoints upstream.

**Moderation**: No built-in moderation is implemented — the system delegates to external services via the prompt guard and zero-shot classification endpoints.

**Plugin/hook points**: The WASM architecture itself is the extensibility mechanism — new guardrails can be added as Envoy filter chains or external agents referenced from the filter config.

> **Editor's note.** Correction: prompt_guards is parsed by the prompt_gateway WASM filter but never passed to its per-request context, so no jailbreak guard call is made at this commit. Guardrails in practice are HTTP agents wired as input_filters/output_filters in brightstaff's filter chains.

Citations: [crates/brightstaff/src/handlers/llm/prompt_caching.rs:1-70](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/prompt_caching.rs#L1-L70) · [crates/hermesllm/src/providers/id.rs:106-228](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/id.rs#L106-L228) · [crates/brightstaff/src/affinity.rs:20-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/affinity.rs#L20-L80) · [crates/common/src/configuration.rs:544-560](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L544-L560)

### How is it observed, deployed and scaled? (answered)

**Plano is observed through OpenTelemetry tracing and structured logging, deployed as a Docker container with supervisord managing three processes, and scales horizontally behind a load balancer.**

**Logs, metrics, and traces**: OpenTelemetry is the primary observability pillar. The `tracing` module in `crates/brightstaff/src/tracing/` initializes a tracer (`init.rs`) with configurable sampling rate and OpenTelemetry gRPC endpoint (`configuration.rs:480-485`). Custom span attributes are injected from HTTP headers via `collect_custom_trace_attributes()`. LLM-specific semantic conventions are defined in `tracing/constants.rs:49` (model name, provider, request/response content length, temperature, token counts). A dedicated `PostHogExporter` (`tracing/posthog_exporter.rs:43`) translates LLM spans into `$ai_generation` events and POSTs them to PostHog's `batch/` API, supporting `distinct_id` from a configurable request header. The `otlp_exporter` gRPC target is also natively supported via Envoy's OpenTelemetry tracing integration (`envoy.template.yaml:47-57`).

**Metrics** are instrumented at two levels: WASM-level counters/gauges via `proxy_wasm::hostcalls` (`crates/common/src/stats.rs`) and application-level metrics collected throughout brightstaff (`crates/brightstaff/src/metrics/`). The `time_to_first_token` histogram is configured in `envoy.template.yaml:7-28`. Access logs are written to `/var/log/access_ingress.log` in standard Envoy format.

**Performance architecture**: The system is written in **Rust** (edition 2021, safe Rust throughout). The two WASM plugins (`prompt_gateway`, `llm_gateway`) are `cdylib` crates compiled for `wasm32-wasip1` and run inside Envoy's WASM sandbox — meaning synchronous, callback-driven execution with no tokio/async, no networking, and no filesystem access. The native `brightstaff` binary runs as a separate async Tokio process for routing, state, and tracing. Communication between Envoy and brightstaff happens over HTTP; brightstaff exposes an orchestrator API (`orchestrator_url`) for routing decisions.

**Deployment**: A single Docker image bundles all components. `supervisord.conf` orchestrates three programs: `config_generator` (Python CLI, runs once to render YAML configs), `brightstaff` (native binary), and `envoy` (the proxy). The config generator runs `envsubst` on rendered templates to inject environment variables. The `planoai` Python CLI (`cli/planoai/main.py`) provides `up`, `down`, `build`, `logs`, `trace` and `init` commands for local development and deployment management.

**High availability**: Session bindings can be stored in Redis (`crates/brightstaff/src/session_cache/redis.rs`) rather than in-memory, enabling multiple brightstaff instances to share session state behind a load balancer. The Redis `ConnectionManager` with automatic reconnection handles connection drops. Session-cache lookups use short timeouts (`connect: 5s, response: 2s`) so a slow backend never blocks routing decisions. The project does not implement active health checks — it relies on Envoy's upstream cluster health monitoring and the response code fallback chain. Key rotation is handled through env-substituted configs and `passthrough_auth` mode.


Citations: [crates/brightstaff/src/tracing/init.rs:1-15](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/init.rs#L1-L15) · [crates/brightstaff/src/tracing/posthog_exporter.rs:1-51](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/posthog_exporter.rs#L1-L51) · [crates/brightstaff/src/tracing/constants.rs:1-61](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/constants.rs#L1-L61) · [crates/brightstaff/src/session_cache/redis.rs:1-35](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/session_cache/redis.rs#L1-L35) · [crates/common/src/configuration.rs:479-528](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L479-L528)
