# How are caching and guardrails implemented?

> LLM gateways — a good answer covers: Exact and semantic caching; PII redaction, moderation and prompt-injection checks; plugin or hook points.

Canonical page: https://llms-technical-reviews.com/llm-gateways/q/caching-guardrails/

## Verdict

[LiteLLM](/p/litellm/) offers the most here: exact and semantic caches on many backends, and about 60 guardrail integrations. [Portkey Gateway](/p/portkey-gateway/) has the most flexible guardrail hooks but only a process-local cache. [Bifrost](/p/bifrost/) has a strong semantic cache and no open-source guardrails.

**Caches and guardrails built in.** LiteLLM caches to memory, Redis, disk, S3, GCS or Azure Blob, with semantic caches on Redis, Qdrant or Valkey. Guardrails run before, during or after the call, including on stream chunks. Integrations include Presidio, Lakera and Bedrock Guardrails. [OmniRoute](/p/omniroute/) checks a SHA-256 exact match first, then embedding similarity on an in-memory or Redis vector store. Its guardrail registry holds a PII masker (off by default), a credential masker and a regex prompt-injection guard. [Higress](/p/higress/)'s `ai-cache` checks Redis, then a vector store. Its default key is the last message's content, without the model or system prompt. `ai-security-guard` sends text to Alibaba Cloud's moderation service.

**One strong half.** Bifrost tries a hashed exact lookup, then cosine similarity (default 0.8) on chromem, Redis, Qdrant, Weaviate or Pinecone. It skips conversations over three messages by default. Guardrails exist only as enterprise schema stubs. Portkey runs guardrails as before- and after-request hooks from about 20 plugin packs, and a failed `deny` check returns HTTP 446. Its cache is exact-match and in memory only. A request with an explicit `stream: false` is never stored, and `SEMANTIC_HIT` is just a label.

**Provider prompt caches, no response cache.** [Plano](/p/plano/) injects provider cache-control markers and keeps a session on the model whose cache is warm. Its `prompt_guards` setting is parsed but never applied. Guardrails work through HTTP agents in input and output filter chains instead. [GPT-Load](/p/gpt-load/) pins requests that share a prompt prefix to one credential. It adds regex redaction that replaces or encrypts text, and an experimental LLM-judge audit.

**Little or nothing.** [One API](/p/one-api/) has no response cache and no content checks. Redis caches only tokens, users and channels. [New API](/p/new-api/) adds one control: an admin word list (Aho-Corasick) that rejects matching prompts before billing. [Agent Router](/p/agent-router/) redacts content only in debug logs.

Pick: LiteLLM for semantic caching plus PII and moderation guardrails in one place.
Pick: Portkey Gateway for configurable per-request guardrails from many vendors.
Pick: Bifrost or Higress when a semantic cache matters more than guardrails.

## Per-project answers

### diegosouzapw/OmniRoute (answered)

Caching uses a dual-layer architecture in `open-sse/services/cache/semanticCacheManager.ts`. Layer 1 is exact-match: messages are normalized and hashed with SHA-256 for O(1) hits. Layer 2 is semantic: conversation text is embedded via `embeddingClient.ts` and matched by cosine similarity against a vector store (in-memory `MemoryVectorStore` or Redis `RedisVectorStore`). Cache entries store full response data and support SSE streaming replay. Cache configuration (`semanticCacheConfig.ts`), TTLs, and namespace isolation are provider-specific. The old deterministic caching also lives in the `proxyDispatcherCache.ts` for HTTP connection reuse. Guardrails are modular in `src/lib/guardrails/`. The `promptInjection.ts` guardrail detects system-override attempts with regex patterns (e.g., `system: override`, markdown code-block system blocks), classifies severity (low/medium/high), and supports block/warn/log modes with configurable block thresholds. `piiMasker.ts` provides PII redaction but is strictly opt-in — disabled by default (Hard Rule #20: both `PII_REDACTION_ENABLED` and `PII_RESPONSE_SANITIZATION` feature flags default to `"false"`). Moderation is handled by `moderationProvider.ts`/`moderationProviderOpenAI.ts`. The guardrail pipeline is assembled by `buildGuardrailDependencyGraph()` from `guardrails/onRequest.ts` (restructured into the current codebase) which sequences execution order. Additional guardrails cover audio/video bridges (`audioBridge.ts`, `videoBridgeHelpers.ts`), vision (`visionBridgeRouter.ts`), and compliance/no-log markers. Custom guardrail/eval/webhook extensions follow documented patterns: guardrails at `src/lib/guardrails/` → `docs/security/GUARDRAILS.md`, evals at `src/lib/evals/` → `docs/frameworks/EVALS.md`. The anti-ReDoS learnings from AGENTS.md mandate strictly bounded regex sequences to prevent catastrophic backtracking on untrusted inputs.

> **Editor's note.** Correction: there is no `buildGuardrailDependencyGraph()` or `guardrails/onRequest.ts`; guardrails are classes registered on a `GuardrailRegistry` by `registerDefaultGuardrails()` in src/lib/guardrails/registry.ts (vision, audio and video bridges, PII masker, credential masker, prompt-injection guard).

Citations: [open-sse/services/cache/semanticCacheManager.ts:1-60](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/cache/semanticCacheManager.ts#L1-L60) · [src/lib/guardrails/promptInjection.ts:1-80](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/src/lib/guardrails/promptInjection.ts#L1-L80) · [open-sse/services/cache/vectorStore.ts:1-5](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/cache/vectorStore.ts#L1-L5) · [open-sse/utils/proxyDispatcherCache.ts:1-20](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/utils/proxyDispatcherCache.ts#L1-L20) · [open-sse/services/cache/embeddingClient.ts:1-5](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/cache/embeddingClient.ts#L1-L5)

### BerriAI/litellm (answered)

**Exact response caching.** The `Cache` class (`litellm/caching/caching.py:74+`) supports multiple backends: in-memory (`InMemoryCache`), Redis (`RedisCache`), Redis Cluster (`RedisClusterCache`), DiskCache, S3, GCS, and Azure Blob. The `DualCache` wrapper maintains a local in-memory LRU + a remote Redis cache for fast local reads with cluster consistency. Response caching is controlled by `CacheMode` (default_on vs default_off) and per-request `ttl`.

**Semantic caching.** Two vector-based semantic caches are available: **QdrantSemanticCache** (`litellm/caching/qdrant_semantic_cache.py`) embeds prompts via a configurable embedding model and queries the Qdrant vector DB for semantically similar requests, returning the cached response if similarity exceeds a threshold. **RedisSemanticCache** (`litellm/caching/redis_semantic_cache.py`) does the same using Redis Stack's vector-similarity-search (VSS) capabilities.

**Guardrails.** The `GuardrailRegistry` in `litellm/proxy/guardrails/guardrail_registry.py:222-386` maintains `guardrail_class_registry` — a dict mapping integration names to `CustomGuardrail` subclasses. ~50 guardrail hooks are auto-discovered from `litellm/proxy/guardrails/guardrail_hooks/` via `get_guardrail_class_from_hooks()` (line 310), which scans subdirectories for `guardrail_class_registry` dicts. Built-in integrations include:
- **PII redaction**: Presidio-based `_OPTIONAL_PresidioPIIMasking` (`guardrail_hooks/presidio.py:172+`) analyzes and anonymizes PII (names, emails, SSNs, credit cards) in request/response content, including SSE stream chunks.
- **Moderation**: Lakera AI (`guardrail_hooks/lakera_ai.py:49`) for prompt injection and content moderation. Also OpenAI Moderation, Bedrock Guardrails, and Guardrails AI.
- **Custom hooks**: Any guardrail can implement `CustomGuardrail` (`litellm/integrations/custom_guardrail.py`) with pre-request (modify/block) and post-response callbacks, plus streaming support for SSE-based anonymization.

**Plugin points.** Guardrails register as standard `litellm.callbacks` via `CustomLogger` hooks, so they also participate in the general callback pipeline (pre-request, post-success, post-failure, streaming).


Citations: [litellm/caching/caching.py:74-100](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/caching/caching.py#L74-L100) · [litellm/caching/qdrant_semantic_cache.py:1-50](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/caching/qdrant_semantic_cache.py#L1-L50) · [litellm/proxy/guardrails/guardrail_registry.py:222-386](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/guardrails/guardrail_registry.py#L222-L386) · [litellm/proxy/guardrails/guardrail_hooks/presidio.py:172-200](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/guardrails/guardrail_hooks/presidio.py#L172-L200)

### QuantumNous/new-api (answered)

**Caching** is minimal in this gateway. The only explicit HTTP cache is the `Cache()` middleware in `middleware/cache.go:7-17`, which sets `Cache-Control: max-age=604800` (one week) on static frontend assets and `no-cache` on the root path. There is **no semantic/response caching** for LLM completions — each relay request is forwarded to the upstream provider without an intermediary cache layer.

**Redis** is used as a shared data store for rate limit counters (`middleware/rate-limit.go`), channel polling state (`model/channel.go:206-290`), and in-memory caches for user/token/channel lookups via the memory cache system (`common` package). But this is operational caching (avoiding repeated DB queries), not semantic response caching.

**Content guardrails** are similarly absent as a core system feature. There is no built-in PII redaction, moderation filter, content safety check, or prompt-injection detection middleware in the request pipeline. The router (`router/relay-router.go:156`) has a single `/v1/moderations` endpoint that relays to OpenAI's moderation API, but this is a passthrough — the gateway does not inject its own moderation.

**What the project provides instead:** (1) **Request policy logging** (`service/request_policy.go`) records every routing decision (channel selection, retries, errors) per request for audit visibility. (2) **Admin security verification** (`service/security_verification.go`) protects sensitive dashboard actions with step-up authentication (2FA, passkey, password re-entry) — this is an admin guardrail, not a content guardrail. (3) **Token-level access controls** — `ModelLimits` (`model/token.go:25`) restricts which models a token can use, `AllowIps` restricts by source IP, and `ExpiredTime` enforces token expiry. (4) **Channel auto-ban** (`model/channel.go:345-350`) — when a channel's upstream keys are all disabled, the channel is automatically marked unavailable.

**Plugin hooks** provide extensibility points: JavaScript task plugins (in `plugins/tasks/`, run via `pkg/jsplugin/`) can implement custom pre-processing and billing logic. The `task_plugin` middleware (`middleware/task_plugin.go`) and plugin protocol routes enable custom endpoint protocols to override the standard relay path. However, these are not content-moderation-specific hooks — they are general-purpose plugin slots for custom request handling.

**In summary**, this gateway focuses on routing, billing, and access control rather than content safety; guardrails like PII filtering, moderation, or injection detection would need to be added at the upstream provider level or through custom plugin development.

> **Editor's note.** Correction: there is one built-in content control: an admin-managed sensitive-word list checked (Aho-Corasick, case-insensitive) in `PrepareRequestBilling` before any upstream call, which rejects matching prompts (service/sensitive.go, relay/request_billing.go).

Citations: [middleware/cache.go:1-17](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/cache.go#L1-L17) · [middleware/rate-limit.go:22-107](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/middleware/rate-limit.go#L22-L107) · [model/token.go:14-33](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/token.go#L14-L33) · [service/security_verification.go:18-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/service/security_verification.go#L18-L80) · [router/relay-router.go:156-158](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/router/relay-router.go#L156-L158)

### songquanpeng/one-api (answered)

**Caching.** There is no semantic caching of LLM responses (no exact-match or semantic cache for chat completions). The `middleware/cache.go` sets only HTTP-level Cache-Control headers (`max-age=604800` for static assets, `no-cache` for the root route) — this is purely for the web frontend, not for API responses. The system caches database entities in Redis: tokens are cached by key with a TTL of `SYNC_FREQUENCY` seconds (`model/cache.go:28-56`), user groups are cached (`cache.go:58-74`), user quotas and enabled status are cached (`cache.go:88-149`), and group→models mappings are cached (`cache.go:151-168`). The in-memory channel index (`group2model2channels`) is refreshed on a timer (`cache.go:170-225`). None of these caches store model outputs.

**PII redaction / prompt injection / moderation.** The system has no built-in PII redaction, prompt-injection detection, or output moderation guardrails. It is a transparent proxy that does not inspect or sanitize request/response content. The `/v1/moderations` endpoint is simply relayed like any other endpoint (`router/relay.go:46`), passing through to the upstream provider's moderation API — not a local check. The only content-level control is the `SystemPrompt` feature on channels (`model/channel.go:40`), which lets an admin override or inject a system prompt into requests, but this is additive, not a guardrail.

**Blacklist.** A simple in-memory user ban list exists (`common/blacklist/main.go`) that blocks banned user IDs from auth. This is operational, not content-based.

**Turnstile.** CAPTCHA (Turnstile) support is available for registration, login, and password reset endpoints via `middleware/turnstile-check.go`, configured through `TurnstileSiteKey` and `TurnstileSecretKey` (`common/config/config.go:89-90`).

**Plugin/hook points.** There are no formal plugin or webhook systems. The adaptor pattern (`relay/adaptor.go`) is the extension mechanism for adding new providers — implementing the Adaptor interface gives you full control over request/response conversion. The message-pusher (`common/message/message-pusher.go`) can send notifications to external systems for channel disable events, but this is one-directional alerting, not an interceptor.


Citations: [common/config/config.go:86-91](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L86-L91) · [model/cache.go:28-56](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/cache.go#L28-L56)

### Portkey-AI/gateway (answered)

**Exact caching** works through two layers. The middleware layer in `src/middlewares/cache/index.ts` intercepts responses after they return from providers. It creates a SHA-256 key from `JSON.stringify(requestBody) + url` (`getCacheKey` at line 14) and stores the response body in an in-memory dict with an optional `maxAge` expiry. The `memoryCache` middleware is conditionally enabled by `conf.cache === true` (`src/index.ts:108`). In `tryPost` (`handlerUtils.ts:372-404`), a `CacheService` (`src/handlers/services/cacheService.ts:16`) checks the cache before making the provider call, passing the request headers and body to `getFromCache`, which returns statuses like `HIT`, `MISS`, `REFRESH`, or `DISABLED`. Force-refresh is triggered by the `x-portkey-cache-force-refresh` header (`cache/index.ts:38`). The second layer is the pluggable cache backed by Redis, file, memory, or Cloudflare KV (`src/shared/services/cache/index.ts:42`), with configurable TTL presets from 1 minute to 30 days (line 29-40). This is initialized in `src/index.ts:49-50` when `REDIS_CONNECTION_STRING` is set.

**Semantic caching** is not implemented in this repository — the `CACHE_STATUS` enum includes `SEMANTIC_HIT` and `SEMANTIC_MISS` labels (`cache/index.ts:6-12`) but no semantic matching logic exists. Cache mode options include `"simple"` but no semantic mode.

**Guardrails** are implemented as a hooks/plugin system. The `HooksManager` in `src/middlewares/hooks/index.ts:202` manages `beforeRequestHooks` and `afterRequestHooks`. Plugins are registered in `plugins/index.ts:71-183` across 18 plugin sources (default, portkey, qualifire, pillar, patronus, bedrock, acuvity, etc.) with ~50 guard check functions. Checks include PII redaction (`qualifire/pii`, `portkey/pii`), content moderation (`qualifire/contentModeration`, `portkey/moderateContent`), prompt injection detection (`qualifire/promptInjections`), hallucinations (`qualifire/hallucinations`), and jailbreak detection (`promptsecurity/protectPrompt`). Guardrails execute as `Check` objects within `HookObject` groups (`src/middlewares/hooks/types.ts:1-27`), supporting `deny`, `sequential`, `async` execution modes. When a verdict is false and `deny` is true, the request is blocked with a 446 status code (`handlerUtils.ts:1337-1341`). The plugin system also supports mutators (request/response transformation) like `addPrefix`, `regexReplace`.

The hook system is configuration-driven: guardrails are specified in the config JSON as `inputGuardrails`/`outputGuardrails` arrays and converted to hook objects via `convertHooksShorthand` at `handlerUtils.ts:238-275`.

> **Editor's note.** Correction: the request-path response cache is only the in-memory `memoryCache` middleware (enabled by `conf.cache`). The Redis/file/KV backends behind `createCacheBackendsRedis` belong to a separate shared cache service that `getFromCache` does not read. Note also that responses to requests with an explicit `stream: false` are not stored, because the write condition reduces to `stream === undefined`.

Citations: [src/middlewares/cache/index.ts:1-113](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/cache/index.ts#L1-L113) · [src/handlers/services/cacheService.ts:16-137](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/cacheService.ts#L16-L137) · [plugins/index.ts:71-183](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/plugins/index.ts#L71-L183) · [src/middlewares/hooks/index.ts:202-556](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/hooks/index.ts#L202-L556)

### higress-group/higress (answered)

**Caching and guardrails are provided by the `ai-cache` and `ai-security-guard` plugins.**

**Exact caching (`ai-cache`)**: The plugin intercepts LLM chat requests and checks Redis for an exact cache hit using the last user message (or all user messages) as the key — `ai-cache/main.go:95-116`. On a hit, the cached response is returned directly as an SSE stream or JSON body — `ai-cache/core.go:80-84`. On a miss, the request is forwarded upstream and the response is cached for future use — `ai-cache/core.go:207-219`. Cache status (hit/miss/skip) is recorded in the AI log — `ai-cache/main.go:142-144`.

**Semantic caching**: When enabled, the `ai-cache` plugin generates text embeddings (using OpenAI, Cohere, DashScope, HuggingFace, Ollama, etc.) and queries a vector database for semantically similar requests — `ai-cache/core.go:88-107`. Vector providers implement `EmbeddingQuerier` or `StringQuerier` interfaces — `ai-cache/core.go:100-107`. An embedding-based similarity search is performed when no exact cache hit is found — `ai-cache/core.go:121-148`. Results are compared against a configurable similarity threshold — `ai-cache/core.go:163-186`. The embedding upload after caching enables building the semantic index progressively — `ai-cache/core.go:222-258`. Multiple embedding providers are supported: Azure, Cohere, DashScope, HuggingFace, Ollama, OpenAI, TextIn, XFYun — `ai-cache/embedding/`.

**PII redaction and moderation (`ai-security-guard`)**: This plugin checks both request and response bodies for harmful content. Two guard modes are available:
 - **MultiModalGuard**: Handles multimodal content with image/text checking — `ai-security-guard/main.go:46-53`.
 - **TextModerationPlus**: Text-based content moderation — `ai-security-guard/main.go:48-49`.
The guard runs on request bodies (user input), response headers, streaming bodies, and full response bodies — `ai-security-guard/main.go:56-105`. Response checking can be disabled independently — `ai-security-guard/main.go:57-58`.

**Additional security extensions**: The `ai-prompt-decorator` and `ai-prompt-template` plugins can inject system prompts for guardrails. The `qwen3guard` extension provides model-specific safety checks. The `waf` extension provides Web Application Firewall capabilities. The `ip-restriction` extension enables IP-based access control — `extensions/ip-restriction/`.

**Prompt injection**: There is no dedicated prompt-injection detection module; the security guard's text moderation plus and multimodal guard serve this purpose. The `ai-security-guard` delegates to the "Lvwang" moderation service for actual content checking — `ai-security-guard/main.go:6-7`.


Citations: [plugins/wasm-go/extensions/ai-cache/main.go:95-116](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/main.go#L95-L116) · [plugins/wasm-go/extensions/ai-cache/core.go:88-107](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L88-L107) · [plugins/wasm-go/extensions/ai-cache/core.go:121-148](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L121-L148) · [plugins/wasm-go/extensions/ai-cache/core.go:163-186](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L163-L186) · [plugins/wasm-go/extensions/ai-cache/core.go:222-258](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-cache/core.go#L222-L258) · [plugins/wasm-go/extensions/ai-security-guard/main.go:43-105](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-security-guard/main.go#L43-L105)

### maximhq/bifrost (answered)

**Caching:** The `semanticcache` plugin (`plugins/semanticcache/main.go`) provides two-tier caching. Tier 1 is direct (exact) caching using xxhash of the deterministic `(provider, model, cacheKey, request_hash, params_hash)` tuple — `performDirectSearch` (`plugins/semanticcache/search.go:23`) does an O(1) point fetch from the vector store by ID. Tier 2 is semantic caching using embedding similarity search — `performSemanticSearch` (`plugins/semanticcache/search.go:54`) generates an embedding of the request text, queries the vector store for nearest neighbors above a configurable cosine similarity `Threshold` (default 0.8), and returns the most similar cached response. The plugin supports TTL, conversation-history thresholds for skipping long threads, and configurable cache keys (by model, provider, excluding system prompt). The `VectorStore` abstraction (`framework/vectorstore/`) backs the cache — Qdrant, Weaviate, and in-process stores are supported. **Guardrails:** There is no separate PII-redaction, moderation, or prompt-injection plugin in the OSS codebase. The `CalculateGuardrailCost` function (`framework/modelcatalog/datasheet/cost.go:228`) prices "judge calls" — internal LLM-as-judge evaluations for guardrail decisions — but the actual guardrail logic (moderation, PII scanning, prompt injection detection) lives in the enterprise build. The OSS plugin architecture exposes hook points (`PreLLMHook`, `PostLLMHook` at `plugins/governance/main.go:100-102`) where such plugins would integrate. The OSS build does support GPT Live integration for realtime moderation via `plugins/logging` redaction and header/config redaction in the OTel plugin's API responses.

> **Editor's note.** Correction: the GPT Live references in the logging and governance plugins concern billing and logging of voice sessions, not moderation. The open-source tree has no moderation or PII guardrail plugin; guardrails_config exists only as an enterprise schema stub. Vector stores also include Redis and Pinecone.

Citations: [plugins/semanticcache/main.go:1-50](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/semanticcache/main.go#L1-L50) · [plugins/semanticcache/search.go:23-85](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/semanticcache/search.go#L23-L85) · [framework/modelcatalog/datasheet/cost.go:228-250](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/framework/modelcatalog/datasheet/cost.go#L228-L250) · [plugins/governance/main.go:95-105](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/governance/main.go#L95-L105)

### katanemo/plano (answered)

**Caching and guardrails are implemented as separate mechanisms: provider prompt-cache marker injection for caching, and filter chains with external guard services for safety checks.**

**Prompt caching** is controlled by the `prompt_caching` config block (`crates/common/src/configuration.rs:270`). When enabled, Plano does two things: (a) auto-injects provider-specific cache-control markers into outbound requests via `inject_cache_markers()` in `crates/brightstaff/src/handlers/llm/prompt_caching.rs:25`, and (b) derives an implicit session key from the stable prompt prefix so follow-up turns reuse the same warm cache. The cache-marking strategy is resolved from the `(gateway × model family × upstream API)` combination in `cache_marker_strategy()` (`crates/hermesllm/src/providers/id.rs:186`). Three strategies exist: `AnthropicMessagesBreakpoints` (injects ephemeral breakpoints on native Anthropic API), `OpenAiContentPartCacheControl` (attaches `cache_control` to OpenAI content parts for Anthropic-family models behind gateways like DigitalOcean/OpenRouter), and `Automatic` (OpenAI-family models cache stably without markers). A `min_prefix_tokens` threshold avoids injecting markers below the provider's minimum cacheable prefix (~1024 tokens).

**Implicit session affinity** (`crates/brightstaff/src/affinity.rs:66`) derives a session key from `hash(system + tools + first_user_message)` using FNV-1a 64-bit with a deterministic salt, so an explicit `X-Model-Affinity` header is not required. The prefix hash (system + tools only) enables drift detection — if the stable prefix changes, the provider cache is already invalidated.

**Exact vs semantic caching**: Plano does not implement a local response cache (exact or semantic). Caching is entirely delegated to upstream provider prompt caches — Plano just optimizes the marker injection to maximize hit rate. (README claim: none found.)

**Guardrails** are divided between the WASM prompt_gateway and the brightstaff backend. `PromptGuards` (`crates/common/src/configuration.rs:545`) configures `input_guards` keyed by `GuardType` (currently only `Jailbreak`), with `on_exception` handling (forward to error target, custom error handler, or message). The `PromptGuardRequest`/`PromptGuardResponse` types (`crates/common/src/api/prompt_guard.rs`) define the API contract for jailbreak and toxicity detection endpoints. `ZeroShotClassificationRequest` (`crates/common/src/api/zero_shot.rs`) provides intent classification via external ML models.

**PII redaction** is limited to auth-header obfuscation in logs (`crates/common/src/pii.rs`). No content-level PII redaction was found.

**Filter chains** are configured per-listener with `input_filters`/`output_filters` (`configuration.rs:123`), resolved to `FilterPipeline` objects that reference named `Agent` entries for executing guard logic. The prompt_gateway WASM filter (`crates/prompt_gateway/src/filter_context.rs:62`) loads config including `prompt_guards` at plugin initialization time and uses `dispatch_http_call()` to invoke guard endpoints upstream.

**Moderation**: No built-in moderation is implemented — the system delegates to external services via the prompt guard and zero-shot classification endpoints.

**Plugin/hook points**: The WASM architecture itself is the extensibility mechanism — new guardrails can be added as Envoy filter chains or external agents referenced from the filter config.

> **Editor's note.** Correction: prompt_guards is parsed by the prompt_gateway WASM filter but never passed to its per-request context, so no jailbreak guard call is made at this commit. Guardrails in practice are HTTP agents wired as input_filters/output_filters in brightstaff's filter chains.

Citations: [crates/brightstaff/src/handlers/llm/prompt_caching.rs:1-70](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/handlers/llm/prompt_caching.rs#L1-L70) · [crates/hermesllm/src/providers/id.rs:106-228](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/hermesllm/src/providers/id.rs#L106-L228) · [crates/brightstaff/src/affinity.rs:20-80](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/affinity.rs#L20-L80) · [crates/common/src/configuration.rs:544-560](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L544-L560)

### tbphp/gpt-load (answered)

**Prompt caching** — the system does not implement a general-purpose response cache. Instead, it provides **prompt affinity routing**: `dialect.inspectPromptAffinityPrefix()` (`internal/dialect/prompt_affinity.go:43`) extracts the first user message content as a cache key. The `affinity.Cache` (`gateway/affinity.go:19-60`) stores a mapping from prompt prefix to the last-used credential ID, enabling session stickiness for providers like Anthropic that support prompt caching natively. Prompt cache keys can also be explicitly set by clients via the `prompt_cache_key` field (`internal/dialect/prompt_cache.go:12`).

**Semantic caching** — not implemented. No vector stores, embedding-based cache lookups, or cache hit/miss logic for responses exist in the codebase.

**PII redaction** — the `requestredact` package (`internal/requestredact/redact.go:1`) applies configurable regex-based redaction to outbound request content before it reaches the upstream provider. Up to 64 rules are supported, with two modes: `ModeReplace` (substitute matching text) and `ModeEncrypt` (encrypt in place with a per-credential cipher). Rules are configured globally or per-access-key. Redaction operates on raw HTTP body bytes with a 128 MB limit (`maxTextBytes`). The `redact.TokenCipher` interface (`redact.go:50-54`) supports token-aware encryption so API keys embedded in text can be redacted without breaking tokenization.

**Moderation and prompt injection** — handled by the `requestaudit` package (`internal/requestaudit/audit.go:1`), described as "experimental Jev guardrails without rewriting content." It defines `Rule` objects with `Instructions` (a prompt for the Jev LLM judge), `Threshold` (0-1 confidence), and `Action` (`block` or `warn`). Request content (up to 24 KB, `MaxRequestBytes`) is sent to the Jev classification engine (`jev` package). The `requestaudit.Cache` (`handler.go:94`) caches audit results. The Jev module (`internal/jev/`) is a mini LLM service specifically for classification decisions.

**Jev Decisions endpoint** — the system exposes a native `/v1/systemone` endpoint (`internal/dialect/decisions.go:16`) for making fast classification decisions. It accepts a `state` (string/object/array) and `questions` object, returning structured decisions. This is used internally for guardrail evaluation.

**Hook points** — the `requestredact` and `requestaudit` configurations are compiled into the immutable `ConfigSnapshot` (`internal/state/snapshot.go:206-207`) and applied per-request. Redaction happens inside the Bifrost executor before upstream dispatch; audit happens during handler processing. No plugin system exists for custom guardrails — both are configuration-driven.


Citations: [internal/requestredact/redact.go:1-80](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/requestredact/redact.go#L1-L80) · [internal/requestaudit/audit.go:1-70](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/requestaudit/audit.go#L1-L70) · [internal/dialect/prompt_affinity.go:43-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/dialect/prompt_affinity.go#L43-L60) · [internal/gateway/affinity.go:19-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/affinity.go#L19-L60)

### theagentrouter/agent-router (answered)

This repository does not implement exact-response caching, semantic caching, PII redaction guardrails, content moderation, or prompt-injection detection. The codebase focuses on routing, translation, and rate limiting — caching and safety guardrails are not present in the data-plane processing path. What does exist: (1) **Response redaction for debug logging**: the `ResponseRedactor` interface allows translators to create redacted copies of responses for safe logging, replacing sensitive content (message text, tool call arguments, audio data) with `[REDACTED LENGTH=n HASH=xxxx]` placeholders while preserving structure for debugging (`internal/translator/openai_openai.go:238-266`). The `RedactString` function uses SHA256 hashing for content correlation in logs (`internal/redaction/redaction.go:29-47`). (2) **Request redaction for debug logging**: `EndpointSpec.RedactSensitiveInfoFromRequest()` creates redacted copies of parsed request bodies, stripping user content while keeping structural fields (tool definitions, response formats) (`internal/endpointspec/endpointspec.go:223-245`). (3) **Token-level caching awareness**: the `TokenUsage` struct tracks `cachedInputTokens` and `cacheCreationInputTokens`, and these values are extracted from provider responses and propagated into dynamic metadata for cost tracking (`internal/metrics/metrics.go:151-152`). The `ExtractTokenUsageFromExplicitCaching` function normalizes Anthropic/AWS Bedrock explicit caching token counts into the unified format (`internal/metrics/metrics.go:298-313`). (4) **Header/body mutation**: routes and backends can configure header and body field removal/modification (e.g., stripping sensitive headers before forwarding to upstream) (`internal/controller/gateway.go:287-379`). These mutations are not automated guardrails — they are declarative per-route configurations.


Citations: [internal/translator/openai_openai.go:238-266](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/translator/openai_openai.go#L238-L266) · [internal/redaction/redaction.go:29-47](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/redaction/redaction.go#L29-L47) · [internal/endpointspec/endpointspec.go:223-245](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/endpointspec/endpointspec.go#L223-L245) · [internal/metrics/metrics.go:143-158](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/metrics/metrics.go#L143-L158) · [internal/metrics/metrics.go:298-313](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/metrics/metrics.go#L298-L313)
