LLMs Technical Reviews

How are caching and guardrails implemented?

Exact and semantic caching; PII redaction, moderation and prompt-injection checks; plugin or hook points.

Verdict

LiteLLM offers the most here: exact and semantic caches on many backends, and about 60 guardrail integrations. Portkey Gateway has the most flexible guardrail hooks but only a process-local cache. Bifrost has a strong semantic cache and no open-source guardrails.

Caches and guardrails built in. LiteLLM caches to memory, Redis, disk, S3, GCS or Azure Blob, with semantic caches on Redis, Qdrant or Valkey. Guardrails run before, during or after the call, including on stream chunks. Integrations include Presidio, Lakera and Bedrock Guardrails. OmniRoute checks a SHA-256 exact match first, then embedding similarity on an in-memory or Redis vector store. Its guardrail registry holds a PII masker (off by default), a credential masker and a regex prompt-injection guard. Higress’s ai-cache checks Redis, then a vector store. Its default key is the last message’s content, without the model or system prompt. ai-security-guard sends text to Alibaba Cloud’s moderation service.

One strong half. Bifrost tries a hashed exact lookup, then cosine similarity (default 0.8) on chromem, Redis, Qdrant, Weaviate or Pinecone. It skips conversations over three messages by default. Guardrails exist only as enterprise schema stubs. Portkey runs guardrails as before- and after-request hooks from about 20 plugin packs, and a failed deny check returns HTTP 446. Its cache is exact-match and in memory only. A request with an explicit stream: false is never stored, and SEMANTIC_HIT is just a label.

Provider prompt caches, no response cache. Plano injects provider cache-control markers and keeps a session on the model whose cache is warm. Its prompt_guards setting is parsed but never applied. Guardrails work through HTTP agents in input and output filter chains instead. GPT-Load pins requests that share a prompt prefix to one credential. It adds regex redaction that replaces or encrypts text, and an experimental LLM-judge audit.

Little or nothing. One API has no response cache and no content checks. Redis caches only tokens, users and channels. New API adds one control: an admin word list (Aho-Corasick) that rejects matching prompts before billing. Agent Router redacts content only in debug logs.

Pick: LiteLLM for semantic caching plus PII and moderation guardrails in one place. Pick: Portkey Gateway for configurable per-request guardrails from many vendors. Pick: Bifrost or Higress when a semantic cache matters more than guardrails.

Per-project answers

diegosouzapw/OmniRoute

answered

Caching uses a dual-layer architecture in open-sse/services/cache/semanticCacheManager.ts. Layer 1 is exact-match: messages are normalized and hashed with SHA-256 for O(1) hits. Layer 2 is semantic: conversation text is embedded via embeddingClient.ts and matched by cosine similarity against a vector store (in-memory MemoryVectorStore or Redis RedisVectorStore). Cache entries store full response data and support SSE streaming replay. Cache configuration (semanticCacheConfig.ts), TTLs, and namespace isolation are provider-specific. The old deterministic caching also lives in the proxyDispatcherCache.ts for HTTP connection reuse. Guardrails are modular in src/lib/guardrails/. The promptInjection.ts guardrail detects system-override attempts with regex patterns (e.g., system: override, markdown code-block system blocks), classifies severity (low/medium/high), and supports block/warn/log modes with configurable block thresholds. piiMasker.ts provides PII redaction but is strictly opt-in — disabled by default (Hard Rule #20: both PII_REDACTION_ENABLED and PII_RESPONSE_SANITIZATION feature flags default to "false"). Moderation is handled by moderationProvider.ts/moderationProviderOpenAI.ts. The guardrail pipeline is assembled by buildGuardrailDependencyGraph() from guardrails/onRequest.ts (restructured into the current codebase) which sequences execution order. Additional guardrails cover audio/video bridges (audioBridge.ts, videoBridgeHelpers.ts), vision (visionBridgeRouter.ts), and compliance/no-log markers. Custom guardrail/eval/webhook extensions follow documented patterns: guardrails at src/lib/guardrails/ → docs/security/GUARDRAILS.md, evals at src/lib/evals/ → docs/frameworks/EVALS.md. The anti-ReDoS learnings from AGENTS.md mandate strictly bounded regex sequences to prevent catastrophic backtracking on untrusted inputs.

Editor's note. Correction: there is no buildGuardrailDependencyGraph() or guardrails/onRequest.ts; guardrails are classes registered on a GuardrailRegistry by registerDefaultGuardrails() in src/lib/guardrails/registry.ts (vision, audio and video bridges, PII masker, credential masker, prompt-injection guard).

BerriAI/litellm

answered

Exact response caching. The Cache class (litellm/caching/caching.py:74+) supports multiple backends: in-memory (InMemoryCache), Redis (RedisCache), Redis Cluster (RedisClusterCache), DiskCache, S3, GCS, and Azure Blob. The DualCache wrapper maintains a local in-memory LRU + a remote Redis cache for fast local reads with cluster consistency. Response caching is controlled by CacheMode (default_on vs default_off) and per-request ttl.

Semantic caching. Two vector-based semantic caches are available: QdrantSemanticCache (litellm/caching/qdrant_semantic_cache.py) embeds prompts via a configurable embedding model and queries the Qdrant vector DB for semantically similar requests, returning the cached response if similarity exceeds a threshold. RedisSemanticCache (litellm/caching/redis_semantic_cache.py) does the same using Redis Stack's vector-similarity-search (VSS) capabilities.

Guardrails. The GuardrailRegistry in litellm/proxy/guardrails/guardrail_registry.py:222-386 maintains guardrail_class_registry — a dict mapping integration names to CustomGuardrail subclasses. ~50 guardrail hooks are auto-discovered from litellm/proxy/guardrails/guardrail_hooks/ via get_guardrail_class_from_hooks() (line 310), which scans subdirectories for guardrail_class_registry dicts. Built-in integrations include:

  • PII redaction: Presidio-based _OPTIONAL_PresidioPIIMasking (guardrail_hooks/presidio.py:172+) analyzes and anonymizes PII (names, emails, SSNs, credit cards) in request/response content, including SSE stream chunks.
  • Moderation: Lakera AI (guardrail_hooks/lakera_ai.py:49) for prompt injection and content moderation. Also OpenAI Moderation, Bedrock Guardrails, and Guardrails AI.
  • Custom hooks: Any guardrail can implement CustomGuardrail (litellm/integrations/custom_guardrail.py) with pre-request (modify/block) and post-response callbacks, plus streaming support for SSE-based anonymization.

Plugin points. Guardrails register as standard litellm.callbacks via CustomLogger hooks, so they also participate in the general callback pipeline (pre-request, post-success, post-failure, streaming).

QuantumNous/new-api

answered

Caching is minimal in this gateway. The only explicit HTTP cache is the Cache() middleware in middleware/cache.go:7-17, which sets Cache-Control: max-age=604800 (one week) on static frontend assets and no-cache on the root path. There is no semantic/response caching for LLM completions — each relay request is forwarded to the upstream provider without an intermediary cache layer.

Redis is used as a shared data store for rate limit counters (middleware/rate-limit.go), channel polling state (model/channel.go:206-290), and in-memory caches for user/token/channel lookups via the memory cache system (common package). But this is operational caching (avoiding repeated DB queries), not semantic response caching.

Content guardrails are similarly absent as a core system feature. There is no built-in PII redaction, moderation filter, content safety check, or prompt-injection detection middleware in the request pipeline. The router (router/relay-router.go:156) has a single /v1/moderations endpoint that relays to OpenAI's moderation API, but this is a passthrough — the gateway does not inject its own moderation.

What the project provides instead: (1) Request policy logging (service/request_policy.go) records every routing decision (channel selection, retries, errors) per request for audit visibility. (2) Admin security verification (service/security_verification.go) protects sensitive dashboard actions with step-up authentication (2FA, passkey, password re-entry) — this is an admin guardrail, not a content guardrail. (3) Token-level access controls — ModelLimits (model/token.go:25) restricts which models a token can use, AllowIps restricts by source IP, and ExpiredTime enforces token expiry. (4) Channel auto-ban (model/channel.go:345-350) — when a channel's upstream keys are all disabled, the channel is automatically marked unavailable.

Plugin hooks provide extensibility points: JavaScript task plugins (in plugins/tasks/, run via pkg/jsplugin/) can implement custom pre-processing and billing logic. The task_plugin middleware (middleware/task_plugin.go) and plugin protocol routes enable custom endpoint protocols to override the standard relay path. However, these are not content-moderation-specific hooks — they are general-purpose plugin slots for custom request handling.

In summary, this gateway focuses on routing, billing, and access control rather than content safety; guardrails like PII filtering, moderation, or injection detection would need to be added at the upstream provider level or through custom plugin development.

Editor's note. Correction: there is one built-in content control: an admin-managed sensitive-word list checked (Aho-Corasick, case-insensitive) in PrepareRequestBilling before any upstream call, which rejects matching prompts (service/sensitive.go, relay/request_billing.go).

songquanpeng/one-api

answered

Caching. There is no semantic caching of LLM responses (no exact-match or semantic cache for chat completions). The middleware/cache.go sets only HTTP-level Cache-Control headers (max-age=604800 for static assets, no-cache for the root route) — this is purely for the web frontend, not for API responses. The system caches database entities in Redis: tokens are cached by key with a TTL of SYNC_FREQUENCY seconds (model/cache.go:28-56), user groups are cached (cache.go:58-74), user quotas and enabled status are cached (cache.go:88-149), and group→models mappings are cached (cache.go:151-168). The in-memory channel index (group2model2channels) is refreshed on a timer (cache.go:170-225). None of these caches store model outputs.

PII redaction / prompt injection / moderation. The system has no built-in PII redaction, prompt-injection detection, or output moderation guardrails. It is a transparent proxy that does not inspect or sanitize request/response content. The /v1/moderations endpoint is simply relayed like any other endpoint (router/relay.go:46), passing through to the upstream provider's moderation API — not a local check. The only content-level control is the SystemPrompt feature on channels (model/channel.go:40), which lets an admin override or inject a system prompt into requests, but this is additive, not a guardrail.

Blacklist. A simple in-memory user ban list exists (common/blacklist/main.go) that blocks banned user IDs from auth. This is operational, not content-based.

Turnstile. CAPTCHA (Turnstile) support is available for registration, login, and password reset endpoints via middleware/turnstile-check.go, configured through TurnstileSiteKey and TurnstileSecretKey (common/config/config.go:89-90).

Plugin/hook points. There are no formal plugin or webhook systems. The adaptor pattern (relay/adaptor.go) is the extension mechanism for adding new providers — implementing the Adaptor interface gives you full control over request/response conversion. The message-pusher (common/message/message-pusher.go) can send notifications to external systems for channel disable events, but this is one-directional alerting, not an interceptor.

Portkey-AI/gateway

answered

Exact caching works through two layers. The middleware layer in src/middlewares/cache/index.ts intercepts responses after they return from providers. It creates a SHA-256 key from JSON.stringify(requestBody) + url (getCacheKey at line 14) and stores the response body in an in-memory dict with an optional maxAge expiry. The memoryCache middleware is conditionally enabled by conf.cache === true (src/index.ts:108). In tryPost (handlerUtils.ts:372-404), a CacheService (src/handlers/services/cacheService.ts:16) checks the cache before making the provider call, passing the request headers and body to getFromCache, which returns statuses like HIT, MISS, REFRESH, or DISABLED. Force-refresh is triggered by the x-portkey-cache-force-refresh header (cache/index.ts:38). The second layer is the pluggable cache backed by Redis, file, memory, or Cloudflare KV (src/shared/services/cache/index.ts:42), with configurable TTL presets from 1 minute to 30 days (line 29-40). This is initialized in src/index.ts:49-50 when REDIS_CONNECTION_STRING is set.

Semantic caching is not implemented in this repository — the CACHE_STATUS enum includes SEMANTIC_HIT and SEMANTIC_MISS labels (cache/index.ts:6-12) but no semantic matching logic exists. Cache mode options include "simple" but no semantic mode.

Guardrails are implemented as a hooks/plugin system. The HooksManager in src/middlewares/hooks/index.ts:202 manages beforeRequestHooks and afterRequestHooks. Plugins are registered in plugins/index.ts:71-183 across 18 plugin sources (default, portkey, qualifire, pillar, patronus, bedrock, acuvity, etc.) with ~50 guard check functions. Checks include PII redaction (qualifire/pii, portkey/pii), content moderation (qualifire/contentModeration, portkey/moderateContent), prompt injection detection (qualifire/promptInjections), hallucinations (qualifire/hallucinations), and jailbreak detection (promptsecurity/protectPrompt). Guardrails execute as Check objects within HookObject groups (src/middlewares/hooks/types.ts:1-27), supporting deny, sequential, async execution modes. When a verdict is false and deny is true, the request is blocked with a 446 status code (handlerUtils.ts:1337-1341). The plugin system also supports mutators (request/response transformation) like addPrefix, regexReplace.

The hook system is configuration-driven: guardrails are specified in the config JSON as inputGuardrails/outputGuardrails arrays and converted to hook objects via convertHooksShorthand at handlerUtils.ts:238-275.

Editor's note. Correction: the request-path response cache is only the in-memory memoryCache middleware (enabled by conf.cache). The Redis/file/KV backends behind createCacheBackendsRedis belong to a separate shared cache service that getFromCache does not read. Note also that responses to requests with an explicit stream: false are not stored, because the write condition reduces to stream === undefined.

higress-group/higress

answered

Caching and guardrails are provided by the ai-cache and ai-security-guard plugins.

Exact caching (ai-cache): The plugin intercepts LLM chat requests and checks Redis for an exact cache hit using the last user message (or all user messages) as the key — ai-cache/main.go:95-116. On a hit, the cached response is returned directly as an SSE stream or JSON body — ai-cache/core.go:80-84. On a miss, the request is forwarded upstream and the response is cached for future use — ai-cache/core.go:207-219. Cache status (hit/miss/skip) is recorded in the AI log — ai-cache/main.go:142-144.

Semantic caching: When enabled, the ai-cache plugin generates text embeddings (using OpenAI, Cohere, DashScope, HuggingFace, Ollama, etc.) and queries a vector database for semantically similar requests — ai-cache/core.go:88-107. Vector providers implement EmbeddingQuerier or StringQuerier interfaces — ai-cache/core.go:100-107. An embedding-based similarity search is performed when no exact cache hit is found — ai-cache/core.go:121-148. Results are compared against a configurable similarity threshold — ai-cache/core.go:163-186. The embedding upload after caching enables building the semantic index progressively — ai-cache/core.go:222-258. Multiple embedding providers are supported: Azure, Cohere, DashScope, HuggingFace, Ollama, OpenAI, TextIn, XFYun — ai-cache/embedding/.

PII redaction and moderation (ai-security-guard): This plugin checks both request and response bodies for harmful content. Two guard modes are available:

  • MultiModalGuard: Handles multimodal content with image/text checking — ai-security-guard/main.go:46-53.
  • TextModerationPlus: Text-based content moderation — ai-security-guard/main.go:48-49. The guard runs on request bodies (user input), response headers, streaming bodies, and full response bodies — ai-security-guard/main.go:56-105. Response checking can be disabled independently — ai-security-guard/main.go:57-58.

Additional security extensions: The ai-prompt-decorator and ai-prompt-template plugins can inject system prompts for guardrails. The qwen3guard extension provides model-specific safety checks. The waf extension provides Web Application Firewall capabilities. The ip-restriction extension enables IP-based access control — extensions/ip-restriction/.

Prompt injection: There is no dedicated prompt-injection detection module; the security guard's text moderation plus and multimodal guard serve this purpose. The ai-security-guard delegates to the "Lvwang" moderation service for actual content checking — ai-security-guard/main.go:6-7.

maximhq/bifrost

answered

Caching: The semanticcache plugin (plugins/semanticcache/main.go) provides two-tier caching. Tier 1 is direct (exact) caching using xxhash of the deterministic (provider, model, cacheKey, request_hash, params_hash) tuple — performDirectSearch (plugins/semanticcache/search.go:23) does an O(1) point fetch from the vector store by ID. Tier 2 is semantic caching using embedding similarity search — performSemanticSearch (plugins/semanticcache/search.go:54) generates an embedding of the request text, queries the vector store for nearest neighbors above a configurable cosine similarity Threshold (default 0.8), and returns the most similar cached response. The plugin supports TTL, conversation-history thresholds for skipping long threads, and configurable cache keys (by model, provider, excluding system prompt). The VectorStore abstraction (framework/vectorstore/) backs the cache — Qdrant, Weaviate, and in-process stores are supported. Guardrails: There is no separate PII-redaction, moderation, or prompt-injection plugin in the OSS codebase. The CalculateGuardrailCost function (framework/modelcatalog/datasheet/cost.go:228) prices "judge calls" — internal LLM-as-judge evaluations for guardrail decisions — but the actual guardrail logic (moderation, PII scanning, prompt injection detection) lives in the enterprise build. The OSS plugin architecture exposes hook points (PreLLMHook, PostLLMHook at plugins/governance/main.go:100-102) where such plugins would integrate. The OSS build does support GPT Live integration for realtime moderation via plugins/logging redaction and header/config redaction in the OTel plugin's API responses.

Editor's note. Correction: the GPT Live references in the logging and governance plugins concern billing and logging of voice sessions, not moderation. The open-source tree has no moderation or PII guardrail plugin; guardrails_config exists only as an enterprise schema stub. Vector stores also include Redis and Pinecone.

katanemo/plano

answered

Caching and guardrails are implemented as separate mechanisms: provider prompt-cache marker injection for caching, and filter chains with external guard services for safety checks.

Prompt caching is controlled by the prompt_caching config block (crates/common/src/configuration.rs:270). When enabled, Plano does two things: (a) auto-injects provider-specific cache-control markers into outbound requests via inject_cache_markers() in crates/brightstaff/src/handlers/llm/prompt_caching.rs:25, and (b) derives an implicit session key from the stable prompt prefix so follow-up turns reuse the same warm cache. The cache-marking strategy is resolved from the (gateway × model family × upstream API) combination in cache_marker_strategy() (crates/hermesllm/src/providers/id.rs:186). Three strategies exist: AnthropicMessagesBreakpoints (injects ephemeral breakpoints on native Anthropic API), OpenAiContentPartCacheControl (attaches cache_control to OpenAI content parts for Anthropic-family models behind gateways like DigitalOcean/OpenRouter), and Automatic (OpenAI-family models cache stably without markers). A min_prefix_tokens threshold avoids injecting markers below the provider's minimum cacheable prefix (~1024 tokens).

Implicit session affinity (crates/brightstaff/src/affinity.rs:66) derives a session key from hash(system + tools + first_user_message) using FNV-1a 64-bit with a deterministic salt, so an explicit X-Model-Affinity header is not required. The prefix hash (system + tools only) enables drift detection — if the stable prefix changes, the provider cache is already invalidated.

Exact vs semantic caching: Plano does not implement a local response cache (exact or semantic). Caching is entirely delegated to upstream provider prompt caches — Plano just optimizes the marker injection to maximize hit rate. (README claim: none found.)

Guardrails are divided between the WASM prompt_gateway and the brightstaff backend. PromptGuards (crates/common/src/configuration.rs:545) configures input_guards keyed by GuardType (currently only Jailbreak), with on_exception handling (forward to error target, custom error handler, or message). The PromptGuardRequest/PromptGuardResponse types (crates/common/src/api/prompt_guard.rs) define the API contract for jailbreak and toxicity detection endpoints. ZeroShotClassificationRequest (crates/common/src/api/zero_shot.rs) provides intent classification via external ML models.

PII redaction is limited to auth-header obfuscation in logs (crates/common/src/pii.rs). No content-level PII redaction was found.

Filter chains are configured per-listener with input_filters/output_filters (configuration.rs:123), resolved to FilterPipeline objects that reference named Agent entries for executing guard logic. The prompt_gateway WASM filter (crates/prompt_gateway/src/filter_context.rs:62) loads config including prompt_guards at plugin initialization time and uses dispatch_http_call() to invoke guard endpoints upstream.

Moderation: No built-in moderation is implemented — the system delegates to external services via the prompt guard and zero-shot classification endpoints.

Plugin/hook points: The WASM architecture itself is the extensibility mechanism — new guardrails can be added as Envoy filter chains or external agents referenced from the filter config.

Editor's note. Correction: prompt_guards is parsed by the prompt_gateway WASM filter but never passed to its per-request context, so no jailbreak guard call is made at this commit. Guardrails in practice are HTTP agents wired as input_filters/output_filters in brightstaff's filter chains.

tbphp/gpt-load

answered

Prompt caching — the system does not implement a general-purpose response cache. Instead, it provides prompt affinity routing: dialect.inspectPromptAffinityPrefix() (internal/dialect/prompt_affinity.go:43) extracts the first user message content as a cache key. The affinity.Cache (gateway/affinity.go:19-60) stores a mapping from prompt prefix to the last-used credential ID, enabling session stickiness for providers like Anthropic that support prompt caching natively. Prompt cache keys can also be explicitly set by clients via the prompt_cache_key field (internal/dialect/prompt_cache.go:12).

Semantic caching — not implemented. No vector stores, embedding-based cache lookups, or cache hit/miss logic for responses exist in the codebase.

PII redaction — the requestredact package (internal/requestredact/redact.go:1) applies configurable regex-based redaction to outbound request content before it reaches the upstream provider. Up to 64 rules are supported, with two modes: ModeReplace (substitute matching text) and ModeEncrypt (encrypt in place with a per-credential cipher). Rules are configured globally or per-access-key. Redaction operates on raw HTTP body bytes with a 128 MB limit (maxTextBytes). The redact.TokenCipher interface (redact.go:50-54) supports token-aware encryption so API keys embedded in text can be redacted without breaking tokenization.

Moderation and prompt injection — handled by the requestaudit package (internal/requestaudit/audit.go:1), described as "experimental Jev guardrails without rewriting content." It defines Rule objects with Instructions (a prompt for the Jev LLM judge), Threshold (0-1 confidence), and Action (block or warn). Request content (up to 24 KB, MaxRequestBytes) is sent to the Jev classification engine (jev package). The requestaudit.Cache (handler.go:94) caches audit results. The Jev module (internal/jev/) is a mini LLM service specifically for classification decisions.

Jev Decisions endpoint — the system exposes a native /v1/systemone endpoint (internal/dialect/decisions.go:16) for making fast classification decisions. It accepts a state (string/object/array) and questions object, returning structured decisions. This is used internally for guardrail evaluation.

Hook points — the requestredact and requestaudit configurations are compiled into the immutable ConfigSnapshot (internal/state/snapshot.go:206-207) and applied per-request. Redaction happens inside the Bifrost executor before upstream dispatch; audit happens during handler processing. No plugin system exists for custom guardrails — both are configuration-driven.

theagentrouter/agent-router

answered

This repository does not implement exact-response caching, semantic caching, PII redaction guardrails, content moderation, or prompt-injection detection. The codebase focuses on routing, translation, and rate limiting — caching and safety guardrails are not present in the data-plane processing path. What does exist: (1) Response redaction for debug logging: the ResponseRedactor interface allows translators to create redacted copies of responses for safe logging, replacing sensitive content (message text, tool call arguments, audio data) with [REDACTED LENGTH=n HASH=xxxx] placeholders while preserving structure for debugging (internal/translator/openai_openai.go:238-266). The RedactString function uses SHA256 hashing for content correlation in logs (internal/redaction/redaction.go:29-47). (2) Request redaction for debug logging: EndpointSpec.RedactSensitiveInfoFromRequest() creates redacted copies of parsed request bodies, stripping user content while keeping structural fields (tool definitions, response formats) (internal/endpointspec/endpointspec.go:223-245). (3) Token-level caching awareness: the TokenUsage struct tracks cachedInputTokens and cacheCreationInputTokens, and these values are extracted from provider responses and propagated into dynamic metadata for cost tracking (internal/metrics/metrics.go:151-152). The ExtractTokenUsageFromExplicitCaching function normalizes Anthropic/AWS Bedrock explicit caching token counts into the unified format (internal/metrics/metrics.go:298-313). (4) Header/body mutation: routes and backends can configure header and body field removal/modification (e.g., stripping sensitive headers before forwarding to upstream) (internal/controller/gateway.go:287-379). These mutations are not automated guardrails — they are declarative per-route configurations.

← How are rate limits, budgets and cost tracking implemented? · How is it observed, deployed and scaled? →