# Portkey-AI/gateway

> Stateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/Portkey-AI/gateway (reviewed at commit `669825cbe89ee51569918b8f78a9db486fd69dd4`, 2026-05-25)
- Stars: 13136 · Language: TypeScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/portkey-gateway/

## Overview

Portkey's AI Gateway is the open-source core of Portkey's hosted LLM gateway. It is a small, stateless TypeScript service built on [Hono](https://hono.dev). It accepts OpenAI-shaped requests (and Anthropic `/v1/messages`), translates them for one of 72 registered providers, and returns OpenAI-shaped responses, streaming included. The same code runs as a Cloudflare Worker or as a Node.js server (`@portkey-ai/gateway` on npm, Docker on port 8787).

The main idea is the per-request **config**. A client sends a JSON routing tree in the `x-portkey-config` header (or plain `x-portkey-provider` plus `Authorization`). That tree can nest fallback, weighted load-balance and conditional nodes, retries, timeouts, caching and guardrail hooks. The gateway stores no provider keys and has no users. Credentials arrive with every request and leave with it.

This makes the open-source gateway a good stateless translation and routing layer. It is not a full self-hosted control plane. Budgets, virtual-key resolution, circuit-breaker state and semantic caching are hook points that the hosted product fills in. In this repository they are empty unless you write the hook yourself.

## Architecture

```mermaid
flowchart LR
  C["Client (OpenAI SDK)"] --> M["Hono middlewares: log, hooks, cache"]
  M --> RV["requestValidator"]
  RV --> H["chatCompletionsHandler"]
  H --> CFG["constructConfigFromRequestHeaders"]
  CFG --> TT["tryTargetsRecursively"]
  TT --> TP["tryPost"]
  TP --> BH["before-request hooks (guardrails)"]
  TP --> TR["transformToProviderRequest"]
  TR --> RR["retryRequest -> provider fetch"]
  RR --> AH["after-request hooks"]
  AH --> RT["responseTransforms -> OpenAI shape"]
```

| Component | Path | Role |
|---|---|---|
| App and routes | `src/index.ts` | Hono app; middleware order; `/v1/*` endpoints and a generic `/v1/*` proxy route |
| Node entry | `src/start-server.ts` | `@hono/node-server`, WebSocket realtime, admin UI, `/log/stream` SSE |
| Router core | `src/handlers/handlerUtils.ts` | Header-to-config parsing, recursive strategy evaluation, `tryPost` pipeline, hook handlers |
| Retries | `src/handlers/retryHandler.ts` | `async-retry` wrapper with `retry-after` support |
| Conditional routing | `src/services/conditionalRouter.ts` | Mongo-style query operators over metadata, params and URL |
| Request mapping | `src/services/transformToProviderRequest.ts` | Maps gateway params to provider params from each provider's config |
| Providers | `src/providers/<name>/` | Per-provider `api` (URL, headers), param configs and response transforms |
| Hooks and plugins | `src/middlewares/hooks/`, `plugins/` | Guardrail and mutator execution; 20+ plugin packs compiled in from `conf.json` |
| Cache | `src/middlewares/cache/`, `src/shared/services/cache/` | In-memory exact cache; pluggable Redis/file/KV backend used mainly for provider token caching |

## How a request flows

Take `POST /v1/chat/completions` with an `x-portkey-config` header that defines a fallback between two targets:

1. **Middlewares.** On Node, `logHandler` records the request. `hooks` creates a `HooksManager`. `memoryCache` is mounted only when `conf.cache` is true ([index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/index.ts#L94-L110)).
2. **Config.** `chatCompletionsHandler` calls `constructConfigFromRequestHeaders`, then `tryTargetsRecursively` with endpoint `chatComplete` ([chatCompletionsHandler.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/chatCompletionsHandler.ts#L16-L31)). The config header is parsed as JSON and merged with provider-specific headers (Azure, Bedrock, Vertex and others) ([handlerUtils.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L1026-L1040)).
3. **Strategy.** `tryTargetsRecursively` merges inherited settings (override params, retry, cache, guardrails) into each node and switches on `strategy.mode`. `fallback` iterates targets until one succeeds or `onStatusCodes` no longer matches. `loadbalance` draws a weighted random target. `conditional` asks `ConditionalRouter`. A leaf calls `tryPost` ([handlerUtils.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L662-L833)).
4. **Before hooks.** `tryPost` builds a `RequestContext` and `ProviderContext`, resolves the provider URL, and runs before-request hooks. A failing guardrail with `deny: true` returns HTTP 446 ([handlerUtils.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L288-L345)).
5. **Transform and cache.** The body is mapped with `transformToProviderRequest`, then the cache is checked, then the optional pre-request validator runs ([handlerUtils.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L362-L440)).
6. **Call and retry.** `recursiveAfterRequestHookHandler` calls `retryRequest`. It retries on 429/500/502/503/504 by default. With `followProviderRetry` it honours `retry-after` headers within a 60-second budget ([retryHandler.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/retryHandler.ts#L84-L150), [globals.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/globals.ts#L5-L39)). After-request hooks can then fail the response and trigger another attempt.
7. **Respond.** The provider's `responseTransforms` entry maps the body or stream chunks back to OpenAI format, and the log object is emitted.

## Key components

### Provider modules

Each provider is a `ProviderConfigs` object: an `api` block (`getBaseURL`, `headers`, `getEndpoint`), a param map per endpoint, and `responseTransforms` keyed by endpoint and stream mode. Anthropic's module is typical ([anthropic/index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/providers/anthropic/index.ts#L19-L32)). `transformUsingProviderConfig` walks the param map and applies renames, defaults, min/max clamps and transforms ([transformToProviderRequest.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/services/transformToProviderRequest.ts#L75-L100)). The registry has 72 entries ([providers/index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/providers/index.ts#L78-L153)).

### Conditional router

`ConditionalRouter.resolveTarget` walks `strategy.conditions` in order and returns the first target whose query matches. Queries support `$eq`, `$regex`, `$in`, `$and`, `$or` and similar operators over `metadata.*`, `params.*` and `url.pathname`. If nothing matches, it uses `strategy.default` ([conditionalRouter.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/services/conditionalRouter.ts#L32-L80)).

### Guardrails as hooks

Guardrails and mutators are hooks with `checks`. `HooksManager.executeHooks` runs them and records verdicts ([hooks/index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/hooks/index.ts#L202-L270)). Check functions come from `plugins/`. `plugins/build.ts` reads `conf.plugins_enabled` and generates `plugins/index.ts` with static imports ([build.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/plugins/build.ts#L1-L40)). The `default` pack (regex, JSON schema, word count, allowed models and others) runs locally. Most other packs (Aporia, Pangea, Patronus, Azure, Bedrock and more) call a vendor API.

### Cache

The open-source cache is exact match only. It hashes the transformed body and URL with SHA-256 and keeps the result in a process-local object with an optional max age ([cache/index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/cache/index.ts#L1-L113)). Streams are never cached. The write condition `requestParams.stream === (false || undefined)` reduces to `=== undefined`, so a request that sends `stream: false` explicitly is not stored either. `SEMANTIC_HIT` exists only as a status label.

### Hook points left for the host

Three request-path features read functions from the Hono context that nothing in this repository sets: `preRequestValidator` (budgets and per-key limits, [preRequestValidatorService.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/preRequestValidatorService.ts#L6-L25)), `handleCircuitBreakerResponse`, and the `isOpen` flag that the circuit-breaker filter reads ([handlerUtils.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L646-L656)). A Redis token-bucket class exists but is not wired in. `initializeSettings.ts` can turn `conf.json` integrations with `rate_limits` into virtual-key objects, but no source file imports it.

## Extending it

- **New provider.** Add `src/providers/<name>/` with an `api` config, param maps and response transforms, then register it in `providers/index.ts`.
- **New guardrail.** Add a plugin folder with `manifest.json` and handler files, list it in `conf.json`, and run `npm run build-plugins`.
- **Routing.** Everything else is config: nested `targets`, `strategy`, `retry`, `cache`, `override_params`, `input_guardrails`/`output_guardrails` and mutators.
- **Host hooks.** Embedders can set `preRequestValidator`, `getFromCache` and circuit-breaker callbacks on the Hono context to add budgets, shared caching or breaker state.

## Running it

- `npx @portkey-ai/gateway` or `npm run dev:node` starts the Node server on port 8787 (`--port=` changes it, `--headless` drops the UI) ([start-server.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/start-server.ts#L18-L40)). The Dockerfile runs `npm run start:node` on the same port.
- `npm run dev:workerd` and `npm run deploy` target Cloudflare Workers through `wrangler.toml`.
- The admin UI and `/log/stream` need `admin_token` in `conf.json` ([adminAuth/index.ts](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/adminAuth/index.ts#L1-L20)). Set `REDIS_CONNECTION_STRING` to back the shared cache service with Redis.
- Requirements: Node.js, or a Workers account. No database.

## Strengths and caveats

- **Strength: small and stateless.** One Hono app, no database, the same code on Workers and Node. It scales horizontally because there is nothing to share.
- **Strength: expressive per-request routing.** Nested fallback, weighted and conditional trees, with retries and guardrails inherited down the tree, cover most multi-provider patterns without gateway redeploys.
- **Strength: guardrail ecosystem.** About 20 vendor packs plus local checks, all on one hook interface with deny or mutate semantics.
- **Caveat: no auth on inference routes.** `/v1/*` does not authenticate callers. Anyone who can reach the gateway can use it with their own provider key. You need a network boundary or an auth proxy in front.
- **Caveat: hosted-only pieces.** Budgets, virtual keys, rate limits, circuit-breaker state and semantic cache are hook points or labels here, not features. The `conf.json` integration block is not loaded by any source file.
- **Caveat: process-local cache.** The exact cache lives in one process's memory and has the explicit `stream: false` quirk described above.
- **Caveat: observability is thin.** Logs go to `console` and an SSE stream for the local UI. There is no metrics or OTLP exporter.

*Sources: code at 669825c, deepwiki-open wiki (12 pages), verified Q&A.*

## How Portkey-AI/gateway answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

Requests are routed through the `tryTargetsRecursively` function in `src/handlers/handlerUtils.ts:476` which evaluates a JSON config (passed via `x-portkey-config` header or the request body). This config supports five strategy modes: **Single** — pick the first target; **Fallback** (`strategy.mode: "fallback"`) — iterate targets sequentially, stopping on the first success (or matching `onStatusCodes`); **Loadbalance** (`mode: "loadbalance"`) — weighted random selection via `selectProviderByWeight` at `handlerUtils.ts:204`, where each target has a numeric `weight` property (defaulting to 1); **Conditional** (`mode: "conditional"`) — a `ConditionalRouter` class at `src/services/conditionalRouter.ts:32` matches conditions against request metadata, URL path, and params using operators like `$eq`, `$regex`, `$in`, `$and`, `$or`, routing to a named target; **Circuit Breaker** — targets with `isOpen` flag are filtered out before selection (`handlerUtils.ts:648-658`).

Retries are handled by `retryRequest` in `src/handlers/retryHandler.ts:65`, which uses the `async-retry` library to retry on configurable `onStatusCodes` (default: 429, 500, 502, 503, 504 per `src/globals.ts:38`). It respects provider `retry-after` headers when `followProviderRetry` is set, with a 60-second max retry budget (`MAX_RETRY_LIMIT_MS`, `globals.ts:5`). Request timeout is configurable per-target via `requestTimeout`.

The gateway does not implement latency-based or cost-based dynamic routing — weights are static. Model aliasing is not built into the router; each provider option specifies its model directly in `overrideParams`. Health checks are passive via the circuit-breaker mechanism (`currentTarget.targets.filter(t => !t.isOpen)` at `handlerUtils.ts:648-653`), which tracks per-target open/close state.

> **Editor's note.** Correction: the circuit-breaker path is a hook point, not a working feature in this repository. `tryTargetsRecursively` filters targets on `isOpen` and calls `c.get('handleCircuitBreakerResponse')`, but nothing in the open-source code sets either, so no target is ever marked open unless an embedding host supplies that callback.

Citations: [src/handlers/handlerUtils.ts:476-834](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L476-L834) · [src/handlers/handlerUtils.ts:204-231](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L204-L231) · [src/services/conditionalRouter.ts:32-156](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/services/conditionalRouter.ts#L32-L156) · [src/handlers/retryHandler.ts:65-220](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/retryHandler.ts#L65-L220) · [src/globals.ts:5-39](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/globals.ts#L5-L39)

### How are different provider APIs unified? (answered)

All providers present an **OpenAI-compatible schema**. Requests arrive in OpenAI format (or Anthropic format at `/v1/messages`). Each provider is a module under `src/providers/` — there are 78 registered providers in `src/providers/index.ts:78-151`, from `openai` and `anthropic` to `ovhcloud`. Each module exports a `ProviderConfigs` object containing parameter configs (`complete`, `chatComplete`, `embed` etc.), an `api` object with `getBaseURL`, `headers`, and `getEndpoint`, and optional `requestHandlers`/`responseTransforms`.

**Request translation** happens in `src/services/transformToProviderRequest.ts:75` via the `transformUsingProviderConfig` function. Each provider's `chatComplete` config (e.g. `src/providers/openai/chatComplete.ts:12`) maps gateway parameter names to provider-specific parameter names (the `param` field). It applies transforms, defaults, min/max clamping, and override merging from `overrideParams`. For example, OpenAI's `max_tokens` maps to `"max_tokens"` with a default of 100, while Anthropic might map it differently via its own config. The `RequestContext` class at `src/handlers/services/requestContext.ts:211` calls `transformToProviderRequest` to produce the provider-specific payload.

**Response translation** flows through `responseHandler` in `src/handlers/responseHandlers.ts:38`, which selects a `responseTransforms` function from the provider config. For streaming, it uses functions like `OpenAIChatCompleteJSONToStreamResponseTransform` (`openai/chatComplete.ts:157`) to break a complete JSON response into `data: ...` SSE chunks with token-level deltas including tool calls, content blocks (text, thinking, data), and citations. Non-OpenAI providers (Anthropic, Bedrock, Vertex AI) each have their own response transforms that normalize to the OpenAI streaming format.

**Streaming** is always normalized to OpenAI SSE format (`data: [DONE]` termination, `chat.completion.chunk` object). **Tool calls** and **multimodal** (image content blocks, audio) are supported through the stream chunk transforms, with tool calls split into name and argument chunks per the OpenAI convention.


Citations: [src/providers/index.ts:78-153](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/providers/index.ts#L78-L153) · [src/services/transformToProviderRequest.ts:75-80](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/services/transformToProviderRequest.ts#L75-L80) · [src/providers/openai/chatComplete.ts:12-133](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/providers/openai/chatComplete.ts#L12-L133) · [src/handlers/services/requestContext.ts:211-224](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/requestContext.ts#L211-L224) · [src/handlers/responseHandlers.ts:38-80](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/responseHandlers.ts#L38-L80)

### How are API keys, users and tenants managed? (answered)

**Upstream credential storage** is handled entirely via request headers, not a server-side secrets store. API keys are passed through the `Authorization: Bearer <key>` header or an `x-portkey-api-key` header (`src/globals.ts:14`). The `constructConfigFromRequestHeaders` function at `src/handlers/handlerUtils.ts:836` extracts the key and spreads it into the provider options object. For provider-specific authentication (Azure, AWS/Bedrock, GCP Vertex, Oracle), additional headers like `x-portkey-azure-resource-name`, `x-portkey-aws-access-key-id`, `x-portkey-vertex-service-account-json` etc. are parsed and merged into the config at lines 839-1179.

**Virtual keys** are referenced via the `x-portkey-virtual-key` header or `virtualKey` field in the config (`src/types/requestBody.ts:49`, `src/globals.ts:27`). They are passed through to the provider config but the repository itself does not store or manage virtual key mappings — they are resolved externally by the Portkey SaaS or a pre-request hook.

**Users and teams** have no built-in model in the gateway code. The `metadata` field (`x-portkey-metadata` header, `src/handlers/services/requestContext.ts:110-116`) can carry user/team identifiers as arbitrary JSON, which is available to hooks and the pre-request validator for budget/rate-limit decisions.

**Authentication for the local admin UI** is handled by `src/middlewares/adminAuth/index.ts:80`. It uses a static `admin_token` from `conf.json`, supporting both Bearer token auth and cookie-based sessions (12-hour expiry, `SESSION_MAX_AGE_SECONDS` at line 5). The login endpoint at `/public/auth` (in `src/start-server.ts:59`) issues an HttpOnly `portkey_admin_session` cookie. This protects `/log/stream` (SSE log streaming) and the UI served at `/public/`.

There is no multi-tenant isolation, user directory, or API-key-based user authentication in the open-source gateway — those are Portkey SaaS features.


Citations: [src/handlers/handlerUtils.ts:836-1179](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/handlerUtils.ts#L836-L1179) · [src/middlewares/adminAuth/index.ts:80-161](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/adminAuth/index.ts#L80-L161) · [src/globals.ts:13-28](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/globals.ts#L13-L28) · [src/types/requestBody.ts:46-51](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/types/requestBody.ts#L46-L51)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limiting** is implemented server-side only when Redis is available, using a token-bucket algorithm in Lua via `src/shared/services/cache/utils/rateLimiter.ts:84`. The `RedisRateLimiter` class takes a `capacity`, `windowSize`, and `key`, and runs a Lua script (`RATE_LIMIT_LUA` at line 5) atomically to decide if a request is allowed. It supports both consume and check-only modes. However, this rate limiter is not wired into the request path by default — it's a utility available for use.

**Per-key/user limits** are delegated to the `preRequestValidator` hook. The `PreRequestValidatorService` at `src/handlers/services/preRequestValidatorService.ts:6` calls `honoContext.get('preRequestValidator')`, which is an externally injected function (from Portkey SaaS, not defined in the open-source core). It returns an optional `response` (to deny the request) and `modelPricingConfig` (to track cost).

**Price tables and cost tracking** are stored per-provider as `modelPricingConfig` (an optional `Record<string, any>` on the provider option, see `src/types/requestBody.ts:179`). This is set by the pre-request validator and attached to the log object (`src/handlers/services/logsService.ts:288-293`). There are no built-in price tables in the gateway code.

**Usage accounting** happens via the logging pipeline. Every request produces a `LogObject` through `LogObjectBuilder` in `src/handlers/services/logsService.ts:215`, capturing `providerOptions`, `transformedRequest`, `requestParams`, `originalResponse` (including usage tokens from provider), `executionTime`, and `modelPricingConfig`. These logs are stored in-memory on the Hono context (`c.get('requestOptions')`) and broadcast via SSE to connected clients at `/log/stream` by `src/middlewares/log/index.ts:80-110`. The logs are not persisted to a database by the gateway — they are ephemeral in-memory structures.

In summary: cost infrastructure exists (modelPricingConfig, usage capture in log objects) but is passive — recording rather than enforcing budgets. Enforcement requires an external pre-request validator.


Citations: [src/shared/services/cache/utils/rateLimiter.ts:5-188](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/shared/services/cache/utils/rateLimiter.ts#L5-L188) · [src/handlers/services/preRequestValidatorService.ts:6-34](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/preRequestValidatorService.ts#L6-L34) · [src/handlers/services/logsService.ts:215-360](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/logsService.ts#L215-L360) · [src/middlewares/log/index.ts:80-166](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/log/index.ts#L80-L166)

### How are caching and guardrails implemented? (answered)

**Exact caching** works through two layers. The middleware layer in `src/middlewares/cache/index.ts` intercepts responses after they return from providers. It creates a SHA-256 key from `JSON.stringify(requestBody) + url` (`getCacheKey` at line 14) and stores the response body in an in-memory dict with an optional `maxAge` expiry. The `memoryCache` middleware is conditionally enabled by `conf.cache === true` (`src/index.ts:108`). In `tryPost` (`handlerUtils.ts:372-404`), a `CacheService` (`src/handlers/services/cacheService.ts:16`) checks the cache before making the provider call, passing the request headers and body to `getFromCache`, which returns statuses like `HIT`, `MISS`, `REFRESH`, or `DISABLED`. Force-refresh is triggered by the `x-portkey-cache-force-refresh` header (`cache/index.ts:38`). The second layer is the pluggable cache backed by Redis, file, memory, or Cloudflare KV (`src/shared/services/cache/index.ts:42`), with configurable TTL presets from 1 minute to 30 days (line 29-40). This is initialized in `src/index.ts:49-50` when `REDIS_CONNECTION_STRING` is set.

**Semantic caching** is not implemented in this repository — the `CACHE_STATUS` enum includes `SEMANTIC_HIT` and `SEMANTIC_MISS` labels (`cache/index.ts:6-12`) but no semantic matching logic exists. Cache mode options include `"simple"` but no semantic mode.

**Guardrails** are implemented as a hooks/plugin system. The `HooksManager` in `src/middlewares/hooks/index.ts:202` manages `beforeRequestHooks` and `afterRequestHooks`. Plugins are registered in `plugins/index.ts:71-183` across 18 plugin sources (default, portkey, qualifire, pillar, patronus, bedrock, acuvity, etc.) with ~50 guard check functions. Checks include PII redaction (`qualifire/pii`, `portkey/pii`), content moderation (`qualifire/contentModeration`, `portkey/moderateContent`), prompt injection detection (`qualifire/promptInjections`), hallucinations (`qualifire/hallucinations`), and jailbreak detection (`promptsecurity/protectPrompt`). Guardrails execute as `Check` objects within `HookObject` groups (`src/middlewares/hooks/types.ts:1-27`), supporting `deny`, `sequential`, `async` execution modes. When a verdict is false and `deny` is true, the request is blocked with a 446 status code (`handlerUtils.ts:1337-1341`). The plugin system also supports mutators (request/response transformation) like `addPrefix`, `regexReplace`.

The hook system is configuration-driven: guardrails are specified in the config JSON as `inputGuardrails`/`outputGuardrails` arrays and converted to hook objects via `convertHooksShorthand` at `handlerUtils.ts:238-275`.

> **Editor's note.** Correction: the request-path response cache is only the in-memory `memoryCache` middleware (enabled by `conf.cache`). The Redis/file/KV backends behind `createCacheBackendsRedis` belong to a separate shared cache service that `getFromCache` does not read. Note also that responses to requests with an explicit `stream: false` are not stored, because the write condition reduces to `stream === undefined`.

Citations: [src/middlewares/cache/index.ts:1-113](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/cache/index.ts#L1-L113) · [src/handlers/services/cacheService.ts:16-137](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/cacheService.ts#L16-L137) · [plugins/index.ts:71-183](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/plugins/index.ts#L71-L183) · [src/middlewares/hooks/index.ts:202-556](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/hooks/index.ts#L202-L556)

### How is it observed, deployed and scaled? (answered)

**Logging** is in-memory and SSE-streaming. The `logHandler` middleware (`src/middlewares/log/index.ts:149`) captures request/response data into a `requestOptions` array on the Hono context. After each request, `processLog` (line 112) sanitizes sensitive headers and broadcasts the log entry to all connected SSE clients at `/log/stream`. The log stream is protected by `adminAuthMiddleware` and includes timing, provider, status, and truncated response bodies (capped at 100KB, line 6).

**Metrics and traces** are minimal. The `LogsService` at `src/handlers/services/logsService.ts:96-162` can construct OTLP-format span objects for tool execution events (with `traceId`, `spanId`, `parentSpanId`, and `gen_ai.*` attributes), but these are not exported to any OTel collector by the gateway — the structure is defined but the exporter is not implemented. The logger is plain `console` (`src/apm/index.ts:1`).

**Performance architecture**: Built on Hono (ultra-fast web framework designed for edge), written in TypeScript. **Concurrency** is handled by Hono's async middleware pipeline and Cloudflare Workers' `waitUntil` for fire-and-forget logging. There is no per-request thread pool — on Node.js it uses the async event loop, and on Workerd it uses the Workers runtime. Hono's sub-millisecond routing overhead is a key design goal (README claim, but the architecture supports it).

**Deployment modes**: Three runtimes are supported: **Cloudflare Workers** (`npm run dev` on wrangler), **Node.js** (`npm run dev:node` or `npm run start:node`), and any Hono-compatible runtime (Bun, Deno via adapter). The `getRuntimeKey()` check in `src/index.ts` and `src/middlewares/log/index.ts` switches behavior per runtime. The Node.js server (`src/start-server.ts`) uses `@hono/node-server` and supports WebSocket upgrades for `/v1/realtime`, SSE log streaming, and a local admin UI at `/public/` that can be hidden with `--headless`. A `conf.json` file controls plugins, caching, and admin credentials.

**High Availability**: The open-source gateway itself has no built-in HA — deployments rely on the underlying platform (Cloudflare Workers' global network, Node.js process manager, Docker orchestrator). Redis provides a shared cache for multi-instance deployments.


Citations: [src/middlewares/log/index.ts:80-166](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/log/index.ts#L80-L166) · [src/handlers/services/logsService.ts:96-162](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/logsService.ts#L96-L162) · [src/index.ts:47-51](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/index.ts#L47-L51) · [src/start-server.ts:1-40](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/start-server.ts#L1-L40) · [src/apm/index.ts:1-1](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/apm/index.ts#L1-L1)
