LLMs Technical Reviews
Home / LLM gateways / gateway

Portkey-AI/gateway

Stateless Hono gateway that routes OpenAI-format calls to 72 providers using per-request fallback, load-balance and guardrail configs.

GitHub ↗★ 13kTypeScriptMITcommit 669825c · 2026-05-25homepage ↗

Overview

Portkey’s AI Gateway is the open-source core of Portkey’s hosted LLM gateway. It is a small, stateless TypeScript service built on Hono. It accepts OpenAI-shaped requests (and Anthropic /v1/messages), translates them for one of 72 registered providers, and returns OpenAI-shaped responses, streaming included. The same code runs as a Cloudflare Worker or as a Node.js server (@portkey-ai/gateway on npm, Docker on port 8787).

The main idea is the per-request config. A client sends a JSON routing tree in the x-portkey-config header (or plain x-portkey-provider plus Authorization). That tree can nest fallback, weighted load-balance and conditional nodes, retries, timeouts, caching and guardrail hooks. The gateway stores no provider keys and has no users. Credentials arrive with every request and leave with it.

This makes the open-source gateway a good stateless translation and routing layer. It is not a full self-hosted control plane. Budgets, virtual-key resolution, circuit-breaker state and semantic caching are hook points that the hosted product fills in. In this repository they are empty unless you write the hook yourself.

Architecture

flowchart LR
  C["Client (OpenAI SDK)"] --> M["Hono middlewares: log, hooks, cache"]
  M --> RV["requestValidator"]
  RV --> H["chatCompletionsHandler"]
  H --> CFG["constructConfigFromRequestHeaders"]
  CFG --> TT["tryTargetsRecursively"]
  TT --> TP["tryPost"]
  TP --> BH["before-request hooks (guardrails)"]
  TP --> TR["transformToProviderRequest"]
  TR --> RR["retryRequest -> provider fetch"]
  RR --> AH["after-request hooks"]
  AH --> RT["responseTransforms -> OpenAI shape"]
Component Path Role
App and routes src/index.ts Hono app; middleware order; /v1/* endpoints and a generic /v1/* proxy route
Node entry src/start-server.ts @hono/node-server, WebSocket realtime, admin UI, /log/stream SSE
Router core src/handlers/handlerUtils.ts Header-to-config parsing, recursive strategy evaluation, tryPost pipeline, hook handlers
Retries src/handlers/retryHandler.ts async-retry wrapper with retry-after support
Conditional routing src/services/conditionalRouter.ts Mongo-style query operators over metadata, params and URL
Request mapping src/services/transformToProviderRequest.ts Maps gateway params to provider params from each provider’s config
Providers src/providers/<name>/ Per-provider api (URL, headers), param configs and response transforms
Hooks and plugins src/middlewares/hooks/, plugins/ Guardrail and mutator execution; 20+ plugin packs compiled in from conf.json
Cache src/middlewares/cache/, src/shared/services/cache/ In-memory exact cache; pluggable Redis/file/KV backend used mainly for provider token caching

How a request flows

Take POST /v1/chat/completions with an x-portkey-config header that defines a fallback between two targets:

  1. Middlewares. On Node, logHandler records the request. hooks creates a HooksManager. memoryCache is mounted only when conf.cache is true (index.ts).
  2. Config. chatCompletionsHandler calls constructConfigFromRequestHeaders, then tryTargetsRecursively with endpoint chatComplete (chatCompletionsHandler.ts). The config header is parsed as JSON and merged with provider-specific headers (Azure, Bedrock, Vertex and others) (handlerUtils.ts).
  3. Strategy. tryTargetsRecursively merges inherited settings (override params, retry, cache, guardrails) into each node and switches on strategy.mode. fallback iterates targets until one succeeds or onStatusCodes no longer matches. loadbalance draws a weighted random target. conditional asks ConditionalRouter. A leaf calls tryPost (handlerUtils.ts).
  4. Before hooks. tryPost builds a RequestContext and ProviderContext, resolves the provider URL, and runs before-request hooks. A failing guardrail with deny: true returns HTTP 446 (handlerUtils.ts).
  5. Transform and cache. The body is mapped with transformToProviderRequest, then the cache is checked, then the optional pre-request validator runs (handlerUtils.ts).
  6. Call and retry. recursiveAfterRequestHookHandler calls retryRequest. It retries on 429/500/502/503/504 by default. With followProviderRetry it honours retry-after headers within a 60-second budget (retryHandler.ts, globals.ts). After-request hooks can then fail the response and trigger another attempt.
  7. Respond. The provider’s responseTransforms entry maps the body or stream chunks back to OpenAI format, and the log object is emitted.

Key components

Provider modules

Each provider is a ProviderConfigs object: an api block (getBaseURL, headers, getEndpoint), a param map per endpoint, and responseTransforms keyed by endpoint and stream mode. Anthropic’s module is typical (anthropic/index.ts). transformUsingProviderConfig walks the param map and applies renames, defaults, min/max clamps and transforms (transformToProviderRequest.ts). The registry has 72 entries (providers/index.ts).

Conditional router

ConditionalRouter.resolveTarget walks strategy.conditions in order and returns the first target whose query matches. Queries support $eq, $regex, $in, $and, $or and similar operators over metadata.*, params.* and url.pathname. If nothing matches, it uses strategy.default (conditionalRouter.ts).

Guardrails as hooks

Guardrails and mutators are hooks with checks. HooksManager.executeHooks runs them and records verdicts (hooks/index.ts). Check functions come from plugins/. plugins/build.ts reads conf.plugins_enabled and generates plugins/index.ts with static imports (build.ts). The default pack (regex, JSON schema, word count, allowed models and others) runs locally. Most other packs (Aporia, Pangea, Patronus, Azure, Bedrock and more) call a vendor API.

Cache

The open-source cache is exact match only. It hashes the transformed body and URL with SHA-256 and keeps the result in a process-local object with an optional max age (cache/index.ts). Streams are never cached. The write condition requestParams.stream === (false || undefined) reduces to === undefined, so a request that sends stream: false explicitly is not stored either. SEMANTIC_HIT exists only as a status label.

Hook points left for the host

Three request-path features read functions from the Hono context that nothing in this repository sets: preRequestValidator (budgets and per-key limits, preRequestValidatorService.ts), handleCircuitBreakerResponse, and the isOpen flag that the circuit-breaker filter reads (handlerUtils.ts). A Redis token-bucket class exists but is not wired in. initializeSettings.ts can turn conf.json integrations with rate_limits into virtual-key objects, but no source file imports it.

Extending it

  • New provider. Add src/providers/<name>/ with an api config, param maps and response transforms, then register it in providers/index.ts.
  • New guardrail. Add a plugin folder with manifest.json and handler files, list it in conf.json, and run npm run build-plugins.
  • Routing. Everything else is config: nested targets, strategy, retry, cache, override_params, input_guardrails/output_guardrails and mutators.
  • Host hooks. Embedders can set preRequestValidator, getFromCache and circuit-breaker callbacks on the Hono context to add budgets, shared caching or breaker state.

Running it

  • npx @portkey-ai/gateway or npm run dev:node starts the Node server on port 8787 (--port= changes it, --headless drops the UI) (start-server.ts). The Dockerfile runs npm run start:node on the same port.
  • npm run dev:workerd and npm run deploy target Cloudflare Workers through wrangler.toml.
  • The admin UI and /log/stream need admin_token in conf.json (adminAuth/index.ts). Set REDIS_CONNECTION_STRING to back the shared cache service with Redis.
  • Requirements: Node.js, or a Workers account. No database.

Strengths and caveats

  • Strength: small and stateless. One Hono app, no database, the same code on Workers and Node. It scales horizontally because there is nothing to share.
  • Strength: expressive per-request routing. Nested fallback, weighted and conditional trees, with retries and guardrails inherited down the tree, cover most multi-provider patterns without gateway redeploys.
  • Strength: guardrail ecosystem. About 20 vendor packs plus local checks, all on one hook interface with deny or mutate semantics.
  • Caveat: no auth on inference routes. /v1/* does not authenticate callers. Anyone who can reach the gateway can use it with their own provider key. You need a network boundary or an auth proxy in front.
  • Caveat: hosted-only pieces. Budgets, virtual keys, rate limits, circuit-breaker state and semantic cache are hook points or labels here, not features. The conf.json integration block is not loaded by any source file.
  • Caveat: process-local cache. The exact cache lives in one process’s memory and has the explicit stream: false quirk described above.
  • Caveat: observability is thin. Logs go to console and an SSE stream for the local UI. There is no metrics or OTLP exporter.

Sources: code at 669825c, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit 669825c. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Requests are routed through the tryTargetsRecursively function in src/handlers/handlerUtils.ts:476 which evaluates a JSON config (passed via x-portkey-config header or the request body). This config supports five strategy modes: Single — pick the first target; Fallback (strategy.mode: "fallback") — iterate targets sequentially, stopping on the first success (or matching onStatusCodes); Loadbalance (mode: "loadbalance") — weighted random selection via selectProviderByWeight at handlerUtils.ts:204, where each target has a numeric weight property (defaulting to 1); Conditional (mode: "conditional") — a ConditionalRouter class at src/services/conditionalRouter.ts:32 matches conditions against request metadata, URL path, and params using operators like $eq, $regex, $in, $and, $or, routing to a named target; Circuit Breaker — targets with isOpen flag are filtered out before selection (handlerUtils.ts:648-658).

Retries are handled by retryRequest in src/handlers/retryHandler.ts:65, which uses the async-retry library to retry on configurable onStatusCodes (default: 429, 500, 502, 503, 504 per src/globals.ts:38). It respects provider retry-after headers when followProviderRetry is set, with a 60-second max retry budget (MAX_RETRY_LIMIT_MS, globals.ts:5). Request timeout is configurable per-target via requestTimeout.

The gateway does not implement latency-based or cost-based dynamic routing — weights are static. Model aliasing is not built into the router; each provider option specifies its model directly in overrideParams. Health checks are passive via the circuit-breaker mechanism (currentTarget.targets.filter(t => !t.isOpen) at handlerUtils.ts:648-653), which tracks per-target open/close state.

Editor's note. Correction: the circuit-breaker path is a hook point, not a working feature in this repository. tryTargetsRecursively filters targets on isOpen and calls c.get('handleCircuitBreakerResponse'), but nothing in the open-source code sets either, so no target is ever marked open unless an embedding host supplies that callback.

How are different provider APIs unified?

answered

All providers present an OpenAI-compatible schema. Requests arrive in OpenAI format (or Anthropic format at /v1/messages). Each provider is a module under src/providers/ — there are 78 registered providers in src/providers/index.ts:78-151, from openai and anthropic to ovhcloud. Each module exports a ProviderConfigs object containing parameter configs (complete, chatComplete, embed etc.), an api object with getBaseURL, headers, and getEndpoint, and optional requestHandlers/responseTransforms.

Request translation happens in src/services/transformToProviderRequest.ts:75 via the transformUsingProviderConfig function. Each provider's chatComplete config (e.g. src/providers/openai/chatComplete.ts:12) maps gateway parameter names to provider-specific parameter names (the param field). It applies transforms, defaults, min/max clamping, and override merging from overrideParams. For example, OpenAI's max_tokens maps to "max_tokens" with a default of 100, while Anthropic might map it differently via its own config. The RequestContext class at src/handlers/services/requestContext.ts:211 calls transformToProviderRequest to produce the provider-specific payload.

Response translation flows through responseHandler in src/handlers/responseHandlers.ts:38, which selects a responseTransforms function from the provider config. For streaming, it uses functions like OpenAIChatCompleteJSONToStreamResponseTransform (openai/chatComplete.ts:157) to break a complete JSON response into data: ... SSE chunks with token-level deltas including tool calls, content blocks (text, thinking, data), and citations. Non-OpenAI providers (Anthropic, Bedrock, Vertex AI) each have their own response transforms that normalize to the OpenAI streaming format.

Streaming is always normalized to OpenAI SSE format (data: [DONE] termination, chat.completion.chunk object). Tool calls and multimodal (image content blocks, audio) are supported through the stream chunk transforms, with tool calls split into name and argument chunks per the OpenAI convention.

How are API keys, users and tenants managed?

answered

Upstream credential storage is handled entirely via request headers, not a server-side secrets store. API keys are passed through the Authorization: Bearer <key> header or an x-portkey-api-key header (src/globals.ts:14). The constructConfigFromRequestHeaders function at src/handlers/handlerUtils.ts:836 extracts the key and spreads it into the provider options object. For provider-specific authentication (Azure, AWS/Bedrock, GCP Vertex, Oracle), additional headers like x-portkey-azure-resource-name, x-portkey-aws-access-key-id, x-portkey-vertex-service-account-json etc. are parsed and merged into the config at lines 839-1179.

Virtual keys are referenced via the x-portkey-virtual-key header or virtualKey field in the config (src/types/requestBody.ts:49, src/globals.ts:27). They are passed through to the provider config but the repository itself does not store or manage virtual key mappings — they are resolved externally by the Portkey SaaS or a pre-request hook.

Users and teams have no built-in model in the gateway code. The metadata field (x-portkey-metadata header, src/handlers/services/requestContext.ts:110-116) can carry user/team identifiers as arbitrary JSON, which is available to hooks and the pre-request validator for budget/rate-limit decisions.

Authentication for the local admin UI is handled by src/middlewares/adminAuth/index.ts:80. It uses a static admin_token from conf.json, supporting both Bearer token auth and cookie-based sessions (12-hour expiry, SESSION_MAX_AGE_SECONDS at line 5). The login endpoint at /public/auth (in src/start-server.ts:59) issues an HttpOnly portkey_admin_session cookie. This protects /log/stream (SSE log streaming) and the UI served at /public/.

There is no multi-tenant isolation, user directory, or API-key-based user authentication in the open-source gateway — those are Portkey SaaS features.

How are rate limits, budgets and cost tracking implemented?

answered

Rate limiting is implemented server-side only when Redis is available, using a token-bucket algorithm in Lua via src/shared/services/cache/utils/rateLimiter.ts:84. The RedisRateLimiter class takes a capacity, windowSize, and key, and runs a Lua script (RATE_LIMIT_LUA at line 5) atomically to decide if a request is allowed. It supports both consume and check-only modes. However, this rate limiter is not wired into the request path by default — it's a utility available for use.

Per-key/user limits are delegated to the preRequestValidator hook. The PreRequestValidatorService at src/handlers/services/preRequestValidatorService.ts:6 calls honoContext.get('preRequestValidator'), which is an externally injected function (from Portkey SaaS, not defined in the open-source core). It returns an optional response (to deny the request) and modelPricingConfig (to track cost).

Price tables and cost tracking are stored per-provider as modelPricingConfig (an optional Record<string, any> on the provider option, see src/types/requestBody.ts:179). This is set by the pre-request validator and attached to the log object (src/handlers/services/logsService.ts:288-293). There are no built-in price tables in the gateway code.

Usage accounting happens via the logging pipeline. Every request produces a LogObject through LogObjectBuilder in src/handlers/services/logsService.ts:215, capturing providerOptions, transformedRequest, requestParams, originalResponse (including usage tokens from provider), executionTime, and modelPricingConfig. These logs are stored in-memory on the Hono context (c.get('requestOptions')) and broadcast via SSE to connected clients at /log/stream by src/middlewares/log/index.ts:80-110. The logs are not persisted to a database by the gateway — they are ephemeral in-memory structures.

In summary: cost infrastructure exists (modelPricingConfig, usage capture in log objects) but is passive — recording rather than enforcing budgets. Enforcement requires an external pre-request validator.

How are caching and guardrails implemented?

answered

Exact caching works through two layers. The middleware layer in src/middlewares/cache/index.ts intercepts responses after they return from providers. It creates a SHA-256 key from JSON.stringify(requestBody) + url (getCacheKey at line 14) and stores the response body in an in-memory dict with an optional maxAge expiry. The memoryCache middleware is conditionally enabled by conf.cache === true (src/index.ts:108). In tryPost (handlerUtils.ts:372-404), a CacheService (src/handlers/services/cacheService.ts:16) checks the cache before making the provider call, passing the request headers and body to getFromCache, which returns statuses like HIT, MISS, REFRESH, or DISABLED. Force-refresh is triggered by the x-portkey-cache-force-refresh header (cache/index.ts:38). The second layer is the pluggable cache backed by Redis, file, memory, or Cloudflare KV (src/shared/services/cache/index.ts:42), with configurable TTL presets from 1 minute to 30 days (line 29-40). This is initialized in src/index.ts:49-50 when REDIS_CONNECTION_STRING is set.

Semantic caching is not implemented in this repository — the CACHE_STATUS enum includes SEMANTIC_HIT and SEMANTIC_MISS labels (cache/index.ts:6-12) but no semantic matching logic exists. Cache mode options include "simple" but no semantic mode.

Guardrails are implemented as a hooks/plugin system. The HooksManager in src/middlewares/hooks/index.ts:202 manages beforeRequestHooks and afterRequestHooks. Plugins are registered in plugins/index.ts:71-183 across 18 plugin sources (default, portkey, qualifire, pillar, patronus, bedrock, acuvity, etc.) with ~50 guard check functions. Checks include PII redaction (qualifire/pii, portkey/pii), content moderation (qualifire/contentModeration, portkey/moderateContent), prompt injection detection (qualifire/promptInjections), hallucinations (qualifire/hallucinations), and jailbreak detection (promptsecurity/protectPrompt). Guardrails execute as Check objects within HookObject groups (src/middlewares/hooks/types.ts:1-27), supporting deny, sequential, async execution modes. When a verdict is false and deny is true, the request is blocked with a 446 status code (handlerUtils.ts:1337-1341). The plugin system also supports mutators (request/response transformation) like addPrefix, regexReplace.

The hook system is configuration-driven: guardrails are specified in the config JSON as inputGuardrails/outputGuardrails arrays and converted to hook objects via convertHooksShorthand at handlerUtils.ts:238-275.

Editor's note. Correction: the request-path response cache is only the in-memory memoryCache middleware (enabled by conf.cache). The Redis/file/KV backends behind createCacheBackendsRedis belong to a separate shared cache service that getFromCache does not read. Note also that responses to requests with an explicit stream: false are not stored, because the write condition reduces to stream === undefined.

How is it observed, deployed and scaled?

answered

Logging is in-memory and SSE-streaming. The logHandler middleware (src/middlewares/log/index.ts:149) captures request/response data into a requestOptions array on the Hono context. After each request, processLog (line 112) sanitizes sensitive headers and broadcasts the log entry to all connected SSE clients at /log/stream. The log stream is protected by adminAuthMiddleware and includes timing, provider, status, and truncated response bodies (capped at 100KB, line 6).

Metrics and traces are minimal. The LogsService at src/handlers/services/logsService.ts:96-162 can construct OTLP-format span objects for tool execution events (with traceId, spanId, parentSpanId, and gen_ai.* attributes), but these are not exported to any OTel collector by the gateway — the structure is defined but the exporter is not implemented. The logger is plain console (src/apm/index.ts:1).

Performance architecture: Built on Hono (ultra-fast web framework designed for edge), written in TypeScript. Concurrency is handled by Hono's async middleware pipeline and Cloudflare Workers' waitUntil for fire-and-forget logging. There is no per-request thread pool — on Node.js it uses the async event loop, and on Workerd it uses the Workers runtime. Hono's sub-millisecond routing overhead is a key design goal (README claim, but the architecture supports it).

Deployment modes: Three runtimes are supported: Cloudflare Workers (npm run dev on wrangler), Node.js (npm run dev:node or npm run start:node), and any Hono-compatible runtime (Bun, Deno via adapter). The getRuntimeKey() check in src/index.ts and src/middlewares/log/index.ts switches behavior per runtime. The Node.js server (src/start-server.ts) uses @hono/node-server and supports WebSocket upgrades for /v1/realtime, SSE log streaming, and a local admin UI at /public/ that can be hidden with --headless. A conf.json file controls plugins, caching, and admin credentials.

High Availability: The open-source gateway itself has no built-in HA — deployments rely on the underlying platform (Cloudflare Workers' global network, Node.js process manager, Docker orchestrator). Redis provides a shared cache for multi-instance deployments.