diegosouzapw/OmniRoute
Self-hosted TypeScript AI gateway that pools provider free tiers, subscription OAuth logins and web sessions behind one OpenAI-style API.
Overview
OmniRoute is a self-hosted AI gateway written in TypeScript on Next.js 16. It puts one OpenAI-compatible endpoint (plus Anthropic /v1/messages, the Responses API, Gemini-style /v1beta and media endpoints) in front of a very large provider catalogue. It is built for people who run coding agents such as Claude Code, Codex, Cline or OpenCode and want them to keep working when one account or provider hits a limit. It ships as an npm CLI (omniroute), a Docker image, an Electron desktop app and a PWA dashboard. State lives in SQLite (better-sqlite3), and Redis is optional.
The project’s headline claim is free access to many models. In the code this is not one trick. It is a mix of five mechanisms: a hand-curated catalogue of providers’ own free tiers; “no-auth” providers that call public or anonymous endpoints; OAuth connections that reuse a user’s existing subscription to a vendor CLI or IDE (Claude Code, Codex, GitHub Copilot, Antigravity, Kiro and others); “web-cookie” providers that drive consumer chat websites with a pasted browser session; and pooling of several accounts per provider with rotation and cooldowns. The first two use access the provider offers to everyone. The last three use credentials or client identities that were issued for a different product, and the codebase itself flags many of them as risky under the upstream terms (see Free access and terms of service).
The codebase is very large: 161 entries in open-sse/executors/, about 200 SQL migrations, and a dashboard with routes for combos, quotas, compression, MCP, A2A, memory, evals and more. TypeScript strict is off. The core is careful about failure handling, but the code is hard to read end to end.
Architecture
flowchart LR
C["Client (agent, SDK, IDE)"] --> R["Next.js route /v1/chat/completions"]
R --> H["handleChat (src/sse)"]
H --> CB["handleComboChat (combos)"]
H --> S["handleSingleModelChat"]
CB --> S
S --> CR["getProviderCredentials"]
CR --> DB["SQLite: provider_connections"]
S --> CC["handleChatCore (open-sse)"]
CC --> T["translateRequest / translateResponse"]
CC --> EX["getExecutor -> provider executor"]
EX --> UP["Upstream: API key, OAuth, no-auth, web cookie"]
S --> AF["markAccountUnavailable (cooldown)"]
| Component | Path | Role |
|---|---|---|
| HTTP routes | src/app/api/v1/ |
Next.js App Router endpoints for chat, messages, responses, embeddings, images, audio, video and more |
| Chat entry | src/sse/handlers/chat.ts |
Admission, model and combo resolution, per-account retry loop |
| Credential selection | src/sse/services/auth.ts |
Picks a connection (account) per provider, applies cooldowns, leases and quotas |
| Combo engine | open-sse/services/combo.ts, combo/ |
Multi-model strategies (priority, weighted, round-robin, p2c, cost, reset-aware, fusion and others) |
| Core pipeline | open-sse/handlers/chatCore.ts |
Compression, guardrails, cache, translation, executor dispatch, streaming |
| Translators | open-sse/translator/ |
Hub-and-spoke format conversion through an OpenAI-shaped middle format |
| Executors | open-sse/executors/ |
One class per upstream family (OAuth CLIs, web sessions, generic OpenAI-compatible) |
| Free-tier catalogue | open-sse/config/freeModelCatalog.data.ts |
Hand-curated free quotas with a per-entry tos verdict |
| Provider catalogue | src/shared/constants/providers/ |
Provider metadata split by auth type: oauth, noauth, web-cookie, apikey, local |
| Persistence | src/lib/db/ |
SQLite tables for keys, connections (encrypted), usage, combos, settings |
How a request flows
Take a POST /v1/chat/completions with model: "cc/claude-sonnet-4-6":
- Route. The App Router handler validates the body shape, applies the prompt-injection guard and API-key policy, initialises translators once, and calls
handleChat(route.ts). - Combo or single model.
handleChatresolves aliases. If the name is a combo, it callshandleComboChat, which orders targets by the combo strategy and calls back intohandleSingleModelChatfor each target (combo.ts). - Pick an account.
handleSingleModelChatruns a loop. Each pass callsgetProviderCredentialsWithQuotaPreflightwith the IDs it has already tried excluded (chat.ts). Inside,getProviderCredentialsloads every active connection for the provider, filters by API-key allowlist, cooldowns and lockouts, and returns one. No-auth providers get a synthetic connection instead (auth.ts). - Translate.
handleChatCoredetects the source format and callstranslateRequest(source, target, …), which normalises thinking budgets and tool-call IDs and converts through the OpenAI middle format when no direct translator exists (translator/index.ts). - Execute.
getExecutor(provider)loads a registered executor or the default OpenAI-compatible one (executors/index.ts).BaseExecutor.executerefreshes OAuth tokens if needed, builds provider headers and sends the request, with per-URL retry (base.ts). - Fail over. On 401/403/429 or quota errors the handler calls
markAccountUnavailable, which sets a cooldown on that connection (or a model lockout), and the loop picks the next account. 5xx errors also feed a per-provider circuit breaker. - Return. The response is translated back to the client’s format, streamed as SSE, and logged with token usage and cost.
Key components
Free access and terms of service
How “free” is achieved, mechanism by mechanism:
- Providers’ own free tiers.
FREE_MODEL_BUDGETSis a hand-curated list of about 490 model entries. Each has afreeType(recurring-daily,recurring-monthly,one-time-initial,keyless,recurring-uncapped,discontinued), a token figure where one is documented, and atosverdict ofok,caution,ambiguous,avoidorunknown(freeModelCatalog.data.ts, freeModelCatalog.ts). The user still signs up for each provider and stores its key. The gateway combines these quotas and fails over between them. - No-auth providers.
NOAUTH_PROVIDERSlists endpoints that need no account, such as DuckDuckGo AI Chat, Cloudflare AI Playground, UncloseAI, AI Horde and “OpenCode Free” (noauth.ts). The OpenCode notice says that tier “can only be used from within OpenCode”. The executor meets that requirement by synthesising the client’s request contract: forcedstream: true, a non-emptytoolsarray, a session ID in OpenCode’s format, and anopencode/<version>User-Agent (opencodeFreeTierContract.ts, opencodeHeaders.ts). It can also rotate several such “accounts” (fingerprints), each with its own proxy (accountRotation.ts). - Subscription OAuth reuse.
OAUTH_PROVIDERSconnects to Claude Code, OpenAI Codex, GitHub Copilot, Antigravity, Kiro, Cursor, Grok Build and others with the user’s own login (oauth.ts). Requests are shaped to match the official client.cliFingerprints.tsreorders headers and JSON fields “to be indistinguishable from the real CLI binary, reducing account flagging risk” (cliFingerprints.ts).claudeIdentity.tsreproduces a pinnedclaude-clirequest shape for the Claude Code OAuth scope (claudeIdentity.ts). Antigravity traffic goes to Google’scloudcode-pav1internalendpoints (antigravityUpstream.ts). - Web-cookie providers.
WEB_COOKIE_PROVIDERS(33 entries) drive consumer chat sites such as chatgpt.com, grok.com and gemini.google.com with a pasted cookie header or Playwright storage state (web-cookie.ts). - Account pooling. Any provider can hold several connections.
selectAccountsupportsfill-first,round-robin,p2candrandom(accountSelector.ts). Quota-aware combos move on when one account’s window is exhausted. An opt-in setting deactivates a connection after a permanent ban (autoDisableBannedAccount.ts). The dashboard can also import public free-proxy lists into a proxy pool (freeProxyProviders/types.ts).
Terms-of-service implications, as the code records them. FREE_TIER_TOS marks 17 providers avoid, described as providers “whose terms PROHIBIT routing through a self-hosted proxy or forbid non-personal use” (freeTierCatalog.ts). The list includes agy (Antigravity), kiro, amazon-q, opencode and duckduckgo-web. In the per-model catalogue, 78 entries carry tos: "avoid". Providers with subscriptionRisk: true show a consent dialog before connecting. For OAuth providers it says the session “is not authorized for proxy/router use” and “the upstream may react by restricting or banning the account” (en.json). The project’s own docs state that the ToS flag “is advisory, not a routing gate”: avoid providers stay in routing and fallback by default (FREE_TIERS.md). The one code-level filter, filterTosAvoidCandidates, applies only to auto-combo candidates and only when excludeTosAvoid is set (strictZeroCostFilter.ts). In short, OmniRoute documents the terms risk and leaves the decision to the operator.
Routing and resilience
There are three layers. Combos order targets across models. Per-connection cooldowns with exponential backoff handle a single account that fails. A per-provider circuit breaker with four states (CLOSED, DEGRADED, OPEN, HALF_OPEN) is persisted to SQLite. 401/403/429 errors go to connection cooldown, and 408/5xx errors trip the breaker. Retry hints in error bodies (“resets in 3h”, retry-after) set the cooldown length.
Translation
The translator registry maps (from, to) format pairs. Formats include Claude, Gemini, Kiro, Cursor, Antigravity and the Responses API. Thinking models get their earlier reasoning_content replayed from a cache to avoid upstream 400 errors.
Keys, limits and cache
Client API keys live in SQLite with model, combo and connection allowlists, per-minute and per-day limits, and daily or weekly USD caps. Upstream credentials are encrypted with AES-256-GCM. A two-layer response cache (SHA-256 exact match, then embedding similarity on an in-memory or Redis vector store) sits in open-sse/services/cache/. Guardrails (PII masking, credential masking, prompt-injection detection, vision/audio/video bridges) register in src/lib/guardrails/registry.ts. PII redaction is off by default.
Extending it
- New provider. Add metadata under
src/shared/constants/providers/and, if the provider is not OpenAI-compatible, an executor inopen-sse/executors/plus a translator pair. - Custom OpenAI-compatible nodes. Add them from the dashboard (
/api/provider-nodes) without code. - Combos. Define named combos with a strategy and targets. Clients then call the combo name as a model.
- MCP and A2A.
omniroute --mcpstarts an MCP server over stdio, and/a2aexposes an agent-to-agent surface. - Guardrails, evals and webhooks. Each is a registry with its own folder under
src/lib/.
Running it
npx omnirouteor a global npm install starts the bundled Next.js server (bin/omniroute.mjs). The default port is 20128 (.env.example).- Docker images, several Compose files and
fly.tomlare included. An Electron wrapper lives inelectron/. - Requirements: Node.js 22 or newer, a writable data directory for SQLite, and optionally Redis for shared rate limits and the vector cache.
Strengths and caveats
- Strength: failure handling. Per-account cooldowns, model lockouts, a persisted breaker and quota-aware combos make it hard for one exhausted account to stop a coding session.
- Strength: honest free-tier bookkeeping. The catalogue separates recurring, one-time and uncapped quotas, needs a cited source for every number, and records a ToS verdict and prompt-training flag per entry.
- Strength: format coverage. OpenAI, Anthropic, Gemini and Responses clients all work against the same pool.
- Caveat: terms-of-service exposure. A large part of the “free” surface depends on subscription OAuth, web cookies, client-identity emulation or account rotation. The code itself marks many of these as prohibited by the upstream and warns of account bans. Its ToS flags do not stop routing by default.
- Caveat: size and churn. A god-file layout (
chatCore.tsabout 3,900 lines,chat.tsabout 2,600), non-strict TypeScript and frequent provider retirements make audits and upgrades expensive. - Caveat: in-process state. Rotation and throttle stores such as the OpenCode egress throttle are in-memory modules, so several replicas do not share them even when Redis is configured.
Sources: code at 8ad6b1c, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the LLM gateways questions
Each answer was drafted by a code-reading agent at commit 8ad6b1c. Its citations were checked mechanically. Compare with the other llm gateways →
How are requests routed across providers and models?
answeredRouting is orchestrated by handleComboChat() in open-sse/services/combo.ts, which implements 19+ strategies (priority, weighted, round-robin, random, least-used, cost-optimized, reset-aware, reset-window, strict-random, auto, fill-first, p2c, lkgp, context-optimized, context-relay, fusion, pipeline, quota-share, chaos). The dispatch passes through a prelude sequence: pinned model → fusion (parallel judge-synthesis) → chaos (multi-model parallel) → pipeline → runtime-unit → round-robin → target iteration with cooldown-aware retry. Target resolution in targetResolution.ts performs provider-wildcard expansion, weighted step-group resolution, prompt-cache affinity, session stickiness, eval scores, and request-compatibility filtering. The auto strategy generates scored candidates via buildAutoCandidates() using 16 factors: cost per token, historical p95 latency, error rate, circuit-breaker state, quota remaining, session affinity, reset-window affinity, OAuth session availability, connection pool size, quality scores, and speed telemetry. Fallbacks operate at three resilience layers (detailed in AGENTS.md): provider circuit breakers (src/shared/utils/circuitBreaker.ts — 4 states: CLOSED/DEGRADED/OPEN/HALF_OPEN, failure-kind-aware thresholds, DB-persisted), connection cooldown (accountFallback.ts — per-credential exponential backoff), and model lockout (per provider+connection+model). Retries use full-jitter backoff with status-code-aware classifiers (408/500/502/503/504 trip the provider breaker; 401/403/429 route to connection cooldown). Model aliases resolve through open-sse/config/providerRegistry.ts where user-supplied names map to provider-specific model IDs. Health checks run via circuit-breaker lazy recovery: expired OPEN states auto-transition to HALF_OPEN on the next getStatus(), and canExecute() gates target selection.
How are different provider APIs unified?
answeredTranslation follows a hub-and-spoke architecture centered on OpenAI format. translateRequest() in open-sse/translator/index.ts first normalizes thinking budget and reasoning-routing directives, then checks for direct translators (e.g., Claude→Gemini). When none exist, it converts source→OpenAI via a registered translator, then OpenAI→target. translateResponse() reverses the pipeline. The translator registry (translator/registry.ts) maps (from, to) format pairs to request/response functions. 12+ request translators exist: claude-to-openai.ts, gemini-to-openai.ts, openai-to-claude.ts, openai-to-gemini.ts, openai-to-kiro.ts, openai-to-cursor.ts, openai-to-clova.ts, antigravity-to-openai.ts, and the openai-responses.ts converter. Streaming is handled by translateStreamChunk() which converts provider-native SSE into unified OpenAI-style deltas. Tool calls are translated via translator/helpers/toolCallHelper.ts (ensureToolCallIds, fixMissingToolResponses, stripOrphanedToolResults) and schema coercion helpers (coerceToolSchemas, sanitizeToolDescriptions). Multimodal content (images, audio, video) passes through the same format pipeline. The responsesTransformer.ts bridges OpenAI Chat Completions and the Responses API via a TransformStream. A critical subsystem is reasoning replay: for thinking models (DeepSeek V4, Kimi K2, Qwen), cached reasoning_content is re-injected on subsequent turns to avoid 400 errors. Reasoning cache uses services/reasoningCache.ts with lookupReasoning()/recordReplay(). Provider support: 325+ providers registered across auth types (OAuth: 25 entries, No-Auth: 10, Local: 14, WebCookie: ~50+), served by 161 executor files in open-sse/executors/. Role normalization handles developer→system mapping, and OpenAI-incompatible echo fields (reasoning_content, refusal, annotations) are stripped for strict upstreams.
How are API keys, users and tenants managed?
answeredAPI keys are managed through src/lib/db/apiKeys.ts which stores keys in a SQLite api_keys table with fields for id, key prefix/hash, provider/model permissions, usage limits, spend tracking, rate limits, access schedules, IP allowlists, and tenant/user associations. Key lookup uses constant-time comparison (timingSafeCompare from src/shared/utils/timingSafeCompare). Virtual keys allow multiple users/teams to share upstream credentials while maintaining distinct rate/spend limits. Upstream credentials (OAuth tokens, API keys for downstream providers) live in the provider_connections table managed by src/lib/db/providers.ts, encrypted at rest with AES-256-GCM via the encryptConnectionFields()/decryptConnectionFields() pipeline. Credentials are fetched by getProviderCredentialsByUserIdAndConnectionId(). Auth methods are diverse: API-key auth (bearer token validated by extractApiKey()/isValidApiKey() in route handlers, with requireApiKey: true flag), OAuth flows via src/lib/oauth/, JWT session tokens (src/lib/auth/jwt.ts with signJwt/verifyJwt), and refresh tokens (src/lib/db/tokens.ts). The user/team/tenant model is multi-tenant with src/lib/db/tenants.ts, src/lib/db/teams.ts, and src/lib/db/users.ts. Each API key carries fine-grained permissions: allowed models (with wildcard support via modelPermissions.ts), allowed combos, allowed connections, allowed quotas, access schedules, per-minute/per-day rate limits, spend limits (daily/weekly USD caps), and key groups (apiKeyGroups.ts). The dashboard at src/app/(dashboard)/ provides a UI for key management, while the /v1/* API routes serve the programmatic interface. Auth method discovery is also driven by src/shared/constants/providers/ which classifies providers as OAuth (25), No-Auth (10), Local (14), API-key (bulk), WebCookie (~50+), and Cloud Agent types.
How are rate limits, budgets and cost tracking implemented?
answeredRate limiting is implemented via src/shared/utils/rateLimiter.ts using a Redis-backed (or in-memory fallback) leaky bucket strategy. API keys carry maxRequestsPerDay, maxRequestsPerMinute, and rich rateLimits rules (parsed by parseRateLimits()). Cost tracking centers on open-sse/services/providerCostData.ts which provides a KNOWN_MODEL_PRICING map covering ~25 key models (e.g., gpt-4o at $2.50/$10 per 1M tokens, claude-sonnet-4.6 at $3/$15) and a getModelPricing() resolver that falls back through: provider-specific KNOWN match → default pricing DB table → generic model name match → free-model check → estimated fallback ($5/$15). The tier resolver (tierResolver.ts) layers operator-configured pricing over the defaults using a DB snapshot from pricing/pricing_synced/models_dev_pricing key-value namespaces and caches the merged result. Usage accounting is stored in two SQLite tables: usage_history (per-request rows with tokens_input/tokens_output and ISO timestamps) and daily_usage_summary (rolled-up aggregates), queried by src/lib/db/usageSummary.ts and src/lib/db/usageLogs.ts. The GET /v1/usage endpoint exposes aggregated usage. Spend limits are enforced by apiKeyPolicyService.ts (enforceApiKeyLimits()) which checks daily/weekly USD caps before allowing requests. Circuit breakers, connection cooldown (exponential backoff on per-credential failures), and model lockout all act as runtime cost-control mechanisms by stopping traffic to failing-expensive or exhausted endpoints. The connectionBillingCatalog.ts classifies connections as subscription (plan-included) vs metered (per-token) to budget routing.
How are caching and guardrails implemented?
answeredCaching uses a dual-layer architecture in open-sse/services/cache/semanticCacheManager.ts. Layer 1 is exact-match: messages are normalized and hashed with SHA-256 for O(1) hits. Layer 2 is semantic: conversation text is embedded via embeddingClient.ts and matched by cosine similarity against a vector store (in-memory MemoryVectorStore or Redis RedisVectorStore). Cache entries store full response data and support SSE streaming replay. Cache configuration (semanticCacheConfig.ts), TTLs, and namespace isolation are provider-specific. The old deterministic caching also lives in the proxyDispatcherCache.ts for HTTP connection reuse. Guardrails are modular in src/lib/guardrails/. The promptInjection.ts guardrail detects system-override attempts with regex patterns (e.g., system: override, markdown code-block system blocks), classifies severity (low/medium/high), and supports block/warn/log modes with configurable block thresholds. piiMasker.ts provides PII redaction but is strictly opt-in — disabled by default (Hard Rule #20: both PII_REDACTION_ENABLED and PII_RESPONSE_SANITIZATION feature flags default to "false"). Moderation is handled by moderationProvider.ts/moderationProviderOpenAI.ts. The guardrail pipeline is assembled by buildGuardrailDependencyGraph() from guardrails/onRequest.ts (restructured into the current codebase) which sequences execution order. Additional guardrails cover audio/video bridges (audioBridge.ts, videoBridgeHelpers.ts), vision (visionBridgeRouter.ts), and compliance/no-log markers. Custom guardrail/eval/webhook extensions follow documented patterns: guardrails at src/lib/guardrails/ → docs/security/GUARDRAILS.md, evals at src/lib/evals/ → docs/frameworks/EVALS.md. The anti-ReDoS learnings from AGENTS.md mandate strictly bounded regex sequences to prevent catastrophic backtracking on untrusted inputs.
buildGuardrailDependencyGraph() or guardrails/onRequest.ts; guardrails are classes registered on a GuardrailRegistry by registerDefaultGuardrails() in src/lib/guardrails/registry.ts (vision, audio and video bridges, PII masker, credential masker, prompt-injection guard).How is it observed, deployed and scaled?
answeredObservability is three-tiered. Structured JSON logging uses pino through src/lib/logger/, with a log-export system (src/lib/logExport/) that can ship call logs to BigQuery or other destinations. OpenTelemetry support is optional and lightweight: open-sse/services/routing/otel.ts exports GenAI semantic-convention spans (gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.*) to an OTLP/HTTP collector. It uses global fetch (no SDK dependency), a bounded buffer with background flush, and is disabled by default unless OMNIROUTE_OTEL_ENDPOINT or OTEL_EXPORTER_OTLP_ENDPOINT is set. Internal metrics (latency, request counts, error rates) are collected in src/lib/logger/metrics.ts. Deployment supports multiple modes: container via Dockerfile (multi-stage Node.js), Fly.io via fly.toml, Docker Compose, and tunnel-based deployment (src/mitm/tunnel/) for local model endpoints. The dashboard at src/app/(dashboard)/ provides visual management of providers, API keys, usage, logs, and settings. Performance architecture is standard Node.js async I/O on Next.js 16, single-threaded event loop with 161 executor modules and connection pooling via Undici (proxyDispatcherCache.ts with round-robin dispatchers to prevent SSE stream monopolization). HA mechanisms include the three resilience layers: provider circuit breakers (persisted to SQLite for restart survival), connection cooldown, and model lockout — all lazily self-recovering. Rate limiting uses Redis when available (falling back to in-memory). The _artifacts/ directory stores temporary working files per CLAUDE.md conventions. TypeScript strict mode is off (target ES2022), with 193 DB migrations for schema evolution.