LLMs Technical Reviews
Home / LLM gateways / bifrost

maximhq/bifrost

Go LLM gateway and embeddable SDK with per-provider worker queues, a hook-based plugin pipeline, virtual-key governance and fallbacks.

GitHub ↗★ 8.6kGoApache-2.0commit 0e9c135 · 2026-10-07homepage ↗

Overview

Bifrost is Maxim’s LLM gateway, written in Go. It comes in two forms. core/ is a Go library: you call bifrost.Init with an Account that supplies provider configs and keys, then call typed methods such as ChatCompletionRequest or EmbeddingRequest. transports/bifrost-http wraps that library in a fasthttp server with an embedded web UI, a config store, and drop-in routes for the OpenAI, Anthropic, Bedrock, GenAI, Cohere and LiteLLM wire formats. The library form is real enough that another gateway in this category, GPT-Load, imports github.com/maximhq/bifrost/core for its protocol conversion.

The project markets itself on speed. The repository description says “50x faster than LiteLLM”, and the README reports 11 µs of added overhead at 5,000 RPS in the maintainers’ own benchmark. This site runs no benchmarks, so those numbers are the project’s claims. What the code does show is a design built to keep the hot path cheap. Each provider gets a buffered Go channel served by a fixed pool of worker goroutines. Request objects, channels and plugin pipelines come from sync.Pools. Plugin lists sit behind atomic.Pointers. Upstream calls use per-provider fasthttp.Clients, and JSON goes through bytedance/sonic. Overhead is measured, too. The core opens named spans (“handle-setup”, “pipeline-pre”, “worker-setup”, “pipeline-post”) so that gateway time can be separated from upstream time.

Most gateway features live in plugins: CEL routing rules, governance (virtual keys, budgets, rate limits), a semantic cache, logging, Prometheus and OpenTelemetry. Some features named in the README, including guardrails, clustering, adaptive load balancing and the custom-plugin docs, are listed under enterprise deployments. The open-source tree has only schema stubs for guardrails.

Architecture

flowchart LR
  C["Client (OpenAI / Anthropic / GenAI SDK)"] --> R["GenericRouter (fasthttp)"]
  R --> MW["HTTP transport hooks: auth, governance"]
  MW --> CORE["Bifrost core: handleRequest"]
  CORE --> PRH["PreRequestHook: routing rules, load balance"]
  PRH --> TRY["tryRequest: PreLLMHooks"]
  TRY -->|"short-circuit"| POST["PostLLMHooks"]
  TRY --> Q["ProviderQueue (buffered chan)"]
  Q --> W["requestWorker goroutines"]
  W --> KEY["Key selection + retries"]
  KEY --> P["Provider impl (fasthttp client)"]
  P --> UP["Upstream LLM API"]
  W --> POST
  POST --> CORE
  CORE -->|"error"| FB["Next fallback"]
  FB --> TRY
Component Path Role
Core engine core/bifrost.go Init, typed request methods, handleRequest / tryRequest, provider queues, workers, retries, plugin pipeline
Schemas core/schemas/ Canonical BifrostRequest / BifrostResponse, plugin interfaces, context keys
Providers core/providers/<name>/ About 30 provider packages that convert to and from each provider’s wire format
Key selection core/keyselectors/, core/sessionaffinity.go Weighted-random key choice and per-session key stickiness
HTTP transport transports/bifrost-http/ fasthttp server, integration routers, admin handlers, embedded UI
Framework framework/ Config store (SQLite/Postgres), log store, vector stores, model catalog and pricing, .so plugin loader
Plugins plugins/ routing, governance, semanticcache, logging, telemetry, otel, compat, jsonparser, mocker, prompts

How a request flows

Take POST /v1/chat/completions on the HTTP server:

  1. Parse. The integration’s route was built by GenericRouter.createHandler. It turns the fasthttp context into a BifrostContext, runs an optional large-payload hook (big uploads stream through without JSON parsing), then parses the body into the integration’s request type and converts it to a BifrostRequest (router.go). It then calls client.ChatCompletionRequest.
  2. Set up the request. handleRequest records the requested route, stamps a request start time for the overhead breakdown, and assigns a request ID (bifrost.go).
  3. Pick a route. RunPreRequestHooks runs once per request. The routing plugin evaluates CEL rules here and may rewrite provider, model, fallbacks and a key pin. Governance’s LoadBalanceProvider then picks a weighted provider when none is set, and fills fallbacks from the remaining weighted providers (governance/main.go). Session affinity has the last word.
  4. Pre-hooks. tryRequest resolves the provider’s queue, merges MCP tool definitions when MCP is configured, and runs RunLLMPreHooks. A plugin can short-circuit with a response or an error. A semantic-cache hit is this case: it skips the queue and goes straight to the post-hooks (bifrost.go).
  5. Enqueue. A pooled ChannelMessage is sent on the provider’s buffered channel. If the channel is full and dropExcessRequests is set, the request fails at once with a queue-full error. Otherwise the caller blocks until there is space, the provider shuts down, or the context is cancelled (bifrost.go).
  6. Work. One of the provider’s requestWorker goroutines dequeues the message, stamps per-attempt context flags, selects a key, and calls executeRequestWithRetries. That function retries with backoff and moves to a different key after a permanent per-key failure (bifrost.go, L6699-L6760).
  7. Post-hooks and fallbacks. The result returns over a pooled response channel, and RunPostLLMHooks runs in reverse order for logging, telemetry and caching. If the attempt failed, shouldTryFallbacks checks for cancellation, a plugin veto (AllowFallbacks=false) and an empty fallback list. If fallbacks are allowed, handleRequest walks them in order and calls tryRequest again for each one (bifrost.go).

Key components

Provider queues and workers

prepareProvider creates one ProviderQueue per provider, sized by ConcurrencyAndBufferSize.BufferSize, and starts Concurrency worker goroutines on it (bifrost.go). This gives per-provider backpressure and isolation: a slow provider fills its own queue and does not slow down the others. The queue channel is never closed. A long comment explains why: closing a channel that other goroutines may still send on can panic. Shutdown uses a separate done channel, and stragglers are drained with errors (bifrost.go). Init prewarms the object pools to InitialPoolSize (bifrost.go). Providers can be added at runtime through getProviderQueue, which creates them lazily under a per-provider lock.

Providers and translation

Each package under core/providers/ turns the canonical request into its own wire format and parses the reply back. Clients are tuned per provider. OpenAI, for example, has separate unary and streaming fasthttp.Clients with a FIFO connection pool, MaxConnsPerHost and keep-alive taken from the provider’s network config (openai.go). Custom providers reuse a built-in base type through CustomProviderConfig.BaseProviderType.

Plugin pipeline

LLMPlugin has three hooks. PreRequestHook runs once per request and decides the route. PreLLMHook runs once per attempt and can short-circuit. PostLLMHook runs in reverse order. Errors returned from the pre-request and pre-LLM hooks are logged but do not stop the request, so a plugin that wants to reject a request has to short-circuit or use the HTTP transport hook (plugin.go). MCP and A2A traffic have their own parallel plugin interfaces.

Routing and governance

The routing engine evaluates CEL rules with scope precedence VirtualKey > User > Team > Customer > Global. Rules can chain: a matched chain rule feeds its new provider and model back into the evaluator, up to a maximum depth (engine.go). Governance owns virtual keys (sk-bf-…), teams, customers, budgets with reset windows, and token and request rate limits. It keeps counters in memory and flushes them to the config store asynchronously.

Semantic cache

plugins/semanticcache first tries a direct lookup by hashed request ID, then an embedding search with a cosine threshold that defaults to 0.8. It skips conversations longer than 3 messages by default (main.go). Backends in framework/vectorstore are chromem (in-process), Redis, Qdrant, Weaviate and Pinecone.

Extending it

  • Go plugins. Implement LLMPlugin (and optionally HTTPTransportPlugin, MCPPlugin) and pass it in BifrostConfig.LLMPlugins when embedding the core. Plugins can be reloaded, removed and reordered at runtime.
  • Native .so plugins. The HTTP server can load Go plugin shared objects from a path or URL. Downloads go through an SSRF-hardened client (soloader.go). Go plugins have to be built with the same toolchain and dependency versions as the host binary.
  • Custom providers and key selectors. Point a custom provider at an OpenAI-compatible base type, or pass your own KeySelector / KeyPoolFilter in BifrostConfig.
  • MCP. The core can connect to MCP servers and inject their tools into requests, and the HTTP server can expose tools as an MCP server.

Running it

  • Binary or container. npx -y @maximhq/bifrost or docker run -p 8080:8080 maximhq/bifrost. The server defaults to localhost:8080 (-host, -port, -app-dir), and automaxprocs sets GOMAXPROCS from cgroup limits (main.go, L150-L165). Configuration comes from config.json in the app directory or from the UI.
  • Storage. The config store is SQLite or PostgreSQL. Logs go to SQLite, PostgreSQL or ClickHouse. The semantic cache needs a vector store and an embedding provider.
  • Kubernetes. Helm charts are in helm-charts/bifrost.
  • As a library. go get github.com/maximhq/bifrost/core, implement schemas.Account, then call bifrost.Init.

Strengths and caveats

  • Strength: a hot path built for low overhead. Bounded per-provider queues, pooled objects, atomic plugin lists and fasthttp keep allocation and locking low. The overhead spans let you measure what the gateway adds in your own setup instead of relying on the README.
  • Strength: embeddable core. The gateway is a Go library first, so you can use it in-process without HTTP. GPT-Load does exactly that.
  • Strength: broad surface. About 30 providers, chat/responses/embeddings/audio/image/video/batch/files, MCP and A2A, plus several SDK-compatible route sets.
  • Caveat: pre-hook errors do not block. Reject requests by short-circuiting or in the transport hook. A routing or policy plugin that only returns an error lets the request through.
  • Caveat: state is per node in OSS. The default KV store behind session affinity is in-memory (kvstore.go), and governance counters are kept per process. The README lists clustering as an enterprise feature.
  • Caveat: guardrails and adaptive load balancing are not in the open tree. The OSS build has only config-schema stubs for guardrails.
  • Caveat: size. core/bifrost.go alone is over 10,000 lines. Expect to read a lot of code before you change core behaviour.

Sources: code at 0e9c135, verified Q&A.

How it answers the LLM gateways questions

Each answer was drafted by a code-reading agent at commit 0e9c135. Its citations were checked mechanically. Compare with the other llm gateways →

How are requests routed across providers and models?

answered

Bifrost uses a three-layer routing pipeline. Layer 1 — CEL rule engine: The routing plugin (plugins/routing/main.go) evaluates Google CEL (Common Expression Language) rules with scope precedence (VirtualKey > User > Team > Customer > Global). Each TableRoutingRule has a CelExpression (e.g. headers["x-region"] == "eu") that matches on model, provider, request_type, headers, query params, complexity_tier, budget_used, tokens_used, and request count. Rules can be terminal or chainable (recursive re-evaluation after a matched rule rewrites provider/model). Layer 2 — Weighted provider load balancing: After rules settle, GovernancePlugin.LoadBalanceProvider (plugins/governance/main.go:430) selects among eligible weighted provider configs using weighted random selection. The eligible set is filtered by model allowance, budget limits, rate limits, and key grants. Fallbacks are automatically derived from remaining weighted providers sorted by weight. Layer 3 — Ordered fallback chain in core: Bifrost.handleRequest (core/bifrost.go:5689) iterates the fallback list sequentially. Each fallback calls tryRequest on the next provider/model, with full tracing spans per attempt. TTFT (time-to-first-token) deadlines can cut short slow streaming attempts via AttemptAbort (core/providers/utils/stream.go:29). Session affinity (core/sessionaffinity.go) remembers which provider/key served a session, routing subsequent requests to the same provider. Health checks are implicit — failures trigger fallback traversal; a provider that fails in tryRequest is skipped for subsequent fallbacks. The model catalog supports aliases via keyconfig.Store.ResolveAlias, allowing one model name to map to different wire names per key or provider.

How are different provider APIs unified?

answered

Bifrost unifies every provider through its internal BifrostRequest/BifrostResponse schema (core/schemas/bifrost.go), which defines type-safe request structs for chat completions, embeddings, responses, rerank, speech, transcription, and image/video generation. Each provider package (30+ in core/providers/) implements a conversion layer: ToBifrostChatRequest normalizes incoming provider-native requests, and ToXxxRequest/ToXxxResponse translates Bifrost's internal format to the provider's wire format. For example, openai/chat.go has ToBifrostChatRequest and ToOpenAIChatRequest that convert between the OpenAI ChatCompletion JSON shape and Bifrost's canonical BifrostChatRequest. Anthropic's anthropic/chat.go similarly converts Anthropic-specific constructs (document blocks, thinking blocks) into the common schema. Streaming is handled separately per provider — the fasthttp-based HTTP clients use SSE parsing for OpenAI-compatible streams and custom chunk logic for Anthropic/other providers. Tool calls, multimodal (images, audio, video), and structured output all flow through the same conversion pattern; the jsonparser plugin (plugins/jsonparser/main.go) handles partial JSON accumulation for streaming responses. The compat plugin (plugins/compat/main.go) bridges LiteLLM compatibility by converting text completions to chat or chat to responses when the target model lacks native support. The 30+ providers span OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Cerebras, Cohere, DeepSeek, ElevenLabs, Fireworks, Gemini, GitHub Copilot, Groq, HuggingFace, Mistral, Nebius, Ollama, OpenRouter, Perplexity, Replicate, Runware, Runway, Sarvam, SGL, Typesafe, vLLM, Wafer, and xAI.

How are API keys, users and tenants managed?

answered

Bifrost implements a virtual-key-based access model. Virtual keys (sk-bf-... prefix) are the primary auth mechanism: each TableVirtualKey (configstore/tables/virtualkey.go) maps to a set of TableVirtualKeyProviderConfig rows that define which providers the key can reach, along with allowed/blacklisted models, weights, and associated upstream API keys in the Keys many-to-many relationship. The Value of each TableKey (configstore/tables/key.go) stores the actual upstream credential (e.g. OpenAI API key) as a SecretVar that supports env-var references. Scope hierarchy: Each virtual key can belong to teams and customers (TableTeam, TableCustomer), forming a hierarchy for budget and rate-limit propagation. Users are identified via UserID on requests and can have their own scoped routing rules. Auth flow: The governance plugin's ResolveAccess (plugins/governance/resolver.go) resolves what a request may reach from its presented virtual key, then evaluates provider gates, model gates, and key restrictions. The HttpTransportPreHook can reject requests before they reach the LLM pipeline. Admin UI: The web UI (React frontend embedded in the Go binary) provides visual configuration of virtual keys, provider keys, teams, customers, and routing rules. OIDC/OAuth2 user provisioning (framework/oauth2/) supports OAuth 2.0 / OIDC login with background directory sync for teams and roles, writing provisioned users into the same governance store. MCP tool authorization (plugins/governance/mcpauthorization.go) extends the virtual key model to MCP tool access, with per-VMCP-tool allow/deny rules.

How are rate limits, budgets and cost tracking implemented?

answered

Rate limits and budgets are enforced per entity via the governance plugin's LocalGovernanceStore. Budgets (TableBudget in configstore/tables/budget.go) carry a MaxLimit (dollars), ResetDuration (e.g. "1d", "1M"), CurrentUsage, and LastReset, plus override support. Budgets are attached at the virtual-key provider-config level, and also at team/customer levels for hierarchical cost control. Rate limits (TableRateLimit) track both token-based (TokenMaxLimit) and request-based (RequestMaxLimit) quotas with independent reset durations. Both budgets and rate limits use a sliding-window reset mechanism computed from LastReset + ResetDuration. The UsageTracker (plugins/governance/tracker.go) processes UsageUpdate events asynchronously through a background goroutine that batches writes to the in-memory store and periodically dumps to the DB. Cost calculation: The datasheet.Store.CalculateCost (framework/modelcatalog/datasheet/cost.go:19) computes response cost from model pricing. The TableModelPricing table (configstore/tables/modelpricing.go) stores per-model per-provider per-mode pricing fields: text token rates (input/output, with batch/priority/fast/ultrafast/flex tiers), image/video/audio pricing, and tiered rates for contexts above 128k/200k/272k tokens. Custom pricing overrides can be scoped globally or per-provider. Usage accounting stores spend in the LocalGovernanceStore in-memory maps and periodically persists to PostgreSQL. The data is exposed via the admin API and used for real-time governance decisions (budget-exceeded/rate-limited denials).

How are caching and guardrails implemented?

answered

Caching: The semanticcache plugin (plugins/semanticcache/main.go) provides two-tier caching. Tier 1 is direct (exact) caching using xxhash of the deterministic (provider, model, cacheKey, request_hash, params_hash) tuple — performDirectSearch (plugins/semanticcache/search.go:23) does an O(1) point fetch from the vector store by ID. Tier 2 is semantic caching using embedding similarity search — performSemanticSearch (plugins/semanticcache/search.go:54) generates an embedding of the request text, queries the vector store for nearest neighbors above a configurable cosine similarity Threshold (default 0.8), and returns the most similar cached response. The plugin supports TTL, conversation-history thresholds for skipping long threads, and configurable cache keys (by model, provider, excluding system prompt). The VectorStore abstraction (framework/vectorstore/) backs the cache — Qdrant, Weaviate, and in-process stores are supported. Guardrails: There is no separate PII-redaction, moderation, or prompt-injection plugin in the OSS codebase. The CalculateGuardrailCost function (framework/modelcatalog/datasheet/cost.go:228) prices "judge calls" — internal LLM-as-judge evaluations for guardrail decisions — but the actual guardrail logic (moderation, PII scanning, prompt injection detection) lives in the enterprise build. The OSS plugin architecture exposes hook points (PreLLMHook, PostLLMHook at plugins/governance/main.go:100-102) where such plugins would integrate. The OSS build does support GPT Live integration for realtime moderation via plugins/logging redaction and header/config redaction in the OTel plugin's API responses.

Editor's note. Correction: the GPT Live references in the logging and governance plugins concern billing and logging of voice sessions, not moderation. The open-source tree has no moderation or PII guardrail plugin; guardrails_config exists only as an enterprise schema stub. Vector stores also include Redis and Pinecone.

How is it observed, deployed and scaled?

answered

Observability: Bifrost has three observability layers. (1) Prometheus metrics via the telemetry plugin (plugins/telemetry/main.go) — exposes counters for upstream requests, success/error rates, input/output tokens, cache hits, and cost; histograms for upstream latency, overhead latency, stream first-token/inter-token latency, and retry counts. Supports push to a Prometheus Push Gateway for multi-node aggregation. (2) OpenTelemetry via the otel plugin (plugins/otel/main.go) — exports traces and metrics to one or more collector profiles (OTLP over HTTP or gRPC). Each profile has independent endpoints, headers, TLS config, and span filters. Traces follow the GenAI semantic convention and include per-provider spans, fallback spans, and overhead breakdowns. Supports distributed tracing via W3C traceparent and session-based trace grouping. (3) Structured logging via the logging plugin (plugins/logging/main.go) — logs all requests/responses to PostgreSQL with search, filter, and pagination via the logstore framework. Performance architecture: Written in Go with fasthttp (a high-performance async HTTP client/server), object pooling for responses, lock-free sync.Map data structures in hot paths, and per-provider connection pools with configurable concurrency. The model catalog uses memoization with generation counters to avoid recomputation. Deployment: Docker images (Alpine-based multi-stage builds embedding a React UI) run the HTTP server on port 8080 by default. Helm charts (helm-charts/bifrost/) deploy to Kubernetes with configurable replica count, secret management, and values examples for various backing stores (PostgreSQL+Weaviate, SQLite+Redis, etc.). The Dockerfile (transports/Dockerfile) builds a static Go binary using CGO. High availability: Horizontal scaling is supported by shared PostgreSQL and vector stores; the KV store (framework/kvstore/) provides distributed state for session affinity and rate-limit coordination across nodes. The OTel circuit breaker (plugins/otel/main.go:56-64) prevents a failed collector from blocking request processing.

Editor's note. Correction: framework/kvstore is an in-memory, per-process store (kvstore.go L25-L33), and governance usage counters are also kept per process, so session affinity and rate limits are not coordinated across nodes in the open-source build; the README lists clustering as an enterprise feature.