How are different provider APIs unified?
OpenAI-compatible (or other) schema; request/response translation; streaming; tool calls and multimodal; provider count.
Verdict
LiteLLM covers the most providers behind an OpenAI-style API. New API and Plano are the strongest when clients speak several formats, such as Claude Code next to an OpenAI SDK. Agent Router is the most careful about cross-provider retries.
OpenAI format in, adapters out. LiteLLM has about 150 provider folders. Each provider is a Config class with transform_request and transform_response, driven by one shared HTTP handler. It also serves /v1/messages and raw pass-through routes. Portkey Gateway registers 72 providers. Each one maps gateway parameters to provider parameters, with renames, defaults and min/max clamps, and streams come back as OpenAI SSE. One API accepts OpenAI format only. GetAdaptor maps 19 API types to a nine-method Adaptor interface, and most other channel types fall through to the OpenAI adaptor.
Many formats in, many out. New API accepts OpenAI Chat, Responses (HTTP and WebSocket), Claude Messages and Gemini. Its relaykit converters are a separate Go module, and they grade each conversion path. Plano’s hermesllm crate maps three client APIs to five upstream shapes, including Bedrock Converse, in one TryFrom match. OmniRoute converts through an OpenAI middle format unless a direct translator exists. It replays cached reasoning_content to thinking models to avoid upstream 400 errors.
A shared schema library. Bifrost converts every request to a canonical BifrostRequest and has about 30 provider packages. Its HTTP server adds routes for the OpenAI, Anthropic, Bedrock, GenAI, Cohere and LiteLLM wire formats. GPT-Load writes no converters of its own. It forwards the client’s native format when the upstream supports it. Converted routes use Bifrost’s Go provider packages, and subscription accounts go through an embedded CLIProxyAPI.
Translation inside Envoy. Higress’s ai-proxy WASM plugin supports 38 provider types. protocol: original skips conversion, and Claude-format requests are converted for providers that lack native support. Agent Router’s ext_proc sidecar translates once per upstream attempt, from the stored original body. A retry from OpenAI to Bedrock therefore gets a fresh, correctly signed request. Its inputs include Anthropic /v1/messages and Cohere rerank, not only OpenAI.
Pick: LiteLLM or Portkey Gateway for the widest OpenAI-compatible provider coverage. Pick: New API, Plano or OmniRoute when Claude, Gemini or Responses clients must reach any backend. Pick: Bifrost to embed the translation layer as a Go library.
Per-project answers
diegosouzapw/OmniRoute
answeredTranslation follows a hub-and-spoke architecture centered on OpenAI format. translateRequest() in open-sse/translator/index.ts first normalizes thinking budget and reasoning-routing directives, then checks for direct translators (e.g., Claude→Gemini). When none exist, it converts source→OpenAI via a registered translator, then OpenAI→target. translateResponse() reverses the pipeline. The translator registry (translator/registry.ts) maps (from, to) format pairs to request/response functions. 12+ request translators exist: claude-to-openai.ts, gemini-to-openai.ts, openai-to-claude.ts, openai-to-gemini.ts, openai-to-kiro.ts, openai-to-cursor.ts, openai-to-clova.ts, antigravity-to-openai.ts, and the openai-responses.ts converter. Streaming is handled by translateStreamChunk() which converts provider-native SSE into unified OpenAI-style deltas. Tool calls are translated via translator/helpers/toolCallHelper.ts (ensureToolCallIds, fixMissingToolResponses, stripOrphanedToolResults) and schema coercion helpers (coerceToolSchemas, sanitizeToolDescriptions). Multimodal content (images, audio, video) passes through the same format pipeline. The responsesTransformer.ts bridges OpenAI Chat Completions and the Responses API via a TransformStream. A critical subsystem is reasoning replay: for thinking models (DeepSeek V4, Kimi K2, Qwen), cached reasoning_content is re-injected on subsequent turns to avoid 400 errors. Reasoning cache uses services/reasoningCache.ts with lookupReasoning()/recordReplay(). Provider support: 325+ providers registered across auth types (OAuth: 25 entries, No-Auth: 10, Local: 14, WebCookie: ~50+), served by 161 executor files in open-sse/executors/. Role normalization handles developer→system mapping, and OpenAI-incompatible echo fields (reasoning_content, refusal, annotations) are stripped for strict upstreams.
BerriAI/litellm
answeredLiteLLM unifies 150+ providers under an OpenAI-compatible schema. The core insight is in litellm/__init__.py (line 1149) and 20 more like deepseek, groq, perplexity, cerebras) that speak OpenAI's wire format directly. Each provider implements a subclass of litellm/constants.py (line ~691): LITELLM_CHAT_PROVIDERS lists ~70 native chat providers (openai, anthropic, gemini, cohere, together_ai, etc.) plus openai_compatible_providers (BaseLLM (litellm/llms/base.py:14-80) that overrides process_response() and validate_environment().
Request translation. Each provider subdirectory under litellm/llms/ (Anthropic, Bedrock, Gemini, Cohere, etc.) contains a provider module that translates the incoming OpenAI-format request (messages, tools, stream, response_format) into the provider's native format and then maps the response back to ModelResponse — LiteLLM's pydantic model that mirrors openai.types.completion.ChatCompletion. For example litellm/llms/anthropic/ handles Anthropic's different tool-call, thinking, and streaming formats.
Streaming. Each provider normalizes its stream chunks into OpenAI-format ChatCompletionChunk objects via a CustomStreamWrapper, so downstream code sees the same SSE event structure. Tool calls and multimodal (image/audio/video inputs mapped to OpenAI content-block arrays) are translated per-provider.
Cost table. The 80,644-line model_prices_and_context_window.json at the repo root is the authoritative price list with input_cost_per_token, output_cost_per_token, max_input_tokens, mode, supports_* flags for every known model — parsed at runtime by cost_calculator.py:1-80.
Pass-through routes. The proxy also supports raw provider-API pass-through for endpoints it hasn't modeled (litellm/proxy/pass_through_endpoints/), allowing the proxy to act as a plain credential/rate-limit gate for any provider API.
Config classes that extend BaseConfig (llms/base_llm/chat/transformation.py) with transform_request/transform_response, called by BaseLLMHTTPHandler; BaseLLM.process_response in llms/base.py is a legacy template.QuantumNous/new-api
answeredThe project unifies 40+ provider APIs (listed in constant/channel.go:4-63) through an OpenAI-compatible schema as the lingua franca. Inbound requests arrive in OpenAI format (chat/completions, embeddings, images, audio, rerank, etc.) and are converted to each provider's native protocol via channel-specific Adaptor implementations.
Each channel type implements the channel.Adaptor interface defined in relay/channel/adapter.go:17-34. The key methods are ConvertOpenAIRequest, ConvertClaudeRequest, ConvertGeminiRequest, DoRequest, and DoResponse. The factory function GetAdaptor in relay/relay_adaptor.go:50-128 maps an APIType constant to a concrete adaptor (e.g., APITypeAnthropic → &claude.Adaptor{}).
Request translation lives in relaykit/relayconvert/internal/, organized by target protocol:
oai_chat/to_claude_messages_req.go— OpenAI chat → Claude Messagesoai_chat/to_gemini_chat_req.go— OpenAI chat → Geminiclaude_messages/to_oai_chat_req.go— Claude → OpenAI chatgemini_chat/to_oai_chat_req.go— Gemini → OpenAI chat
Response translation uses the reverse direction (*_resp.go files in the same directories). The service/convert.go wrapper exposes ResponseOpenAI2Claude, StreamResponseOpenAI2Gemini, etc. which call into relaykit conversion functions.
Streaming is handled by translating each SSE chunk from the upstream format back into OpenAI-compatible ChatCompletionsStreamResponse chunks. The DoResponse method on each adaptor receives the raw HTTP response and parses/extracts usage. The helper functions in relay/helper/common.go handle streaming headers (SetEventStreamHeaders) and format-specific chunk serialization (ClaudeData, etc.).
Tool calls and multimodal are fully converted: videos, images, audio, and tool_use blocks are translated between schemas in the internal converters.
Relay format selection happens in the router (router/relay-router.go:97-153): the controller receives a RelayFormat enum (OpenAI, Claude, Gemini, Embedding, etc.) based on the URL path structure. GetAndValidateRequest in relay/helper/valid_request.go:22 dispatches to the appropriate request parser per format. The relay_handler in controller/relay.go:33 then dispatches to a mode-specific relay helper (TextHelper, ImageHelper, AudioHelper, ResponsesHelper, GeminiHelper, etc.).
songquanpeng/one-api
answeredArchitecture. Every provider has an adaptor implementing the adaptor.Adaptor interface (5 methods: Init, GetRequestURL, SetupRequestHeader, ConvertRequest/ConvertImageRequest, DoRequest, DoResponse, GetModelList, GetChannelName). The dispatcher relay/adaptor.go:27-68 maps APIType → adaptor instance (27 provider adaptors registered). channeltype/helper.go:5-47 maps channel type → API type.
OpenAI-compatible schema. All incoming requests are parsed into relay/model/general.go's GeneralOpenAIRequest struct — a superset of the OpenAI chat/completions, embeddings, images, and audio schemas. This is the internal canonical form. The relaymode/helper.go dispatches by URL path to decide relay mode (chat, completions, embeddings, images, audio, proxy).
Per-provider translation. Adaptors like Anthropic's (relay/adaptor/anthropic/main.go:39-146) convert the OpenAI-shaped request to the provider's native format: messages become Anthropic Messages, system prompts are extracted, tool definitions are reshaped to Claude's input_schema, and multi-part content (text + image) is converted to Claude's content blocks. Responses flow back through reverse converters: ResponseClaude2OpenAI (anthropic/main.go:210-247) maps Claude's Response → OpenAI TextResponse, including tool_use → function_call translation. Streaming uses StreamResponseClaude2OpenAI (anthropic/main.go:149-208) which translates each Claude SSE event (message_start, content_block_start, content_block_delta, message_delta) into OpenAI SSE chunks.
Streaming and tools. The OpenAI adaptor (openai/adaptor.go) passes through for pure OpenAI-compatible providers. For others, streaming is re-encoded event-by-event. Tool calls are converted bidirectionally: e.g., Anthropic tool_use content blocks → OpenAI tool_calls with function type. The anthropic/main.go:21-37 stopReasonClaude2OpenAI maps Claude stop reasons (end_turn, tool_use) to OpenAI equivalents (stop, tool_calls).
Multimodal. Image content is handled via Message.ParseContent() (relay/model/message.go:41-80) which extracts text and image_url parts from the OpenAI content array. The Anthropic adaptor fetches and base64-encodes images (anthropic/main.go:132-139). The relay/controller/image.go handles image generation relay with size/quality validation and cost ratio computation.
Provider count. The channel type enum (relay/channeltype/define.go) defines 56 channel types. The adaptor registry (relay/adaptor.go) explicitly maps 27 API types; the remaining channel types (Azure, OpenAI-compatible, etc.) fall through to the OpenAI adaptor or Custom handling.
GetAdaptor maps 19 API types (not 27) and the Adaptor interface has nine methods; the other channel types fall through to the OpenAI adaptor.Portkey-AI/gateway
answeredAll providers present an OpenAI-compatible schema. Requests arrive in OpenAI format (or Anthropic format at /v1/messages). Each provider is a module under src/providers/ — there are 78 registered providers in src/providers/index.ts:78-151, from openai and anthropic to ovhcloud. Each module exports a ProviderConfigs object containing parameter configs (complete, chatComplete, embed etc.), an api object with getBaseURL, headers, and getEndpoint, and optional requestHandlers/responseTransforms.
Request translation happens in src/services/transformToProviderRequest.ts:75 via the transformUsingProviderConfig function. Each provider's chatComplete config (e.g. src/providers/openai/chatComplete.ts:12) maps gateway parameter names to provider-specific parameter names (the param field). It applies transforms, defaults, min/max clamping, and override merging from overrideParams. For example, OpenAI's max_tokens maps to "max_tokens" with a default of 100, while Anthropic might map it differently via its own config. The RequestContext class at src/handlers/services/requestContext.ts:211 calls transformToProviderRequest to produce the provider-specific payload.
Response translation flows through responseHandler in src/handlers/responseHandlers.ts:38, which selects a responseTransforms function from the provider config. For streaming, it uses functions like OpenAIChatCompleteJSONToStreamResponseTransform (openai/chatComplete.ts:157) to break a complete JSON response into data: ... SSE chunks with token-level deltas including tool calls, content blocks (text, thinking, data), and citations. Non-OpenAI providers (Anthropic, Bedrock, Vertex AI) each have their own response transforms that normalize to the OpenAI streaming format.
Streaming is always normalized to OpenAI SSE format (data: [DONE] termination, chat.completion.chunk object). Tool calls and multimodal (image content blocks, audio) are supported through the stream chunk transforms, with tool calls split into name and argument chunks per the OpenAI convention.
higress-group/higress
answeredProvider API unification is the central function of the ai-proxy plugin: all requests enter the gateway as OpenAI-format requests and are translated to each provider's native format.
API Name Resolution: Incoming request paths are mapped to ApiName constants (e.g. /v1/chat/completions → ApiNameChatCompletion) via suffix and regex matching — ai-proxy/main.go:56-104. Over 30 API types are defined across OpenAI, Anthropic, Cohere, Gemini, and Qwen formats — ai-proxy/provider/provider.go:37-82.
OpenAI-compatible schema is the internal standard: chatCompletionRequest, chatCompletionResponse, embeddingsRequest, imageGenerationRequest, etc. are defined in model.go with proper JSON field names — ai-proxy/provider/model.go:43-78. The protocol field can be set to "openai" (default) or "original" to bypass conversion — ai-proxy/provider/provider.go:411.
Request/response translation happens per-provider via the Provider interface with TransformRequestHeadersHandler, TransformRequestBodyHandler, and TransformResponseBodyHandler — ai-proxy/provider/provider.go:302-322. Each provider file (48 provider implementations) implements these interfaces. For example, openai.go passes through with path rewriting to api.openai.com, while claude.go translates to Anthropic's Messages API format (api.anthropic.com) — ai-proxy/provider/claude.go:18-20. The claude_to_openai.go converter handles bidirectional protocol translation for providers that don't natively support Claude's protocol — ai-proxy/provider/claude_to_openai.go:13-35.
Auto protocol detection converts Claude-format requests to OpenAI-format when the upstream provider doesn't natively support Claude — ai-proxy/main.go:254-266. The response is converted back to Claude format on streaming/non-streaming response callbacks — ai-proxy/main.go:688-771.
Streaming is handled by StreamingResponseBodyHandler or StreamingEventHandler interfaces, with SSE framing and event extraction — ai-proxy/provider/provider.go:290-296. The ExtractStreamingEvents function parses raw SSE chunks into structured StreamEvent objects, handling partial chunk buffering — ai-proxy/provider/provider.go:1130-1206.
Tool calls and multimodal are fully supported: the chatMessage struct includes tool_calls, ToolCallId, FunctionCall, and multimodal content types (text, image_url, input_audio, file) — ai-proxy/provider/model.go:207-224,348-370.
Provider count: 38 provider types are registered in the providerInitializers map — ai-proxy/provider/provider.go:236-275.
maximhq/bifrost
answeredBifrost unifies every provider through its internal BifrostRequest/BifrostResponse schema (core/schemas/bifrost.go), which defines type-safe request structs for chat completions, embeddings, responses, rerank, speech, transcription, and image/video generation. Each provider package (30+ in core/providers/) implements a conversion layer: ToBifrostChatRequest normalizes incoming provider-native requests, and ToXxxRequest/ToXxxResponse translates Bifrost's internal format to the provider's wire format. For example, openai/chat.go has ToBifrostChatRequest and ToOpenAIChatRequest that convert between the OpenAI ChatCompletion JSON shape and Bifrost's canonical BifrostChatRequest. Anthropic's anthropic/chat.go similarly converts Anthropic-specific constructs (document blocks, thinking blocks) into the common schema. Streaming is handled separately per provider — the fasthttp-based HTTP clients use SSE parsing for OpenAI-compatible streams and custom chunk logic for Anthropic/other providers. Tool calls, multimodal (images, audio, video), and structured output all flow through the same conversion pattern; the jsonparser plugin (plugins/jsonparser/main.go) handles partial JSON accumulation for streaming responses. The compat plugin (plugins/compat/main.go) bridges LiteLLM compatibility by converting text completions to chat or chat to responses when the target model lacks native support. The 30+ providers span OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Cerebras, Cohere, DeepSeek, ElevenLabs, Fireworks, Gemini, GitHub Copilot, Groq, HuggingFace, Mistral, Nebius, Ollama, OpenRouter, Perplexity, Replicate, Runware, Runway, Sarvam, SGL, Typesafe, vLLM, Wafer, and xAI.
katanemo/plano
answeredProvider API unification is done through a layered type system that normalizes all request shapes into one of 4 internal formats and translates responses back to the client's expected API.
The hermesllm crate owns this layer. Three client-facing APIs are accepted: OpenAI Chat Completions (/v1/chat/completions), Anthropic Messages (/v1/messages), and OpenAI Responses (/v1/responses), enumerated in SupportedAPIsFromClient (crates/hermesllm/src/clients/endpoints.rs:7). Upstream, the system translates to 5 possible shapes: OpenAI Chat Completions, Anthropic Messages API, Amazon Bedrock Converse (streaming and non-streaming), and OpenAI Responses API (SupportedUpstreamAPIs, same file lines 14-20).
Request translation is handled by ProviderRequestType (crates/hermesllm/src/providers/request.rs:17), an enum that wraps ChatCompletionsRequest, MessagesRequest, BedrockConverse, BedrockConverseStream, or ResponsesAPIRequest. Conversion between formats is chained via TryFrom<(ProviderRequestType, &SupportedUpstreamAPIs)> (line 319), which contains ~30 match arms covering every cross-format pair — for example ResponsesAPI → ChatCompletions → Anthropic Messages is done by chaining two conversions. The ProviderRequest trait provides a uniform interface (model(), set_messages(), get_messages(), is_streaming(), to_bytes()) across all types.
Provider-specific normalization is applied in normalize_for_upstream() (line 81) — the system strips xAI's deprecated web_search_options from chat completions, normalizes Moonshot's Kimi models (strips fixed sampling fields for first-party K3, only reasoning_effort kept), and adds mandatory instructions for ChatGPT Codex.
Response translation (crates/hermesllm/src/providers/response.rs:103) is symmetric: ProviderResponseType wraps ChatCompletionsResponse, MessagesResponse, or ResponsesAPIResponse, and conversion between formats goes through transform modules in crates/hermesllm/src/transforms/response/to_openai.rs and to_anthropic.rs. Bedrock responses are first converted to the canonical format, then to the client's target. Token usage is extracted via the TokenUsage trait with support for cached input tokens (cached_input_tokens()), cache creation tokens (cache_creation_tokens()), and reasoning tokens (reasoning_tokens()).
Streaming uses separate transform modules (crates/hermesllm/src/transforms/response_streaming/) that convert between provider streaming formats (SSE chunk processors). The system supports 27+ providers via the ProviderId enum, covering OpenAI, Anthropic, Deepseek, Gemini, Mistral, Groq, Amazon Bedrock, Moonshot AI, Zhipu, Qwen, DigitalOcean, OpenRouter, and more.
Tool calls and multimodal are supported through the OpenAI-style Message type which includes tool_calls, function fields, and content parts including image_url for images. These are preserved across format conversions.
tbphp/gpt-load
answeredProtocol translation is handled by the dialect package plus the execution/bifrost executor. The system defines 11 protocol.Protocol values (internal/protocol/protocol.go:6-19): OpenAICompletions, OpenAIResponses, OpenAIImages, OpenAIEmbeddings, Anthropic, Gemini, GeminiEmbeddings, Mistral, CodexLive, Rerank, and Decisions.
Dialect pattern — each protocol implements dialect.Dialect (internal/dialect/dialect.go:36), which has two methods: Protocol() returns its identity, and InspectRequest() parses a raw HTTP request into RequestMetadata. Implementations: dialect.OpenAI (internal/dialect/openai.go), dialect.Anthropic (internal/dialect/anthropic.go), dialect.Gemini (internal/dialect/gemini.go), dialect.Mistral (internal/dialect/mistral.go), and others for images, embeddings, and rerank. All JSON-based protocols share the inspectJSONRequestFields() helper.
Two execution modes — RouteNative sends the request using the client's original protocol wire format to an upstream that supports it. RouteConverted translates to a provider-neutral format via execution/bifrost (the Bifrost executor, internal/execution/bifrost/). The conversion layer includes chat_conversion.go, compatible.go, compatible_stream.go, responses_passthrough.go, etc. For example, an OpenAI-format chat request can be converted to Anthropic's format and vice versa.
Provider adapter registry — provideradapter.Registry (internal/provideradapter/registry.go:41-45) compiles ProviderKind-to-execution.Executor bindings. Each adapter implements RouteCapabilityValidator and declares which execution.Operation values it can handle natively vs converted. The execution/bifrost/executor.go is the primary multi-protocol executor.
Streaming — abstracted via StreamEvent (internal/dialect/stream_event.go:5), a provider-neutral SSE event representation. StreamEventClassifier (stream_event.go:38) classifies events as continue/completed/failed. UsageStreamEventObserver captures token usage from stream events. Multiple SSE transformers exist for native event gates (bifrost/model_alias.go).
Tool calls and multimodal — handled in the Bifrost conversion layer: tool_compatibility.go, compatible_conversion.go, and images_conversion_test.go cover function-calling conversion across wire formats. Multimodal (images in chat) goes through openai_images.go, openai_images_multipart.go for the dialect layer, and gemini_images_test.go for Gemini-specific image handling.
Provider count — the channel.ID type lists about 40 upstream provider channels (internal/channel/spec/definition.go:15-51), including OpenAI, Anthropic, Claude, Gemini, Grok, DeepSeek, Cohere, HuggingFace, Azure, AWS Bedrock, Google Vertex, Mistral, and various Chinese providers (ZhipuAI, Alibaba, SiliconFlow, MoonshotAI).
theagentrouter/agent-router
answeredProvider APIs are unified through a multi-layer translation architecture. The Translator interface (internal/translator/translator.go:48-83) defines RequestBody, ResponseBody, ResponseHeaders, and ResponseError methods, parameterized with request/response types (ReqT/SpanT). Ten endpoint-specific translator types are defined, covering chat completions, completions, embeddings, image generation, responses, speech, transcription, translation, rerank, tokenize, and count-tokens (internal/translator/translator.go:132-160). Nine API schemas are supported: OpenAI, Cohere, AWSBedrock, AzureOpenAI, GCPVertexAI, GCPAnthropic, Anthropic, AWSAnthropic, AWSOpenAI, and TypeSafe (internal/apischema/typesafe/systemone.go). The EndpointSpec.GetTranslator() method selects the correct translator based on the backend's API schema — for example, ChatCompletionsEndpointSpec.GetTranslator() has seven schema branches, translating OpenAI input to each backend's native format (internal/endpointspec/endpointspec.go:202-221). OpenAI-compatible schema is the canonical input format: all requests enter as OpenAI-shaped JSON and are translated to the backend's format. Streaming is handled per-endpoint: the upstream processor sets streaming mode in response headers when stream=true and processes streamed SSE chunks (internal/extproc/processor_impl.go:517-520,538-666). Tool calls and multimodal content are supported through the full OpenAI chat completions schema including tool_calls, ToolCalls, ContentPart, image URL, audio data, and file parts — all of which are preserved through translation (internal/endpointspec/endpointspec.go:689-859). Each translator pair (e.g., openai_awsbedrock.go, anthropic_gcpanthropic.go, openai_azureopenai.go) handles the specific request/response format conversion between OpenAI input and the target backend API.
← How are requests routed across providers and models? · How are API keys, users and tenants managed? →