# Which models are supported and how are they called?

> Open-source coding agents — a good answer covers: Providers; local models; tool calling vs text formats; per-model prompt tuning; cost tracking.

Canonical page: https://llms-technical-reviews.com/coding-agents/q/models/

## Verdict

[Aider](/p/aider/) is the only coding agent that works without native tool calling, so it suits small and local models best. [Codex](/p/codex/) is the narrowest, with one wire protocol. Among the app builders, real model choice is thinner than the docs suggest.

**Text edit formats, any model.** [Aider](/p/aider/) calls everything through LiteLLM. 357 entries in `model-settings.yml` tune the edit format and helper models per model.

**Native tool calling over broad gateways.**
- [Cline](/p/cline/): its own `@cline/llms` gateway with 211 generated provider IDs. It has no text-format fallback.
- [OpenCode](/p/opencode/): the models.dev catalog plus AI SDK packages, with per-family system prompts.
- [Codewhale](/p/codewhale/): about 50 provider kinds and three wire clients. `model = "auto"` picks a big or cheap model, and its decision router can call a hosted third-party service.
- [Qwen Code](/p/qwen-code/): five adapters, all translated into `@google/genai` shapes, and 13 provider presets.
- [PI-Desktop](/p/pi-desktop/): seven wire styles through pi-ai, including the ChatGPT/Codex subscription endpoint.
- [DeepSeek-Reasonix](/p/deepseek-reasonix/): three provider kinds. It picks the reasoning parameters from the base URL and ships per-vendor price tables.
- [Open Interpreter](/p/openinterpreter/): picks a harness per provider and model, for example `claude-code-bare` for DeepSeek. Most harnesses need a Chat Completions endpoint.

**One protocol.** [Codex](/p/codex/) speaks only the OpenAI Responses API. The built-in providers are OpenAI, Bedrock, Ollama and LM Studio, and any other provider must speak Responses.

**App builders.**
- [bolt.diy](/p/bolt.diy/): 22 provider classes, but only OpenRouter, Anthropic, OpenAI and Google are registered. Ollama and LM Studio need a source edit despite the README.
- [Open Lovable](/p/open-lovable/): four providers, no tool calling, one prompt for every model, and no local models.
- [VibeSDK](/p/vibesdk/): the default Think engine is hard-wired to a Gemini Flash model through AI Gateway. The per-action model settings only reach the legacy engines.

Pick: Aider for local or weak models that cannot call functions.
Pick: OpenCode, Cline or Codewhale for the widest choice of tool-capable providers. DeepSeek-Reasonix is the pick if cost tracking for DeepSeek-class pricing matters.
Pick: Codex if you are on OpenAI models and want the best-tuned path.

More projects in this category are being researched.

## Per-project answers

### anomalyco/opencode (answered)

## Model Support

OpenCode uses the **Vercel AI SDK** (`ai` package) as its primary LLM abstraction layer, with an opt-in native LLM runtime (`@opencode-ai/llm`). The `Provider` service (`provider/provider.ts`) loads model configurations from a catalog (`@opencode-ai/core/models-dev.ts`), which defines models with fields for context limits, output limits, cost tiers, reasoning options, modalities (text/image/audio/pdf), temperature support, tool-call support, and per-provider npm packages.

**Providers** are resolved from the catalog and configured via `@ai-sdk/*` npm packages. The `sdkKey()` function in `provider/transform.ts` (lines 41-98) maps npm packages to AI SDK provider option keys: `@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/google`, `@ai-sdk/amazon-bedrock`, `@ai-sdk/mistral`, `@ai-sdk/groq`, `@ai-sdk/xai`, `@ai-sdk/cohere`, `@ai-sdk/perplexity`, `@ai-sdk/togetherai`, `@ai-sdk/azure`, `@ai-sdk/alibaba`, `@ai-sdk/cerebras`, `@ai-sdk/deepinfra`, `@ai-sdk/vercel`, `@ai-sdk/gateway`, plus venice, openrouter, gitlab-ai-provider, and OpenAI-compatible providers via `@ai-sdk/openai-compatible`. Plugins add auth support for GitHub Copilot, Codex, GitLab, Poe, Cloudflare, Azure, DigitalOcean, Snowflake Cortex, xAI, Cerebras, and Modal.

**Local models** are supported via OpenAI-compatible endpoints (`@ai-sdk/openai-compatible`), enabling any local LLM server (like Ollama, LM Studio, or vLLM) that exposes an OpenAI-compatible API.

**Tool-calling vs text formats**: Models with `tool_call: true` in the catalog definition get full tool definitions. The AI SDK handles tool-calling natively; OpenCode normalizes the stream events through `LLMAISDK.toLLMEvents` (`session/llm/ai-sdk.ts`). For providers without tool-calling, the model would only produce text.

**Per-model prompt tuning** is in `session/system.ts` (lines 28-51): Anthropic models get one system prompt (`anthropic.txt`), GPT-4/o1/o3 get `beast.txt`, newer GPTs get `astra.txt` or `codex.txt`, Gemini gets `gemini.txt`, Kimi gets `kimi.txt`, Meta/Muse gets `meta.txt`, and everything else gets `default.txt`.

**Provider-specific message transforms** in `provider/transform.ts` handle quirks: Anthropic models filter empty content parts and normalize tool call IDs (lines 170-194, 224-251), Bedrock uses signature-based reasoning, Mistral pads/scrambles tool call IDs (lines 253-277), and OpenAI uses `include` for encrypted reasoning (line 23).

**Cost tracking** uses `Session.getUsage()` (`session.ts` lines 338-405), which reads cost tiers from the model catalog, handles cache read/write pricing, and accounts for provider-specific metadata keys (Anthropic vertex, Bedrock, Venice).


Citations: [packages/opencode/src/provider/transform.ts:41-98](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/provider/transform.ts#L41-L98) · [packages/opencode/src/session/system.ts:28-51](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/session/system.ts#L28-L51) · [packages/opencode/src/provider/transform.ts:100-280](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/provider/transform.ts#L100-L280) · [packages/opencode/src/session/session.ts:338-405](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/session/session.ts#L338-L405) · [packages/opencode/src/provider/provider.ts:1-33](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/provider/provider.ts#L1-L33) · [packages/opencode/src/session/llm.ts:226-278](https://github.com/anomalyco/opencode/blob/4ac0d9c3d169bbe81d9570013effdda3fe24d36e/packages/opencode/src/session/llm.ts#L226-L278)

### openai/codex (answered)

Codex CLI connects primarily to **OpenAI's Responses API** as the default provider, with model catalog fetched from `api.openai.com` and cached in `models_cache.json` (managed by `codex-rs/models-manager/src/manager.rs:34`). The `ModelProvider` trait in `codex-rs/model-provider/src/provider.rs:108` defines the abstraction for model backends, with implementations for: OpenAI's production API (the default, via `CoreAuthProvider` in lib.rs:27), **Amazon Bedrock** (`amazon_bedrock` in lib.rs:17), **LM Studio** (`codex-rs/lmstudio/`), and **Ollama** (`codex-rs/ollama/`). The `ModelsEndpointClient` trait (manager.rs:42) handles remote catalog fetching, while `static` and `OpenAI` models managers serve the bundled catalog (`bundled_models_response()` in lib.rs:13 from a bundled `models.json`). Model selection uses a **pinned model** per session, with the option to switch mid-conversation. The `ModelInfo` struct carries per-model metadata: context window size, reasoning-effort defaults, input/output modalities, auto-compact token limits, and `ToolMode` (standard, code-mode-only, or code-mode — `core/src/tools/mod.rs:75`). Models are called through the OpenAI Responses API streaming endpoints via the `ModelClientSession` in `core/src/client.rs:1`, which manages WebSocket connections for sticky routing, retries with exponential backoff (`core/src/responses_retry.rs:33`), and transport fallback. Per-model prompts are rendered via `render_model_instructions()` in `prompts/src/model_instructions.rs:8`, which uses the model's instruction template from its metadata. Cost tracking uses `TokenUsageInfo` collected per turn and aggregated in analytics events. Reasonable efforts (`ReasoningEffort`) controls are configurable per-turn.


Citations: [codex-rs/model-provider/src/provider.rs:108-150](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/model-provider/src/provider.rs#L108-L150) · [codex-rs/models-manager/src/manager.rs:34-72](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/models-manager/src/manager.rs#L34-L72) · [codex-rs/model-provider/src/lib.rs:17-36](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/model-provider/src/lib.rs#L17-L36) · [codex-rs/core/src/client.rs:1-26](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/core/src/client.rs#L1-L26) · [codex-rs/core/src/responses_retry.rs:27-46](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/core/src/responses_retry.rs#L27-L46) · [codex-rs/prompts/src/model_instructions.rs:1-18](https://github.com/openai/codex/blob/ccde2fc8b7fd08216757d3d3f67362847ab8341b/codex-rs/prompts/src/model_instructions.rs#L1-L18)

### cline/cline (answered)

The **`@cline/llms`** package provides a gateway-based provider architecture (`gateway.ts`). A central **GatewayRegistry** routes requests to registered provider handlers. Supported vendors include **Anthropic** (`vendors/anthropic.ts`), **OpenAI** (`vendors/openai.ts`), **Google** (`vendors/google.ts`), **Mistral** (`vendors/mistral.ts`), **AWS Bedrock** (`vendors/bedrock.ts`), **Vertex AI** (`vendors/vertex.ts`), **Ollama** (local models, `vendors/ollama.ts`), **OpenAI-compatible** APIs (`vendors/openai-compatible.ts`), a **community provider** system (`vendors/community.ts`), the **Cline** managed API (`vendors/cline.ts`), and **MiniMax Thinking** (`vendors/minimax-thinking.ts`). There is also a generic **ai-sdk** provider (`providers/ai-sdk.ts`). Each provider implements a common `ApiHandler` interface that can operate in **tool-calling mode** (when the model supports it) or **text-format mode** (shimmed via structured prompts). A generated provider spec file (`providers.generated.ts`) and model catalog (`catalog.generated.ts`) register known model IDs, capabilities, and operation support. The **model-tools** system (`model-tools.ts`) detects which models support native tool definitions vs. requiring tool-use prompt shims. **Cost tracking** is handled through billing module (`billing.ts`) and telemetry events. **Per-model prompt tuning** is implemented via routing rules: reasoning effort, cache point placement, and max-token budgeting are adjusted per-provider/model in the routing modules (`routing/anthropic-compatible.ts`, `routing/bedrock-cache-point.ts`, `routing/reasoning-options.ts`).

> **Editor's note.** Correction: there is no text-format tool shim; all calls use AI SDK native tool calling. `model-tools.ts` resolves provider-side model tools such as `web_search`, not native tool support.

Citations: [sdk/packages/llms/src/providers/gateway.ts:1-40](https://github.com/cline/cline/blob/cd80a20e96481f5f5d413789f6847accf846487b/sdk/packages/llms/src/providers/gateway.ts#L1-L40) · [sdk/packages/llms/src/providers/vendors/anthropic.ts:1-10](https://github.com/cline/cline/blob/cd80a20e96481f5f5d413789f6847accf846487b/sdk/packages/llms/src/providers/vendors/anthropic.ts#L1-L10) · [sdk/packages/llms/src/providers/vendors/openai-compatible.ts:1-10](https://github.com/cline/cline/blob/cd80a20e96481f5f5d413789f6847accf846487b/sdk/packages/llms/src/providers/vendors/openai-compatible.ts#L1-L10) · [sdk/packages/llms/src/providers/builtins.ts:1-60](https://github.com/cline/cline/blob/cd80a20e96481f5f5d413789f6847accf846487b/sdk/packages/llms/src/providers/builtins.ts#L1-L60) · [sdk/packages/llms/src/providers/routing/anthropic-compatible.ts:1-30](https://github.com/cline/cline/blob/cd80a20e96481f5f5d413789f6847accf846487b/sdk/packages/llms/src/providers/routing/anthropic-compatible.ts#L1-L30)

### openinterpreter/openinterpreter (answered)

Models use a provider abstraction: ModelProvider (model-provider/src/provider.rs:42-73) with ProviderCapabilities (namespace_tools, image_generation, web_search, remote_compaction). Catalog: models-manager/models.json (1663 lines) with slugs, context_window, max_context_window, reasoning levels (low/medium/high/xhigh/max), tool_mode, input_modalities, comp_hash. ModelsManager (models-manager/src/manager.rs) provides OpenAiModelsManager and StaticModelsManager backed by ModelsCache. Provider resolution via create_model_provider dispatches to Amazon Bedrock, OpenAI/Anthropic, or combined auth (model-provider/src/provider.rs:39). bundled_provider_catalog (model-provider-info/src/bundled_provider_catalog.rs) serves offline data. Model calls go through ModelClient (core/src/client.rs); turn-scoped ModelClientSession caches WebSocket connections. Per-turn settings (model, reasoning_effort, service_tier) pass to client_session.stream(). Tool calling: Prompt.tools = step_context.tool_router.model_visible_specs() (turn.rs:1566). Cost tracking via SessionTelemetry and AnalyticsEventsClient.

> **Editor's note.** Correction: the answer misses the harness layer: `resolve_stream_transport_route` maps (wire_api = responses|chat|messages, harness) to a transport and per-harness request builder (core/src/harness/routing.rs), and `default_harness_for_provider_model` picks claude-code, kimi-code, qwen-code, zcode or claude-code-bare from the provider/model.

Citations: [codex-rs/model-provider/src/provider.rs:42-73](https://github.com/openinterpreter/openinterpreter/blob/2767e5f20d6927500b8f1938c773c61afb823245/codex-rs/model-provider/src/provider.rs#L42-L73) · [codex-rs/models-manager/models.json:1-60](https://github.com/openinterpreter/openinterpreter/blob/2767e5f20d6927500b8f1938c773c61afb823245/codex-rs/models-manager/models.json#L1-L60) · [codex-rs/core/src/client.rs:1-17](https://github.com/openinterpreter/openinterpreter/blob/2767e5f20d6927500b8f1938c773c61afb823245/codex-rs/core/src/client.rs#L1-L17) · [codex-rs/core/src/session/turn.rs:1558-1580](https://github.com/openinterpreter/openinterpreter/blob/2767e5f20d6927500b8f1938c773c61afb823245/codex-rs/core/src/session/turn.rs#L1558-L1580) · [codex-rs/model-provider/src/lib.rs:1-40](https://github.com/openinterpreter/openinterpreter/blob/2767e5f20d6927500b8f1938c773c61afb823245/codex-rs/model-provider/src/lib.rs#L1-L40)

### Aider-AI/aider (answered)

**Provider-agnostic via litellm.** All model calls go through `litellm.completion()` (`aider/llm.py` lazily imports litellm). This means any model that litellm supports — OpenAI, Anthropic, Gemini, DeepSeek, Ollama, Together, OpenRouter, AWS Bedrock, Azure, Cohere, Replicate, etc. — is usable. API keys are set via environment variables (e.g. `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) or the `--api-key` flag (`aider/main.py:600-630`).

**Built-in known model lists.** `aider/models.py:33-123` defines `OPENAI_MODELS` (30+ variants from gpt-3.5-turbo through o1, o3-mini, gpt-5.5), `ANTHROPIC_MODELS` (claude-2 through claude-sonnet-4-6, haiku-4-5, opus-4-7) and `MODEL_ALIASES` (e.g. "sonnet" → "claude-sonnet-4-6").

**Per-model prompt tuning.** The `ModelSettings` dataclass (`models.py:128-151`) carries per-model parameters: `edit_format` (diff/whole/udiff/etc), `use_repo_map`, `use_system_prompt`, `use_temperature`, `reminder` style ("user" or "sys"), `cache_control`, `examples_as_sys_msg`, `extra_params`, and `accepts_settings` (e.g. "thinking_tokens"). Settings are loaded from `aider/resources/model-settings.yml` (`models.py:153-158`) and can be extended with user-provided `.aider.model.settings.yml` files registered via `register_models()` (`main.py:335-358`). Additional per-model heuristics in `Model.__init_post__()` (`models.py:490-598`) set defaults based on model name patterns — DeepSeek R1 gets `reasoning_tag="think"` and `use_temperature=False`, GPT-4-turbo gets `edit_format="udiff"`, o1 models disable `use_system_prompt`.

**Tool calling vs text formats.** When a model has a function-based coder (like `WholeFileFunctionCoder`), the LLM receives `tools` in the litellm API call (`models.py:1006-1009`). Otherwise edits are text-based and parsed from the model's plain-text response. The `edit_format` field on the model determines which.

**Cost tracking.** The `ModelInfoManager` (`models.py:161`) fetches `model_prices_and_context_window.json` from the litellm GitHub repo (cached for 24h). The OpenRouter model database is also cached locally. Costs are calculated after each completion (`base_coder.py:1811`).


Citations: [aider/models.py:33-123](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/models.py#L33-L123) · [aider/models.py:128-200](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/models.py#L128-L200) · [aider/models.py:985-1038](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/models.py#L985-L1038) · [aider/llm.py:1-47](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/llm.py#L1-L47) · [aider/models.py:490-600](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/models.py#L490-L600) · [aider/main.py:600-630](https://github.com/Aider-AI/aider/blob/5dc9490bb35f9729ef2c95d00a19ccd30c26339c/aider/main.py#L600-L630)

### codewhale-hq/Codewhale (answered)

Codewhale is **provider-neutral** and supports 46+ provider kinds through a unified `LlmClient` async trait.

**Supported providers** (`config/src/provider_kind.rs`): DeepSeek, OpenAI, Anthropic (native Messages API), Google (Gemini), Mistral AI, Ollama, Ollama Cloud, Hugging Face, Together, Groq, Fireworks, Novita, OpenRouter, Grok (xAI), Meta (Muse Spark), and many more. Custom providers use the `Custom` kind for arbitrary OpenAI-compatible endpoints. Dual-wire variants (`DeepSeekAnthropic`, `MinimaxAnthropic`, `ModelstudioTokenPlanAnthropic`) select the Anthropic Messages wire protocol instead of OpenAI Chat Completions.

**Wire protocols:** Three client implementations in `crates/tui/src/client/`:
- **`chat.rs`** — OpenAI Chat Completions (the production path for most providers): streaming, request building, SSE parsing.
- **`anthropic.rs`** — Native Anthropic Messages API: handles adaptive thinking, `cache_control` breakpoints, signed-thinking signatures, usage normalization.
- **`responses.rs`** — OpenAI Responses API (for newer OpenAI endpoints).
The `LlmClient` trait (`llm_client/mod.rs:54`) defines `create_message`, `create_message_stream`, and `provider_name`/`model()` with retry logic (`with_retry`), exponential backoff, jitter, and error classification.

**Model selection:**
- **Auto model detection** (`auto_model.rs`): For DeepSeek, a rule-based classifier scores the user's prompt to choose between "deepseek-v4-flash" (simple tasks) and "deepseek-v4-pro" (complex).
- **Model catalog** (`catalog.rs:1`): Layered resolution — bundled Models.dev snapshot < live Models.dev < corrections < signed cloud facts < live provider `/v1/models` < user config overrides.
- **Fleet roster**: Role-based model assignment for sub-agents via `FleetRoster` and `FleetRole`, where each role can pin a specific provider/model.
- **Route resolver** in `config/src/route/` compiles catalog rows into executable routes with resolved `base_url`, wire protocol, and limits.

**Per-model prompt tuning:** System prompt shape adapts per mode (Plan/Act/Operate). Provider-specific request shaping includes: `max_completion_tokens` vs `max_tokens` selection (`chat.rs:46`), reasoning effort parameters (`chat.rs:67`), token limits per route, and provider-specific tool-choice modes (`strict_tool_mode` for DeepSeek beta).

**Cost tracking:** `Usage` struct with input/output/cache tokens, `CostScopeToken` for billing attribution, and route-level `RouteLimits` with per-provider pricing.

> **Editor's note.** Correction: `crates/config/src/auto_model.rs` is a legacy DeepSeek-only scorer kept for API compatibility and is not used for `model = "auto"`; auto routing is done in `crates/tui/src/model_routing.rs`, which picks a big/cheap pair for the active provider via a chat router or a TypeSafe/OpenRouter decision router.

Citations: [crates/config/src/provider_kind.rs:1-303](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/provider_kind.rs#L1-L303) · [crates/tui/src/llm_client/mod.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/llm_client/mod.rs#L1-L80) · [crates/tui/src/client/chat.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/client/chat.rs#L1-L80) · [crates/tui/src/client/anthropic.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/client/anthropic.rs#L1-L80) · [crates/config/src/auto_model.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/auto_model.rs#L1-L60) · [crates/config/src/catalog.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/catalog.rs#L1-L60)

### esengine/DeepSeek-Reasonix (answered)

Models are supported through two provider implementations registered in a `provider.Registry`. **OpenAI-compatible** (`internal/model/openai/openai.go:1`): registered as kind `"openai"`, calls `/chat/completions` with SSE streaming. Per-vendor wire-shape detection selects different reasoning parameters: `api.deepseek.com` uses `thinking.type=enabled` with `reasoning_effort`; `api.minimaxi.com` uses `thinking.type=adaptive|disabled` (M3's binary knob); `open.bigmodel.cn` / `api.z.ai` (Zhipu GLM) have documented depth fields; `api.longcat.chat` uses `thinking.type=enabled|disabled` without effort scale; `ollama.com` accepts hosted Ollama Cloud's effort scale including `max`; Kimi K3 preserves complete messages and uses `max_completion_tokens`. Everything else uses vanilla `reasoning_effort`. **Anthropic** (`internal/model/anthropic/anthropic.go:1`): registered as kind `"anthropic"`, calls POST `/v1/messages` with SSE streaming. Supports extended thinking via signed reasoning blocks that are replayed on multi-turn. No temperature/top_p (Claude rejects sampling params). DeepSeek's compatible Anthropic endpoint uses unsigned thinking blocks with binary enabled|disabled and output_config.effort. Configuration is via `reasonix.toml` `[provider]` blocks with `kind`, `base_url`, `model`, `api_key`, and kind-specific extras. Providers resolve via `config.go` (`internal/contract/provider/config.go:6`) and the provider broker serves them. `local.go` (`internal/model/providerbroker/local.go`) provides no local-model fallback; there is no Ollama-hosted local inference path — the Ollama integration reaches hosted cloud API. Cost tracking uses `pricing.rates.go` (`internal/contract/pricing/rates.go:36`) with official rates for DeepSeek, Anthropic Claude, OpenAI GPT, MiniMax, LongCat, Moonshot (Kimi), and others, stored as oldest-first generations per vendor. `make pricecheck` reads vendor pages to verify rates.

> **Editor's note.** Correction: there are three provider kinds, not two; `responses` (also `dashscope-responses`, internal/model/responses/responses.go) targets the OpenAI Responses API alongside `openai` and `anthropic`.

Citations: [internal/model/openai/openai.go:1-80](https://github.com/esengine/DeepSeek-Reasonix/blob/7f30fcbdb371585cd684947723097ec67ba4cf4f/internal/model/openai/openai.go#L1-L80) · [internal/model/anthropic/anthropic.go:1-80](https://github.com/esengine/DeepSeek-Reasonix/blob/7f30fcbdb371585cd684947723097ec67ba4cf4f/internal/model/anthropic/anthropic.go#L1-L80) · [internal/contract/provider/vendor.go:1-80](https://github.com/esengine/DeepSeek-Reasonix/blob/7f30fcbdb371585cd684947723097ec67ba4cf4f/internal/contract/provider/vendor.go#L1-L80) · [internal/contract/pricing/rates.go:1-80](https://github.com/esengine/DeepSeek-Reasonix/blob/7f30fcbdb371585cd684947723097ec67ba4cf4f/internal/contract/pricing/rates.go#L1-L80)

### firecrawl/open-lovable (answered)

Four providers are supported via the Vercel AI SDK (`@ai-sdk/*`): Anthropic, OpenAI, Google Generative AI, and Groq. All are wrapped in `lib/ai/provider-manager.ts`. The `getProviderForModel()` function (`provider-manager.ts:79-117`) routes based on model name prefixes: `anthropic/` → Anthropic SDK, `openai/` → OpenAI SDK, `google/` → Google Gemini SDK, and `moonshotai/kimi-k2-instruct-0905` → Groq SDK. Without a prefix it defaults to Groq. Models listed in `config/app.config.ts:56-62`: `openai/gpt-5`, `moonshotai/kimi-k2-instruct-0905`, `anthropic/claude-sonnet-4-20250514`, `google/gemini-3-pro-preview` (the default). Custom model routing can be configured via `modelApiConfig` in the app config (`provider-manager.ts:81-86`). Vercel AI Gateway is supported as a unified proxy — all requests go to `https://ai-gateway.vercel.sh/v1` when `AI_GATEWAY_API_KEY` is set (`provider-manager.ts:21-23`). **Tool/function calling is not used** — the code at `generate-ai-code-stream/route.ts:1310-1311` explicitly notes no support in the available models. Instead, XML tags in text output are parsed. `generateObject()` (Vercel AI SDK) is used solely for the edit-intent search-plan endpoint with Zod schemas (`analyze-edit-intent/route.ts:34-60,126`). Temperature defaults to 0.7 for non-reasoning models and is unset for GPT-5. Max tokens is 8192. GPT-5 uses `reasoningEffort: 'high'` (`route.ts:1320-1326`). No local models are supported. No cost tracking or per-model prompt tuning exists — the same massive system prompt template is used for all models.


Citations: [lib/ai/provider-manager.ts:79-117](https://github.com/firecrawl/open-lovable/blob/69bd93bae7a9c97ef989eb70aabe6797fb3dac89/lib/ai/provider-manager.ts#L79-L117) · [config/app.config.ts:53-71](https://github.com/firecrawl/open-lovable/blob/69bd93bae7a9c97ef989eb70aabe6797fb3dac89/config/app.config.ts#L53-L71) · [app/api/generate-ai-code-stream/route.ts:1310-1326](https://github.com/firecrawl/open-lovable/blob/69bd93bae7a9c97ef989eb70aabe6797fb3dac89/app/api/generate-ai-code-stream/route.ts#L1310-L1326) · [app/api/analyze-edit-intent/route.ts:34-60](https://github.com/firecrawl/open-lovable/blob/69bd93bae7a9c97ef989eb70aabe6797fb3dac89/app/api/analyze-edit-intent/route.ts#L34-L60)

### QwenLM/qwen-code (answered)

Models are accessed through a **multi-protocol** architecture. `AuthType` (from `core/contentGenerator.ts`) defines supported protocols: `openai`, `openai-responses`, `anthropic`, `gemini`, `vertex-ai`, and `qwen-oauth`. Constants in `packages/core/src/models/constants.ts` (lines 74–105) map each protocol to environment variables (e.g., `OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, `GEMINI_API_KEY`).

There are **13 provider presets** registered in `packages/core/src/providers/all-providers.ts` (lines 55–69): Alibaba (coding plan, token plan, standard), DeepSeek, Grok, Minimax, Z.ai, Moonshot, IdeaLab, ModelScope, OpenRouter, Requesty, and a generic Custom provider. Each provider defines `envKey`, `baseUrl`, `modelNamePrefix`, and optional per-model generation config (thinking, context window, modalities) — built by `provider-config.ts`.

**Content generators** translate the protocol into actual API calls: `OpenAIContentGenerator` (`core/openaiContentGenerator/`) for all OpenAI-compatible backends, `AnthropicContentGenerator` (`core/anthropicContentGenerator/`), and `GeminiChat`/`geminiContentGenerator` for Gemini. The model registry (`models/modelRegistry.ts`) resolves a provider ID + model config into the correct `AuthType` via `resolveModelProtocol()`, supporting `wireApi` routing between `chat-completions` and `responses` (lines 69–92). Per-provider tuning lives in `openaiContentGenerator/provider/` (DeepSeek, Fireworks, Grok, Mistral, etc.). Cost tracking is in `telemetry/loggers.ts`. Local models (Ollama/vLLM) connect via the Custom provider.


Citations: [packages/core/src/models/constants.ts:74-105](https://github.com/QwenLM/qwen-code/blob/ac81c07dc8cc13a7a6d22d3af7b9fcc3f4858607/packages/core/src/models/constants.ts#L74-L105) · [packages/core/src/providers/all-providers.ts:1-70](https://github.com/QwenLM/qwen-code/blob/ac81c07dc8cc13a7a6d22d3af7b9fcc3f4858607/packages/core/src/providers/all-providers.ts#L1-L70) · [packages/core/src/providers/provider-config.ts:52-101](https://github.com/QwenLM/qwen-code/blob/ac81c07dc8cc13a7a6d22d3af7b9fcc3f4858607/packages/core/src/providers/provider-config.ts#L52-L101) · [packages/core/src/models/modelRegistry.ts:69-92](https://github.com/QwenLM/qwen-code/blob/ac81c07dc8cc13a7a6d22d3af7b9fcc3f4858607/packages/core/src/models/modelRegistry.ts#L69-L92)

### stackblitz-labs/bolt.diy (answered)

Bolt supports 20+ providers registered in `app/lib/modules/llm/registry.ts`. Only 4 are enabled by default via `ENABLED_PROVIDERS` in the manager (`OpenRouter`, `Anthropic`, `OpenAI`, `Google`), but all 22 registered providers (Anthropic, OpenAI, Google, Groq, Cohere, DeepSeek, HuggingFace, Mistral, Ollama, OpenRouter, LMStudio, AmazonBedrock, Perplexity, Together, xAI, Fireworks, Hyperbolic, Cerebras, GitHub, Moonshot, OpenAILike, Z-AI) can be enabled. Each provider extends `BaseProvider` (`app/lib/modules/llm/base-provider.ts`) which provides common infrastructure: getting base URL and API key from environment variables, cookie settings, or server env; Docker URL rewriting; and model caching with cache-key generation. The `getModelInstance` method creates an SDK model instance per provider — for example, `AnthropicProvider` uses `createAnthropic()` from `@ai-sdk/anthropic`, `OllamaProvider` uses `createOllama()` from `ollama-ai-provider`. Models can be static (hardcoded in the provider class, e.g., Anthropic's three fallback models) or dynamic (fetched at runtime via `getDynamicModels` — Ollama fetches from `/api/tags`, Anthropic from `api.anthropic.com/v1/models`). The `LLMManager` singleton (`app/lib/modules/llm/manager.ts`) orchestrates provider registration and model list updates. It caches dynamic models based on a cache key derived from API keys and settings. The Vercel AI SDK serves as the unified interface: `streamText` in the SDK handles streaming, tool calling, token counting, and error handling. Reasoning models (o1, o3, GPT-5) get special handling — they use `maxCompletionTokens` instead of `maxTokens`, and have unsupported parameters (temperature, topP, etc.) filtered out (stream-text.ts lines 240-280). Per-model prompt tuning is handled by the `PromptLibrary` which supports swapping system prompts (`default`, `original`, `optimized`). No cost tracking is implemented in this codebase.

> **Editor's note.** Correction: at this commit LLMManager registers only OpenRouter, Anthropic, OpenAI and Google (ENABLED_PROVIDERS in manager.ts). The other 18 provider classes, including Ollama and LM Studio, are skipped and need a source edit to become usable.

Citations: [app/lib/modules/llm/manager.ts:14-239](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/modules/llm/manager.ts#L14-L239) · [app/lib/modules/llm/registry.ts:1-47](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/modules/llm/registry.ts#L1-L47) · [app/lib/modules/llm/base-provider.ts:10-176](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/modules/llm/base-provider.ts#L10-L176) · [app/lib/modules/llm/providers/anthropic.ts:1-136](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/modules/llm/providers/anthropic.ts#L1-L136) · [app/lib/modules/llm/providers/ollama.ts:30-134](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/modules/llm/providers/ollama.ts#L30-L134) · [app/lib/.server/llm/stream-text.ts:228-295](https://github.com/stackblitz-labs/bolt.diy/blob/db479e84ac7ad905bbaafc6aed097f7c0a87c849/app/lib/.server/llm/stream-text.ts#L228-L295)

### vastsa/PI-Desktop (answered)

**PI-Desktop supports multiple model providers through the `pi-ai` library**, with a universal provider model (ADR 0012).

**Wire protocols** — `provider-binding.ts:129-176` maps user-configured `apiStyle` values to pi-ai wire adapters: `chat_completions` (OpenAI-compatible), `responses` (OpenAI Responses), `anthropic_messages` (Claude), `openai_codex_responses` (ChatGPT/Codex), `pi_messages` (Pi service), `google_generative_ai` (Gemini), and `opencode_go` (OpenCode Go API). If no explicit style is set, it defaults to OpenAI Chat Completions (`provider-binding.ts:170-175`).

**Model resolution** — Providers are stored with `apiStyle`, `baseUrl`, `modelId`, `apiKey`, and optional `modelConfig` with per-model overrides. The `apiBindingForProviderModel()` function (`provider-binding.ts:194-196`) resolves the correct wire protocol considering both provider-level and model-level API styles. Vendor accounts (GitHub Copilot, Pi, custom) are handled with special auth paths including OAuth token refresh (`provider-binding.ts:76`).

**Thinking/reasoning** — Models can be configured with thinking levels via `SessionThinkingLevel` / `SubagentThinkingLevel`. The `thinking-level.ts` module clamps and normalizes these. Claude models with adaptive thinking vs. budget-token thinking are handled differently (`provider-binding.ts:298+`). DeepSeek Flash gets special tool-declaration optimizations (`fixed-tool-declarations.ts:29-42`).

**Cost tracking** — Usage is tracked via `MessageUsage` objects accumulating tokens and costs per turn, surfaced through `request-usage.ts`. The `provider-retry.ts` module handles rate limits (up to 5 retries) and transient errors (up to 3 retries) at the provider level.

**Subagent models** — Delegations can pin their own model via `SubagentDefinition`. Electron main resolves credentials; subagent model keys must be explicitly opted-in for Task.model overrides via `subagentModelKeys` (`runtime.ts:1037-1039`). Missing or failed model bindings fall through a fallback chain.


Citations: [packages/agent-runtime/src/provider-binding.ts:129-176](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/provider-binding.ts#L129-L176) · [packages/agent-runtime/src/provider-binding.ts:47-77](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/provider-binding.ts#L47-L77) · [packages/agent-runtime/src/provider-binding.ts:194-196](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/provider-binding.ts#L194-L196) · [packages/agent-runtime/src/runtime.ts:1037-1039](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/runtime.ts#L1037-L1039) · [packages/agent-runtime/src/fixed-tool-declarations.ts:29-42](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/fixed-tool-declarations.ts#L29-L42) · [packages/agent-runtime/src/provider-binding.ts:298-299](https://github.com/vastsa/PI-Desktop/blob/d403c96030381009c266140b7fba97edb6c1ca5d/packages/agent-runtime/src/provider-binding.ts#L298-L299)

### cloudflare/vibesdk (answered)

VibeSDK supports models across multiple providers through Cloudflare AI Gateway routing. The model master list is defined in `worker/agents/inferutils/config.types.ts` with over 20+ models classified by size (Lite, Regular, Large), each with a `provider`, `creditCost`, and `contextSize`.

**Supported providers:** Google AI Studio (Gemini 2.5 Flash/Pro, Gemini 3 Flash/Pro Preview), Anthropic (Claude 3.7 Sonnet, Claude 4/4.5 Sonnet/Opus/Haiku), OpenAI (GPT-5, GPT-5.1, GPT-5.2, GPT-5 Mini), Grok (Grok 4 Fast, Grok 4.1 Fast, Grok Code Fast 1), Vertex AI (GPT-OSS 120B, Kimi K2 Thinking, Qwen 3 Coder 480B). Several models (Cerebras, OpenAI OSS/Codex variants) are commented out in the master list.

**ThinkAgent model:** Configured in `worker/agents/think/model-config.ts` with `THINK_MODEL_ID = 'google-ai-studio/gemini-3.6-flash'`. The default config in `worker/agents/inferutils/config.ts` selects a Gemini-only set; when `PLATFORM_MODEL_PROVIDERS` env var is set, it uses a multi-provider config with Gemini 3 Pro Preview for blueprints, Grok 4.1 Fast for conversational response and deep debugging, and OpenAI 5 Mini as fallback.

**How models are called:** All model calls go through AI Gateway. The `getConfigurationForModel()` function in `worker/agents/inferutils/core.ts` resolves a model's baseURL, apiKey, and headers by checking: (1) runtime overrides (SDK BYOK keys), (2) user's encrypted gateway token, (3) user's custom gateway override, (4) environment variable per provider (`${PROVIDER}_API_KEY`), (5) platform AI Gateway binding. The OpenAI SDK (`openai` package) is used as the inference client with `client.chat.completions.create()`. The `ThinkAgent` uses `@ai-sdk/openai` via `createOpenAI()` with a custom `fetch` wrapper for Gemini compatibility patching.

**Per-model prompt tuning:** The Think agent has tuned system prompts per model family in `worker/agents/think/prompts/`: `gemini.txt`, `gpt.txt`, `anthropic.txt`, `trinity.txt`, `kimi.txt`, `codex.txt`, `beast.txt` (for GPT-4/o1/o3), and `default.txt`. These differ in tone, detail, and instructions; e.g., the Gemini prompt emphasizes security and tool-use patterns, while the Anthropic prompt is more concise and direct.

**Cost tracking:** Each model has a `creditCost` relative to GPT-5 Mini baseline (1.0 credit = $0.25/1M input). Rate limiting is enforced via `RateLimitService.enforceLLMCallsRateLimit()` which checks configurable budgets per user. Multi-provider configs have per-action model assignments with fallback chains (e.g., blueprint: Gemini 3 Pro Preview → Gemini 2.5 Flash).

**No local model support:** All models are called remotely through AI Gateway. There is no Ollama, llama.cpp, or local model infrastructure.


Citations: [worker/agents/inferutils/config.types.ts:1-340](https://github.com/cloudflare/vibesdk/blob/9da158d82c597a0e8f4bf033cdccd1053fb6fb15/worker/agents/inferutils/config.types.ts#L1-L340) · [worker/agents/inferutils/config.ts:58-186](https://github.com/cloudflare/vibesdk/blob/9da158d82c597a0e8f4bf033cdccd1053fb6fb15/worker/agents/inferutils/config.ts#L58-L186) · [worker/agents/inferutils/core.ts:424-497](https://github.com/cloudflare/vibesdk/blob/9da158d82c597a0e8f4bf033cdccd1053fb6fb15/worker/agents/inferutils/core.ts#L424-L497) · [worker/agents/think/prompts.ts:1-50](https://github.com/cloudflare/vibesdk/blob/9da158d82c597a0e8f4bf033cdccd1053fb6fb15/worker/agents/think/prompts.ts#L1-L50) · [worker/agents/think/ThinkAgent.ts:225-262](https://github.com/cloudflare/vibesdk/blob/9da158d82c597a0e8f4bf033cdccd1053fb6fb15/worker/agents/think/ThinkAgent.ts#L225-L262)
