Which models are supported and how are they called?
Providers; local models; tool calling vs text formats; per-model prompt tuning; cost tracking.
Verdict
Aider is the only coding agent that works without native tool calling, so it suits small and local models best. Codex is the narrowest, with one wire protocol. Among the app builders, real model choice is thinner than the docs suggest.
Text edit formats, any model. Aider calls everything through LiteLLM. 357 entries in model-settings.yml tune the edit format and helper models per model.
Native tool calling over broad gateways.
- Cline: its own
@cline/llmsgateway with 211 generated provider IDs. It has no text-format fallback. - OpenCode: the models.dev catalog plus AI SDK packages, with per-family system prompts.
- Codewhale: about 50 provider kinds and three wire clients.
model = "auto"picks a big or cheap model, and its decision router can call a hosted third-party service. - Qwen Code: five adapters, all translated into
@google/genaishapes, and 13 provider presets. - PI-Desktop: seven wire styles through pi-ai, including the ChatGPT/Codex subscription endpoint.
- DeepSeek-Reasonix: three provider kinds. It picks the reasoning parameters from the base URL and ships per-vendor price tables.
- Open Interpreter: picks a harness per provider and model, for example
claude-code-barefor DeepSeek. Most harnesses need a Chat Completions endpoint.
One protocol. Codex speaks only the OpenAI Responses API. The built-in providers are OpenAI, Bedrock, Ollama and LM Studio, and any other provider must speak Responses.
App builders.
- bolt.diy: 22 provider classes, but only OpenRouter, Anthropic, OpenAI and Google are registered. Ollama and LM Studio need a source edit despite the README.
- Open Lovable: four providers, no tool calling, one prompt for every model, and no local models.
- VibeSDK: the default Think engine is hard-wired to a Gemini Flash model through AI Gateway. The per-action model settings only reach the legacy engines.
Pick: Aider for local or weak models that cannot call functions. Pick: OpenCode, Cline or Codewhale for the widest choice of tool-capable providers. DeepSeek-Reasonix is the pick if cost tracking for DeepSeek-class pricing matters. Pick: Codex if you are on OpenAI models and want the best-tuned path.
More projects in this category are being researched.
Per-project answers
anomalyco/opencode
answeredModel Support
OpenCode uses the Vercel AI SDK (ai package) as its primary LLM abstraction layer, with an opt-in native LLM runtime (@opencode-ai/llm). The Provider service (provider/provider.ts) loads model configurations from a catalog (@opencode-ai/core/models-dev.ts), which defines models with fields for context limits, output limits, cost tiers, reasoning options, modalities (text/image/audio/pdf), temperature support, tool-call support, and per-provider npm packages.
Providers are resolved from the catalog and configured via @ai-sdk/* npm packages. The sdkKey() function in provider/transform.ts (lines 41-98) maps npm packages to AI SDK provider option keys: @ai-sdk/anthropic, @ai-sdk/openai, @ai-sdk/google, @ai-sdk/amazon-bedrock, @ai-sdk/mistral, @ai-sdk/groq, @ai-sdk/xai, @ai-sdk/cohere, @ai-sdk/perplexity, @ai-sdk/togetherai, @ai-sdk/azure, @ai-sdk/alibaba, @ai-sdk/cerebras, @ai-sdk/deepinfra, @ai-sdk/vercel, @ai-sdk/gateway, plus venice, openrouter, gitlab-ai-provider, and OpenAI-compatible providers via @ai-sdk/openai-compatible. Plugins add auth support for GitHub Copilot, Codex, GitLab, Poe, Cloudflare, Azure, DigitalOcean, Snowflake Cortex, xAI, Cerebras, and Modal.
Local models are supported via OpenAI-compatible endpoints (@ai-sdk/openai-compatible), enabling any local LLM server (like Ollama, LM Studio, or vLLM) that exposes an OpenAI-compatible API.
Tool-calling vs text formats: Models with tool_call: true in the catalog definition get full tool definitions. The AI SDK handles tool-calling natively; OpenCode normalizes the stream events through LLMAISDK.toLLMEvents (session/llm/ai-sdk.ts). For providers without tool-calling, the model would only produce text.
Per-model prompt tuning is in session/system.ts (lines 28-51): Anthropic models get one system prompt (anthropic.txt), GPT-4/o1/o3 get beast.txt, newer GPTs get astra.txt or codex.txt, Gemini gets gemini.txt, Kimi gets kimi.txt, Meta/Muse gets meta.txt, and everything else gets default.txt.
Provider-specific message transforms in provider/transform.ts handle quirks: Anthropic models filter empty content parts and normalize tool call IDs (lines 170-194, 224-251), Bedrock uses signature-based reasoning, Mistral pads/scrambles tool call IDs (lines 253-277), and OpenAI uses include for encrypted reasoning (line 23).
Cost tracking uses Session.getUsage() (session.ts lines 338-405), which reads cost tiers from the model catalog, handles cache read/write pricing, and accounts for provider-specific metadata keys (Anthropic vertex, Bedrock, Venice).
openai/codex
answeredCodex CLI connects primarily to OpenAI's Responses API as the default provider, with model catalog fetched from api.openai.com and cached in models_cache.json (managed by codex-rs/models-manager/src/manager.rs:34). The ModelProvider trait in codex-rs/model-provider/src/provider.rs:108 defines the abstraction for model backends, with implementations for: OpenAI's production API (the default, via CoreAuthProvider in lib.rs:27), Amazon Bedrock (amazon_bedrock in lib.rs:17), LM Studio (codex-rs/lmstudio/), and Ollama (codex-rs/ollama/). The ModelsEndpointClient trait (manager.rs:42) handles remote catalog fetching, while static and OpenAI models managers serve the bundled catalog (bundled_models_response() in lib.rs:13 from a bundled models.json). Model selection uses a pinned model per session, with the option to switch mid-conversation. The ModelInfo struct carries per-model metadata: context window size, reasoning-effort defaults, input/output modalities, auto-compact token limits, and ToolMode (standard, code-mode-only, or code-mode — core/src/tools/mod.rs:75). Models are called through the OpenAI Responses API streaming endpoints via the ModelClientSession in core/src/client.rs:1, which manages WebSocket connections for sticky routing, retries with exponential backoff (core/src/responses_retry.rs:33), and transport fallback. Per-model prompts are rendered via render_model_instructions() in prompts/src/model_instructions.rs:8, which uses the model's instruction template from its metadata. Cost tracking uses TokenUsageInfo collected per turn and aggregated in analytics events. Reasonable efforts (ReasoningEffort) controls are configurable per-turn.
cline/cline
answeredThe @cline/llms package provides a gateway-based provider architecture (gateway.ts). A central GatewayRegistry routes requests to registered provider handlers. Supported vendors include Anthropic (vendors/anthropic.ts), OpenAI (vendors/openai.ts), Google (vendors/google.ts), Mistral (vendors/mistral.ts), AWS Bedrock (vendors/bedrock.ts), Vertex AI (vendors/vertex.ts), Ollama (local models, vendors/ollama.ts), OpenAI-compatible APIs (vendors/openai-compatible.ts), a community provider system (vendors/community.ts), the Cline managed API (vendors/cline.ts), and MiniMax Thinking (vendors/minimax-thinking.ts). There is also a generic ai-sdk provider (providers/ai-sdk.ts). Each provider implements a common ApiHandler interface that can operate in tool-calling mode (when the model supports it) or text-format mode (shimmed via structured prompts). A generated provider spec file (providers.generated.ts) and model catalog (catalog.generated.ts) register known model IDs, capabilities, and operation support. The model-tools system (model-tools.ts) detects which models support native tool definitions vs. requiring tool-use prompt shims. Cost tracking is handled through billing module (billing.ts) and telemetry events. Per-model prompt tuning is implemented via routing rules: reasoning effort, cache point placement, and max-token budgeting are adjusted per-provider/model in the routing modules (routing/anthropic-compatible.ts, routing/bedrock-cache-point.ts, routing/reasoning-options.ts).
model-tools.ts resolves provider-side model tools such as web_search, not native tool support.openinterpreter/openinterpreter
answeredModels use a provider abstraction: ModelProvider (model-provider/src/provider.rs:42-73) with ProviderCapabilities (namespace_tools, image_generation, web_search, remote_compaction). Catalog: models-manager/models.json (1663 lines) with slugs, context_window, max_context_window, reasoning levels (low/medium/high/xhigh/max), tool_mode, input_modalities, comp_hash. ModelsManager (models-manager/src/manager.rs) provides OpenAiModelsManager and StaticModelsManager backed by ModelsCache. Provider resolution via create_model_provider dispatches to Amazon Bedrock, OpenAI/Anthropic, or combined auth (model-provider/src/provider.rs:39). bundled_provider_catalog (model-provider-info/src/bundled_provider_catalog.rs) serves offline data. Model calls go through ModelClient (core/src/client.rs); turn-scoped ModelClientSession caches WebSocket connections. Per-turn settings (model, reasoning_effort, service_tier) pass to client_session.stream(). Tool calling: Prompt.tools = step_context.tool_router.model_visible_specs() (turn.rs:1566). Cost tracking via SessionTelemetry and AnalyticsEventsClient.
resolve_stream_transport_route maps (wire_api = responses|chat|messages, harness) to a transport and per-harness request builder (core/src/harness/routing.rs), and default_harness_for_provider_model picks claude-code, kimi-code, qwen-code, zcode or claude-code-bare from the provider/model.Aider-AI/aider
answeredProvider-agnostic via litellm. All model calls go through litellm.completion() (aider/llm.py lazily imports litellm). This means any model that litellm supports — OpenAI, Anthropic, Gemini, DeepSeek, Ollama, Together, OpenRouter, AWS Bedrock, Azure, Cohere, Replicate, etc. — is usable. API keys are set via environment variables (e.g. ANTHROPIC_API_KEY, OPENAI_API_KEY) or the --api-key flag (aider/main.py:600-630).
Built-in known model lists. aider/models.py:33-123 defines OPENAI_MODELS (30+ variants from gpt-3.5-turbo through o1, o3-mini, gpt-5.5), ANTHROPIC_MODELS (claude-2 through claude-sonnet-4-6, haiku-4-5, opus-4-7) and MODEL_ALIASES (e.g. "sonnet" → "claude-sonnet-4-6").
Per-model prompt tuning. The ModelSettings dataclass (models.py:128-151) carries per-model parameters: edit_format (diff/whole/udiff/etc), use_repo_map, use_system_prompt, use_temperature, reminder style ("user" or "sys"), cache_control, examples_as_sys_msg, extra_params, and accepts_settings (e.g. "thinking_tokens"). Settings are loaded from aider/resources/model-settings.yml (models.py:153-158) and can be extended with user-provided .aider.model.settings.yml files registered via register_models() (main.py:335-358). Additional per-model heuristics in Model.__init_post__() (models.py:490-598) set defaults based on model name patterns — DeepSeek R1 gets reasoning_tag="think" and use_temperature=False, GPT-4-turbo gets edit_format="udiff", o1 models disable use_system_prompt.
Tool calling vs text formats. When a model has a function-based coder (like WholeFileFunctionCoder), the LLM receives tools in the litellm API call (models.py:1006-1009). Otherwise edits are text-based and parsed from the model's plain-text response. The edit_format field on the model determines which.
Cost tracking. The ModelInfoManager (models.py:161) fetches model_prices_and_context_window.json from the litellm GitHub repo (cached for 24h). The OpenRouter model database is also cached locally. Costs are calculated after each completion (base_coder.py:1811).
codewhale-hq/Codewhale
answeredCodewhale is provider-neutral and supports 46+ provider kinds through a unified LlmClient async trait.
Supported providers (config/src/provider_kind.rs): DeepSeek, OpenAI, Anthropic (native Messages API), Google (Gemini), Mistral AI, Ollama, Ollama Cloud, Hugging Face, Together, Groq, Fireworks, Novita, OpenRouter, Grok (xAI), Meta (Muse Spark), and many more. Custom providers use the Custom kind for arbitrary OpenAI-compatible endpoints. Dual-wire variants (DeepSeekAnthropic, MinimaxAnthropic, ModelstudioTokenPlanAnthropic) select the Anthropic Messages wire protocol instead of OpenAI Chat Completions.
Wire protocols: Three client implementations in crates/tui/src/client/:
chat.rs— OpenAI Chat Completions (the production path for most providers): streaming, request building, SSE parsing.anthropic.rs— Native Anthropic Messages API: handles adaptive thinking,cache_controlbreakpoints, signed-thinking signatures, usage normalization.responses.rs— OpenAI Responses API (for newer OpenAI endpoints). TheLlmClienttrait (llm_client/mod.rs:54) definescreate_message,create_message_stream, andprovider_name/model()with retry logic (with_retry), exponential backoff, jitter, and error classification.
Model selection:
- Auto model detection (
auto_model.rs): For DeepSeek, a rule-based classifier scores the user's prompt to choose between "deepseek-v4-flash" (simple tasks) and "deepseek-v4-pro" (complex). - Model catalog (
catalog.rs:1): Layered resolution — bundled Models.dev snapshot < live Models.dev < corrections < signed cloud facts < live provider/v1/models< user config overrides. - Fleet roster: Role-based model assignment for sub-agents via
FleetRosterandFleetRole, where each role can pin a specific provider/model. - Route resolver in
config/src/route/compiles catalog rows into executable routes with resolvedbase_url, wire protocol, and limits.
Per-model prompt tuning: System prompt shape adapts per mode (Plan/Act/Operate). Provider-specific request shaping includes: max_completion_tokens vs max_tokens selection (chat.rs:46), reasoning effort parameters (chat.rs:67), token limits per route, and provider-specific tool-choice modes (strict_tool_mode for DeepSeek beta).
Cost tracking: Usage struct with input/output/cache tokens, CostScopeToken for billing attribution, and route-level RouteLimits with per-provider pricing.
crates/config/src/auto_model.rs is a legacy DeepSeek-only scorer kept for API compatibility and is not used for model = "auto"; auto routing is done in crates/tui/src/model_routing.rs, which picks a big/cheap pair for the active provider via a chat router or a TypeSafe/OpenRouter decision router.esengine/DeepSeek-Reasonix
answeredModels are supported through two provider implementations registered in a provider.Registry. OpenAI-compatible (internal/model/openai/openai.go:1): registered as kind "openai", calls /chat/completions with SSE streaming. Per-vendor wire-shape detection selects different reasoning parameters: api.deepseek.com uses thinking.type=enabled with reasoning_effort; api.minimaxi.com uses thinking.type=adaptive|disabled (M3's binary knob); open.bigmodel.cn / api.z.ai (Zhipu GLM) have documented depth fields; api.longcat.chat uses thinking.type=enabled|disabled without effort scale; ollama.com accepts hosted Ollama Cloud's effort scale including max; Kimi K3 preserves complete messages and uses max_completion_tokens. Everything else uses vanilla reasoning_effort. Anthropic (internal/model/anthropic/anthropic.go:1): registered as kind "anthropic", calls POST /v1/messages with SSE streaming. Supports extended thinking via signed reasoning blocks that are replayed on multi-turn. No temperature/top_p (Claude rejects sampling params). DeepSeek's compatible Anthropic endpoint uses unsigned thinking blocks with binary enabled|disabled and output_config.effort. Configuration is via reasonix.toml [provider] blocks with kind, base_url, model, api_key, and kind-specific extras. Providers resolve via config.go (internal/contract/provider/config.go:6) and the provider broker serves them. local.go (internal/model/providerbroker/local.go) provides no local-model fallback; there is no Ollama-hosted local inference path — the Ollama integration reaches hosted cloud API. Cost tracking uses pricing.rates.go (internal/contract/pricing/rates.go:36) with official rates for DeepSeek, Anthropic Claude, OpenAI GPT, MiniMax, LongCat, Moonshot (Kimi), and others, stored as oldest-first generations per vendor. make pricecheck reads vendor pages to verify rates.
responses (also dashscope-responses, internal/model/responses/responses.go) targets the OpenAI Responses API alongside openai and anthropic.firecrawl/open-lovable
answeredFour providers are supported via the Vercel AI SDK (@ai-sdk/*): Anthropic, OpenAI, Google Generative AI, and Groq. All are wrapped in lib/ai/provider-manager.ts. The getProviderForModel() function (provider-manager.ts:79-117) routes based on model name prefixes: anthropic/ → Anthropic SDK, openai/ → OpenAI SDK, google/ → Google Gemini SDK, and moonshotai/kimi-k2-instruct-0905 → Groq SDK. Without a prefix it defaults to Groq. Models listed in config/app.config.ts:56-62: openai/gpt-5, moonshotai/kimi-k2-instruct-0905, anthropic/claude-sonnet-4-20250514, google/gemini-3-pro-preview (the default). Custom model routing can be configured via modelApiConfig in the app config (provider-manager.ts:81-86). Vercel AI Gateway is supported as a unified proxy — all requests go to https://ai-gateway.vercel.sh/v1 when AI_GATEWAY_API_KEY is set (provider-manager.ts:21-23). Tool/function calling is not used — the code at generate-ai-code-stream/route.ts:1310-1311 explicitly notes no support in the available models. Instead, XML tags in text output are parsed. generateObject() (Vercel AI SDK) is used solely for the edit-intent search-plan endpoint with Zod schemas (analyze-edit-intent/route.ts:34-60,126). Temperature defaults to 0.7 for non-reasoning models and is unset for GPT-5. Max tokens is 8192. GPT-5 uses reasoningEffort: 'high' (route.ts:1320-1326). No local models are supported. No cost tracking or per-model prompt tuning exists — the same massive system prompt template is used for all models.
QwenLM/qwen-code
answeredModels are accessed through a multi-protocol architecture. AuthType (from core/contentGenerator.ts) defines supported protocols: openai, openai-responses, anthropic, gemini, vertex-ai, and qwen-oauth. Constants in packages/core/src/models/constants.ts (lines 74–105) map each protocol to environment variables (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY).
There are 13 provider presets registered in packages/core/src/providers/all-providers.ts (lines 55–69): Alibaba (coding plan, token plan, standard), DeepSeek, Grok, Minimax, Z.ai, Moonshot, IdeaLab, ModelScope, OpenRouter, Requesty, and a generic Custom provider. Each provider defines envKey, baseUrl, modelNamePrefix, and optional per-model generation config (thinking, context window, modalities) — built by provider-config.ts.
Content generators translate the protocol into actual API calls: OpenAIContentGenerator (core/openaiContentGenerator/) for all OpenAI-compatible backends, AnthropicContentGenerator (core/anthropicContentGenerator/), and GeminiChat/geminiContentGenerator for Gemini. The model registry (models/modelRegistry.ts) resolves a provider ID + model config into the correct AuthType via resolveModelProtocol(), supporting wireApi routing between chat-completions and responses (lines 69–92). Per-provider tuning lives in openaiContentGenerator/provider/ (DeepSeek, Fireworks, Grok, Mistral, etc.). Cost tracking is in telemetry/loggers.ts. Local models (Ollama/vLLM) connect via the Custom provider.
stackblitz-labs/bolt.diy
answeredBolt supports 20+ providers registered in app/lib/modules/llm/registry.ts. Only 4 are enabled by default via ENABLED_PROVIDERS in the manager (OpenRouter, Anthropic, OpenAI, Google), but all 22 registered providers (Anthropic, OpenAI, Google, Groq, Cohere, DeepSeek, HuggingFace, Mistral, Ollama, OpenRouter, LMStudio, AmazonBedrock, Perplexity, Together, xAI, Fireworks, Hyperbolic, Cerebras, GitHub, Moonshot, OpenAILike, Z-AI) can be enabled. Each provider extends BaseProvider (app/lib/modules/llm/base-provider.ts) which provides common infrastructure: getting base URL and API key from environment variables, cookie settings, or server env; Docker URL rewriting; and model caching with cache-key generation. The getModelInstance method creates an SDK model instance per provider — for example, AnthropicProvider uses createAnthropic() from @ai-sdk/anthropic, OllamaProvider uses createOllama() from ollama-ai-provider. Models can be static (hardcoded in the provider class, e.g., Anthropic's three fallback models) or dynamic (fetched at runtime via getDynamicModels — Ollama fetches from /api/tags, Anthropic from api.anthropic.com/v1/models). The LLMManager singleton (app/lib/modules/llm/manager.ts) orchestrates provider registration and model list updates. It caches dynamic models based on a cache key derived from API keys and settings. The Vercel AI SDK serves as the unified interface: streamText in the SDK handles streaming, tool calling, token counting, and error handling. Reasoning models (o1, o3, GPT-5) get special handling — they use maxCompletionTokens instead of maxTokens, and have unsupported parameters (temperature, topP, etc.) filtered out (stream-text.ts lines 240-280). Per-model prompt tuning is handled by the PromptLibrary which supports swapping system prompts (default, original, optimized). No cost tracking is implemented in this codebase.
vastsa/PI-Desktop
answeredPI-Desktop supports multiple model providers through the pi-ai library, with a universal provider model (ADR 0012).
Wire protocols — provider-binding.ts:129-176 maps user-configured apiStyle values to pi-ai wire adapters: chat_completions (OpenAI-compatible), responses (OpenAI Responses), anthropic_messages (Claude), openai_codex_responses (ChatGPT/Codex), pi_messages (Pi service), google_generative_ai (Gemini), and opencode_go (OpenCode Go API). If no explicit style is set, it defaults to OpenAI Chat Completions (provider-binding.ts:170-175).
Model resolution — Providers are stored with apiStyle, baseUrl, modelId, apiKey, and optional modelConfig with per-model overrides. The apiBindingForProviderModel() function (provider-binding.ts:194-196) resolves the correct wire protocol considering both provider-level and model-level API styles. Vendor accounts (GitHub Copilot, Pi, custom) are handled with special auth paths including OAuth token refresh (provider-binding.ts:76).
Thinking/reasoning — Models can be configured with thinking levels via SessionThinkingLevel / SubagentThinkingLevel. The thinking-level.ts module clamps and normalizes these. Claude models with adaptive thinking vs. budget-token thinking are handled differently (provider-binding.ts:298+). DeepSeek Flash gets special tool-declaration optimizations (fixed-tool-declarations.ts:29-42).
Cost tracking — Usage is tracked via MessageUsage objects accumulating tokens and costs per turn, surfaced through request-usage.ts. The provider-retry.ts module handles rate limits (up to 5 retries) and transient errors (up to 3 retries) at the provider level.
Subagent models — Delegations can pin their own model via SubagentDefinition. Electron main resolves credentials; subagent model keys must be explicitly opted-in for Task.model overrides via subagentModelKeys (runtime.ts:1037-1039). Missing or failed model bindings fall through a fallback chain.
cloudflare/vibesdk
answeredVibeSDK supports models across multiple providers through Cloudflare AI Gateway routing. The model master list is defined in worker/agents/inferutils/config.types.ts with over 20+ models classified by size (Lite, Regular, Large), each with a provider, creditCost, and contextSize.
Supported providers: Google AI Studio (Gemini 2.5 Flash/Pro, Gemini 3 Flash/Pro Preview), Anthropic (Claude 3.7 Sonnet, Claude 4/4.5 Sonnet/Opus/Haiku), OpenAI (GPT-5, GPT-5.1, GPT-5.2, GPT-5 Mini), Grok (Grok 4 Fast, Grok 4.1 Fast, Grok Code Fast 1), Vertex AI (GPT-OSS 120B, Kimi K2 Thinking, Qwen 3 Coder 480B). Several models (Cerebras, OpenAI OSS/Codex variants) are commented out in the master list.
ThinkAgent model: Configured in worker/agents/think/model-config.ts with THINK_MODEL_ID = 'google-ai-studio/gemini-3.6-flash'. The default config in worker/agents/inferutils/config.ts selects a Gemini-only set; when PLATFORM_MODEL_PROVIDERS env var is set, it uses a multi-provider config with Gemini 3 Pro Preview for blueprints, Grok 4.1 Fast for conversational response and deep debugging, and OpenAI 5 Mini as fallback.
How models are called: All model calls go through AI Gateway. The getConfigurationForModel() function in worker/agents/inferutils/core.ts resolves a model's baseURL, apiKey, and headers by checking: (1) runtime overrides (SDK BYOK keys), (2) user's encrypted gateway token, (3) user's custom gateway override, (4) environment variable per provider (${PROVIDER}_API_KEY), (5) platform AI Gateway binding. The OpenAI SDK (openai package) is used as the inference client with client.chat.completions.create(). The ThinkAgent uses @ai-sdk/openai via createOpenAI() with a custom fetch wrapper for Gemini compatibility patching.
Per-model prompt tuning: The Think agent has tuned system prompts per model family in worker/agents/think/prompts/: gemini.txt, gpt.txt, anthropic.txt, trinity.txt, kimi.txt, codex.txt, beast.txt (for GPT-4/o1/o3), and default.txt. These differ in tone, detail, and instructions; e.g., the Gemini prompt emphasizes security and tool-use patterns, while the Anthropic prompt is more concise and direct.
Cost tracking: Each model has a creditCost relative to GPT-5 Mini baseline (1.0 credit = $0.25/1M input). Rate limiting is enforced via RateLimitService.enforceLLMCallsRateLimit() which checks configurable budgets per user. Multi-provider configs have per-action model assignments with fallback chains (e.g., blueprint: Gemini 3 Pro Preview → Gemini 2.5 Flash).
No local model support: All models are called remotely through AI Gateway. There is no Ollama, llama.cpp, or local model infrastructure.
← How are shell commands and file writes kept safe? · How can it be extended and customised? →