# Which models are supported and how are they called?

> Browser & computer control — a good answer covers: Providers; vision requirement; structured output / tool calling; small or specialised models.

Canonical page: https://llms-technical-reviews.com/browser-control/q/models/

## Verdict

Both projects require structured output. They differ in how many providers they reach and how.

[browser-use](/p/browser-use/) defines a small `BaseChatModel` protocol (`ainvoke(messages, output_format)`) with native adapters for about fifteen backends. These include OpenAI, Anthropic, Google, Azure, Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, Cerebras, LiteLLM and its own hosted `ChatBrowserUse`, which is the default when no model is given. The agent always asks for a Pydantic `AgentOutput` whose action field is a discriminated union. Vision is on by default and switched off for DeepSeek and some Grok models. Timeouts, screenshot size and coordinate clicking are tuned by model name.

[Stagehand](/p/stagehand/) calls models from inside its browser extension through the Vercel AI SDK. Five providers are supported (OpenAI, Anthropic, Google, Groq, Cerebras), with allow-listed model ids. Every call uses JSON Schema structured output generated from Zod. Without a provider key, calls go through the Browserbase Model Gateway. A client-side `generate` hook lets you route inference through any model you control. Vision is only used for optional `extract` screenshots.

Pick browser-use for the widest provider choice, including local Ollama. Pick Stagehand when a short list of frontier providers, or your own `generate` function, is enough.

More projects in this category are being researched.

## Per-project answers

### browser-use/browser-use (answered)

**Supported providers.** The library supports 14+ LLM providers, each in its own `browser_use/llm/<provider>/` package: OpenAI (`openai/chat.py`), Anthropic (`anthropic/chat.py`), Google Gemini (`google/chat.py`), Azure OpenAI (`azure/chat.py`), AWS Bedrock (both Anthropic via `aws/chat_anthropic.py` and native via `aws/chat_bedrock.py`), DeepSeek (`deepseek/chat.py`), Groq (`groq/chat.py`), Mistral (`mistral/chat.py`), Ollama (`ollama/chat.py`), OpenRouter (`openrouter/chat.py`), OrcaRouter (`orcarouter/chat.py`), Cerebras (`cerebras/chat.py`), Vercel (`vercel/chat.py`), lite llm (`litellm/chat.py`), OCI Raw (`oci_raw/chat.py`), and a custom fine-tuned browser-use model (`browser_use/llm/browser_use/chat.py`). Lazy imports via `__getattr__` keep startup fast (`browser_use/llm/__init__.py:82-98`).

**Base protocol.** All LLMs implement the `BaseChatModel` Protocol (`browser_use/llm/base.py:32-75`) which requires `model` attribute, `provider` property, and `ainvoke(messages, output_format=None, **kwargs) -> ChatInvokeCompletion`. The `ainvoke` method is overloaded: without `output_format` it returns a string; with a Pydantic model class it returns a validated structured output. The `Agent` always calls with `output_format=AgentOutput` (a Pydantic model with `action` as a discriminated union).

**Vision support.** Controlled by `use_vision` parameter (default `True` or `'auto'`). When enabled, screenshots are sent as image content parts alongside the DOM text. Some models auto-disable vision: DeepSeek (doesn't support it, `browser_use/agent/service.py:476-478`), older Grok variants (grok-3/grok-code, line 483-485). `vision_detail_level` (`'auto' | 'low' | 'high'`) controls image quality. Claude Sonnet models auto-configure 1400×850 screenshot resizing to fit vision context windows.

**Fine-tuned model.** `ChatBrowserUse` (`browser_use/llm/browser_use/chat.py`) is a first-party fine-tuned model that defaults when no LLM is specified. When used, `flash_mode=True` is auto-set, stripping planning fields from the output and disabling `use_thinking` — yielding faster, more token-efficient inference (`browser_use/agent/service.py:238-243`).

**Structured output / tool calling.** All providers return Pydantic-validated structured outputs. The mechanism differs per provider: OpenAI uses JSON mode / response_format; Anthropic uses tool_use with a single tool; Google uses response_schema; fine-tuned models return raw JSON matching the schema. The `output_format` parameter is passed through `ainvoke()`. For the agent loop, `AgentOutput` is dynamically generated via `AgentOutput.type_with_custom_actions(ActionModel)` which creates a discriminated union of all registered action types.

**Model-to-timeout mapping.** LLM timeouts are auto-configured per model family (`browser_use/agent/service.py:261-277`): Gemini 3 Pro gets 90s, Gemini other gets 75s, Groq gets 30s (fast inference), o3/Claude/DeepSeek get 90s.

**Pre-configured model instances.** `browser_use/llm/models.py` provides shorthands like `openai_gpt_4o`, `google_gemini_2_5_pro`, `anthropic_claude_sonnet_4` etc., each constructing the appropriate Chat class with model string and optional API key from environment.


Citations: [browser_use/llm/__init__.py:82-134](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/llm/__init__.py#L82-L134) · [browser_use/agent/service.py:237-278](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L237-L278) · [browser_use/agent/service.py:476-486](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L476-L486) · [browser_use/agent/service.py:1943-1981](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1943-L1981) · [browser_use/agent/service.py:327-333](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L327-L333)

### browserbase/stagehand (answered)

Stagehand supports **five providers** with explicitly allowlisted model IDs defined in `packages/protocol/schemas.ts`:

- **OpenAI**: gpt-4.1 family, gpt-4o, o1/o3/o4-mini, gpt-5 family (gpt-5 through gpt-5.6)
- **Anthropic**: claude-3-haiku, claude-haiku-4-5, claude-opus-4 through 4.8, claude-sonnet-4 through 4.6, claude-fable-5, claude-sonnet-5
- **Google (Gemini)**: gemini-2.0-flash, gemini-2.5-pro/flash, gemini-3 variants, gemma-3
- **Groq**: llama-3.x, gemma2, mixtral, deepseek-r1-distill, qwen, kimi-k2
- **Cerebras**: llama3.1-8b, gpt-oss-120b, qwen-3 variants, zai-glm

**Provider resolution** (`LLMProvider.ts`): Model names use the format `{provider}/{modelId}` (e.g., `openai/gpt-4o`, `anthropic/claude-sonnet-4-5`). The provider is extracted from the prefix, and the model is created via the Vercel AI SDK (`@ai-sdk/openai`, `@ai-sdk/anthropic`, `@ai-sdk/google`, `@ai-sdk/groq`, `@ai-sdk/cerebras`).

**Structured output / tool calling**: All LLM calls use JSON Schema structured output (`responseFormat: { type: "json_schema", name, schema }`). The Zod schemas are converted to JSON Schema via `z.toJSONSchema()`. This is not tool calling (function calling) — it's native JSON Schema mode supported by OpenAI, Anthropic, and Gemini.

**Vision**: Vision is optional. When `extract(options.screenshot: true)`, the screenshot is sent as an image content block alongside the DOM text. The `vision` requirement is per-model (e.g., Gemini models generally support it). The `LLMClient` abstract class tracks `hasVision: boolean`.

**Model Gateway / Auto-selection**: When using Browserbase cloud sessions without an explicit API key, `createGatewayLanguageModel` routes through Browserbase's Model Gateway (an OpenAI-compatible proxy at `{apiUrl}/llm`). With `modelName: "auto"`, Browserbase selects the model automatically.

**Client-side LLM**: The SDK client can register a `model.generate` function. If a model with `source: "client"` is configured, `clientLlmClient.ts` delegates generation to the connected client rather than a built-in provider.

**Model resolution fallback**: In `stagehandController.ts`, the per-call `options.model` overrides the instance-level `initParams.model`. If neither is set, the Browserbase gateway must be available.


Citations: [packages/protocol/schemas.ts:9-76](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/protocol/schemas.ts#L9-L76) · [packages/protocol/schemas.ts:78-100](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/protocol/schemas.ts#L78-L100) · [packages/protocol/schemas.ts:182-213](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/protocol/schemas.ts#L182-L213) · [packages/extension/llm/LLMProvider.ts:22-28](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/llm/LLMProvider.ts#L22-L28) · [packages/extension/llm/LLMProvider.ts:78-113](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/llm/LLMProvider.ts#L78-L113) · [packages/extension/services/llmService.ts:13-36](https://github.com/browserbase/stagehand/blob/c9c8a41778b2000c9a9bdfc4b68e6c0c4866ab1a/packages/extension/services/llmService.ts#L13-L36)
