Which models are supported and how are they called?
Providers; vision requirement; structured output / tool calling; small or specialised models.
Verdict
Both projects require structured output. They differ in how many providers they reach and how.
browser-use defines a small BaseChatModel protocol (ainvoke(messages, output_format)) with native adapters for about fifteen backends. These include OpenAI, Anthropic, Google, Azure, Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, Cerebras, LiteLLM and its own hosted ChatBrowserUse, which is the default when no model is given. The agent always asks for a Pydantic AgentOutput whose action field is a discriminated union. Vision is on by default and switched off for DeepSeek and some Grok models. Timeouts, screenshot size and coordinate clicking are tuned by model name.
Stagehand calls models from inside its browser extension through the Vercel AI SDK. Five providers are supported (OpenAI, Anthropic, Google, Groq, Cerebras), with allow-listed model ids. Every call uses JSON Schema structured output generated from Zod. Without a provider key, calls go through the Browserbase Model Gateway. A client-side generate hook lets you route inference through any model you control. Vision is only used for optional extract screenshots.
Pick browser-use for the widest provider choice, including local Ollama. Pick Stagehand when a short list of frontier providers, or your own generate function, is enough.
More projects in this category are being researched.
Per-project answers
browser-use/browser-use
answeredSupported providers. The library supports 14+ LLM providers, each in its own browser_use/llm/<provider>/ package: OpenAI (openai/chat.py), Anthropic (anthropic/chat.py), Google Gemini (google/chat.py), Azure OpenAI (azure/chat.py), AWS Bedrock (both Anthropic via aws/chat_anthropic.py and native via aws/chat_bedrock.py), DeepSeek (deepseek/chat.py), Groq (groq/chat.py), Mistral (mistral/chat.py), Ollama (ollama/chat.py), OpenRouter (openrouter/chat.py), OrcaRouter (orcarouter/chat.py), Cerebras (cerebras/chat.py), Vercel (vercel/chat.py), lite llm (litellm/chat.py), OCI Raw (oci_raw/chat.py), and a custom fine-tuned browser-use model (browser_use/llm/browser_use/chat.py). Lazy imports via __getattr__ keep startup fast (browser_use/llm/__init__.py:82-98).
Base protocol. All LLMs implement the BaseChatModel Protocol (browser_use/llm/base.py:32-75) which requires model attribute, provider property, and ainvoke(messages, output_format=None, **kwargs) -> ChatInvokeCompletion. The ainvoke method is overloaded: without output_format it returns a string; with a Pydantic model class it returns a validated structured output. The Agent always calls with output_format=AgentOutput (a Pydantic model with action as a discriminated union).
Vision support. Controlled by use_vision parameter (default True or 'auto'). When enabled, screenshots are sent as image content parts alongside the DOM text. Some models auto-disable vision: DeepSeek (doesn't support it, browser_use/agent/service.py:476-478), older Grok variants (grok-3/grok-code, line 483-485). vision_detail_level ('auto' | 'low' | 'high') controls image quality. Claude Sonnet models auto-configure 1400×850 screenshot resizing to fit vision context windows.
Fine-tuned model. ChatBrowserUse (browser_use/llm/browser_use/chat.py) is a first-party fine-tuned model that defaults when no LLM is specified. When used, flash_mode=True is auto-set, stripping planning fields from the output and disabling use_thinking — yielding faster, more token-efficient inference (browser_use/agent/service.py:238-243).
Structured output / tool calling. All providers return Pydantic-validated structured outputs. The mechanism differs per provider: OpenAI uses JSON mode / response_format; Anthropic uses tool_use with a single tool; Google uses response_schema; fine-tuned models return raw JSON matching the schema. The output_format parameter is passed through ainvoke(). For the agent loop, AgentOutput is dynamically generated via AgentOutput.type_with_custom_actions(ActionModel) which creates a discriminated union of all registered action types.
Model-to-timeout mapping. LLM timeouts are auto-configured per model family (browser_use/agent/service.py:261-277): Gemini 3 Pro gets 90s, Gemini other gets 75s, Groq gets 30s (fast inference), o3/Claude/DeepSeek get 90s.
Pre-configured model instances. browser_use/llm/models.py provides shorthands like openai_gpt_4o, google_gemini_2_5_pro, anthropic_claude_sonnet_4 etc., each constructing the appropriate Chat class with model string and optional API key from environment.
browserbase/stagehand
answeredStagehand supports five providers with explicitly allowlisted model IDs defined in packages/protocol/schemas.ts:
- OpenAI: gpt-4.1 family, gpt-4o, o1/o3/o4-mini, gpt-5 family (gpt-5 through gpt-5.6)
- Anthropic: claude-3-haiku, claude-haiku-4-5, claude-opus-4 through 4.8, claude-sonnet-4 through 4.6, claude-fable-5, claude-sonnet-5
- Google (Gemini): gemini-2.0-flash, gemini-2.5-pro/flash, gemini-3 variants, gemma-3
- Groq: llama-3.x, gemma2, mixtral, deepseek-r1-distill, qwen, kimi-k2
- Cerebras: llama3.1-8b, gpt-oss-120b, qwen-3 variants, zai-glm
Provider resolution (LLMProvider.ts): Model names use the format {provider}/{modelId} (e.g., openai/gpt-4o, anthropic/claude-sonnet-4-5). The provider is extracted from the prefix, and the model is created via the Vercel AI SDK (@ai-sdk/openai, @ai-sdk/anthropic, @ai-sdk/google, @ai-sdk/groq, @ai-sdk/cerebras).
Structured output / tool calling: All LLM calls use JSON Schema structured output (responseFormat: { type: "json_schema", name, schema }). The Zod schemas are converted to JSON Schema via z.toJSONSchema(). This is not tool calling (function calling) — it's native JSON Schema mode supported by OpenAI, Anthropic, and Gemini.
Vision: Vision is optional. When extract(options.screenshot: true), the screenshot is sent as an image content block alongside the DOM text. The vision requirement is per-model (e.g., Gemini models generally support it). The LLMClient abstract class tracks hasVision: boolean.
Model Gateway / Auto-selection: When using Browserbase cloud sessions without an explicit API key, createGatewayLanguageModel routes through Browserbase's Model Gateway (an OpenAI-compatible proxy at {apiUrl}/llm). With modelName: "auto", Browserbase selects the model automatically.
Client-side LLM: The SDK client can register a model.generate function. If a model with source: "client" is configured, clientLlmClient.ts delegates generation to the connected client rather than a built-in provider.
Model resolution fallback: In stagehandController.ts, the per-call options.model overrides the instance-level initParams.model. If neither is set, the Browserbase gateway must be available.
← How are failures, retries and self-healing handled? · How are browser sessions, profiles, auth and anti-bot handled? →