LLMs Technical Reviews

Skyvern-AI/skyvern

Self-hosted LLM browser agent that picks actions from a scraped DOM tree plus screenshots, runs them in Playwright, and chains workflows.

GitHub ↗★ 23kPythonAGPL-3.0commit 44de8cd · 2026-10-05homepage ↗

Overview

Skyvern is a self-hostable server that runs browser tasks described in natural language. You send a prompt and a start URL to its REST API, and a background worker drives a Playwright browser until an LLM decides the goal is met. Around that core sit a workflow engine (around 30 block types: navigation, extraction, login, loops, conditionals, code, HTTP, email, PDF), a React UI with live browser streaming, a Python SDK, a CLI and an MCP server with over a hundred tools.

The page is shown to the model as a DOM-derived element tree, not as raw pixels. An injected script tags every interactable element with a unique_id attribute and serialises the tree as compact HTML. That HTML goes to the LLM together with a few scrolled screenshots. The model answers with a JSON list of actions that name elements by id, and Python handlers execute them with Playwright locators. Despite the “computer vision” in the marketing, screenshots are context only. The set-of-marks box overlay is switched off and marked deprecated in the scraper.

At this commit, “Skyvern” means four different agents behind one API, chosen by the engine field: skyvern-1.0 (the default per-step engine), skyvern-2.0 (a planner that writes and runs workflow blocks), skyvern-3.0 (a single persistent tool-calling conversation), and third-party computer-use models (OpenAI CUA, Anthropic CUA, UI-TARS, Yutori Navigator). The repository is also the open core of a hosted product. Many behaviours, such as CAPTCHA solving, cross-run caches and A/B engine routing, are hooks that are no-ops here and are filled in by the cloud.

Architecture

flowchart LR
  C["Client: SDK / CLI / MCP / UI"] --> API["FastAPI routes"]
  API --> EX["BackgroundTaskExecutor"]
  EX --> AG["ForgeAgent.execute_step"]
  API --> T2["task_v2_service planner"]
  T2 --> WF["Workflow blocks"]
  WF --> AG
  AG --> V3["Task V3 tool loop"]
  AG --> SC["Scraper + domUtils.js"]
  AG --> LLM["LLMAPIHandlerFactory (LiteLLM)"]
  AG --> AH["ActionHandler"]
  SC --> BM["RealBrowserManager"]
  AH --> BM
  V3 --> BM
  BM --> PW["Playwright Chromium or CDP"]
  AG --> DB["PostgreSQL"]
Component Path Role
API routes skyvern/forge/sdk/routes/agent_protocol.py /v1/run/tasks, workflow runs, engine dispatch
Executor skyvern/forge/sdk/executor/background_task_executor.py Creates the first step and schedules execute_step as a FastAPI background task
Step engine skyvern/forge/agent.py ForgeAgent: scrape, prompt, LLM, parse, act, verify, recurse
Planner (2.0) skyvern/services/task_v2_service.py Iterative “plan next block” loop that builds and runs a workflow
Task V3 skyvern/forge/taskv3/ Persistent tool-use conversation (engine.py, loop.py, tools.py)
Scraper skyvern/webeye/scraper/ scraper.py + 4,000-line domUtils.js: element tree, frames, split screenshots
Actions skyvern/webeye/actions/ Action models, parse_actions.py, the 16k-line handler.py
Browser skyvern/webeye/real_browser_manager.py, browser_factory.py Browser acquisition, persistent sessions, profiles, proxies
LLM layer skyvern/forge/sdk/api/llm/ Config registry, LiteLLM handlers and routers, LLMCaller for chat-history engines
Workflows skyvern/forge/sdk/workflow/ Block models, parameters, service.py
Script generation skyvern/core/script_generations/ Turns recorded runs into Playwright-style Python (libcst)
Client surfaces skyvern/library/, skyvern/cli/, skyvern-frontend/ SDK, Typer CLI + FastMCP server, React UI

How a request flows

Take POST /v1/run/tasks with engine: skyvern-1.0:

  1. Accept. run_task validates the webhook URL and rate limit, rejects options the v1 engine cannot honour, and builds a legacy TaskRequest. If no URL was given, an LLM call (generate_task) derives one first (agent_protocol.py).
  2. Persist and dispatch. task_v1_service.run_task creates the task and a task_runs row, asks the AGENT_FUNCTION hook to resolve the engine (OSS returns the requested one), and calls the executor (task_v1_service.py). BackgroundTaskExecutor.execute_task creates step 0, marks the task running, maps the run type back to an engine, and schedules app.agent.execute_step (background_task_executor.py).
  3. Step entry. execute_step sets context, turns off completion verification for CUA engines, and either hands the whole task to Task V3 or calls agent_step (agent.py, L3950-L3970).
  4. Scrape. build_and_record_step_prompt walks a ladder of scrape strategies (normal, stop-loading, reload) through _scrape_with_type (agent.py, L7030-L7073). scrape_web_unsafe builds the element tree for the main frame and every child frame, trims it, counts tokens, and takes up to MAX_NUM_SCREENSHOTS scrolled screenshots, or just one if the tree is too large (scraper.py, L908-L956).
  5. Decide. _generate_step_actions branches by engine. For skyvern-1.0 it calls the org-aware LLM handler with the extract-action prompt and the screenshots, then parse_actions maps each JSON entry, by id/element_id, onto the scraped element (agent.py, parse_actions.py).
  6. Act. _execute_step_actions runs each action through ActionHandler.handle_action, which dispatches to a handler registered per ActionType (agent.py, handler.py). A click resolves the element by id, checks the desired toggle state and the disabled state, scrolls it into view and clicks with Playwright. Coordinate clicks go to page.mouse (handler.py).
  7. Verify and loop. A failed step goes to handle_failed_step, which retries up to MAX_RETRIES_PER_STEP (5 unless the org overrides it). A successful one goes to handle_completed_step, which can run complete_verify (a second LLM call over a fresh scrape) and creates the next step. execute_step then calls itself recursively (agent.py, L10181-L10200).
  8. Finish. clean_up_task uploads video, HAR and artifacts, closes the browser unless it belongs to a persistent session, and fires the webhook.

Key components

Element tree perception

domUtils.js runs in each frame, decides what is interactable (inputs, hover styles, ARIA roles), stamps unique_id, and returns elements plus a nested tree. Python side, build_element_dict turns each id into a CSS selector [unique_id='…']. json_to_html renders the trimmed tree for the prompt. Placeholder nodes stand in for frames the filter skipped, so the model sees a cross-origin CAPTCHA frame but can never target it (scraper.py). When an ARIA popup is open, scroll capture can be suppressed so the dropdown survives into the next action.

Four engines

The RunEngine enum lists them in one place (run_enums.py):

  • skyvern-1.0 is the step engine above: one scrape, one LLM call and a batch of actions per step.
  • skyvern-2.0 is a planner. On each iteration run_task_v2_helper scrapes, asks the task_v2 prompt for user_goal_achieved and a task_type (navigate, extract, loop, compute or goto_url), generates a matching workflow block and runs it with block.execute_safe (task_v2_service.py, L1262-L1272). The result is a reusable workflow, not only an answer.
  • skyvern-3.0 (Task V3) drops the per-step prompt. A single conversation calls tools such as observe, get_html, look, click, type, select_option, navigate and file_upload until it calls finish (tools.py). Backstops are 80 turns, 300 tool calls, 1,800 seconds and 1.5M tokens (engine.py). The module docstring says why: compact observe snapshots matched raw DOM at equal success with a tighter tail (engine.py).
  • CUA engines receive one screenshot without scrolling and return coordinate actions through _generate_cua_actions, _generate_anthropic_actions, _generate_ui_tars_actions or _generate_yutori_navigator_actions.

Action handlers

ActionHandler keeps three registries (setup, main, teardown) keyed by ActionType (handler.py). The 16,000-line handler.py is where most site-specific hardening lives: custom dropdowns, auto-complete inputs, date pickers, multi-field TOTP, downloads and file choosers. In the open-source build, SOLVE_CAPTCHA just logs “Please solve the captcha” and sleeps 30 seconds (handler.py). The solving ladder in captcha_solver.py calls AGENT_FUNCTION hooks that return False outside the cloud (captcha_solver.py).

LLM layer

LLMAPIHandlerFactory.get_llm_api_handler looks up a key in LLMConfigRegistry and returns a LiteLLM-backed handler or a LiteLLM Router for fallback groups (api_handler_factory.py). Providers are switched on with ENABLE_* settings: OpenAI, Anthropic, Azure, Bedrock, Gemini, Vertex, Groq, xAI, OpenRouter, Ollama, Novita, Moonshot, Volcengine, Yutori and any OpenAI-compatible endpoint. Screenshots are silently dropped for models without supports_vision, except for prompts that require vision, which raise an error (api_handler_factory.py). LLM_KEY and SECONDARY_LLM_KEY split the main reasoning model from cheaper helper calls.

Browser management

The OSS app wires RealBrowserManager as its browser manager (forge_app.py). BrowserContextFactory registers three launch modes: chromium-headless, chromium-headful and cdp-connect for attaching to an existing Chrome (browser_factory.py). Persistent browser sessions, saved browser profiles, cookie restore and per-run proxy location are first-class. A per-run engine-selection seam lets a future image ship drivers other than stock Playwright (browser_engine.py).

Caches and code generation

Extraction results are cached in process, per workflow run, keyed by a hash of the element tree, page text, URL, goal, schema and model (extraction_cache.py). The cross-run Redis tier sits behind cloud hooks. Separately, run_with: "code" replays a workflow as a generated Python script (generate_workflow_script_python_code, built with libcst), with AI fallbacks per action. Repeat runs can then skip most LLM calls (runs.py).

Extending it

  • Workflows. Compose YAML/JSON workflows from the block types in BlockType, with parameters, credentials, loops and conditionals. This is the main customisation surface, and the UI has a visual editor for it.
  • SDK. skyvern.library wraps a Playwright page with page.act(...), ai_click, ai_extract, ai_validate and an agent.run_task/run_workflow, so AI steps can be mixed with normal Playwright code (skyvern_browser_page.py).
  • MCP. skyvern/cli/mcp_tools exposes sessions, tabs, network, credentials, workflows, schedules and scripts as FastMCP tools. It also has deterministic skyvern_observe/skyvern_execute primitives next to the LLM-backed skyvern_act and skyvern_run_task (mcp_tools/init.py).
  • Models. Add a custom LLM through OPENAI_COMPATIBLE_* settings or a per-org custom LLM record. Change prompts in skyvern/forge/prompts/skyvern/*.j2.
  • Platform hooks. AgentFunction has about 180 overridable methods, from engine routing to CAPTCHA solving and audit logging. The OSS versions are mostly pass-throughs (agent_functions.py). Subclass it to change behaviour without forking the agent.

Running it

  • pip. pip install skyvern, then skyvern quickstart (or skyvern init + skyvern run all). It supports Python 3.11 to 3.14 and needs Playwright’s Chromium and PostgreSQL (the SDK’s `Skyvern.local()` can use an in-memory database instead). skyvern run server and skyvern run ui start the pieces separately.
  • Docker Compose. docker-compose.yml runs postgres:14, the skyvern API image (API plus a VNC WebSocket for live view) and skyvern-ui, with artifacts, videos, HAR files and logs on mounted volumes (docker-compose.yml). Provider keys and LLM_KEY go in .env.
  • Kubernetes. Manifests live in kubernetes-deployment/ and k8s/.
  • Licence. AGPL-3.0, which matters if you embed the server in a hosted product.

Strengths and caveats

  • Strength: grounded actions. The model selects element ids from a real DOM tree that includes iframes and shadow roots. Execution uses Playwright locators, not pixel guesses, so runs are inspectable and replayable.
  • Strength: more than a loop. Workflows, credentials (Bitwarden, 1Password), TOTP, downloads, scheduled runs, video and HAR artifacts, and code generation for cheap reruns cover what a production form-filling job needs.
  • Strength: engine choice. The same task API can run a cheap step engine, a planner, a tool-loop agent or a vendor CUA model, which makes side-by-side comparisons easy.
  • Caveat: size and churn. agent.py is 11,600 lines, handler.py 16,400 and taskv3/tools.py 20,500. Behaviour often hangs on experiment flags (DISABLE_TASK_V3, FORCE_TASK_V1) whose real values only exist in the cloud. Reading the code to predict behaviour is hard.
  • Caveat: open core. CAPTCHA solving, cross-run extraction caching, engine A/B routing and some stealth or browser engines are hooks that do nothing in the OSS build. Expect a self-hosted instance to stall on CAPTCHAs.
  • Caveat: heavy per-step prompts. In skyvern-1.0, every step re-scrapes and sends the HTML tree plus up to ten screenshots, and completion checks add calls. Task V3 exists largely to cut that cost.
  • Caveat: infrastructure. The normal deployment is a server process plus PostgreSQL. `Skyvern.local()` can run the same stack in-process, optionally on an in-memory database, but it is still the whole platform, not a thin library.

Sources: code at 44de8cd, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (22 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit 44de8cd. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

The page is represented to the model as a combination of a scraped interactive element tree and full-page screenshots (base64 PNG, potentially multiple scroll-segments). Both are sent to the LLM on every step.

DOM serialisation + accessibility tree. The scraper (skyvern/webeye/scraper/scraper.py:908-956, get_interactable_element_tree) runs a JavaScript DOM inspector that walks the page's accessible, visible, interactable elements. It annotates each with attributes (type, role, aria-*, value, placeholder, disabled, readonly, text, href, etc.), bounding rects, and a unique_id Skyvern-injected attribute (SKYVERN_ID_ATTR). The tree includes all child frames recursively (add_frame_interactable_elements, scraper.py:827–904). The resulting list is rendered as HTML (via json_to_html in scraped_page.py:63-132) and sent to the LLM — typically the "trimmed" tree where non-essential attributes and non-interactable nodes are pruned (trim_element_tree, scraper.py:1247-1250). Three tree variants exist: the full tree, the economy tree (strips SVG branches), and the lean tree (optionally compresses long hrefs/srcs).

Size limits and pruning. The tree is token-counted after rendering (scraper.py:623): if it exceeds DEFAULT_MAX_TOKENS (a tunable ceiling), the screenshot count is capped to 1. Individual elements have their attributes filtered to a reserved set (RESERVED_ATTRIBUTES, scraper.py:113–143); enriched trees include additional accessibility attributes. Hashed URLs >150 chars are replaced with a SHA-256 placeholder (json_to_html, scraped_page.py:82-88).

Screenshots. Screenshots are taken after the tree build via SkyvernFrame.take_split_screenshots (scraper.py:679-686): the page is scrolled in viewport-sized increments, capturing each segment. Screenshot count is bounded by MAX_NUM_SCREENSHOTS (default 10). When a transient popup is detected, scrolling is suppressed so the overlay survives into the next action (scraper.py:632-641). The agent can also suppress screenshots entirely via the take_screenshots=False flag.

Set-of-marks/element index. Skyvern injects a unique_id attribute into each interactable element's DOM node at scrape time. The model refers to elements by this ID (the "element_id" in actions like click/input), NOT by coordinates or CSS selectors in the prompt. The scraped page maintains an id_to_css_dict mapping these IDs to CSS selectors for Playwright resolution (scraper.py:266-282). Text content and extracted text (from get_frame_text, scraper.py:703) are also included.

How are actions executed and how are elements targeted?

answered

Actions are executed through Playwright (async API) against the live browser page. The codebase defines ~30 action types (skyvern/webeye/actions/action_types.py:4-37), including click, input_text, select_option, upload_file, scroll, keypress, move, drag, goto_url, go_back, go_forward, close_page, new_tab, switch_tab, solve_captcha, wait, hover, execute_js, reload_page, and the decisive actions complete/terminate.

Element targeting uses three strategies, ordered by preference:

  1. Skyvern element_id — The model names elements by their injected unique_id attribute. The action handler resolves this to a CSS selector (skyvern/webeye/scraper/scraper.py:275, id_to_css_dict[element_id] = f"[{SKYVERN_ID_ATTR}='{element_id}']") and uses Playwright's frame.locator() to find it. This is the primary targeting method.
  2. Coordinate-based — click actions can include x/y pixel coordinates (e.g. for image maps). The handler uses page.mouse.click(x, y, button) directly (handler.py:7473-7480).
  3. XPath fallback — Each stored action can carry an xpath field derived from skyelement data (actions.py:291-298).

Click execution is the most complex handler (handler.py:7428-7524+): it resolves the element by ID, checks disabled state (retargeting to a child if the parent is inert), validates desired state for toggle controls, scrolls into view, optionally moves the cursor via events, then calls locator.click() or page.mouse.click() with support for left/right button, double-click, and triple-click (via repeat).

Other action mechanics: INPUT_TEXT fills by clearing then typing via Playwright's fill() or type(), with TOTP code injection support. SELECT_OPTION uses locator.select_option() by label/value/index. UPLOAD_FILE uses Playwright's file chooser. SCROLL uses page.evaluate('window.scrollTo(...)'). CLOSE_PAGE closes the current Playwright page; NEW_TAB/SWITCH_TAB manage page objects. EXECUTE_JS runs arbitrary JavaScript. SOLVE_CAPTCHA routes through the AGENT_FUNCTION seam (the OSS base returns False; cloud deployments solve via vendor handlers). The Task V3 loop (forge/taskv3/tools.py) offers raw-browser tools — click, type, select_option, press_key, file_upload, navigate, scroll, wait, solve_captcha, reload_page, and observe — each as a ToolSpec the model calls via tool_calls.

How is the agent loop / planning implemented?

answered

Skyvern has three major agent loop implementations, selectable per task:

Legacy step engine (V1/V2). Located in skyvern/forge/agent.py. The loop (execute_step at line 3653) is: scrape the page → build prompt with element tree + screenshots → call LLM → parse actions → execute each action via the action handler → verify goal completion → repeat or terminate. Each iteration is one "step" with scraped page state. The LLM response is parsed into a list of Action objects (via parse_actions in parse_actions.py) which are then dispatched to ActionHandler. After action execution, the loop runs complete_verify — a separate LLM call that checks whether the user's goal was achieved. Stop conditions: COMPLETE or TERMINATE action emitted by the model, or budget exhaustion (max steps, max retries). The V2 variant includes message history across steps via llm_messages_builder_with_history.

Task V3 engine (persistent-conversation tool-loop). Located in skyvern/forge/taskv3/engine.py and loop.py. This is fundamentally different: it runs one persistent LLM conversation where the model calls tools (not a fixed prompt + actions) in a tool-use loop. Perception is a tool the model calls (observe), not automatically injected. Tools are defined as ToolSpec objects in tools.py (click, type, navigate, scroll, etc.). The loop calls run_agent_tool_loop (loop.py:1-13), which threads tool results back as tool-role messages. It is capped by DEFAULT_MAX_TURNS (80), DEFAULT_MAX_TOOL_CALLS (300), a wall-clock deadline (1800s), and a token budget (1.5M). Goal checking runs via GoalJudge (goal_check.py) with configurable verification. An unlisted_reask mechanism lets the model ask the user about unprompted decisions.

Third-party engines. CUA (OpenAI/Anthropic computer-use) and UI-TARS are supported as alternative engines. CUA tasks pass an existing OpenAIResponse object into execution (agent.py:3669). Yutori Navigator is a separate LLM caller for navigation-specific applications.

Across all loops, the SkyvernContext object carries step state, error codes, secrets, TOTP state, and multi-field TOTP tracking (multi-fill OTP splitting). The step/action history is persisted to the database for debugging and retry.

Editor's note. Correction: skyvern-2.0 is not a variant of the step engine. It is a planner (task_v2_service.run_task_v2_helper) that asks the task_v2 prompt for the next navigate/extract/loop/compute block, generates that workflow block and runs it with block.execute_safe. The step engine itself is skyvern-1.0.

How are failures, retries and self-healing handled?

answered

Reliability is handled at multiple layers with a rich exception hierarchy and explicit retry mechanics.

Error classes. Skyvern defines dozens of exception types (skyvern/exceptions.py), including MissingElement, MultipleElementsFound, InteractWithDisabledElement, FailToClick, ScrapingFailed, ActionExecutionTimeout, CaptchaSolveError, LLMProviderError, and SkyvernPageAnalysisTimeout. Each action handler wraps its operation and raises typed exceptions, which are caught by the step loop.

Retries within a step. The action handler returns ActionResult objects (responses.py), which have a success flag and optional exception_type. The step engine evaluates these: on failure, it retries the step (up to max_retries_per_step, default 3). The step has an explicit retry_index field (set in agent.py:3676). The entire step retry loop re-scrapes the page and re-calls the LLM. Scraping itself has retry logic (scrape_website, scraper.py:285-410): up to MAX_SCRAPING_RETRIES (currently 0 in prod) with an intermediate 3s wait and page reload.

Replanning. On step failure, the agent can escalate to a "recovery" step or, in workflow context, the WorkflowService triggers replanning with different credential parameters (e.g. a fresh login). The failure_classifier.py module categorises errors into FailureCategory values (LLM_ERROR, CAPTCHA_SOLVE_FAILURE, etc.), which feed into the retry/recovery decision.

Caching. The extraction cache (skyvern/forge/sdk/cache/) caches LLM extraction responses for up to 4 weeks (cache expiry = timedelta(weeks=4), base.py:6), keyed by page content hash, so identical pages reuse prior extraction results. It supports both local in-memory and Redis backends. The extraction shadow maintains a parallel write-through cache for consistency verification.

Self-heal tracking (services/self_heal_reliability_service.py) tracks workflow reliability metrics: it collects recent workflow run outcomes and "heal episodes" (instances where a recovery action was taken), computing a WorkflowReliability score per workflow over the RELIABILITY_WINDOW. This is a reporting/observability system rather than an automatic healing mechanism.

Timeouts. Multiple timeout layers exist: BROWSER_LOADING_TIMEOUT_MS for page loads, SkyvernPageAnalysisTimeout for page analysis, ActionExecutionTimeout for individual actions, task-level timeouts (TaskAlreadyTimeout), and workflow-level timeout. The Task V3 loop adds wall-clock, token, and tool-call budget caps (engine.py:91-94).

Editor's note. Correction: MAX_RETRIES_PER_STEP defaults to 5, not 3. The OSS extraction cache is in-process and scoped per workflow run (LRU/FIFO caps); the 4-week TTL is the generic local cache constant, and a cross-run Redis tier exists only behind cloud AgentFunction hooks.

Which models are supported and how are they called?

answered

Skyvern supports all major LLM providers through LiteLLM, with a plugin-style configuration registry (skyvern/forge/sdk/api/llm/config_registry.py).

Supported providers and models are registered via _register_builtin_config under gate-keeping env variable flags (e.g. ENABLE_OPENAI, ENABLE_ANTHROPIC, ENABLE_BEDROCK, ENABLE_GEMINI, ENABLE_AZURE, ENABLE_YUTORI). The built-in keys include:

  • OpenAI: GPT-4o, GPT-4.1, GPT-4.5, GPT-5 family (GPT-5 mini/nano, GPT-5.1/5.2/5.4/5.5/5.6), GPT-6 family (Astra/Sol/Luna), o3-mini, o4-mini (through standard, flex, and nano tiers)
  • Anthropic: Claude 4 Opus/Sonnet, 4.5 Sonnet/Haiku/Opus, 4.6 Opus/Sonnet, 4.7 Opus, 4.8 Opus, 5 Fable, 5.1 Fable, 5 Opus, 5.5 Opus, 5.5 Sonnet
  • Azure OpenAI: Deployments for all the above models via Azure endpoints
  • AWS Bedrock: Amazon Nova Pro/Lite, Anthropic Claude inference profiles (4 through 5.5)
  • Gemini (direct): Flash 2.0, Flash 2.0 Lite, Pro, 2.5 Pro, 2.5 Pro Preview, 2.5 Pro Exp, 3.0 models
  • Yutori Navigator: Custom navigation-focused model
  • OpenRouter: Dynamic model resolution with supports_vision
  • Custom LLM: OPENAI_COMPATIBLE_API_BASE for any OpenAI-compatible endpoint (Ollama, vLLM, etc.) and fully custom providers via CUSTOM_LLM_KEY
  • xAI: grok-4.5 (registered with its own cost table)

Vision requirement. The LLMConfig has a supports_vision boolean. Screenshots are stripped from messages sent to non-vision models (api_handler_factory.py:658-665). The extract-actions prompt is the primary vision user.

Structured output / tool calling. All models are called via LiteLLM's acompletion with OpenAI-format messages and tools. The router handler (api_handler_factory.py:1867-1920+) builds a LiteLLM Router with configurable main/fallback model groups and retry policies. Thinking budgets (budget_tokens for Claude, reasoning_effort for GPT-5, thinking_level for Gemini 3) are applied per-prompt-name. Prompt caching is injected for OpenAI and Vertex models.

Small/specialised models. Secondary models (configured via SECONDARY_LLM_KEY) can be used for lightweight operations like goal verification. The UItarsLLMCaller and YutoriNavigatorLLMCaller provide specialised inference paths. Custom LLMs can be registered per-organization with their own API base, key, and model family.

How are browser sessions, profiles, auth and anti-bot handled?

answered

Skyvern manages browser sessions through a layered architecture built on Playwright, supporting both local ephemeral browsers and remote persistent sessions.

Local vs remote/cloud browsers. The BrowserManager protocol (skyvern/webeye/browser_manager.py:23-100) defines get_or_create_for_task/get_or_create_for_workflow_run to acquire a browser. The default implementation creates Playwright browser instances locally. Cloud deployments use RealBrowserManager (real_browser_manager.py) backed by CDP-connected remote browsers. The BrowserAcquisitionSample module records the acquisition mode (create vs attach vs reuse). The engine selection determines whether to use local Playwright, a remote CDP endpoint, or a cloud persistent session.

Persistent browser sessions. The PersistentSessionsManager protocol (persistent_sessions_manager.py:54-80) manages a pool of pre-warmed, reusable browser sessions. These are long-lived browser instances (launched headlessly on remote machines) that can be leased by workflows and tasks. The protocol supports session watching, a reaper for idle/expired sessions, eviction/reconnection, and startup timeout configuration. Sessions are tracked in the database through PersistentBrowserSession records.

Profiles and cookies. The browser factory (browser_factory.py) orchestrates profile-based browsing. At startup, it can:

  • Apply a browser_profile_id (pre-saved profile) — a snapshot including cookies and local storage
  • Restore session cookies from a stored directory (browser_factory.py:826-827)
  • Restore "banked" cookies for verified-login healing (restore_banked_cookies, browser_factory.py:831-832)
  • Restore sign-in cookies for OAuth recovery
  • Snapshot seed profile state at start (_capture_seed_profile_state, browser_factory.py:277-291) to track cookie changes for freshness guards

Stealth and anti-bot. The project includes anti-detection measures: it monitors CDP connections, detects challenge vendors via URL signatures (challenge_signature.py:12-14: captcha, turnstile, cloudflare, arkoselabs, funcaptcha, datadome, perimeterx), and runs a captcha-solving ladder (captcha_solver.py) that probes for DOM checkbox challenges, reCAPTCHA anchor frames, solver extensions, and token-based routes. The DialogHandler auto-accepts browser dialogs. CDP download interception mediates network requests and proxy auth challenges.

Proxies. Proxy location is a first-class concept (ProxyLocation, ProxyLocationInput). The browser factory configures Playwright's proxy settings and sets timezone info based on the proxy location. CDP interceptors pass proxy credentials for authenticated proxies.

CAPTCHA handling. The captcha solver (utils/captcha_solver.py) provides a multi-arm solving ladder: DOM checkbox click, reCAPTCHA anchor iframe click, solver extension, and reCAPTCHA token route. The vendor challenge signature (CHALLENGE_VENDOR_SIGNATURE) detects cloudflare Turnstile, DataDome, PerimeterX, reCAPTCHA, hCaptcha, and FUNCAPTCHA.

Editor's note. Correction: RealBrowserManager is the browser manager the open-source app wires in (forge_app.py), not a cloud-only implementation. It launches local chromium-headless/headful browsers or attaches over CDP (cdp-connect).