# Skyvern-AI/skyvern

> Self-hosted LLM browser agent that picks actions from a scraped DOM tree plus screenshots, runs them in Playwright, and chains workflows.

- Category: [Browser & computer control](https://llms-technical-reviews.com/browser-control/)
- Repository: https://github.com/Skyvern-AI/skyvern (reviewed at commit `44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1`, 2026-10-05)
- Stars: 23142 · Language: Python · License: AGPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/skyvern/

## Overview

Skyvern is a self-hostable server that runs browser tasks described in natural language. You send a prompt and a start URL to its REST API, and a background worker drives a Playwright browser until an LLM decides the goal is met. Around that core sit a workflow engine (around 30 block types: navigation, extraction, login, loops, conditionals, code, HTTP, email, PDF), a React UI with live browser streaming, a Python SDK, a CLI and an MCP server with over a hundred tools.

The page is shown to the model as a DOM-derived element tree, not as raw pixels. An injected script tags every interactable element with a `unique_id` attribute and serialises the tree as compact HTML. That HTML goes to the LLM together with a few scrolled screenshots. The model answers with a JSON list of actions that name elements by id, and Python handlers execute them with Playwright locators. Despite the "computer vision" in the marketing, screenshots are context only. The set-of-marks box overlay is switched off and marked deprecated in the scraper.

At this commit, "Skyvern" means four different agents behind one API, chosen by the `engine` field: `skyvern-1.0` (the default per-step engine), `skyvern-2.0` (a planner that writes and runs workflow blocks), `skyvern-3.0` (a single persistent tool-calling conversation), and third-party computer-use models (OpenAI CUA, Anthropic CUA, UI-TARS, Yutori Navigator). The repository is also the open core of a hosted product. Many behaviours, such as CAPTCHA solving, cross-run caches and A/B engine routing, are hooks that are no-ops here and are filled in by the cloud.

## Architecture

```mermaid
flowchart LR
  C["Client: SDK / CLI / MCP / UI"] --> API["FastAPI routes"]
  API --> EX["BackgroundTaskExecutor"]
  EX --> AG["ForgeAgent.execute_step"]
  API --> T2["task_v2_service planner"]
  T2 --> WF["Workflow blocks"]
  WF --> AG
  AG --> V3["Task V3 tool loop"]
  AG --> SC["Scraper + domUtils.js"]
  AG --> LLM["LLMAPIHandlerFactory (LiteLLM)"]
  AG --> AH["ActionHandler"]
  SC --> BM["RealBrowserManager"]
  AH --> BM
  V3 --> BM
  BM --> PW["Playwright Chromium or CDP"]
  AG --> DB["PostgreSQL"]
```

| Component | Path | Role |
|---|---|---|
| API routes | `skyvern/forge/sdk/routes/agent_protocol.py` | `/v1/run/tasks`, workflow runs, engine dispatch |
| Executor | `skyvern/forge/sdk/executor/background_task_executor.py` | Creates the first step and schedules `execute_step` as a FastAPI background task |
| Step engine | `skyvern/forge/agent.py` | `ForgeAgent`: scrape, prompt, LLM, parse, act, verify, recurse |
| Planner (2.0) | `skyvern/services/task_v2_service.py` | Iterative "plan next block" loop that builds and runs a workflow |
| Task V3 | `skyvern/forge/taskv3/` | Persistent tool-use conversation (`engine.py`, `loop.py`, `tools.py`) |
| Scraper | `skyvern/webeye/scraper/` | `scraper.py` + 4,000-line `domUtils.js`: element tree, frames, split screenshots |
| Actions | `skyvern/webeye/actions/` | Action models, `parse_actions.py`, the 16k-line `handler.py` |
| Browser | `skyvern/webeye/real_browser_manager.py`, `browser_factory.py` | Browser acquisition, persistent sessions, profiles, proxies |
| LLM layer | `skyvern/forge/sdk/api/llm/` | Config registry, LiteLLM handlers and routers, `LLMCaller` for chat-history engines |
| Workflows | `skyvern/forge/sdk/workflow/` | Block models, parameters, `service.py` |
| Script generation | `skyvern/core/script_generations/` | Turns recorded runs into Playwright-style Python (libcst) |
| Client surfaces | `skyvern/library/`, `skyvern/cli/`, `skyvern-frontend/` | SDK, Typer CLI + FastMCP server, React UI |

## How a request flows

Take `POST /v1/run/tasks` with `engine: skyvern-1.0`:

1. **Accept.** `run_task` validates the webhook URL and rate limit, rejects options the v1 engine cannot honour, and builds a legacy `TaskRequest`. If no URL was given, an LLM call (`generate_task`) derives one first ([agent_protocol.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/routes/agent_protocol.py#L310-L450)).
2. **Persist and dispatch.** `task_v1_service.run_task` creates the task and a `task_runs` row, asks the `AGENT_FUNCTION` hook to resolve the engine (OSS returns the requested one), and calls the executor ([task_v1_service.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/services/task_v1_service.py#L111-L182)). `BackgroundTaskExecutor.execute_task` creates step 0, marks the task running, maps the run type back to an engine, and schedules `app.agent.execute_step` ([background_task_executor.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/executor/background_task_executor.py#L697-L760)).
3. **Step entry.** `execute_step` sets context, turns off completion verification for CUA engines, and either hands the whole task to Task V3 or calls `agent_step` ([agent.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L3653-L3700), [L3950-L3970](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L3950-L3970)).
4. **Scrape.** `build_and_record_step_prompt` walks a ladder of scrape strategies (normal, stop-loading, reload) through `_scrape_with_type` ([agent.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L7430-L7460), [L7030-L7073](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L7030-L7073)). `scrape_web_unsafe` builds the element tree for the main frame and every child frame, trims it, counts tokens, and takes up to `MAX_NUM_SCREENSHOTS` scrolled screenshots, or just one if the tree is too large ([scraper.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L539-L700), [L908-L956](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L908-L956)).
5. **Decide.** `_generate_step_actions` branches by engine. For skyvern-1.0 it calls the org-aware LLM handler with the `extract-action` prompt and the screenshots, then `parse_actions` maps each JSON entry, by `id`/`element_id`, onto the scraped element ([agent.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L5497-L5615), [parse_actions.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/parse_actions.py#L73-L100)).
6. **Act.** `_execute_step_actions` runs each action through `ActionHandler.handle_action`, which dispatches to a handler registered per `ActionType` ([agent.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L5250-L5262), [handler.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L11622-L11648)). A click resolves the element by id, checks the desired toggle state and the disabled state, scrolls it into view and clicks with Playwright. Coordinate clicks go to `page.mouse` ([handler.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L7428-L7530)).
7. **Verify and loop.** A failed step goes to `handle_failed_step`, which retries up to `MAX_RETRIES_PER_STEP` (5 unless the org overrides it). A successful one goes to `handle_completed_step`, which can run `complete_verify` (a second LLM call over a fresh scrape) and creates the next step. `execute_step` then calls itself recursively ([agent.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L4060-L4160), [L10181-L10200](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L10181-L10200)).
8. **Finish.** `clean_up_task` uploads video, HAR and artifacts, closes the browser unless it belongs to a persistent session, and fires the webhook.

## Key components

### Element tree perception

`domUtils.js` runs in each frame, decides what is interactable (inputs, hover styles, ARIA roles), stamps `unique_id`, and returns elements plus a nested tree. Python side, `build_element_dict` turns each id into a CSS selector `[unique_id='…']`. `json_to_html` renders the trimmed tree for the prompt. Placeholder nodes stand in for frames the filter skipped, so the model sees a cross-origin CAPTCHA frame but can never target it ([scraper.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L908-L956)). When an ARIA popup is open, scroll capture can be suppressed so the dropdown survives into the next action.

### Four engines

The `RunEngine` enum lists them in one place ([run_enums.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/schemas/run_enums.py#L15-L26)):

- **skyvern-1.0** is the step engine above: one scrape, one LLM call and a batch of actions per step.
- **skyvern-2.0** is a planner. On each iteration `run_task_v2_helper` scrapes, asks the `task_v2` prompt for `user_goal_achieved` and a `task_type` (navigate, extract, loop, compute or goto_url), generates a matching workflow block and runs it with `block.execute_safe` ([task_v2_service.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/services/task_v2_service.py#L1025-L1075), [L1262-L1272](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/services/task_v2_service.py#L1262-L1272)). The result is a reusable workflow, not only an answer.
- **skyvern-3.0** (Task V3) drops the per-step prompt. A single conversation calls tools such as `observe`, `get_html`, `look`, `click`, `type`, `select_option`, `navigate` and `file_upload` until it calls `finish` ([tools.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/tools.py#L19868-L19900)). Backstops are 80 turns, 300 tool calls, 1,800 seconds and 1.5M tokens ([engine.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/engine.py#L86-L96)). The module docstring says why: compact `observe` snapshots matched raw DOM at equal success with a tighter tail ([engine.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/engine.py#L1-L20)).
- **CUA engines** receive one screenshot without scrolling and return coordinate actions through `_generate_cua_actions`, `_generate_anthropic_actions`, `_generate_ui_tars_actions` or `_generate_yutori_navigator_actions`.

### Action handlers

`ActionHandler` keeps three registries (setup, main, teardown) keyed by `ActionType` ([handler.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L4718-L4770)). The 16,000-line `handler.py` is where most site-specific hardening lives: custom dropdowns, auto-complete inputs, date pickers, multi-field TOTP, downloads and file choosers. In the open-source build, `SOLVE_CAPTCHA` just logs "Please solve the captcha" and sleeps 30 seconds ([handler.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L6588-L6600)). The solving ladder in `captcha_solver.py` calls `AGENT_FUNCTION` hooks that return False outside the cloud ([captcha_solver.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/utils/captcha_solver.py#L1-L13)).

### LLM layer

`LLMAPIHandlerFactory.get_llm_api_handler` looks up a key in `LLMConfigRegistry` and returns a LiteLLM-backed handler or a LiteLLM Router for fallback groups ([api_handler_factory.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/api_handler_factory.py#L2700-L2720)). Providers are switched on with `ENABLE_*` settings: OpenAI, Anthropic, Azure, Bedrock, Gemini, Vertex, Groq, xAI, OpenRouter, Ollama, Novita, Moonshot, Volcengine, Yutori and any OpenAI-compatible endpoint. Screenshots are silently dropped for models without `supports_vision`, except for prompts that require vision, which raise an error ([api_handler_factory.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/api_handler_factory.py#L650-L680)). `LLM_KEY` and `SECONDARY_LLM_KEY` split the main reasoning model from cheaper helper calls.

### Browser management

The OSS app wires `RealBrowserManager` as its browser manager ([forge_app.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/forge_app.py#L145-L160)). `BrowserContextFactory` registers three launch modes: `chromium-headless`, `chromium-headful` and `cdp-connect` for attaching to an existing Chrome ([browser_factory.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/browser_factory.py#L1480-L1482)). Persistent browser sessions, saved browser profiles, cookie restore and per-run proxy location are first-class. A per-run engine-selection seam lets a future image ship drivers other than stock Playwright ([browser_engine.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/browser_engine.py#L1-L20)).

### Caches and code generation

Extraction results are cached in process, per workflow run, keyed by a hash of the element tree, page text, URL, goal, schema and model ([extraction_cache.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/cache/extraction_cache.py#L1-L54)). The cross-run Redis tier sits behind cloud hooks. Separately, `run_with: "code"` replays a workflow as a generated Python script (`generate_workflow_script_python_code`, built with libcst), with AI fallbacks per action. Repeat runs can then skip most LLM calls ([runs.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/schemas/runs.py#L327-L331)).

## Extending it

- **Workflows.** Compose YAML/JSON workflows from the block types in `BlockType`, with parameters, credentials, loops and conditionals. This is the main customisation surface, and the UI has a visual editor for it.
- **SDK.** `skyvern.library` wraps a Playwright page with `page.act(...)`, `ai_click`, `ai_extract`, `ai_validate` and an `agent.run_task`/`run_workflow`, so AI steps can be mixed with normal Playwright code ([skyvern_browser_page.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/library/skyvern_browser_page.py#L146-L163)).
- **MCP.** `skyvern/cli/mcp_tools` exposes sessions, tabs, network, credentials, workflows, schedules and scripts as FastMCP tools. It also has deterministic `skyvern_observe`/`skyvern_execute` primitives next to the LLM-backed `skyvern_act` and `skyvern_run_task` ([mcp_tools/__init__.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/cli/mcp_tools/__init__.py#L348-L349)).
- **Models.** Add a custom LLM through `OPENAI_COMPATIBLE_*` settings or a per-org custom LLM record. Change prompts in `skyvern/forge/prompts/skyvern/*.j2`.
- **Platform hooks.** `AgentFunction` has about 180 overridable methods, from engine routing to CAPTCHA solving and audit logging. The OSS versions are mostly pass-throughs ([agent_functions.py](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent_functions.py#L988-L1017)). Subclass it to change behaviour without forking the agent.

## Running it

- **pip.** `pip install skyvern`, then `skyvern quickstart` (or `skyvern init` + `skyvern run all`). It supports Python 3.11 to 3.14 and needs Playwright's Chromium and PostgreSQL (the SDK's \`Skyvern.local()\` can use an in-memory database instead). `skyvern run server` and `skyvern run ui` start the pieces separately.
- **Docker Compose.** `docker-compose.yml` runs `postgres:14`, the `skyvern` API image (API plus a VNC WebSocket for live view) and `skyvern-ui`, with artifacts, videos, HAR files and logs on mounted volumes ([docker-compose.yml](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/docker-compose.yml#L20-L40)). Provider keys and `LLM_KEY` go in `.env`.
- **Kubernetes.** Manifests live in `kubernetes-deployment/` and `k8s/`.
- **Licence.** AGPL-3.0, which matters if you embed the server in a hosted product.

## Strengths and caveats

- **Strength: grounded actions.** The model selects element ids from a real DOM tree that includes iframes and shadow roots. Execution uses Playwright locators, not pixel guesses, so runs are inspectable and replayable.
- **Strength: more than a loop.** Workflows, credentials (Bitwarden, 1Password), TOTP, downloads, scheduled runs, video and HAR artifacts, and code generation for cheap reruns cover what a production form-filling job needs.
- **Strength: engine choice.** The same task API can run a cheap step engine, a planner, a tool-loop agent or a vendor CUA model, which makes side-by-side comparisons easy.
- **Caveat: size and churn.** `agent.py` is 11,600 lines, `handler.py` 16,400 and `taskv3/tools.py` 20,500. Behaviour often hangs on experiment flags (`DISABLE_TASK_V3`, `FORCE_TASK_V1`) whose real values only exist in the cloud. Reading the code to predict behaviour is hard.
- **Caveat: open core.** CAPTCHA solving, cross-run extraction caching, engine A/B routing and some stealth or browser engines are hooks that do nothing in the OSS build. Expect a self-hosted instance to stall on CAPTCHAs.
- **Caveat: heavy per-step prompts.** In skyvern-1.0, every step re-scrapes and sends the HTML tree plus up to ten screenshots, and completion checks add calls. Task V3 exists largely to cut that cost.
- **Caveat: infrastructure.** The normal deployment is a server process plus PostgreSQL. \`Skyvern.local()\` can run the same stack in-process, optionally on an in-memory database, but it is still the whole platform, not a thin library.

*Sources: code at 44de8cd, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (22 pages), verified Q&A.*

## How Skyvern-AI/skyvern answers the Browser & computer control questions

### How is the page represented to the model? (answered)

The page is represented to the model as a combination of a **scraped interactive element tree** and **full-page screenshots** (base64 PNG, potentially multiple scroll-segments). Both are sent to the LLM on every step.

**DOM serialisation + accessibility tree.** The scraper (`skyvern/webeye/scraper/scraper.py:908-956`, `get_interactable_element_tree`) runs a JavaScript DOM inspector that walks the page's accessible, visible, interactable elements. It annotates each with attributes (type, role, aria-*, value, placeholder, disabled, readonly, text, href, etc.), bounding rects, and a `unique_id` Skyvern-injected attribute (`SKYVERN_ID_ATTR`). The tree includes all child frames recursively (`add_frame_interactable_elements`, scraper.py:827–904). The resulting list is rendered as **HTML** (via `json_to_html` in `scraped_page.py:63-132`) and sent to the LLM — typically the "trimmed" tree where non-essential attributes and non-interactable nodes are pruned (`trim_element_tree`, scraper.py:1247-1250). Three tree variants exist: the full tree, the economy tree (strips SVG branches), and the lean tree (optionally compresses long hrefs/srcs).

**Size limits and pruning.** The tree is token-counted after rendering (scraper.py:623): if it exceeds `DEFAULT_MAX_TOKENS` (a tunable ceiling), the screenshot count is capped to 1. Individual elements have their attributes filtered to a reserved set (`RESERVED_ATTRIBUTES`, scraper.py:113–143); enriched trees include additional accessibility attributes. Hashed URLs >150 chars are replaced with a SHA-256 placeholder (`json_to_html`, scraped_page.py:82-88).

**Screenshots.** Screenshots are taken after the tree build via `SkyvernFrame.take_split_screenshots` (scraper.py:679-686): the page is scrolled in viewport-sized increments, capturing each segment. Screenshot count is bounded by `MAX_NUM_SCREENSHOTS` (default 10). When a transient popup is detected, `scrolling is suppressed` so the overlay survives into the next action (scraper.py:632-641). The agent can also suppress screenshots entirely via the `take_screenshots=False` flag.

**Set-of-marks/element index.** Skyvern injects a `unique_id` attribute into each interactable element's DOM node at scrape time. The model refers to elements by this ID (the "element_id" in actions like click/input), NOT by coordinates or CSS selectors in the prompt. The scraped page maintains an `id_to_css_dict` mapping these IDs to CSS selectors for Playwright resolution (scraper.py:266-282). Text content and extracted text (from `get_frame_text`, scraper.py:703) are also included.


Citations: [skyvern/webeye/scraper/scraper.py:908-956](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L908-L956) · [skyvern/webeye/scraper/scraper.py:113-164](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L113-L164) · [skyvern/webeye/scraper/scraper.py:597-700](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L597-L700) · [skyvern/webeye/scraper/scraped_page.py:63-132](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraped_page.py#L63-L132) · [skyvern/webeye/scraper/scraper.py:266-282](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L266-L282)

### How are actions executed and how are elements targeted? (answered)

Actions are executed **through Playwright** (async API) against the live browser page. The codebase defines ~30 action types (`skyvern/webeye/actions/action_types.py:4-37`), including click, input_text, select_option, upload_file, scroll, keypress, move, drag, goto_url, go_back, go_forward, close_page, new_tab, switch_tab, solve_captcha, wait, hover, execute_js, reload_page, and the decisive actions complete/terminate.

**Element targeting** uses three strategies, ordered by preference:
1. **Skyvern element_id** — The model names elements by their injected `unique_id` attribute. The action handler resolves this to a CSS selector (`skyvern/webeye/scraper/scraper.py:275`, `id_to_css_dict[element_id] = f"[{SKYVERN_ID_ATTR}='{element_id}']"`) and uses Playwright's `frame.locator()` to find it. This is the primary targeting method.
2. **Coordinate-based** — `click` actions can include `x`/`y` pixel coordinates (e.g. for image maps). The handler uses `page.mouse.click(x, y, button)` directly (`handler.py:7473-7480`).
3. **XPath fallback** — Each stored action can carry an `xpath` field derived from skyelement data (`actions.py:291-298`).

**Click execution** is the most complex handler (`handler.py:7428-7524+`): it resolves the element by ID, checks disabled state (retargeting to a child if the parent is inert), validates desired state for toggle controls, scrolls into view, optionally moves the cursor via events, then calls `locator.click()` or `page.mouse.click()` with support for left/right button, double-click, and triple-click (via `repeat`).

**Other action mechanics:** `INPUT_TEXT` fills by clearing then typing via Playwright's `fill()` or `type()`, with TOTP code injection support. `SELECT_OPTION` uses `locator.select_option()` by label/value/index. `UPLOAD_FILE` uses Playwright's file chooser. `SCROLL` uses `page.evaluate('window.scrollTo(...)')`. `CLOSE_PAGE` closes the current Playwright page; `NEW_TAB`/`SWITCH_TAB` manage page objects. `EXECUTE_JS` runs arbitrary JavaScript. `SOLVE_CAPTCHA` routes through the AGENT_FUNCTION seam (the OSS base returns False; cloud deployments solve via vendor handlers). The Task V3 loop (`forge/taskv3/tools.py`) offers raw-browser tools — click, type, select_option, press_key, file_upload, navigate, scroll, wait, solve_captcha, reload_page, and observe — each as a ToolSpec the model calls via tool_calls.


Citations: [skyvern/webeye/actions/action_types.py:4-37](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/action_types.py#L4-L37) · [skyvern/webeye/actions/handler.py:7428-7480](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L7428-L7480) · [skyvern/webeye/actions/handler.py:288-299](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/handler.py#L288-L299) · [skyvern/webeye/actions/actions.py:146-290](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/actions.py#L146-L290) · [skyvern/forge/taskv3/tools.py:1-11](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/tools.py#L1-L11) · [skyvern/webeye/scraper/scraper.py:272-278](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L272-L278)

### How is the agent loop / planning implemented? (answered)

Skyvern has **three major agent loop implementations**, selectable per task:

**Legacy step engine (V1/V2).** Located in `skyvern/forge/agent.py`. The loop (`execute_step` at line 3653) is: scrape the page → build prompt with element tree + screenshots → call LLM → parse actions → execute each action via the action handler → verify goal completion → repeat or terminate. Each iteration is one "step" with scraped page state. The LLM response is parsed into a list of `Action` objects (via `parse_actions` in `parse_actions.py`) which are then dispatched to `ActionHandler`. After action execution, the loop runs `complete_verify` — a separate LLM call that checks whether the user's goal was achieved. Stop conditions: `COMPLETE` or `TERMINATE` action emitted by the model, or budget exhaustion (max steps, max retries). The V2 variant includes message history across steps via `llm_messages_builder_with_history`.

**Task V3 engine (persistent-conversation tool-loop).** Located in `skyvern/forge/taskv3/engine.py` and `loop.py`. This is fundamentally different: it runs one persistent LLM conversation where the model calls **tools** (not a fixed prompt + actions) in a tool-use loop. Perception is a tool the model calls (`observe`), not automatically injected. Tools are defined as `ToolSpec` objects in `tools.py` (click, type, navigate, scroll, etc.). The loop calls `run_agent_tool_loop` (loop.py:1-13), which threads tool results back as tool-role messages. It is capped by `DEFAULT_MAX_TURNS` (80), `DEFAULT_MAX_TOOL_CALLS` (300), a wall-clock deadline (1800s), and a token budget (1.5M). Goal checking runs via `GoalJudge` (`goal_check.py`) with configurable verification. An `unlisted_reask` mechanism lets the model ask the user about unprompted decisions.

**Third-party engines.** CUA (OpenAI/Anthropic computer-use) and UI-TARS are supported as alternative engines. CUA tasks pass an existing `OpenAIResponse` object into execution (`agent.py:3669`). Yutori Navigator is a separate LLM caller for navigation-specific applications.

Across all loops, the `SkyvernContext` object carries step state, error codes, secrets, TOTP state, and multi-field TOTP tracking (multi-fill OTP splitting). The step/action history is persisted to the database for debugging and retry.

> **Editor's note.** Correction: skyvern-2.0 is not a variant of the step engine. It is a planner (task_v2_service.run_task_v2_helper) that asks the task_v2 prompt for the next navigate/extract/loop/compute block, generates that workflow block and runs it with block.execute_safe. The step engine itself is skyvern-1.0.

Citations: [skyvern/forge/agent.py:3653-3690](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/agent.py#L3653-L3690) · [skyvern/forge/taskv3/engine.py:1-100](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/engine.py#L1-L100) · [skyvern/forge/taskv3/loop.py:1-165](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/loop.py#L1-L165) · [skyvern/forge/taskv3/goal_check.py:1-10](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/goal_check.py#L1-L10) · [skyvern/webeye/actions/parse_actions.py:1-10](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/actions/parse_actions.py#L1-L10)

### How are failures, retries and self-healing handled? (answered)

Reliability is handled at multiple layers with a rich exception hierarchy and explicit retry mechanics.

**Error classes.** Skyvern defines dozens of exception types (`skyvern/exceptions.py`), including `MissingElement`, `MultipleElementsFound`, `InteractWithDisabledElement`, `FailToClick`, `ScrapingFailed`, `ActionExecutionTimeout`, `CaptchaSolveError`, `LLMProviderError`, and `SkyvernPageAnalysisTimeout`. Each action handler wraps its operation and raises typed exceptions, which are caught by the step loop.

**Retries within a step.** The action handler returns `ActionResult` objects (`responses.py`), which have a `success` flag and optional `exception_type`. The step engine evaluates these: on failure, it retries the step (up to `max_retries_per_step`, default 3). The step has an explicit `retry_index` field (set in `agent.py:3676`). The entire step retry loop re-scrapes the page and re-calls the LLM. Scraping itself has retry logic (`scrape_website`, scraper.py:285-410): up to `MAX_SCRAPING_RETRIES` (currently 0 in prod) with an intermediate 3s wait and page reload.

**Replanning.** On step failure, the agent can escalate to a "recovery" step or, in workflow context, the WorkflowService triggers replanning with different credential parameters (e.g. a fresh login). The `failure_classifier.py` module categorises errors into `FailureCategory` values (LLM_ERROR, CAPTCHA_SOLVE_FAILURE, etc.), which feed into the retry/recovery decision.

**Caching.** The extraction cache (`skyvern/forge/sdk/cache/`) caches LLM extraction responses for up to 4 weeks (cache expiry = timedelta(weeks=4), `base.py:6`), keyed by page content hash, so identical pages reuse prior extraction results. It supports both local in-memory and Redis backends. The extraction shadow maintains a parallel write-through cache for consistency verification.

**Self-heal tracking** (`services/self_heal_reliability_service.py`) tracks workflow reliability metrics: it collects recent workflow run outcomes and "heal episodes" (instances where a recovery action was taken), computing a `WorkflowReliability` score per workflow over the `RELIABILITY_WINDOW`. This is a reporting/observability system rather than an automatic healing mechanism.

**Timeouts.** Multiple timeout layers exist: `BROWSER_LOADING_TIMEOUT_MS` for page loads, `SkyvernPageAnalysisTimeout` for page analysis, `ActionExecutionTimeout` for individual actions, task-level timeouts (`TaskAlreadyTimeout`), and workflow-level timeout. The Task V3 loop adds wall-clock, token, and tool-call budget caps (engine.py:91-94).

> **Editor's note.** Correction: MAX_RETRIES_PER_STEP defaults to 5, not 3. The OSS extraction cache is in-process and scoped per workflow run (LRU/FIFO caps); the 4-week TTL is the generic local cache constant, and a cross-run Redis tier exists only behind cloud AgentFunction hooks.

Citations: [skyvern/webeye/scraper/scraper.py:285-410](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/scraper/scraper.py#L285-L410) · [skyvern/forge/sdk/cache/base.py:1-60](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/cache/base.py#L1-L60) · [skyvern/services/self_heal_reliability_service.py:1-59](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/services/self_heal_reliability_service.py#L1-L59) · [skyvern/forge/failure_classifier.py:1-10](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/failure_classifier.py#L1-L10) · [skyvern/forge/taskv3/engine.py:89-94](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/taskv3/engine.py#L89-L94)

### Which models are supported and how are they called? (answered)

Skyvern supports **all major LLM providers** through LiteLLM, with a plugin-style configuration registry (`skyvern/forge/sdk/api/llm/config_registry.py`).

**Supported providers and models** are registered via `_register_builtin_config` under gate-keeping env variable flags (e.g. `ENABLE_OPENAI`, `ENABLE_ANTHROPIC`, `ENABLE_BEDROCK`, `ENABLE_GEMINI`, `ENABLE_AZURE`, `ENABLE_YUTORI`). The built-in keys include:
- **OpenAI**: GPT-4o, GPT-4.1, GPT-4.5, GPT-5 family (GPT-5 mini/nano, GPT-5.1/5.2/5.4/5.5/5.6), GPT-6 family (Astra/Sol/Luna), o3-mini, o4-mini (through standard, flex, and nano tiers)
- **Anthropic**: Claude 4 Opus/Sonnet, 4.5 Sonnet/Haiku/Opus, 4.6 Opus/Sonnet, 4.7 Opus, 4.8 Opus, 5 Fable, 5.1 Fable, 5 Opus, 5.5 Opus, 5.5 Sonnet
- **Azure OpenAI**: Deployments for all the above models via Azure endpoints
- **AWS Bedrock**: Amazon Nova Pro/Lite, Anthropic Claude inference profiles (4 through 5.5)
- **Gemini** (direct): Flash 2.0, Flash 2.0 Lite, Pro, 2.5 Pro, 2.5 Pro Preview, 2.5 Pro Exp, 3.0 models
- **Yutori Navigator**: Custom navigation-focused model
- **OpenRouter**: Dynamic model resolution with supports_vision
- **Custom LLM**: `OPENAI_COMPATIBLE_API_BASE` for any OpenAI-compatible endpoint (Ollama, vLLM, etc.) and fully custom providers via `CUSTOM_LLM_KEY`
- **xAI**: grok-4.5 (registered with its own cost table)

**Vision requirement.** The `LLMConfig` has a `supports_vision` boolean. Screenshots are stripped from messages sent to non-vision models (`api_handler_factory.py:658-665`). The `extract-actions` prompt is the primary vision user.

**Structured output / tool calling.** All models are called via LiteLLM's `acompletion` with OpenAI-format messages and tools. The router handler (`api_handler_factory.py:1867-1920+`) builds a LiteLLM Router with configurable main/fallback model groups and retry policies. Thinking budgets (budget_tokens for Claude, reasoning_effort for GPT-5, thinking_level for Gemini 3) are applied per-prompt-name. Prompt caching is injected for OpenAI and Vertex models.

**Small/specialised models.** Secondary models (configured via `SECONDARY_LLM_KEY`) can be used for lightweight operations like goal verification. The UItarsLLMCaller and YutoriNavigatorLLMCaller provide specialised inference paths. Custom LLMs can be registered per-organization with their own API base, key, and model family.


Citations: [skyvern/forge/sdk/api/llm/config_registry.py:238-550](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/config_registry.py#L238-L550) · [skyvern/forge/sdk/api/llm/config_registry.py:594-750](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/config_registry.py#L594-L750) · [skyvern/forge/sdk/api/llm/api_handler_factory.py:658-665](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/api_handler_factory.py#L658-L665) · [skyvern/forge/sdk/api/llm/api_handler_factory.py:1867-1950](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/api_handler_factory.py#L1867-L1950) · [skyvern/forge/sdk/api/llm/config_registry.py:1390-1470](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/forge/sdk/api/llm/config_registry.py#L1390-L1470)

### How are browser sessions, profiles, auth and anti-bot handled? (answered)

Skyvern manages browser sessions through a layered architecture built on Playwright, supporting both local ephemeral browsers and remote persistent sessions.

**Local vs remote/cloud browsers.** The `BrowserManager` protocol (`skyvern/webeye/browser_manager.py:23-100`) defines `get_or_create_for_task`/`get_or_create_for_workflow_run` to acquire a browser. The default implementation creates Playwright browser instances locally. Cloud deployments use `RealBrowserManager` (`real_browser_manager.py`) backed by CDP-connected remote browsers. The `BrowserAcquisitionSample` module records the acquisition mode (`create` vs `attach` vs `reuse`). The engine selection determines whether to use local Playwright, a remote CDP endpoint, or a cloud persistent session.

**Persistent browser sessions.** The `PersistentSessionsManager` protocol (`persistent_sessions_manager.py:54-80`) manages a pool of pre-warmed, reusable browser sessions. These are long-lived browser instances (launched headlessly on remote machines) that can be leased by workflows and tasks. The protocol supports session watching, a reaper for idle/expired sessions, eviction/reconnection, and startup timeout configuration. Sessions are tracked in the database through `PersistentBrowserSession` records.

**Profiles and cookies.** The browser factory (`browser_factory.py`) orchestrates profile-based browsing. At startup, it can:
  - Apply a `browser_profile_id` (pre-saved profile) — a snapshot including cookies and local storage
  - Restore session cookies from a stored directory (`browser_factory.py:826-827`)
  - Restore "banked" cookies for verified-login healing (`restore_banked_cookies`, browser_factory.py:831-832)
  - Restore sign-in cookies for OAuth recovery
  - Snapshot seed profile state at start (`_capture_seed_profile_state`, browser_factory.py:277-291) to track cookie changes for freshness guards

**Stealth and anti-bot.** The project includes anti-detection measures: it monitors CDP connections, detects challenge vendors via URL signatures (`challenge_signature.py:12-14`: captcha, turnstile, cloudflare, arkoselabs, funcaptcha, datadome, perimeterx), and runs a captcha-solving ladder (`captcha_solver.py`) that probes for DOM checkbox challenges, reCAPTCHA anchor frames, solver extensions, and token-based routes. The `DialogHandler` auto-accepts browser dialogs. CDP download interception mediates network requests and proxy auth challenges.

**Proxies.** Proxy location is a first-class concept (`ProxyLocation`, `ProxyLocationInput`). The browser factory configures Playwright's proxy settings and sets timezone info based on the proxy location. CDP interceptors pass proxy credentials for authenticated proxies.

**CAPTCHA handling.** The captcha solver (`utils/captcha_solver.py`) provides a multi-arm solving ladder: DOM checkbox click, reCAPTCHA anchor iframe click, solver extension, and reCAPTCHA token route. The vendor challenge signature (`CHALLENGE_VENDOR_SIGNATURE`) detects cloudflare Turnstile, DataDome, PerimeterX, reCAPTCHA, hCaptcha, and FUNCAPTCHA.

> **Editor's note.** Correction: RealBrowserManager is the browser manager the open-source app wires in (forge_app.py), not a cloud-only implementation. It launches local chromium-headless/headful browsers or attaches over CDP (cdp-connect).

Citations: [skyvern/webeye/browser_manager.py:23-100](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/browser_manager.py#L23-L100) · [skyvern/webeye/persistent_sessions_manager.py:54-80](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/persistent_sessions_manager.py#L54-L80) · [skyvern/webeye/browser_factory.py:815-833](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/browser_factory.py#L815-L833) · [skyvern/webeye/utils/captcha_solver.py:1-85](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/utils/captcha_solver.py#L1-L85) · [skyvern/webeye/utils/challenge_signature.py:1-28](https://github.com/Skyvern-AI/skyvern/blob/44de8cd1af2150a8cf0b28fe5cf9ec1f2f5728e1/skyvern/webeye/utils/challenge_signature.py#L1-L28)
