Skyvern-AI/skyvern
Self-hosted LLM browser agent that picks actions from a scraped DOM tree plus screenshots, runs them in Playwright, and chains workflows.
Overview
Skyvern is a self-hostable server that runs browser tasks described in natural language. You send a prompt and a start URL to its REST API, and a background worker drives a Playwright browser until an LLM decides the goal is met. Around that core sit a workflow engine (around 30 block types: navigation, extraction, login, loops, conditionals, code, HTTP, email, PDF), a React UI with live browser streaming, a Python SDK, a CLI and an MCP server with over a hundred tools.
The page is shown to the model as a DOM-derived element tree, not as raw pixels. An injected script tags every interactable element with a unique_id attribute and serialises the tree as compact HTML. That HTML goes to the LLM together with a few scrolled screenshots. The model answers with a JSON list of actions that name elements by id, and Python handlers execute them with Playwright locators. Despite the “computer vision” in the marketing, screenshots are context only. The set-of-marks box overlay is switched off and marked deprecated in the scraper.
At this commit, “Skyvern” means four different agents behind one API, chosen by the engine field: skyvern-1.0 (the default per-step engine), skyvern-2.0 (a planner that writes and runs workflow blocks), skyvern-3.0 (a single persistent tool-calling conversation), and third-party computer-use models (OpenAI CUA, Anthropic CUA, UI-TARS, Yutori Navigator). The repository is also the open core of a hosted product. Many behaviours, such as CAPTCHA solving, cross-run caches and A/B engine routing, are hooks that are no-ops here and are filled in by the cloud.
Architecture
flowchart LR
C["Client: SDK / CLI / MCP / UI"] --> API["FastAPI routes"]
API --> EX["BackgroundTaskExecutor"]
EX --> AG["ForgeAgent.execute_step"]
API --> T2["task_v2_service planner"]
T2 --> WF["Workflow blocks"]
WF --> AG
AG --> V3["Task V3 tool loop"]
AG --> SC["Scraper + domUtils.js"]
AG --> LLM["LLMAPIHandlerFactory (LiteLLM)"]
AG --> AH["ActionHandler"]
SC --> BM["RealBrowserManager"]
AH --> BM
V3 --> BM
BM --> PW["Playwright Chromium or CDP"]
AG --> DB["PostgreSQL"]
| Component | Path | Role |
|---|---|---|
| API routes | skyvern/forge/sdk/routes/agent_protocol.py |
/v1/run/tasks, workflow runs, engine dispatch |
| Executor | skyvern/forge/sdk/executor/background_task_executor.py |
Creates the first step and schedules execute_step as a FastAPI background task |
| Step engine | skyvern/forge/agent.py |
ForgeAgent: scrape, prompt, LLM, parse, act, verify, recurse |
| Planner (2.0) | skyvern/services/task_v2_service.py |
Iterative “plan next block” loop that builds and runs a workflow |
| Task V3 | skyvern/forge/taskv3/ |
Persistent tool-use conversation (engine.py, loop.py, tools.py) |
| Scraper | skyvern/webeye/scraper/ |
scraper.py + 4,000-line domUtils.js: element tree, frames, split screenshots |
| Actions | skyvern/webeye/actions/ |
Action models, parse_actions.py, the 16k-line handler.py |
| Browser | skyvern/webeye/real_browser_manager.py, browser_factory.py |
Browser acquisition, persistent sessions, profiles, proxies |
| LLM layer | skyvern/forge/sdk/api/llm/ |
Config registry, LiteLLM handlers and routers, LLMCaller for chat-history engines |
| Workflows | skyvern/forge/sdk/workflow/ |
Block models, parameters, service.py |
| Script generation | skyvern/core/script_generations/ |
Turns recorded runs into Playwright-style Python (libcst) |
| Client surfaces | skyvern/library/, skyvern/cli/, skyvern-frontend/ |
SDK, Typer CLI + FastMCP server, React UI |
How a request flows
Take POST /v1/run/tasks with engine: skyvern-1.0:
- Accept.
run_taskvalidates the webhook URL and rate limit, rejects options the v1 engine cannot honour, and builds a legacyTaskRequest. If no URL was given, an LLM call (generate_task) derives one first (agent_protocol.py). - Persist and dispatch.
task_v1_service.run_taskcreates the task and atask_runsrow, asks theAGENT_FUNCTIONhook to resolve the engine (OSS returns the requested one), and calls the executor (task_v1_service.py).BackgroundTaskExecutor.execute_taskcreates step 0, marks the task running, maps the run type back to an engine, and schedulesapp.agent.execute_step(background_task_executor.py). - Step entry.
execute_stepsets context, turns off completion verification for CUA engines, and either hands the whole task to Task V3 or callsagent_step(agent.py, L3950-L3970). - Scrape.
build_and_record_step_promptwalks a ladder of scrape strategies (normal, stop-loading, reload) through_scrape_with_type(agent.py, L7030-L7073).scrape_web_unsafebuilds the element tree for the main frame and every child frame, trims it, counts tokens, and takes up toMAX_NUM_SCREENSHOTSscrolled screenshots, or just one if the tree is too large (scraper.py, L908-L956). - Decide.
_generate_step_actionsbranches by engine. For skyvern-1.0 it calls the org-aware LLM handler with theextract-actionprompt and the screenshots, thenparse_actionsmaps each JSON entry, byid/element_id, onto the scraped element (agent.py, parse_actions.py). - Act.
_execute_step_actionsruns each action throughActionHandler.handle_action, which dispatches to a handler registered perActionType(agent.py, handler.py). A click resolves the element by id, checks the desired toggle state and the disabled state, scrolls it into view and clicks with Playwright. Coordinate clicks go topage.mouse(handler.py). - Verify and loop. A failed step goes to
handle_failed_step, which retries up toMAX_RETRIES_PER_STEP(5 unless the org overrides it). A successful one goes tohandle_completed_step, which can runcomplete_verify(a second LLM call over a fresh scrape) and creates the next step.execute_stepthen calls itself recursively (agent.py, L10181-L10200). - Finish.
clean_up_taskuploads video, HAR and artifacts, closes the browser unless it belongs to a persistent session, and fires the webhook.
Key components
Element tree perception
domUtils.js runs in each frame, decides what is interactable (inputs, hover styles, ARIA roles), stamps unique_id, and returns elements plus a nested tree. Python side, build_element_dict turns each id into a CSS selector [unique_id='…']. json_to_html renders the trimmed tree for the prompt. Placeholder nodes stand in for frames the filter skipped, so the model sees a cross-origin CAPTCHA frame but can never target it (scraper.py). When an ARIA popup is open, scroll capture can be suppressed so the dropdown survives into the next action.
Four engines
The RunEngine enum lists them in one place (run_enums.py):
- skyvern-1.0 is the step engine above: one scrape, one LLM call and a batch of actions per step.
- skyvern-2.0 is a planner. On each iteration
run_task_v2_helperscrapes, asks thetask_v2prompt foruser_goal_achievedand atask_type(navigate, extract, loop, compute or goto_url), generates a matching workflow block and runs it withblock.execute_safe(task_v2_service.py, L1262-L1272). The result is a reusable workflow, not only an answer. - skyvern-3.0 (Task V3) drops the per-step prompt. A single conversation calls tools such as
observe,get_html,look,click,type,select_option,navigateandfile_uploaduntil it callsfinish(tools.py). Backstops are 80 turns, 300 tool calls, 1,800 seconds and 1.5M tokens (engine.py). The module docstring says why: compactobservesnapshots matched raw DOM at equal success with a tighter tail (engine.py). - CUA engines receive one screenshot without scrolling and return coordinate actions through
_generate_cua_actions,_generate_anthropic_actions,_generate_ui_tars_actionsor_generate_yutori_navigator_actions.
Action handlers
ActionHandler keeps three registries (setup, main, teardown) keyed by ActionType (handler.py). The 16,000-line handler.py is where most site-specific hardening lives: custom dropdowns, auto-complete inputs, date pickers, multi-field TOTP, downloads and file choosers. In the open-source build, SOLVE_CAPTCHA just logs “Please solve the captcha” and sleeps 30 seconds (handler.py). The solving ladder in captcha_solver.py calls AGENT_FUNCTION hooks that return False outside the cloud (captcha_solver.py).
LLM layer
LLMAPIHandlerFactory.get_llm_api_handler looks up a key in LLMConfigRegistry and returns a LiteLLM-backed handler or a LiteLLM Router for fallback groups (api_handler_factory.py). Providers are switched on with ENABLE_* settings: OpenAI, Anthropic, Azure, Bedrock, Gemini, Vertex, Groq, xAI, OpenRouter, Ollama, Novita, Moonshot, Volcengine, Yutori and any OpenAI-compatible endpoint. Screenshots are silently dropped for models without supports_vision, except for prompts that require vision, which raise an error (api_handler_factory.py). LLM_KEY and SECONDARY_LLM_KEY split the main reasoning model from cheaper helper calls.
Browser management
The OSS app wires RealBrowserManager as its browser manager (forge_app.py). BrowserContextFactory registers three launch modes: chromium-headless, chromium-headful and cdp-connect for attaching to an existing Chrome (browser_factory.py). Persistent browser sessions, saved browser profiles, cookie restore and per-run proxy location are first-class. A per-run engine-selection seam lets a future image ship drivers other than stock Playwright (browser_engine.py).
Caches and code generation
Extraction results are cached in process, per workflow run, keyed by a hash of the element tree, page text, URL, goal, schema and model (extraction_cache.py). The cross-run Redis tier sits behind cloud hooks. Separately, run_with: "code" replays a workflow as a generated Python script (generate_workflow_script_python_code, built with libcst), with AI fallbacks per action. Repeat runs can then skip most LLM calls (runs.py).
Extending it
- Workflows. Compose YAML/JSON workflows from the block types in
BlockType, with parameters, credentials, loops and conditionals. This is the main customisation surface, and the UI has a visual editor for it. - SDK.
skyvern.librarywraps a Playwright page withpage.act(...),ai_click,ai_extract,ai_validateand anagent.run_task/run_workflow, so AI steps can be mixed with normal Playwright code (skyvern_browser_page.py). - MCP.
skyvern/cli/mcp_toolsexposes sessions, tabs, network, credentials, workflows, schedules and scripts as FastMCP tools. It also has deterministicskyvern_observe/skyvern_executeprimitives next to the LLM-backedskyvern_actandskyvern_run_task(mcp_tools/init.py). - Models. Add a custom LLM through
OPENAI_COMPATIBLE_*settings or a per-org custom LLM record. Change prompts inskyvern/forge/prompts/skyvern/*.j2. - Platform hooks.
AgentFunctionhas about 180 overridable methods, from engine routing to CAPTCHA solving and audit logging. The OSS versions are mostly pass-throughs (agent_functions.py). Subclass it to change behaviour without forking the agent.
Running it
- pip.
pip install skyvern, thenskyvern quickstart(orskyvern init+skyvern run all). It supports Python 3.11 to 3.14 and needs Playwright’s Chromium and PostgreSQL (the SDK’s `Skyvern.local()` can use an in-memory database instead).skyvern run serverandskyvern run uistart the pieces separately. - Docker Compose.
docker-compose.ymlrunspostgres:14, theskyvernAPI image (API plus a VNC WebSocket for live view) andskyvern-ui, with artifacts, videos, HAR files and logs on mounted volumes (docker-compose.yml). Provider keys andLLM_KEYgo in.env. - Kubernetes. Manifests live in
kubernetes-deployment/andk8s/. - Licence. AGPL-3.0, which matters if you embed the server in a hosted product.
Strengths and caveats
- Strength: grounded actions. The model selects element ids from a real DOM tree that includes iframes and shadow roots. Execution uses Playwright locators, not pixel guesses, so runs are inspectable and replayable.
- Strength: more than a loop. Workflows, credentials (Bitwarden, 1Password), TOTP, downloads, scheduled runs, video and HAR artifacts, and code generation for cheap reruns cover what a production form-filling job needs.
- Strength: engine choice. The same task API can run a cheap step engine, a planner, a tool-loop agent or a vendor CUA model, which makes side-by-side comparisons easy.
- Caveat: size and churn.
agent.pyis 11,600 lines,handler.py16,400 andtaskv3/tools.py20,500. Behaviour often hangs on experiment flags (DISABLE_TASK_V3,FORCE_TASK_V1) whose real values only exist in the cloud. Reading the code to predict behaviour is hard. - Caveat: open core. CAPTCHA solving, cross-run extraction caching, engine A/B routing and some stealth or browser engines are hooks that do nothing in the OSS build. Expect a self-hosted instance to stall on CAPTCHAs.
- Caveat: heavy per-step prompts. In skyvern-1.0, every step re-scrapes and sends the HTML tree plus up to ten screenshots, and completion checks add calls. Task V3 exists largely to cut that cost.
- Caveat: infrastructure. The normal deployment is a server process plus PostgreSQL. `Skyvern.local()` can run the same stack in-process, optionally on an in-memory database, but it is still the whole platform, not a thin library.
Sources: code at 44de8cd, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (22 pages), verified Q&A.
How it answers the Browser & computer control questions
Each answer was drafted by a code-reading agent at commit 44de8cd. Its citations were checked mechanically. Compare with the other browser & computer control →
How is the page represented to the model?
answeredThe page is represented to the model as a combination of a scraped interactive element tree and full-page screenshots (base64 PNG, potentially multiple scroll-segments). Both are sent to the LLM on every step.
DOM serialisation + accessibility tree. The scraper (skyvern/webeye/scraper/scraper.py:908-956, get_interactable_element_tree) runs a JavaScript DOM inspector that walks the page's accessible, visible, interactable elements. It annotates each with attributes (type, role, aria-*, value, placeholder, disabled, readonly, text, href, etc.), bounding rects, and a unique_id Skyvern-injected attribute (SKYVERN_ID_ATTR). The tree includes all child frames recursively (add_frame_interactable_elements, scraper.py:827–904). The resulting list is rendered as HTML (via json_to_html in scraped_page.py:63-132) and sent to the LLM — typically the "trimmed" tree where non-essential attributes and non-interactable nodes are pruned (trim_element_tree, scraper.py:1247-1250). Three tree variants exist: the full tree, the economy tree (strips SVG branches), and the lean tree (optionally compresses long hrefs/srcs).
Size limits and pruning. The tree is token-counted after rendering (scraper.py:623): if it exceeds DEFAULT_MAX_TOKENS (a tunable ceiling), the screenshot count is capped to 1. Individual elements have their attributes filtered to a reserved set (RESERVED_ATTRIBUTES, scraper.py:113–143); enriched trees include additional accessibility attributes. Hashed URLs >150 chars are replaced with a SHA-256 placeholder (json_to_html, scraped_page.py:82-88).
Screenshots. Screenshots are taken after the tree build via SkyvernFrame.take_split_screenshots (scraper.py:679-686): the page is scrolled in viewport-sized increments, capturing each segment. Screenshot count is bounded by MAX_NUM_SCREENSHOTS (default 10). When a transient popup is detected, scrolling is suppressed so the overlay survives into the next action (scraper.py:632-641). The agent can also suppress screenshots entirely via the take_screenshots=False flag.
Set-of-marks/element index. Skyvern injects a unique_id attribute into each interactable element's DOM node at scrape time. The model refers to elements by this ID (the "element_id" in actions like click/input), NOT by coordinates or CSS selectors in the prompt. The scraped page maintains an id_to_css_dict mapping these IDs to CSS selectors for Playwright resolution (scraper.py:266-282). Text content and extracted text (from get_frame_text, scraper.py:703) are also included.
How are actions executed and how are elements targeted?
answeredActions are executed through Playwright (async API) against the live browser page. The codebase defines ~30 action types (skyvern/webeye/actions/action_types.py:4-37), including click, input_text, select_option, upload_file, scroll, keypress, move, drag, goto_url, go_back, go_forward, close_page, new_tab, switch_tab, solve_captcha, wait, hover, execute_js, reload_page, and the decisive actions complete/terminate.
Element targeting uses three strategies, ordered by preference:
- Skyvern element_id — The model names elements by their injected
unique_idattribute. The action handler resolves this to a CSS selector (skyvern/webeye/scraper/scraper.py:275,id_to_css_dict[element_id] = f"[{SKYVERN_ID_ATTR}='{element_id}']") and uses Playwright'sframe.locator()to find it. This is the primary targeting method. - Coordinate-based —
clickactions can includex/ypixel coordinates (e.g. for image maps). The handler usespage.mouse.click(x, y, button)directly (handler.py:7473-7480). - XPath fallback — Each stored action can carry an
xpathfield derived from skyelement data (actions.py:291-298).
Click execution is the most complex handler (handler.py:7428-7524+): it resolves the element by ID, checks disabled state (retargeting to a child if the parent is inert), validates desired state for toggle controls, scrolls into view, optionally moves the cursor via events, then calls locator.click() or page.mouse.click() with support for left/right button, double-click, and triple-click (via repeat).
Other action mechanics: INPUT_TEXT fills by clearing then typing via Playwright's fill() or type(), with TOTP code injection support. SELECT_OPTION uses locator.select_option() by label/value/index. UPLOAD_FILE uses Playwright's file chooser. SCROLL uses page.evaluate('window.scrollTo(...)'). CLOSE_PAGE closes the current Playwright page; NEW_TAB/SWITCH_TAB manage page objects. EXECUTE_JS runs arbitrary JavaScript. SOLVE_CAPTCHA routes through the AGENT_FUNCTION seam (the OSS base returns False; cloud deployments solve via vendor handlers). The Task V3 loop (forge/taskv3/tools.py) offers raw-browser tools — click, type, select_option, press_key, file_upload, navigate, scroll, wait, solve_captcha, reload_page, and observe — each as a ToolSpec the model calls via tool_calls.
How is the agent loop / planning implemented?
answeredSkyvern has three major agent loop implementations, selectable per task:
Legacy step engine (V1/V2). Located in skyvern/forge/agent.py. The loop (execute_step at line 3653) is: scrape the page → build prompt with element tree + screenshots → call LLM → parse actions → execute each action via the action handler → verify goal completion → repeat or terminate. Each iteration is one "step" with scraped page state. The LLM response is parsed into a list of Action objects (via parse_actions in parse_actions.py) which are then dispatched to ActionHandler. After action execution, the loop runs complete_verify — a separate LLM call that checks whether the user's goal was achieved. Stop conditions: COMPLETE or TERMINATE action emitted by the model, or budget exhaustion (max steps, max retries). The V2 variant includes message history across steps via llm_messages_builder_with_history.
Task V3 engine (persistent-conversation tool-loop). Located in skyvern/forge/taskv3/engine.py and loop.py. This is fundamentally different: it runs one persistent LLM conversation where the model calls tools (not a fixed prompt + actions) in a tool-use loop. Perception is a tool the model calls (observe), not automatically injected. Tools are defined as ToolSpec objects in tools.py (click, type, navigate, scroll, etc.). The loop calls run_agent_tool_loop (loop.py:1-13), which threads tool results back as tool-role messages. It is capped by DEFAULT_MAX_TURNS (80), DEFAULT_MAX_TOOL_CALLS (300), a wall-clock deadline (1800s), and a token budget (1.5M). Goal checking runs via GoalJudge (goal_check.py) with configurable verification. An unlisted_reask mechanism lets the model ask the user about unprompted decisions.
Third-party engines. CUA (OpenAI/Anthropic computer-use) and UI-TARS are supported as alternative engines. CUA tasks pass an existing OpenAIResponse object into execution (agent.py:3669). Yutori Navigator is a separate LLM caller for navigation-specific applications.
Across all loops, the SkyvernContext object carries step state, error codes, secrets, TOTP state, and multi-field TOTP tracking (multi-fill OTP splitting). The step/action history is persisted to the database for debugging and retry.
How are failures, retries and self-healing handled?
answeredReliability is handled at multiple layers with a rich exception hierarchy and explicit retry mechanics.
Error classes. Skyvern defines dozens of exception types (skyvern/exceptions.py), including MissingElement, MultipleElementsFound, InteractWithDisabledElement, FailToClick, ScrapingFailed, ActionExecutionTimeout, CaptchaSolveError, LLMProviderError, and SkyvernPageAnalysisTimeout. Each action handler wraps its operation and raises typed exceptions, which are caught by the step loop.
Retries within a step. The action handler returns ActionResult objects (responses.py), which have a success flag and optional exception_type. The step engine evaluates these: on failure, it retries the step (up to max_retries_per_step, default 3). The step has an explicit retry_index field (set in agent.py:3676). The entire step retry loop re-scrapes the page and re-calls the LLM. Scraping itself has retry logic (scrape_website, scraper.py:285-410): up to MAX_SCRAPING_RETRIES (currently 0 in prod) with an intermediate 3s wait and page reload.
Replanning. On step failure, the agent can escalate to a "recovery" step or, in workflow context, the WorkflowService triggers replanning with different credential parameters (e.g. a fresh login). The failure_classifier.py module categorises errors into FailureCategory values (LLM_ERROR, CAPTCHA_SOLVE_FAILURE, etc.), which feed into the retry/recovery decision.
Caching. The extraction cache (skyvern/forge/sdk/cache/) caches LLM extraction responses for up to 4 weeks (cache expiry = timedelta(weeks=4), base.py:6), keyed by page content hash, so identical pages reuse prior extraction results. It supports both local in-memory and Redis backends. The extraction shadow maintains a parallel write-through cache for consistency verification.
Self-heal tracking (services/self_heal_reliability_service.py) tracks workflow reliability metrics: it collects recent workflow run outcomes and "heal episodes" (instances where a recovery action was taken), computing a WorkflowReliability score per workflow over the RELIABILITY_WINDOW. This is a reporting/observability system rather than an automatic healing mechanism.
Timeouts. Multiple timeout layers exist: BROWSER_LOADING_TIMEOUT_MS for page loads, SkyvernPageAnalysisTimeout for page analysis, ActionExecutionTimeout for individual actions, task-level timeouts (TaskAlreadyTimeout), and workflow-level timeout. The Task V3 loop adds wall-clock, token, and tool-call budget caps (engine.py:91-94).
Which models are supported and how are they called?
answeredSkyvern supports all major LLM providers through LiteLLM, with a plugin-style configuration registry (skyvern/forge/sdk/api/llm/config_registry.py).
Supported providers and models are registered via _register_builtin_config under gate-keeping env variable flags (e.g. ENABLE_OPENAI, ENABLE_ANTHROPIC, ENABLE_BEDROCK, ENABLE_GEMINI, ENABLE_AZURE, ENABLE_YUTORI). The built-in keys include:
- OpenAI: GPT-4o, GPT-4.1, GPT-4.5, GPT-5 family (GPT-5 mini/nano, GPT-5.1/5.2/5.4/5.5/5.6), GPT-6 family (Astra/Sol/Luna), o3-mini, o4-mini (through standard, flex, and nano tiers)
- Anthropic: Claude 4 Opus/Sonnet, 4.5 Sonnet/Haiku/Opus, 4.6 Opus/Sonnet, 4.7 Opus, 4.8 Opus, 5 Fable, 5.1 Fable, 5 Opus, 5.5 Opus, 5.5 Sonnet
- Azure OpenAI: Deployments for all the above models via Azure endpoints
- AWS Bedrock: Amazon Nova Pro/Lite, Anthropic Claude inference profiles (4 through 5.5)
- Gemini (direct): Flash 2.0, Flash 2.0 Lite, Pro, 2.5 Pro, 2.5 Pro Preview, 2.5 Pro Exp, 3.0 models
- Yutori Navigator: Custom navigation-focused model
- OpenRouter: Dynamic model resolution with supports_vision
- Custom LLM:
OPENAI_COMPATIBLE_API_BASEfor any OpenAI-compatible endpoint (Ollama, vLLM, etc.) and fully custom providers viaCUSTOM_LLM_KEY - xAI: grok-4.5 (registered with its own cost table)
Vision requirement. The LLMConfig has a supports_vision boolean. Screenshots are stripped from messages sent to non-vision models (api_handler_factory.py:658-665). The extract-actions prompt is the primary vision user.
Structured output / tool calling. All models are called via LiteLLM's acompletion with OpenAI-format messages and tools. The router handler (api_handler_factory.py:1867-1920+) builds a LiteLLM Router with configurable main/fallback model groups and retry policies. Thinking budgets (budget_tokens for Claude, reasoning_effort for GPT-5, thinking_level for Gemini 3) are applied per-prompt-name. Prompt caching is injected for OpenAI and Vertex models.
Small/specialised models. Secondary models (configured via SECONDARY_LLM_KEY) can be used for lightweight operations like goal verification. The UItarsLLMCaller and YutoriNavigatorLLMCaller provide specialised inference paths. Custom LLMs can be registered per-organization with their own API base, key, and model family.
How are browser sessions, profiles, auth and anti-bot handled?
answeredSkyvern manages browser sessions through a layered architecture built on Playwright, supporting both local ephemeral browsers and remote persistent sessions.
Local vs remote/cloud browsers. The BrowserManager protocol (skyvern/webeye/browser_manager.py:23-100) defines get_or_create_for_task/get_or_create_for_workflow_run to acquire a browser. The default implementation creates Playwright browser instances locally. Cloud deployments use RealBrowserManager (real_browser_manager.py) backed by CDP-connected remote browsers. The BrowserAcquisitionSample module records the acquisition mode (create vs attach vs reuse). The engine selection determines whether to use local Playwright, a remote CDP endpoint, or a cloud persistent session.
Persistent browser sessions. The PersistentSessionsManager protocol (persistent_sessions_manager.py:54-80) manages a pool of pre-warmed, reusable browser sessions. These are long-lived browser instances (launched headlessly on remote machines) that can be leased by workflows and tasks. The protocol supports session watching, a reaper for idle/expired sessions, eviction/reconnection, and startup timeout configuration. Sessions are tracked in the database through PersistentBrowserSession records.
Profiles and cookies. The browser factory (browser_factory.py) orchestrates profile-based browsing. At startup, it can:
- Apply a
browser_profile_id(pre-saved profile) — a snapshot including cookies and local storage - Restore session cookies from a stored directory (
browser_factory.py:826-827) - Restore "banked" cookies for verified-login healing (
restore_banked_cookies, browser_factory.py:831-832) - Restore sign-in cookies for OAuth recovery
- Snapshot seed profile state at start (
_capture_seed_profile_state, browser_factory.py:277-291) to track cookie changes for freshness guards
Stealth and anti-bot. The project includes anti-detection measures: it monitors CDP connections, detects challenge vendors via URL signatures (challenge_signature.py:12-14: captcha, turnstile, cloudflare, arkoselabs, funcaptcha, datadome, perimeterx), and runs a captcha-solving ladder (captcha_solver.py) that probes for DOM checkbox challenges, reCAPTCHA anchor frames, solver extensions, and token-based routes. The DialogHandler auto-accepts browser dialogs. CDP download interception mediates network requests and proxy auth challenges.
Proxies. Proxy location is a first-class concept (ProxyLocation, ProxyLocationInput). The browser factory configures Playwright's proxy settings and sets timezone info based on the proxy location. CDP interceptors pass proxy credentials for authenticated proxies.
CAPTCHA handling. The captcha solver (utils/captcha_solver.py) provides a multi-arm solving ladder: DOM checkbox click, reCAPTCHA anchor iframe click, solver extension, and reCAPTCHA token route. The vendor challenge signature (CHALLENGE_VENDOR_SIGNATURE) detects cloudflare Turnstile, DataDome, PerimeterX, reCAPTCHA, hCaptcha, and FUNCAPTCHA.