LLMs Technical Reviews

browser-use/workflow-use

Turns a recorded or agent-generated browser session into a step workflow that replays with text-based element finding, not an agent.

GitHub ↗★ 4.2kPythonAGPL-3.0commit 5d2d19f · 2026-08-27homepage ↗

Overview

Workflow Use is Browser Use’s attempt at “RPA 2.0”: capture a browser task once, save it as a typed list of steps, and replay it later without an agent deciding every click. A workflow is a Pydantic WorkflowDefinitionSchema (name, version, input_schema, steps) stored as JSON or YAML. Steps are mostly deterministic (navigation, click, input, select_change, key_press, scroll, go_back, extract), and an agent step type exists for parts that really need reasoning.

There are two ways to get a workflow. The recorder opens Chromium with a bundled MV3 extension that captures your clicks and typing, then an LLM cleans the raw event log into a parameterised workflow. Generation mode runs a browser-use Agent on a natural-language task, captures the element text it touched, and converts that history into a workflow. Either way, the result runs from the CLI, as an LLM “tool” that parses inputs from a prompt, or as an MCP server where each saved workflow becomes one tool.

The README calls it “very early development” and that is fair. The package is a single large Python tree (semantic_executor.py alone is about 3,200 lines), several features are stubs or switched off, and the promised “fallback to Browser Use when a step fails” is commented out at this commit. What does work is a solid idea: record visible text, not brittle CSS, and find the element again by text at run time.

Architecture

flowchart LR
  EXT["Recorder extension (rrweb + listeners)"] --> REC["RecordingService :7331"]
  REC --> BLD["BuilderService (LLM)"]
  TASK["Task prompt"] --> HEAL["HealingService"]
  HEAL --> AGENT["browser-use Agent"]
  HEAL --> WF["Workflow file (JSON/YAML)"]
  BLD --> WF
  WF --> STORE["WorkflowStorageService"]
  WF --> RUN["Workflow.run / run_with_no_ai"]
  RUN --> CTRL["WorkflowController (selectors)"]
  RUN --> SEM["SemanticWorkflowExecutor (text)"]
  RUN --> EXTR["Extraction LLM"]
  CTRL --> BR["browser-use Browser (CDP)"]
  SEM --> BR
  WF --> MCP["FastMCP server"]
  MCP --> RUN
Component Path Role
CLI workflows/cli.py Typer app: create-workflow, generate-workflow, run-workflow, run-workflow-no-ai, run-as-tool, run-workflow-csv, mcp-server, launch-gui and storage commands
Schema workflows/workflow_use/schema/views.py Pydantic step union, input schema, “must end with extract” validator
Runner workflows/workflow_use/workflow/service.py Workflow: placeholder resolution, step dispatch, context, output model
Semantic executor workflows/workflow_use/workflow/semantic_executor.py Text-to-element mapping, typed step handlers, retries, form-error detection
Element finder workflows/workflow_use/workflow/element_finder.py Tries selectorStrategies against browser-use’s selector map
Controller workflows/workflow_use/controller/ browser-use Controller subclass with selector-based click/input/navigation actions
Recorder workflows/workflow_use/recorder/, extension/ FastAPI event sink plus a WXT/React extension built on rrweb
Builder workflows/workflow_use/builder/service.py LLM turns a raw recording into a workflow (optionally with screenshots)
Healing / generation workflows/workflow_use/healing/ Agent-driven generation, selector strategies, variable detection, optional AI validation
Storage workflows/workflow_use/storage/service.py Files plus a metadata.json index
MCP workflows/workflow_use/mcp/service.py Registers each workflow file as a FastMCP tool
Backend + UI workflows/backend/, ui/ FastAPI service and a React Flow visualiser

How a request flows

Take python cli.py run-workflow my.workflow.json:

  1. Load. Workflow.load_from_file reads the file with yaml.safe_load (so JSON works too), builds the schema and a Workflow with a browser-use Browser, a WorkflowController and an ElementFinder (service.py). The CLI prompts for each declared input.
  2. Validate the shape. WorkflowDefinitionSchema rejects any workflow whose last step is not extract or extract_page_content (views.py). Every run therefore ends with at least one LLM call.
  3. Loop. Workflow.run starts the browser and walks schema.steps, sleeping wait_time between steps, checking an optional cancel_event, and filling {placeholders} from the context dict (service.py).
  4. Dispatch. _execute_step branches on the step (service.py): extraction steps go to _run_extraction_step; a click/input/key/select step with target_text but no cssSelector goes to the SemanticWorkflowExecutor; everything else goes to _run_deterministic_step. A failure raises. There is no agent fallback.
  5. Deterministic path. _run_deterministic_step strips workflow-only fields, tries ElementFinder on selectorStrategies to get a browser-use element index, then builds a one-action model and calls controller.act, followed by a networkidle wait for navigation-like actions (service.py). The controller’s click walks get_best_element_handle: the recorded selector, attribute-based fallbacks, then XPath alternatives, each with a 1-second visibility wait (utils.py, controller/service.py).
  6. Extract. _run_extraction_step takes the page as clean markdown, truncates it to 10,000 characters and sends one plain prompt to the extraction LLM (service.py).
  7. Store and return. Each result is saved into the context under the step’s output key, and an optional output_model triggers a final LLM pass that maps all extracted text into a Pydantic object. The browser is closed in a finally block.

run_with_no_ai is the same loop, but every step goes through the semantic executor and agent steps raise (service.py).

Key components

Semantic executor

Before every step, execute_step refreshes a mapping of visible text to element metadata (semantic_executor.py). SemanticExtractor.extract_semantic_mapping runs page JavaScript over interactive elements and keys each one by its text, label or a generated fallback, adding container and position context to break ties (semantic_extractor.py). Steps then resolve target_text against that map. _execute_with_verification_and_retry allows 3 retries per step, stops after 3 consecutive or 5 total failures, and re-scans the page for form validation messages after each attempt (semantic_executor.py). The StepVerifier that would check outcomes is off by default (enable_step_verification=False, L31-L57). The file also carries domain helpers for calendars, dropdowns and even flight options, which says something about the demos it was tuned on.

Recorder

RecordingService.capture_workflow starts a local FastAPI server on 127.0.0.1:7331 and launches a headed browser with the built extension loaded from extension/.output/chrome-mv3 and a persistent profile (recorder/service.py). The content script runs rrweb.record for scroll and navigation events and adds capture-phase listeners for click, input, change, keydown and focus (each handler also records an XPath, an enhanced CSS selector and semantic info for its target) (content.ts). The background script posts the converted steps to the Python server (background.ts). BuilderService.build_workflow then asks the LLM, with structured output, to turn the raw steps plus your stated goal into a cleaned, parameterised workflow (builder/service.py).

Generation mode

HealingService.generate_workflow_from_prompt wraps a browser-use Agent (vision on, memory off, up to 10 failures) in a CapturingController that snapshots the selector map before each action, so it knows the text and XPath of every element touched. The prompt asks the agent to tag targets as [ELEMENT: "text"] (healing/service.py, L800-L898). The history is converted either by an LLM call that returns a WorkflowDefinitionSchema (the default) or by a DeterministicWorkflowConverter, followed by pattern-based variable detection and optional AI validation (L288-L342). The SelectorGenerator documents a 10-rung ladder from IDs to fuzzy text, but HealingService caps it at 2 strategies per step (L41-L88, selector_generator.py).

Controller

WorkflowController subclasses browser-use’s Controller, removes 25 default actions (tabs, scrolling helpers, sheets, done and so on) and registers its own selector-driven actions (controller/service.py). Navigation is page.goto plus a fixed 2-second sleep.

Storage and MCP

WorkflowStorageService.save_workflow writes <uuid>.workflow.yaml and updates a metadata.json index; there is no database (storage/service.py). get_mcp_server globs *.workflow.{json,yaml,yml} in a folder, builds a Python Signature from each workflow’s input model, and registers a FastMCP tool that calls Workflow.run (mcp/service.py).

Extending it

  • Write workflows by hand. The schema is plain Pydantic with extra: allow; prefer target_text plus container_hint/position_hint over cssSelector, which the schema marks as legacy.
  • Mix in agent steps. An agent step runs a browser-use Agent with the neighbouring steps as context, for parts that cannot be scripted.
  • Typed results. Pass output_model to run() to get a Pydantic object back.
  • Expose as tools. run_as_tool(prompt) parses inputs with the LLM; cli.py mcp-server serves a folder of workflows over MCP.
  • Bring your own model. The library takes any browser-use BaseChatModel. The CLI does not: it hardcodes ChatBrowserUse(model='bu-latest') and ignores its --agent-model/--extraction-model/--workflow-model flags (cli.py).

Running it

  • Build the extension (cd extension && npm install && npm run build); the recorder looks for its .output/chrome-mv3 folder.
  • In workflows/: uv sync, playwright install chromium, and an API key. With the CLI as shipped that means BROWSER_USE_API_KEY, since every CLI path uses the Browser Use cloud model (cli.py).
  • --use-cloud swaps the local Chromium for a Browser Use cloud browser. launch-gui starts the FastAPI backend and the Vite UI side by side.
  • mcp-server serves SSE on 0.0.0.0:8008 by default with no authentication (cli.py), so bind it yourself before exposing it.

Strengths and caveats

  • Strength: the text-first targeting idea. Recording visible text and resolving it fresh on each step survives class and layout churn far better than replaying CSS selectors.
  • Strength: cheap replays. Navigation, clicks and typing make no model calls. Only extract steps (and optional output mapping) hit an LLM.
  • Strength: a reusable artefact. A workflow is a readable YAML file with typed inputs that also becomes an MCP tool for free.
  • Caveat: no self-healing at run time. Despite the README and the fallback_to_agent=True flag, the agent fallback is commented out; a failed step simply raises.
  • Caveat: “no AI” is not quite no AI. The schema forces a final extract step, which calls an LLM when one is configured.
  • Caveat: two overlapping execution paths. Selector-based controller steps and text-based semantic steps coexist, chosen per step by which fields are present, and behave differently on failure.
  • Caveat: early-stage code. Debug prints, dead parameters (CLI model flags, an unused faiss-cpu dependency), verification switched off by default, and no CAPTCHA, proxy or stealth handling of its own.
  • Caveat: AGPL-3.0. Embedding it in a hosted product carries copyleft obligations.

Sources: code at 5d2d19f, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (19 pages), verified Q&A.

How it answers the Browser & computer control questions

Each answer was drafted by a code-reading agent at commit 5d2d19f. Its citations were checked mechanically. Compare with the other browser & computer control →

How is the page represented to the model?

answered

The page is NOT represented to a model as a DOM tree or screenshot in this project's main deterministic path. On the semantic execution path, the SemanticExtractor runs JavaScript via CDP (page.evaluate) to query document.querySelectorAll across 50+ interactive selectors (button, input, select, textarea, [role="button"], [role="combobox"], [data-testid], etc.) and extracts bounding-box-visible elements into a dict (element_info) with fields like text_content, css_selector, hierarchical_selector (a DOM path like form > div:nth-of-type(2) > input[name="email"]), aria_label, placeholder, label_text, container_context, and position (SemanticExtractor lines 206–775). The SemanticExtractor.extract_semantic_mapping method (line 816) collects these into a mapping of visible text string → element metadata.

On the agent fallback path (which uses browser-use's agent), the model sees the full DOM via browser-use's internal EnhancedDOM representation — not built in this repo. Browser-use provides an element-index-based selector map (browser_session.get_selector_map(), used at element_finder.py:69–74) that maps integer indexes to DOM nodes with text, role, tag, and visibility flags.

CSS selectors from the recorded workflow are stored at the schema level in SelectorWorkflowSteps.cssSelector and xpath fields (schema/views.py:32–36), but are considered legacy — the preferred approach is target_text (text-based semantic targeting) plus selectorStrategies (a list of ordered fallback strategies like text_exact, role_text, xpath with priority scores, generated by SelectorGenerator). There are no screenshots or set-of-marks visual representations used; the system is entirely text/selector-based.

How are actions executed and how are elements targeted?

answered

Actions execute through two parallel systems. Path 1 — Deterministic Controller: WorkflowController (controller/service.py) registers custom Playwright-based actions via browser-use's @self.registry.action decorator: navigation (page.goto), click (locator via get_best_element_handle which tries multiple stability-ranked CSS selectors, then XPath), input (locator.fill), select_change (locator.select_option by label), key_press (locator.press), and scroll (page.evaluate JavaScript). Elements are targeted by CSS selector with 10+ fallback strategies generated by generate_stable_selectors — attribute-based ([placeholder], [aria-label], [name]), tag+class, stripping dynamic IDs/state classes, and elementTag:has-text() (controller/utils.py:51–101). XPath with stable alternatives is tried as last resort (utils.py:33–48).

Path 2 — Semantic Executor: SemanticWorkflowExecutor.execute_step (semantic_executor.py:1863) dispatches to typed step handlers (NavigationStep → page.goto + wait for input, button, form, ClickStep → multi-strategy element finder via ElementFinder.find_element_with_strategies trying text_exact, role_text, aria_label, placeholder, title, alt_text, text_fuzzy, and XPath via JavaScript document.evaluate), InputStep → element.fill, SelectChangeStep → element.select_option, KeyPressStep → CDP dispatchKeyEvent. For clicking, there is a direct text-based JavaScript click method (_click_element_by_text_direct) that searches all visible interactive elements by text content with 4-priority matching (exact → word → contains → fuzzy). File uploads and tab management are NOT implemented in this workflow system — the WorkflowController explicitly disables the browser-use actions for switch_tab, open_tab, close_tab, drag_drop, and get_sheet_contents (controller/service.py:26–52). Calendars, dropdowns, and booking widgets have specialised extraction in the JavaScript component (semantic_extractor JS, lines 259–424).

How is the agent loop / planning implemented?

answered

The primary execution model is a deterministic step loop, not an LLM agent loop. In Workflow.run() (service.py:851), the workflow iterates over self.schema.steps, resolving placeholder variables from a context dict using Python str.format(), then executing each step via _execute_step. Each step is a Pydantic discriminated union (WorkflowStep = DeterministicWorkflowStep | AgenticWorkflowStep). Deterministic steps are dispatched by type field (navigation, click, input, extract, etc.) and run through either the WorkflowController or the SemanticWorkflowExecutor — no LLM involvement. In run_with_no_ai() (service.py:1047), only semantic execution is used.

Agent steps: When a step has type: 'agent', _run_agent_step (service.py:329) constructs a browser_use Agent with the step's task as the task and a message_context showing the last 2, current, and next 2 steps for contextual awareness. An override_system_message from agent_step_system_prompt.md controls agent behavior. The agent runs with its own WorkflowStepAgentController and returns an AgentHistoryList.

Tool-call schema: For deterministic steps, the controller uses create_action_model from browser-use's registry to build a Pydantic action model, then calls controller.act(action_model, browser, ...) (service.py:235–244). No LLM tool calling is involved — the schema is purely for validation.

Stop conditions: The loop supports a cancel_event (asyncio.Event) for external cancellation (service.py:904–906). An output_model parameter with a Pydantic class drives an LLM-based conversion pass at the end (_convert_results_to_output_model) that aggregates all extracted_content from step results and uses the LLM to parse them into the target model (service.py:723–769). The schema validator requires the last step to be an extract or extract_page_content type (schema/views.py:240–256).

Memory between steps: A plain self.context: dict[str, Any] is maintained across the run. Steps can declare an output key (a string naming a context variable), and _store_output (service.py:586) stores either the JSON-parsed extracted_content from ActionResult or the last AgentHistoryList content into that key. Downstream steps reference values via Python format placeholders like {output_variable_name}.

How are failures, retries and self-healing handled?

answered

Error categories and retries: The SemanticWorkflowExecutor._execute_with_verification_and_retry method (semantic_executor.py:1901–2117) implements a retry loop with configurable max_retries=3, max_global_failures=5, and max_verification_failures=3. It tracks consecutive_failures (after 3, raises immediately) and consecutive_verification_failures (after max_verification_failures, raises). Before each retry attempt, the semantic mapping is refreshed and a 1-second delay is applied (line 1928–1930). Errors caught include Playwright Element not found, Timeout, No such element patterns, form validation errors via _detect_form_validation_errors (which inspects the page for patterns like "required", "invalid", "must"), and general execution exceptions. After all retries exhaust, failure counters increment and a comprehensive ErrorReport is generated.

Self-healing — multi-strategy element finding: The ElementFinder (element_finder.py:20) tries strategies in priority order: semantic strategies (text_exact, role_text, aria_label, placeholder, title, alt_text, text_fuzzy) against browser-use's selector map (a dict of index → DOM node), then XPath via JavaScript document.evaluate as fallback. The SelectorGenerator (healing/selector_generator.py:47) generates these strategies at workflow-creation time, ordered by reliability: ID selectors → data attributes → name attributes → exact text → ARIA labels → role+text → placeholder → class+text → fuzzy text → XPath. XPath alternatives are further optimized by XPathOptimizer.

Caching and verification: The _execute_with_verification_and_retry method uses an optional StepVerifier (step_verifier.py:66) with deterministic checks (DOM inspection, URL changes) and optional AI-assisted verification. On success, all failure counters reset. The ErrorReporter (error_reporter.py:79) logs each failure with strategy attempt records, page screenshots (saved to .workflow_screenshots/), and generates root-cause analysis and suggestions (e.g., if all semantic strategies failed, XPath might work). Workflow-level retries that happen in _execute_step (service.py) are disabled — the fallback-to-agent code path is commented out. However, the CLI's run-workflow and run-stored-workflow commands do not wrap the entire run in retry logic; the retries are per-step only.

Editor's note. Correction: StepVerifier is off by default (SemanticWorkflowExecutor enable_step_verification=False), so post-step verification does not run unless enabled; HealingService also caps SelectorGenerator output at 2 strategies per step.

Which models are supported and how are they called?

answered

Models are abstracted entirely through browser-use's BaseChatModel interface. The project does not import any model directly — it receives pre-configured BaseChatModel instances from the caller. The CLI demonstrates two concrete providers: ChatOpenAI (OpenAI models like gpt-4.1-mini, gpt-4o, gpt-4.1) and ChatAnthropic (Claude models like claude-3-5-sonnet-20241022), both from browser_use.llm. The default CLI model factory is ChatBrowserUse(model='bu-latest'), the Browser-Use Cloud model (cli.py:37–44). Model selection in the CLI is per-task: --agent-model, --extraction-model, and --workflow-model flags (cli.py:2306–2308) defaulting to gpt-4.1-mini, gpt-4.1-mini, and gpt-4.1 respectively.

Vision requirement: Only the agent generation path uses use_vision=True (healing/service.py:826). The semantic extraction and deterministic execution paths do NOT send screenshots to the LLM — they use text-only.

Where models are used: 1) Workflow generation (HealingService.generate_workflow_from_prompt) uses two LLMs: an agent_llm for browser agent decision-making and an extraction_llm for page content extraction (healing/service.py:432). 2) The _run_extraction_step method (service.py:354) calls the LLM directly with page markdown text truncated to 10,000 characters, using ainvoke([UserMessage(...)]). 3) The _run_agent_step creates a full browser-use Agent for complex tasks. 4) The run_as_tool method (service.py:1013) uses LLM to parse natural language prompts into structured workflow inputs via Pydantic output models. 5) Optional AI validation (WorkflowValidator) can review generated workflows.

Structured output: The LLM is called with Pydantic schemas via browser-use's output_format parameter for workflow definition generation (WorkflowDefinitionSchema) and input parsing. The base model interface provides ainvoke(messages, output_format=ModelClass) which forces structured JSON responses from the LLM.

Editor's note. Correction: the CLI's --agent-model, --extraction-model and --workflow-model flags are printed but ignored; generate-workflow and every other CLI path hardcode ChatBrowserUse(model='bu-latest'). Custom models only work through the Python API.

How are browser sessions, profiles, auth and anti-bot handled?

answered

Local vs cloud: The Browser object from browser-use supports use_cloud boolean parameter (service.py:76). Cloud mode sends browser sessions to Browser-Use Cloud (Browser(use_cloud=True)), which runs remote browser instances behind an API. Local mode uses Playwright Chromium installed via playwright install chromium. CLI commands consistently expose --use-cloud flags (cli.py:1137, 1192, 1344, 1777, 2311).

Persistent profiles and cookies: The RecordingService (recorder/service.py:104–140) creates a persistent BrowserProfile with headless=False, a user_data_dir for Chrome profile persistence, and extension-loading arguments. The profile has keep_alive=True so the browser stays running between steps until explicitly stopped. In the Workflow runner, self.browser.browser_profile.keep_alive = True is set at init (service.py:79) and set to False only when close_browser_at_end=True (service.py:940–943). Each run call opens the browser and closes it in a finally block.

Stealth and proxies: Neither stealth mode nor proxy configuration is implemented in this project. The Browser and BrowserProfile classes are imported from browser-use, which may support these, but the workflow_use code never passes stealth or proxy arguments. The extension's Chrome launch args include --no-default-browser-check and --no-first-run for cleaner startup (recorder/service.py:120–125).

CAPTCHA: Not handled. There is no CAPTCHA detection, solving, or bypass mechanism. The README notes the project is "in very early development" and not production-ready.

MCP integration: The WorkflowController disables 22 of browser-use's default actions (including switch_tab, open_tab, close_tab) for deterministic execution (controller/service.py:26–52), limiting the session to the workflow's predefined step sequence. The MCP service (mcp/service.py) exposes recorded workflow files as callable MCP tools via FastMCP, with dynamically-typed signatures derived from each workflow's input schema.