LLMs Technical Reviews

How are failures, retries and self-healing handled?

Error classes caught; retries; replanning; caching of successful actions or workflows; timeouts.

Verdict

Both projects bound each operation with timeouts. They differ in where recovery happens.

browser-use recovers inside the loop. Steps time out after 180 s and actions after 180 s by default. LLM calls get a per-model timeout. An empty action list gets one clarification retry. Rate-limit, auth, 5xx and truncation errors can switch the run once to a fallback_llm. Dropped CDP WebSockets get three reconnect attempts. An ActionLoopDetector watches repeated actions and unchanged pages and adds escalating hints to the prompt. After max_failures consecutive failed steps, one final recovery call is made. Nothing is cached between runs. rerun_history can replay a saved run and re-match elements.

Stagehand recovers per call. With selfHeal, a failed action triggers a fresh snapshot and a new inference. TimeoutError is always passed to the caller. Before each snapshot it waits for DOM and network quiet. Its main reliability tool is a server-side action cache keyed on the accessibility tree. A hit replays the stored actions without the LLM and falls back to inference if replay fails. The cache needs a Browserbase API key and session.

Use browser-use for long unattended runs that must keep going. Use Stagehand’s cache when the same flows repeat often on Browserbase.

More projects in this category are being researched.

Per-project answers

browser-use/browser-use

answered

Error classification. Agent._handle_step_error() (browser_use/agent/service.py:1258-1314) categorizes exceptions into: InterruptedError (user pause/stop — silent), ConnectionError/WebSocket-closed errors (browser disconnect — triggers reconnection), and all other exceptions (counted as consecutive failures). _is_connection_like_error() checks error strings for patterns like 'websocket connection closed', 'browser has been closed'. Browser-closed errors set self.state.stopped = True and exit the loop.

Retries for empty actions. _get_model_output_with_retry() (browser_use/agent/service.py:1670-1704) wraps the LLM call: if the model returns an empty or all-empty action list, it sends a clarification message and retries once. If the retry also returns empty, it inserts a synthetic done(success=false) action.

Consecutive failure counting. Single-action steps that error increment consecutive_failures (capped at max_failures, default 5). Multi-action steps with errors are handled by loop detection instead (not counted as failures). When consecutive_failures >= max_failures, a final recovery response is attempted (if final_response_after_failure=True), then the agent stops.

Step-level timeout. _execute_step() (browser_use/agent/service.py:2438-2498) wraps the entire step() in asyncio.wait_for(timeout=self.settings.step_timeout) (default 180s). On timeout, it increments failures and continues. LLM calls have their own llm_timeout (auto-detected per model: 30s for Groq, 90s for Claude/DeepSeek/o3, 75s default).

Action timeout. Tools.act() applies asyncio.wait_for(timeout=action_timeout) (env BROWSER_USE_ACTION_TIMEOUT_S, default 180s) to prevent hung CDP calls. The timeout is above the 120s extract action LLM cap (browser_use/tools/service.py:99-110).

Fallback LLM. If the primary LLM fails with a rate limit (429), auth error (401/402), server error (500-504), or truncated output, and a fallback_llm is configured, the agent switches providers mid-run (browser_use/agent/service.py:1983-2025). This is a one-way switch; once on fallback, the agent stays there.

Connection recovery. BrowserSession has WebSocket reconnection logic (browser_use/browser/session.py:585-593) with 3 attempts, delays (1+2+4s), 54s total wait timeout. When _handle_step_error detects a connection error during reconnection, it waits for the reconnect event. If reconnection succeeds, the step is retried; otherwise, it's treated as a terminal browser-closed error.

Loop detection. ActionLoopDetector (browser_use/agent/views.py:157-248) tracks the last 20 action hashes for repetition and page fingerprints for stagnation. At thresholds (5/8/12 repeated actions, 5 stagnant pages), escalating nudges are injected into the LLM context. Wait/done/go_back actions are exempt from tracking. This is advisory only — never blocks actions.

CAPTCHA handling. CaptchaWatchdog (browser_use/browser/watchdogs/captcha_watchdog.py) listens for CDP captchaSolverStarted/Finished events from the browser proxy. The agent's step() checks wait_if_captcha_solving() at the start and blocks until resolution or timeout. The result (success/failed/timeout) is injected into the action result for the LLM. Cloud browsers include proxy rotation and stealth fingerprinting to reduce CAPTCHA incidence.

Message compaction. Old conversation history is summarized by an LLM every 25 steps when the prompt exceeds ~10K tokens (MessageCompactionSettings). This prevents context windows from overflowing while preserving key information.

browserbase/stagehand

answered

Self-healing: When a deterministic action fails (any non-TimeoutError exception), and selfHeal: true is set at init, actService.ts re-captures the full page snapshot and re-asks the LLM for a new action decision (selfHealAction). This handles cases where the DOM changed between capture and execution. TimeoutError is NOT caught — it propagates up to the caller as a fatal timeout.

Error classes (errors.ts): TimeoutError, StagehandProtocolCompatibilityError, DuplicatePageEventSubscriptionError, ShadowRootEvaluationError, ShadowRootEvaluationUnavailableError.

Caching (cacheService.ts): Uses the Browserbase API's stateless cache routes. On each act()/observe()/extract():

  1. The raw CDP accessibility tree(s) are collected (collectCdpTree)
  2. A cache key is computed server-side from the CDP tree + request params
  3. On hit: the cached action array is replayed deterministically — if replay fails (e.g., stale selectors), it falls back to execution
  4. On miss: the LLM inference runs and the result is persisted
  5. The cache has a hit-count threshold and supports bypass for requests with locator scoping
  6. Cache read/write failures are best-effort: they log warnings but never break the operation
  7. No caching of successful workflows; caching is per single act/observe/extract

Timeouts: Each act()/observe()/extract() has a configurable timeout (passed through options.timeout). The Progress class in progress.ts enforces a deadline and provides an abort signal. waitForDomNetworkQuiet uses a 500ms quiet-after-last-network-request heuristic with a 2s stale-request sweep. The timeoutConfig.ts and DEFAULT_LOCATOR_TIMEOUT_MS set baseline defaults.

DOM settle: Before each snapshot capture, waitForDomNetworkQuiet uses CDP Network events (Network.requestWillBeSent, Network.loadingFinished, etc.) to wait until no requests are in-flight for 500ms. Stalled requests older than 2s are force-completed.

← How is the agent loop / planning implemented? · Which models are supported and how are they called? →