LLMs Technical Reviews

How is the agent loop / planning implemented?

Step loop; planner vs executor; tool-call schema; stop conditions; memory between steps.

Verdict

This is where the two designs differ most.

browser-use is an agent. Agent.run loops through step(): capture state, call the LLM for an AgentOutput (thinking, evaluation of the last goal, memory, next goal, optional plan update, action list), execute, then post-process. Planning is optional. It keeps a plan list with statuses, and nudges the model to replan after repeated failures or to make a plan after five steps without one. Memory is the message history, compacted by an LLM summary on long runs, plus a free-text memory field. The run stops on done, at max_steps (where the schema is forced to done), after max_failures consecutive failures, or on a user stop.

Stagehand has no loop at this commit, and its SDK has no agent() method. act, observe and extract each make one or two structured LLM calls and keep no memory between calls. Multi-step behaviour is left to your own code, or to an agent harness that uses Stagehand’s integration tools (run, snapshot, screenshot) over MCP or natively.

Choose browser-use for “give it a goal and let it go” tasks. Choose Stagehand if you already have an agent framework or prefer scripted control flow with AI only at the fuzzy steps.

More projects in this category are being researched.

Per-project answers

browser-use/browser-use

answered

Main loop. Agent.run(max_steps=500) (browser_use/agent/service.py:2501-2727) is the entry point. It starts the browser session, dispatches session/task events, runs initial actions (URL extraction from task), then enters a while self.state.n_steps <= max_steps loop calling _execute_step() which internally calls step().

Step loop. step() (browser_use/agent/service.py:1033-1085) has four phases:

  1. CAPTCHA wait: checks if a CAPTCHA solver is in progress and blocks if needed (browser_use/agent/service.py:1044-1061).
  2. Context preparation (_prepare_context): captures browser state (DOM + screenshot), checks downloads, checks stop/pause signals, updates action models for the current page, builds plan description, creates LLM messages with the state, runs optional message compaction, injects loop-detection/replan/exploration/budget nudges.
  3. LLM call + execution (_get_next_action + _execute_actions): calls LLM with the prepared messages, parses structured AgentOutput, then calls multi_act() to run all returned actions sequentially.
  4. Post-processing (_post_process): updates plan state, records actions for loop detection, increments consecutive failures counter on error.

Tool-call schema. LLM output is parsed into AgentOutput (browser_use/agent/views.py:388-398) containing thinking, evaluation_previous_goal, memory, next_goal, current_plan_item, plan_update, and action (list of ActionModel). The ActionModel is a discriminated union of all registered action types. In flash mode, planning fields are stripped. The output goes through AgentOutput.type_with_custom_actions() which dynamically creates a Pydantic model with the registry's current actions as a discriminated union.

Planning. When enable_planning=True, the LLM can output plan_update (list of step strings) to create/replace a plan stored as PlanItem objects with status tracking (pending/current/done/skipped). _render_plan_description() formats it for re-injection. Replan nudges trigger after 3 consecutive failures (planning_replan_on_stall). An exploration nudge suggests creating a plan after 5 steps without one.

Memory between steps. The MessageManager maintains message history. _maybe_compact_messages() summarizes old history every 25 steps (configurable) when the prompt exceeds ~10K tokens, using a compaction LLM. memory field in AgentOutput allows the LLM to persist structured notes. ActionResult.long_term_memory lets actions inject persistent context.

Stop conditions. max_steps reached (forces done action with last-step warning), max_failures exceeded (default 5, with optional final recovery attempt), manual stop/pause via signal handler (SIGINT toggles pause/resume; second SIGINT force-exits). The _force_done_after_last_step() method replaces the action model with a done-only schema on the final step.

History. Each step records an AgentHistory item with model output, results, browser state (URL, title, tabs, screenshot path). After completion, an optional LLM judge evaluates the trace.

browserbase/stagehand

answered

Stagehand does not implement a traditional autonomous agent loop (no planner/executor loop within the SDK itself). Instead, it provides a tool-call API designed to be used from external agent frameworks. The core operations are act(), observe(), and extract(), each making one (or two) LLM calls per invocation.

act() service flow (actService.ts):

  1. Wait for DOM/network quiet via waitForDomNetworkQuiet (CDP Network events)
  2. Capture a hybrid snapshot (page.captureSnapshot)
  3. Call the LLM with buildActPrompt + the snapshot — the LLM returns an ActInferenceSchema (elementId, method, args, twoStep flag)
  4. Execute the action deterministically via CDP (takeDeterministicAction)
  5. If twoStep is true, diff the new snapshot against the old one (diffCombinedTrees), make a second LLM call (buildStepTwoPrompt), and execute a follow-up action

LLM calling (inference.ts): All calls use structured output via responseFormat: { type: "json_schema", name, schema }. The schemas are defined as Zod objects — ActInferenceSchema, ObservationSchema, ExtractMetadataSchema — and serialized to JSON Schema via z.toJSONSchema(). This is JSON Schema-driven structured generation (not tool calling).

observe() captures a snapshot and asks the LLM to return an array of candidate actions (element + method + args) matching the user's instruction.

extract() captures a snapshot, calls the LLM for structured data extraction against a user-provided Zod schema, then makes a second LLM call for a metadata judgment on whether extraction is complete.

External agent loop: The prompt.ts file contains buildOperatorSystemPrompt and buildGoogleCUASystemPrompt — system prompts for an external agent that calls act()/extract()/goto() as tools in a loop. But Stagehand itself is not that loop; it's the tool layer beneath it.

Memory between steps: There is no built-in memory system — each act()/observe()/extract() call is stateless with respect to the LLM. The full snapshot is re-captured on each call. Variables (%varName% placeholders) can be substituted into action arguments before execution.

← How are actions executed and how are elements targeted? · How are failures, retries and self-healing handled? →