LLMs Technical Reviews

How is the agent loop implemented?

Planner vs single loop; tool-call schema; turn structure; stop conditions; sub-agents.

Verdict

For long autonomous runs, Codex has the most mature turn engine among the coding agents. Aider is the pick if you want to approve every step. Among the prompt-to-app builders the loops are much thinner, and one has no loop at all.

Human-driven, no tool loop. Aider parses edit blocks from plain text in Coder.run_one(). Only “reflections” (a failed SEARCH block, lint errors) loop back, and they are capped at three per message.

One tool-calling loop, with sub-agents. This is how the other coding agents work.

  • Codex: run_turn over the Responses API. Stop hooks can veto a stop and force another pass.
  • Open Interpreter: the same engine, as a Codex fork. Each turn goes through an emulated “harness” (Claude Code, Qwen Code, Kimi Code…) with that agent’s prompt and tool names.
  • Cline: ends on no tool calls, a terminal tool such as submit_and_exit, or maxIterations.
  • Codewhale: one run_turn with prepare, model, continue and tool-batch phases. A guard test fails if a second loop is added.
  • Qwen Code: capped at MAX_TURNS = 100. The front ends execute tools, not the core client.
  • OpenCode: reloads history from SQLite every step. task child sessions nest one level by default.
  • PI-Desktop: every tool call crosses into a Rust host process. All tools except Task run sequentially.
  • DeepSeek-Reasonix: the only optional planner/executor split. The planner must answer through submit_plan, and the executor runs alone if it fails.

App-builder pipelines.

  • VibeSDK: the default Think engine is a model-and-tool loop with bash disabled. The older Phasic and Agentic engines are still in the tree.
  • bolt.diy: one streamText call per request, plus up to two continuation segments when output is cut off.
  • Open Lovable: one streamed completion parsed for <file> tags. Truncation recovery exists but is off by default, so only two retries on rate-limit or service errors remain.

Pick: Aider for step-by-step pair programming. Pick: Codex, Codewhale or Cline for unattended multi-step work, and DeepSeek-Reasonix for plan-then-execute with two models. Pick: VibeSDK for an app builder that iterates on its own deploy errors.

More projects in this category are being researched.

Per-project answers

anomalyco/opencode

answered

Agent Loop

The agent loop is a single-loop architecture — there is no separate Planner; the model calls tools directly until it decides to stop. The main loop is runLoop in packages/opencode/src/session/prompt.ts (line 1081-1341), which runs inside a while(true) controlled by SessionRunState.ensureRunning (line 1343-1347). Each iteration (step): loads messages from the database, resolves the model and agent, assembles tools via SessionTools.resolve, prepares system prompts (environment, instructions, skills, MCP), and calls LLM.stream which delegates to either the native LLM runtime or the AI SDK (session/llm.ts lines 226-269). The stream events are processed by SessionProcessor.process (processor.ts lines 641-697), which handles tool calls, reasoning, text deltas, and step completion via an event handler (handleEvent, lines 278-551).

Tool-call schema uses the Vercel AI SDK's Tool type (ai package) — tools emit tool-input-start/delta/end, tool-call, tool-result, and tool-error events. Tools execute via SessionTools.resolve (session/tools.ts), which iterates the ToolRegistry to get built-in tools (edit, write, bash, read, grep, glob, task, etc.) and MCP-provided tools. MCP tools' schemas are normalized through ProviderTransform.schema (provider/transform.ts line 1575).

Stop conditions: the loop breaks when the model returns a finish reason other than "tool-calls" or "unknown" (line 1111-1116), or when the processor returns "stop" (line 1319), or when compaction creates a loop (line 1166-1168). If the processor returns "compact", a compaction step is created and the loop continues (line 1321-1328).

Sub-agents use the task tool (tool/task.ts), which spawns a new session with a different agent via handleSubtask (prompt.ts lines 255-449). Agents like general, explore, and custom agents are defined in agent/agent.ts (lines 140-265) with distinct permission rulesets and system prompts. Background sub-agents use BackgroundJob service for async execution.

openai/codex

answered

Codex CLI uses a single-loop architecture without a separate planner; the model itself plans and acts within each turn. The core loop lives in run_turn() in core/src/session/turn.rs:164 which accepts TurnInput (user messages, function-call outputs, inter-agent communications) and runs a loop { ... } (line 427) that repeats until the model produces a final assistant message with no pending input. Inside the loop, the session builds a prompt from its cloned history (line 527), invokes the OpenAI Responses API via run_sampling_request() (line 535), streams the response, and executes any tool calls the model makes through the ToolCallRuntime (line 1633). After each sampling round the loop checks for pending user input (line 431), token-limit thresholds for auto-compaction (line 612), and hook stop/block requests (line 663–719). Tool calls are dispatched through try_run_sampling_request() (line 1679) which invokes the ToolRouter — a registry that maps tool names to CoreToolRuntime implementations in core/src/tools/registry.rs. Sub-agents are supported via a multi-agent V2 system in core/src/tools/handlers/multi_agents_v2.rs:1 where tools like spawn_agent, send_message, and wait let the model create and communicate with child threads, coordinated by the AgentControl trait in core/src/agent/api.rs:41. The turn's tool-call schema is defined by the Responses API's native function-calling format, with tool definitions served by the ToolRouter and filtered per-step based on permissions and feature flags.

cline/cline

answered

Cline implements a single-turn agent loop — there is no Planner/executor split. The core loop lives in AgentRuntime.execute() in @cline/agents (agent-runtime.ts:783). It iterates: prepare the model request → call the LLM provider → extract tool calls → execute each tool → collect results → repeat. Stop conditions are: (1) the assistant returns no tool calls and no completion reminders fire (agent-runtime.ts:936-957), (2) a terminal/completion tool was called (agent-runtime.ts:977-993), (3) maxIterations is exceeded (agent-runtime.ts:831-833), or (4) an error/abort occurs (agent-runtime.ts:1000-1024). The tool-call schema uses AgentToolCallPart objects with typed toolCallId, toolName, and input fields; the model's response is parsed into these structured parts. Sub-agents are supported via the spawn-agent tool (spawn-agent-tool.ts), creating child AgentRuntime instances inside separate session contexts. Each turn goes through beforeModel/afterModel and beforeTool/afterTool hooks for extension points. Output-token-limit recovery and provider-error retry with exponential backoff are built in (agent-runtime.ts:72-91).

openinterpreter/openinterpreter

answered

The agent loop is two-layer: an outer turn loop and an inner sampling loop. RegularTask (core/src/tasks/regular.rs:31-126) is the SessionTask implementation. Its run() emits TurnStarted, runs turn lifecycle, then enters a loop calling run_turn() while sess.input_queue.has_pending_input() returns true (line 120). Inside run_turn() (core/src/session/turn.rs:163-817), the inner loop builds a StepContext with a ToolRouter and calls run_sampling_request() (line 1592). The Prompt struct (line 1558) contains cloned history, model_visible_specs(), base_instructions, and output_schema. The model streams through ModelClientSession.stream() (line 2491). There is no separate planner -- plan mode (ModeKind::Plan) is a tool-mode flag on the ToolRouter enabling PlanDelta streaming within the same request. Sub-agents use multi-agent collaboration (core/src/tools/handlers/multi_agents.rs:1-78): spawn_agent/send_message create child threads via LocalAgentControl (core/src/agent/control.rs). AgentRegistry (core/src/agent/registry.rs:26-28) tracks agents and enforces depth/session limits. Stop conditions: terminal errors (line 117), no pending input (line 120), turn-stop hooks (line 646-691), context-window compaction (line 602).

Aider-AI/aider

answered

Single loop with coder subclasses, not a separate planner. The main loop lives in Coder.run() (aider/coders/base_coder.py:876-891). It calls self.get_input() to read a user message, then self.run_one() to process it. run_one() (base_coder.py:924-944) calls send_message() synchronously, then optionally loops on self.reflected_message (up to max_reflections=3) to handle multi-turn self-correction (e.g., file-mention follow-ups, lint-error feedback).

No tool-call schema by default — edits are text-based. The LLM returns SEARCH/REPLACE blocks in plain text for the "diff" edit format (via EditBlockCoder.get_edits(), aider/coders/editblock_coder.py:21-36). There is an alternative function-calling path: WholeFileFunctionCoder and EditBlockFunctionCoder define JSON-schema functions lists (e.g. write_file with path/content params), which are sent as tools in the litellm request (aider/models.py:1006-1009). The architect mode (ArchitectCoder, aider/coders/architect_coder.py) uses a two-stage sub-agent: the architect model plans, then spawns a separate editor sub-coder via Coder.create() to apply changes.

Turn structure. Each turn: user input is preprocessed (URL scraping, file-mention detection via check_for_file_mentions at base_coder.py:1761), formatted into structured message chunks (system + examples + done-messages + repo-map + readonly files + chat files + cur + reminder — see format_chat_chunks() at base_coder.py:1226-1331), sent to the LLM via model.send_completion() (aider/models.py:985-1037), and the response is parsed for edits or shell commands. After the response, edits are applied, auto-commits happen, lint runs, shell commands execute with user confirmation, and tests optionally run.

Stop conditions. The loop runs until EOF (Ctrl-D), double Ctrl-C (keyboard_interrupt() at base_coder.py:986-1000), or /exit. Within a turn, the reflection loop stops when msg is empty or max_reflections is hit.

Sub-agents. Only in architect mode: ArchitectCoder.reply_completed() creates a new Coder instance to execute edits (architect_coder.py:37-44). There is no general sub-agent framework.

codewhale-hq/Codewhale

answered

Codewhale uses a single, unified turn loop — no planner/solver separation. The authoritative turn loop is Engine::run_turn in crates/tui/src/core/engine/turn_loop.rs:1358, and a guard test (crates/core/tests/single_turn_loop.rs) fails if a second loop is introduced.

Turn structure: Each call to run_turn iterates through four sequential phases, looped until a stop condition:

  1. prepare_model_step (preparation.rs) — loads the tool catalog, refreshes MCP servers, drains sub-agent completions, processes pending steers (user mid-turn nudges), and applies soft-landing notices at ~80% of the step budget.
  2. run_model_step (model_step.rs) — dispatches the model request (OpenAI Chat Completions, Anthropic Messages API, or OpenAI Responses API depending on provider), streams the response, decodes SSE events, and handles retries/stream resumes.
  3. continue_model_step (continuation.rs) — handles output-limit truncation, RLM (reasoning language model) continuation, and prompts the model to continue from where it left off.
  4. run_tool_batch_phase (tool_batch.rs) — plans tool executions (resolves tool schemas, applies Auto-Review gates, checks permission policy) and then executes them in parallel, collecting results into the session.

Tool-call schema: Tools are registered as ToolSpec implementations with JSON Schema inputs (tools/spec/spec.rs). The catalog for a turn is built from built-in "active native tools" (read, write, edit, bash, agent, workflow, todo_write) plus deferred/discovery tools accessed via tool_search. Tool calls follow the OpenAI/Anthropic function-calling wire format, with aliases for cross-harness parameter spellings (old_string→search, file_path→path, etc.).

Stop conditions: The loop exits when: step budget exhausted (max_steps), per-turn wall-clock timeout, user cancellation, tool-call budget exhausted, consecutive empty model responses, reasoning-only retry limit, no-progress detection from repeated permission denials, or the model stops emitting tool calls.

Sub-agents: Spawned via the agent tool (tools/subagent/mod.rs:11506). The model calls agent with an action (Start, Roster, Status, Message, Followup, etc.), and a background agent session runs with its own filtered toolset and workspace context. Foreground (turn-owned) sub-agents are tracked via ForegroundChildRegistry and cancelled when the turn ends. Background sub-agents complete asynchronously and deliver results via runtime messages. A fleet roster (FleetRoster) provides role-based model assignment for spawned agents.

esengine/DeepSeek-Reasonix

answered

The agent loop is a tool-round loop inside Agent.runToolLoop (internal/runtime/agent/run_loop.go:261). There is no separate planner model by default — the single model both plans and executes. An optional two-model Coordinator (internal/runtime/coordinator/coordinator.go:97) runs a planner model first (via submit_plan, planner tools are read-only), then hands its plan to an executor Agent. The loop in runToolLoop iterates steps up to maxSteps (default from config). Each step: (a) freezes the provider request via prepareSamplingRequest — tool schemas are captured once (ProviderSchemas), prefix shape recorded; (b) streams the model response with streamWithSamplingRecovery (run_loop.go:407), which retries up to maxSamplingAttempts=6 on interruption or context overflow; (c) classifies the response boundary via classifyResponseBoundary (response_boundary.go:27) to drop half-written tool calls; (d) commits assistant message to conversation; (e) for tool calls, runs handleToolRound → executeOne (execute_one.go:95), which resolves the tool (parse, policy, permission gate, sandbox extend, execute, post-write receipt); (f) loops back for another step. Stop conditions: no tool calls → final response handled by handleFinalResponse; max steps reached → armFinalizationRound/gracePause (finalization.go:40,55); task budget exhausted → taskBudgetPause; perseveration detected → cut-and-retry with nudge (settlePerseveration). Sub-agents are spawned via RunSubAgentWithSession (child_run.go:36), each getting its own Agent instance, session, event sink, and session-private temp directory.

firecrawl/open-lovable

answered

There is no iterative agent loop with tool-calling; it is a linear single-shot pipeline per user request. The flow: (1) user prompt hits /api/generate-ai-code-stream (app/api/generate-ai-code-stream/route.ts:91), (2) for edits, it first calls /api/analyze-edit-intent which uses Vercel AI SDK's generateObject() with a Zod schema to produce a structured search plan, (3) the search plan is executed locally by executeSearchPlan() in lib/file-search-executor.ts:42 which greps file contents line-by-line, (4) a single LLM completion is called via streamText() from the Vercel AI SDK (lib/ai/provider-manager.ts wraps the providers), (5) the text stream is parsed in real-time for XML tags (<file>, <package>, <command>) in generate-ai-code-stream/route.ts:1397-1504, and (6) parsed files are sent to /api/apply-ai-code-stream (app/api/apply-ai-code-stream/route.ts:264) which writes them to the sandbox. The system does NOT use tool/function calling — the code explicitly states: "Neither Groq nor Anthropic models support tool/function calling in this context" (route.ts:1310-1311). Instead it relies on XML tags in text output. Stop conditions are simple: up to 2 retries on rate-limit/service errors with exponential backoff (route.ts:1329-1382), and truncation recovery that fires focused API calls to complete files missing closing braces or tags (route.ts:1658-1810). No sub-agents, no multi-turn planning loop.

Editor's note. Correction: truncation recovery is disabled by default (codeApplication.enableTruncationRecovery = false in config/app.config.ts), so in a default install the only stop/retry logic is the two retries on rate-limit, timeout or service-unavailable errors.

QwenLM/qwen-code

answered

The agent loop centres on LlmClient.sendMessageStream() in packages/core/src/core/client.ts, an async generator that orchestrates one complete model interaction turn at a time. The file caps turns at MAX_TURNS = 100 (line 212) and supports multiple message types — UserQuery, ToolResult, Retry, Cron, Notification, Teammate, and Goal (lines 218–237). Each turn: the model generates text plus tool calls, the CoreToolScheduler (core/coreToolScheduler.ts) executes them, and results are fed back as function responses. The loop ends on model text (no tool calls), turn limit, session token limit, loop detection (LoopDetectionService), user interrupt, or stop hooks. There is no separate Planner loop — plan mode is gated by enterPlanMode/exitPlanMode tools that the model itself calls.

Sub-agents use two modes: forked agent (packages/core/src/agents/forkedAgent.ts) — a single-turn cache-sharing call with or without tools, used for auto-memory tasks — and headless/multi-turn via AgentCore (packages/core/src/agents/runtime/agent-core.ts), which owns runReasoningLoop(), tool scheduling, stats, and event emission. AgentInteractive (packages/core/src/agents/runtime/agent-interactive.ts) composes AgentCore with persistent state (message queue, three-level cancellation). Workflows orchestrate many sub-agents in parallel via workflow-orchestrator.ts. Tool-call schemas follow @google/genai's FunctionDeclaration format throughout.

stackblitz-labs/bolt.diy

answered

There is no Planner sub-agent — Bolt uses a single-turn streaming LLM loop orchestrated from the server. The entry point is app/routes/api.chat.ts, the server-side action handler for /api/chat. The client (Chat.client.tsx) calls useChat from the @ai-sdk/react SDK which sends a POST to that route. On the server, chatAction creates a SwitchableStream and runs a createDataStream block with an execute callback (line 93). Inside that callback, it first processes MCP tool invocations with mcpService.processToolInvocations (line 103), then optionally runs createSummary (line 122) and selectContext (line 164) — two separate LLM calls that create a chat summary and choose relevant files. The main generation is streamText (line 311) which calls the Vercel AI SDK's streamText with model, system prompt, and messages. Tool calls are exposed to the LLM via mcpService.toolsWithoutExecute set as options.tools (line 214), with toolChoice: 'auto' and maxSteps configurable per request. When the LLM response hits finishReason === 'length', the loop constructs a new user message with CONTINUE_PROMPT and calls streamText again, up to MAX_RESPONSE_SEGMENTS = 2 times (lines 251-298). The client uses @ai-sdk/react's useChat which handles streaming state, abort via stop(), and renders tool-call annotations in the UI.

vastsa/PI-Desktop

answered

PI-Desktop has a turn-structured agent loop built on the pi-agent-core library's Agent class, not a separate Planner. The SessionRuntime class in packages/agent-runtime/src/runtime.ts wraps an Agent instance and drives the conversation turn by turn.

Turn structure — Each user message enters through SessionRuntime.prompt() (runtime.ts:8460), which appends user content, checks context budget (auto-compacting if needed), then calls this.agent.prompt(...) (runtime.ts:8555-8557). The agent streams the model response, collects tool calls, executes them sequentially via toolExecution: "sequential" (subagent.ts:296), and loops back for the next assistant turn via waitForIdle() + model .continue(). Stop conditions include: the model emits no tool call and a stop reason (stop/end_turn); the user aborts (abort(), runtime.ts:8735); a graceful stop is requested (requestGracefulStop(), runtime.ts:8751); or a silent/progress-only turn is auto-recovered once then stopped (SILENT_TURN_NUDGE, PROGRESS_TURN_NUDGE, runtime.ts:828-849).

Tool-call schema — Tools are declared to the model via JSON Schema using pi-ai's Type.Object() helpers. The core tools (Read, Write, Edit, Bash, Glob, Grep, ToolSearch, Task, TaskWait, etc.) are defined in runtime.ts:3154-3263 with schema parameters. The Edit tool uses a line-anchored format with tag, ops (PUT/CUT/REM/MV), and optional legacy old_string/new_string. Bash accepts command + optional timeout.

Sub-agents — The Task tool (SUBAGENT_TOOL_NAME, subagent.ts:83) spawns a SubagentRun class that creates its own Agent instance inside the same sidecar process (subagent.ts:278-298). Delegates have their own system prompt, tool set, and model (possibly pinned). They share the host connection so all permission checks go through host-core. The parent only sees the final report; intermediate messages are filtered out of model context (session-context.ts:33-41). TaskWait converges on running delegates, TaskList reports status, TaskStop stops them. Delegates are bounded: max captures, max tokens, and keep the parent from overflowing.

Editor's note. Correction: the runtime class is DesktopAgentRuntime, not SessionRuntime, and the session agent uses toolExecution: "parallel" with every tool except Task marked sequential (runtime.ts L2170-L2179); the sequential setting cited is the subagent's.

cloudflare/vibesdk

answered

VibeSDK has three distinct agent-loop architectures, selected per app at creation: Think (model-and-tool loop via @cloudflare/think), Agentic (recursive inference with tool calling), and Phasic (deterministic phase orchestration).

The Think loop is the primary/canonical loop, implemented in ThinkAgent (extends @cloudflare/think) at worker/agents/think/ThinkAgent.ts. It is a single model-and-tool loop, not a planner+executor split. Think.getTools() returns a ToolSet of SpaceDO-backed workspace tools (read, write, edit, list, find, grep, delete) from @cloudflare/think/tools/workspace, plus VibeSDK-specific tools: commit, deploy_space, set_title, ask_questions, get_browser_console_logs. Bash is disabled (workspaceBash = false). Each turn is driven by ThinkAgent.chat(), called from ThinkCodingBehavior.runPrompt() which submits user text to the DO and translates streamed UIMessageChunk frames into WebSocket events. Tools use the OpenAI-compatible function-calling schema.

The Agentic loop (worker/agents/core/behaviors/agentic.ts) uses the infer() function in worker/agents/inferutils/core.ts — a generic inference engine that assembles tool-call requests, streams responses, accumulates tool-call deltas from chat-completions chunks, executes tools via executeToolCallsWithDependencies(), and recurses until a depth limit or CompletionDetector signal fires. Stop conditions include: max tool-calling depth per action key, a max message limit (64000 in constants), and completion signals from specific tools. The Agentic loop also includes conversation compactification triggers every 9 tool calls.

The Phasic loop (worker/agents/core/behaviors/phasic.ts) is a deterministic orchestrator that runs blueprint generation, phase generation, phase implementation, and review cycles in a fixed sequence, then iterates with the user.

Think has no sub-agents in the traditional sense; it issues tool calls to the SpaceDO workspace and other services. The infer() function does support recursive tool calling (model calls tool, tool result feeds back to model), which is its form of sub-loop.

Turn structure (Think): beforeTurn() selects context messages, beforeStep() enforces rate limits and injects a max-steps prompt on the final step. getModel() wraps an OpenAI-compatible provider over AI Gateway with custom SSE patching for Google Gemini compatibility.

How is repository context gathered and kept within the context window? →