Open-source coding agents: opencode vs codex vs cline vs openinterpreter vs aider vs Codewhale vs DeepSeek-Reasonix vs open-lovable vs qwen-code vs bolt.diy vs PI-Desktop vs vibesdk
Terminal, IDE and desktop agents that read a codebase, edit files and run commands — how they gather context, apply edits and stay safe. This page puts every verdict for the category on one page. Each question links to the full per-project answers and their code citations.
At a glance
● answered from code · — not applicable (the project does not do this) · ? insufficient evidence
How is the agent loop implemented?
For long autonomous runs, Codex has the most mature turn engine among the coding agents. Aider is the pick if you want to approve every step. Among the prompt-to-app builders the loops are much thinner, and one has no loop at all.
Human-driven, no tool loop. Aider parses edit blocks from plain text in Coder.run_one(). Only "reflections" (a failed SEARCH block, lint errors) loop back, and they are capped at three per message.
One tool-calling loop, with sub-agents. This is how the other coding agents work.
- Codex:
run_turnover the Responses API. Stop hooks can veto a stop and force another pass. - Open Interpreter: the same engine, as a Codex fork. Each turn goes through an emulated "harness" (Claude Code, Qwen Code, Kimi Code…) with that agent's prompt and tool names.
- Cline: ends on no tool calls, a terminal tool such as
submit_and_exit, ormaxIterations. - Codewhale: one
run_turnwith prepare, model, continue and tool-batch phases. A guard test fails if a second loop is added. - Qwen Code: capped at
MAX_TURNS = 100. The front ends execute tools, not the core client. - OpenCode: reloads history from SQLite every step.
taskchild sessions nest one level by default. - PI-Desktop: every tool call crosses into a Rust host process. All tools except
Taskrun sequentially. - DeepSeek-Reasonix: the only optional planner/executor split. The planner must answer through
submit_plan, and the executor runs alone if it fails.
App-builder pipelines.
- VibeSDK: the default Think engine is a model-and-tool loop with bash disabled. The older Phasic and Agentic engines are still in the tree.
- bolt.diy: one
streamTextcall per request, plus up to two continuation segments when output is cut off. - Open Lovable: one streamed completion parsed for
<file>tags. Truncation recovery exists but is off by default, so only two retries on rate-limit or service errors remain.
Pick: Aider for step-by-step pair programming. Pick: Codex, Codewhale or Cline for unattended multi-step work, and DeepSeek-Reasonix for plan-then-execute with two models. Pick: VibeSDK for an app builder that iterates on its own deploy errors.
More projects in this category are being researched.
How is repository context gathered and kept within the context window?
Aider is the only coding agent that builds a global view of the repository up front. The others let the model search on demand and differ mainly in how they compact long sessions. None of the twelve uses embeddings.
Precomputed repo map. Aider ranks tree-sitter definitions and references with personalised PageRank. The map fits a budget of 1/8 of the input window, clamped to 1K–4K tokens.
On-demand tools plus compaction.
- Cline: reads up to 2,000 lines, caps each tool result at 8,000 characters, and compacts at 90% of the window down to 70%.
- Codewhale: prunes old tool output before any LLM summary and keeps the last user round word for word.
- DeepSeek-Reasonix: a tree-sitter
code_indexgives outlines without a language server. At 85% it writes a structured digest of at most 16K tokens, keeping the prompt prefix cache-stable. - Qwen Code: replaces old tool outputs with
[Old tool result content cleared]and runs background auto-memory under~/.qwen/memories/. - OpenCode: prunes tool outputs while protecting the last 40K tokens, then hands off to a dedicated
compactionagent. - PI-Desktop: a model-written summary by default. It throws rather than send an oversized request.
- Codex: no dedicated read or search tools. The model reads with shell commands, each going through approval and the sandbox.
- Open Interpreter: Codex's compaction, keeping up to 20,000 tokens of recent user messages beside the summary.
Per-request selection in the app builders.
- bolt.diy: two extra LLM calls per request, a chat summary and a file selection. Only the last three messages are kept, and the five-file limit is a prompt instruction, not code.
- Open Lovable: a structured search plan drives a local grep. Context files are cut to 2,000 characters, and nothing is summarised.
- VibeSDK: Think trims history but keeps tool pairs around
deploy_spaceandcommit.
Pick: Aider for a cheap, global map of a large repository. Pick: DeepSeek-Reasonix or Codewhale for long sessions where prompt-cache hits drive cost. Pick: bolt.diy if you want context chosen per request without hand-picking files.
More projects in this category are being researched.
How are code edits applied?
Aider gives the most reviewable result, with one git commit per edit. OpenCode and Codewhale forgive the most model mistakes, and PI-Desktop is the hardest to apply to stale text. The app builders mostly rewrite whole files.
Edit blocks parsed from text. Aider uses SEARCH/REPLACE with exact matching, because the fuzzy matcher is disabled in code. It lints after each edit and offers /undo for its own commits.
Exact search/replace tools.
- Cline:
editorneeds a unique match, and GPT/Codex models getapply_patchinstead. Git checkpoints are stored under private refs. - Qwen Code: refuses an edit unless the file was read this session and is unchanged on disk.
- DeepSeek-Reasonix:
multi_editapplies a batch atomically. A per-turn rewind store sits outside git and does not cover bash.
Forgiving matchers.
- OpenCode: nine fallback replacers, LSP diagnostics after each write, and snapshots in a separate git directory.
- Codewhale: exact, then indentation-tolerant, then punctuation-tolerant matching, with an
expected_hashguard. A side git repository backs/restore N.
Patch grammars and anchors.
- Codex: its own
*** Begin Patchgrammar and no lint step. The editor found that the git-snapshotundofeature has been removed at this commit and is kept only as a no-op flag, so rollback is your own git. - Open Interpreter:
apply_patchin native mode. Emulated harnesses edit through alias tools such as a Claude-styleEdit. - PI-Desktop: line-range ops guarded by a 4-hex tag of the file. A stale tag fails the edit.
Whole files (app builders).
- bolt.diy: full files in
<boltAction>tags.saveFileHistoryhas no callers, so there is no file history. - Open Lovable: full
<file>blocks. Snippet merges happen only with a Morph API key, andbuild-validator.tsis never called. - VibeSDK: Think's exact
editon a git-backed workspace. The model choosescommitrestore points, and rollback re-commits old content.
Pick: Aider for every change as a reviewable commit. Pick: OpenCode or Codewhale for weaker models that slip on whitespace. Pick: VibeSDK for an app builder with real restore points.
More projects in this category are being researched.
How are shell commands and file writes kept safe?
Codex, and Open Interpreter, which inherits its engine, put every shell command in an OS sandbox on macOS, Linux and Windows. DeepSeek-Reasonix also jails bash and fails closed. The app builders get their safety from where the code runs.
OS sandbox as the default path.
- Codex: approval, then sandbox selection (Seatbelt, bubblewrap plus seccomp, or a restricted token), then a retry with an escalated sandbox. The default approval policy is
on-request. - DeepSeek-Reasonix: Seatbelt or bubblewrap, and it refuses to run bash where there is no backend (Windows). However, an
Askrule resolves to allow when no approver is attached.
Sandbox available, but conditional.
- Qwen Code: bwrap or Landlock for shell commands, and Seatbelt, Docker or Podman for the whole CLI.
automode puts a regex filter and a fail-closed LLM classifier in front of risky calls. - Codewhale: Seatbelt on macOS. Bubblewrap on Linux is opt-in, and Windows has no sandbox.
Approval only, with no OS sandbox.
- Aider: shell prompts need an explicit yes even with
--yes-always. - Cline: the CLI auto-approves every tool by default.
- OpenCode: the default ruleset is
"*": "allow". Only repeated identical calls, external directories and.envreads ask. - PI-Desktop: a Rust host process enforces risk tiers and concurrency budgets, but bash runs as you, and
automode allows everything.
Isolated by architecture (app builders).
- bolt.diy: code runs in a WebContainer in the browser tab, and MCP tools run only after an explicit approval.
/api/chathas no rate limiting. - Open Lovable: an E2B or Vercel cloud sandbox with no approvals at all. Creating a sandbox kills every other one.
- VibeSDK: the default Think engine has no shell (
workspaceBash = false), and each app gets its own Durable Objects.preDeploySafetyGatechecks React render-loop hazards, not secrets, and only on the legacy engines.
Pick: Codex for kernel-level isolation on all three desktop platforms. Pick: DeepSeek-Reasonix or Qwen Code for open-model agents with a real shell jail. Pick: bolt.diy if generated code must never leave the user's browser.
More projects in this category are being researched.
Which models are supported and how are they called?
Aider is the only coding agent that works without native tool calling, so it suits small and local models best. Codex is the narrowest, with one wire protocol. Among the app builders, real model choice is thinner than the docs suggest.
Text edit formats, any model. Aider calls everything through LiteLLM. 357 entries in model-settings.yml tune the edit format and helper models per model.
Native tool calling over broad gateways.
- Cline: its own
@cline/llmsgateway with 211 generated provider IDs. It has no text-format fallback. - OpenCode: the models.dev catalog plus AI SDK packages, with per-family system prompts.
- Codewhale: about 50 provider kinds and three wire clients.
model = "auto"picks a big or cheap model, and its decision router can call a hosted third-party service. - Qwen Code: five adapters, all translated into
@google/genaishapes, and 13 provider presets. - PI-Desktop: seven wire styles through pi-ai, including the ChatGPT/Codex subscription endpoint.
- DeepSeek-Reasonix: three provider kinds. It picks the reasoning parameters from the base URL and ships per-vendor price tables.
- Open Interpreter: picks a harness per provider and model, for example
claude-code-barefor DeepSeek. Most harnesses need a Chat Completions endpoint.
One protocol. Codex speaks only the OpenAI Responses API. The built-in providers are OpenAI, Bedrock, Ollama and LM Studio, and any other provider must speak Responses.
App builders.
- bolt.diy: 22 provider classes, but only OpenRouter, Anthropic, OpenAI and Google are registered. Ollama and LM Studio need a source edit despite the README.
- Open Lovable: four providers, no tool calling, one prompt for every model, and no local models.
- VibeSDK: the default Think engine is hard-wired to a Gemini Flash model through AI Gateway. The per-action model settings only reach the legacy engines.
Pick: Aider for local or weak models that cannot call functions. Pick: OpenCode, Cline or Codewhale for the widest choice of tool-capable providers. DeepSeek-Reasonix is the pick if cost tracking for DeepSeek-class pricing matters. Pick: Codex if you are on OpenAI models and want the best-tuned path.
More projects in this category are being researched.
How can it be extended and customised?
Codex, Qwen Code and Cline have the fullest extension surfaces, with MCP, hooks, plugins, skills and an embeddable API. Aider and Open Lovable have no MCP client and no plugins.
Full stack: MCP, hooks, plugins, skills and SDKs.
- Codex: twelve hook events from
config.tomlorhooks.json, and skills and plugins as separate mechanisms.codex app-serverexposes the full JSON-RPC protocol. - Open Interpreter: the same surface, plus new harnesses (a prompt, tool JSON and routing entry each).
- Qwen Code: one extension, installed from git, npm or a GitHub release, can ship MCP servers, skills, sub-agents, hooks and workflows. SDKs exist for TypeScript, Python and Java.
- Cline: JS/TS plugins run in a subprocess sandbox.
@cline/coreembeds the agent in your own code. - OpenCode: plugins from npm, plus custom tools you add by dropping files into a
tool/directory. Every UI is a client of one server API.
Host-mediated extensions.
- DeepSeek-Reasonix: Extension Protocol v2 sidecars can intercept tool calls and host provider streams.
sdk/gohelps you write them. - PI-Desktop:
.piplugpackages whose tools pass the same Rust permission gate as built-ins. - Codewhale: project hooks run only after
/hooks approve <digest>for those exact bytes. The TypeScript extension host is experimental and behind a feature flag.
Configuration and scripting. Aider offers --lint-cmd, --test-cmd, --watch-files comments and the Python Coder.create() API, but no hooks.
App builders.
- bolt.diy: MCP over stdio, SSE or HTTP, with per-call approval. It does not read instruction files such as
AGENTS.md. - VibeSDK: a TypeScript client SDK and four built-in
SKILL.mdskills. Its MCP manager ships with an empty server list, and thecli/tuiscripts point to a directory that does not exist. - Open Lovable: extension stops at the sandbox-provider class and config files.
Pick: Codex or Qwen Code for a hook-and-plugin ecosystem with SDKs. Pick: Cline or OpenCode to embed an agent in your own product. Pick: bolt.diy for an app builder you can wire to MCP tools.
More projects in this category are being researched.
The projects
- anomalyco/opencode: Client/server TypeScript coding agent on Bun and Effect: any AI SDK provider, permission-ruled tools, TUI, desktop and SDK clients.
- openai/codex: OpenAI's Rust terminal coding agent: a Responses-API turn loop with apply_patch edits and OS-level sandboxes for every shell command.
- cline/cline: TypeScript agent SDK behind a CLI, VS Code and desktop app: tool-calling loop, per-tool approval policies, git checkpoints, MCP and plugins.
- openinterpreter/openinterpreter: Rust fork of OpenAI Codex that re-creates other agents' prompts, tools and wire formats ("harnesses") to get more out of open models.
- Aider-AI/aider: Python terminal pair-programmer: plain-text SEARCH/REPLACE edits, a tree-sitter repo map ranked by PageRank, and git auto-commits.
- codewhale-hq/Codewhale: Rust terminal coding agent with ~50 provider kinds, Plan/Act/Operate modes, approval postures, OS sandboxing and sub-agents.
- esengine/DeepSeek-Reasonix: Single-binary Go coding agent built around prefix-cache-stable prompts, an optional planner/executor split, OS sandboxing and per-turn rewind.
- firecrawl/open-lovable: Next.js prompt-to-app builder that clones sites via Firecrawl into a React/Vite app previewed in a cloud sandbox.
- QwenLM/qwen-code: Terminal coding agent forked from Gemini CLI, with a multi-provider model layer, bwrap/Landlock shell sandboxing and an LLM auto-approver.
- stackblitz-labs/bolt.diy: Open-source bolt.new: a Remix prompt-to-app builder that streams LLM file and shell actions into an in-browser WebContainer.
- vastsa/PI-Desktop: Electron desktop agent running the pi engine in a Node sidecar, with every tool call gated and executed by a Rust host-core.
- cloudflare/vibesdk: Self-hostable prompt-to-app platform on Cloudflare Workers; a Durable Object agent writes, deploys and previews whole web apps.