# codewhale-hq/Codewhale

> Rust terminal coding agent with ~50 provider kinds, Plan/Act/Operate modes, approval postures, OS sandboxing and sub-agents.

- Category: [Open-source coding agents](https://llms-technical-reviews.com/coding-agents/)
- Repository: https://github.com/codewhale-hq/Codewhale (reviewed at commit `ecbf2869ab9110486a9e1d6d75481d1d7adf8acf`, 2026-10-05)
- Stars: 41059 · Language: Rust · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/codewhale/

## Overview

Codewhale is a terminal coding agent written in Rust. It works with many model providers. One `codewhale` binary gives you an interactive TUI, a headless `exec` mode for scripts and CI, an MCP server, an ACP/app server with an HTTP/SSE runtime API, and a small web client. The same engine runs behind all of them. The npm package and the crates.io package install that same binary.

The project is large. The workspace has 28 crates plus two test crates, and almost all of the code is in `crates/tui`, which holds well over a million lines of Rust: the engine, tools, providers, sandbox, MCP, hooks, sub-agents and the UI. The smaller crates hold configuration (`config`), command policy (`execpolicy`), protocol and state types, hooks, memory and workflow support. The code is heavily commented with issue numbers and "known limitations" notes, which makes it easier to audit than its size suggests.

Codewhale is in the Claude Code / Codex family, and the resemblance is deliberate: it accepts other harnesses' tool parameter spellings, reads `CLAUDE.md` and `AGENTS.md`, and copies the deferred `tool_search` pattern. On top of that it adds a wide provider matrix (about 50 provider kinds), Plan/Act/Operate modes, git-backed restore points, multi-agent "fleets" and a decision-model auto-router.

## Architecture

```mermaid
flowchart LR
  CLI["codewhale CLI"] --> TUI["TUI (ratatui)"]
  CLI --> EXEC["exec / serve / app-server"]
  TUI --> ENG["Engine event loop"]
  EXEC --> ENG
  ENG --> LOOP["run_turn: 4 phases"]
  LOOP --> ROUTE["Model routing + LlmClient"]
  ROUTE --> PROV["Chat / Messages / Responses wires"]
  LOOP --> GATE["Plan tools: hooks, approval, Auto-Review"]
  GATE --> TOOLS["Native tools + tool_search"]
  TOOLS --> SBX["Seatbelt / bwrap sandbox"]
  TOOLS --> MCP["MCP pool"]
  TOOLS --> SUB["agent tool: sub-agents"]
  LOOP --> COMP["Compaction"]
  LOOP --> SNAP["Side-git snapshots"]
```

| Component | Path | Role |
|---|---|---|
| CLI entry | `crates/cli/` | Argument parsing and subcommands (`run`, `exec`, `serve`, `mcp-server`, `app-server`, `doctor`, ...) |
| Engine | `crates/tui/src/core/engine.rs` | Long-lived event loop; receives `Op`s, owns the session, MCP boot and sub-agent completions |
| Turn loop | `crates/tui/src/core/engine/turn_loop.rs`, `turn_loop/` | `run_turn` and its prepare / model / continue / tool-batch phases |
| Tool catalog | `crates/tui/src/core/engine/tool_catalog.rs` | Eager native tools, deferred tools, `tool_search` |
| Tools | `crates/tui/src/tools/` | File read/write/edit, `apply_patch`, shell, sub-agents, workflow, memory and more |
| Policy | `crates/execpolicy/` | Approval modes and command prefix rules |
| Sandbox | `crates/tui/src/sandbox/` | Seatbelt on macOS, opt-in bubblewrap on Linux, sandbox policies |
| Providers | `crates/config/src/provider_kind.rs`, `crates/tui/src/client/`, `llm_client/` | Provider kinds, wire clients, retries |
| Routing | `crates/tui/src/model_routing.rs` | `model = "auto"` big/cheap routing |
| Compaction | `crates/tui/src/compaction.rs`, `compaction/` | Context pruning and summarisation with a survival contract |
| Hooks and extensions | `crates/tui/src/hooks/`, `extension_host/` | Subprocess hooks; experimental TypeScript plugin host |

## How a request flows

1. **Start.** `main` resets SIGPIPE and calls `codewhale_cli::run_cli()`. The interactive and headless paths share the parser and the engine entry point ([main.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/cli/src/main.rs#L44-L61)). `spawn_engine` builds an `Engine` and runs it as a supervised task ([engine.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine.rs#L8948-L8961)).
2. **Receive.** `Engine::run_owned` boots MCP, starts a stall watchdog, and loops over inputs: sub-agent completions, MCP updates, shell wake-ups and operations. A user message arrives as `Op::SendMessage` ([engine.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine.rs#L3348-L3420)).
3. **Set up the turn.** `run_turn` starts the wall-clock budget, revalidates tools activated earlier in the conversation, checks whether the tool surface changed against the pinned prefix fingerprint (for prompt caching), and creates the retry budgets for streams, empty replies and reasoning-only replies ([turn_loop.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop.rs#L1358-L1530)).
4. **Loop.** Each iteration runs four phases, and any of them can retry, break or return ([turn_loop.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop.rs#L1532-L1591)). `prepare_model_step` builds the request (catalog, steers, sub-agent results). `run_model_step` dispatches and streams it ([model_step.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop/model_step.rs#L7-L47)). `continue_model_step` handles truncated output, and `run_tool_batch_phase` runs any tool calls.
5. **Tools.** `run_tool_batch_phase` connects the MCP pool if needed, then calls `plan_tool_calls` (hooks, approval policy, Auto-Review, call budget, deferred-tool activation), `execute_planned_tools` and `process_tool_results`. Steers that arrived during the batch are added as user messages ([tool_batch.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop/tool_batch.rs#L6-L150)).
6. **Finish.** When the model stops calling tools, or a budget or cancel ends the loop, `run_turn` reports Completed, Failed or Interrupted. Turn-owned sub-agents keep running in the background, and the session gets a note that they will report back ([turn_loop.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop.rs#L1591-L1648)).

## Key components

### Tool catalog

Only a few tools go on the wire eagerly: `read`, `write`, `edit`, `bash`, `agent`, `workflow`, `todo_write`, the goal tools and `load_skill`. Everything else, including MCP tools, is found through `tool_search` and activated for the rest of the conversation ([tool_catalog.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/tool_catalog.rs#L27-L66)). This keeps the prompt prefix small and cache-stable. The comments justify each eager entry in bytes.

### File editing

`edit` tries an exact match first, in a CRLF-normalised view. If that fails, it tries an indentation-tolerant match, then a typographic-punctuation match (smart quotes, dashes, NBSP). It refuses ambiguous matches and shows the start of the search text when nothing matches ([file.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L2630-L2680)). Writes and edits accept an optional `expected_hash` and refuse to write if the file changed since it was read ([file.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L89-L108)). Aliases such as `old_string` and `file_path` are mapped to the canonical names, and conflicting duplicates are rejected ([file.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L228-L255)). After a write, a syntax guard and a formatter pass run, and LSP diagnostics are attached to the result. `apply_patch` handles unified diffs with a default fuzz of 3 (maximum 50) and needs at least four anchor lines before it will move a hunk ([apply_patch.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/apply_patch.rs#L27-L38)).

### Modes, approvals and sandbox

There are two separate controls. The mode (`Plan`, `Agent` shown as Act, `Operate`) sets what kind of work is allowed. The mode cycle is defined in [app_mode.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/app_mode.rs#L8-L33). Plan has no file-mutation authority, and the `write`/`edit` fallbacks tell the model to switch modes first ([tool_catalog.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/tool_catalog.rs#L67-L90)). The approval posture (`Suggest`/Ask by default, `Auto` review, `Bypass`, `Never`) sets when to ask, and Shift+Tab cycles through the first three ([approval_mode.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/execpolicy/src/approval_mode.rs#L1-L64)). Containment depends on the platform: Seatbelt on macOS when the probe succeeds, bubblewrap on Linux only if you opt in and `/usr/bin/bwrap` exists, and nothing on Windows yet ([sandbox/mod.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/sandbox/mod.rs#L1-L80)). The code says plainly when no OS sandbox is active, and that honesty is to its credit.

### Providers and auto-routing

`ProviderKind` has roughly 50 variants: DeepSeek (the default), OpenAI, Anthropic, OpenRouter, Ollama, vLLM, SGLang, many Chinese clouds, and dual-wire variants that speak Anthropic Messages to non-Anthropic backends ([provider_kind.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/provider_kind.rs#L13-L93)). Three wire clients cover Chat Completions, Anthropic Messages and OpenAI Responses. With `model = "auto"`, `model_routing.rs` chooses between a big and a cheap model for the active provider ([model_routing.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/model_routing.rs#L26-L40)). It decides either with a chat router model or with a "decision" router that asks TypeSafe's hosted Jev model, directly or through OpenRouter ([system_one.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/client/system_one.rs#L1-L27)). The older keyword scorer in `auto_model.rs` is kept only for compatibility and is not used for provider-neutral auto routing ([auto_model.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/auto_model.rs#L1-L25)).

### Context and restore points

There is no embedding index. The model gathers context with read and search tools. Compaction is on by default. It prunes old tool output first, then summarises older turns while keeping the last user round word for word (the "survival contract"). Real thresholds come from each model's window, with an 800K-token fallback ([compaction.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/compaction.rs#L150-L170)). Before a turn, around each tool call and after the turn, the workspace is snapshotted into a side git repository. `/restore N` and the `revert_turn` tool use those snapshots ([turn.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/turn.rs#L1-L15)).

## Extending it

- **Hooks.** Subprocesses run on events such as `SessionStart`, `MessageSubmit`, `ToolCallBefore`/`After`, `ModeChange`, `SubagentSpawn`, `ShellEnv` (injects env vars) and `SessionIdle` ([hooks/config.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/hooks/config.rs#L27-L75)). Exit code 2 blocks the action ([executor.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/hooks/executor.rs#L1882-L1889)). Project hooks run only in a trusted workspace, after `/hooks review` and `/hooks approve <digest>` for those exact bytes ([hooks/config.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/hooks/config.rs#L428-L436)).
- **MCP.** Supports stdio, SSE, streamable HTTP and OAuth. Tools from connected servers join the deferred catalog. `codewhale mcp-server` exposes Codewhale itself over MCP.
- **Skills.** `SKILL.md` folders are listed in the prompt and loaded with `load_skill`.
- **Extension host (experimental).** A Bun or Node process runs TypeScript plugins behind `[features] extension_host`. Plugins can add tools, slash commands, admission hooks and prompt sections, all behind the core's approval gate ([extension_host/mod.rs](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/extension_host/mod.rs#L1-L60)).
- **Automation.** `codewhale exec` runs headless. `app-server` serves the `/v1` runtime API (sessions, threads, turns) and a mobile control page. Sub-agents and workflows can be scripted through the `agent` and `workflow` tools.

## Running it

- **Install.** Use the install script, `npm install -g codewhale` (a wrapper around the release binaries) or `cargo install codewhale-cli --locked`. Building from source needs Rust 1.89 or later.
- **Configure.** Add a provider key or OAuth login, or point a provider at a local Ollama/vLLM/SGLang endpoint. `codewhale doctor` checks the setup.
- **Optional.** Install `bwrap` and opt in to get a Linux sandbox. Install `git` for restore points. Bun or Node is needed only for the extension host.

## Strengths and caveats

- **Strength: careful edit tooling.** Hash-guarded writes, alias normalisation, staged fuzzy matching and post-edit syntax and LSP checks deal directly with the usual ways models break edits.
- **Strength: provider breadth.** About 50 provider kinds, three wire protocols and local engines, with a layered model catalog.
- **Strength: honest safety reporting.** Modes, approval postures and sandboxing are separate controls, and the code says when a sandbox is not in effect.
- **Caveat: no sandbox by default on Linux or Windows.** On Linux, approval is your only guard unless you opt in to bubblewrap.
- **Caveat: sheer size.** More than a million lines in one crate and many feature flags. Understanding or forking it is a serious investment, and behaviour changes quickly between releases.
- **Caveat: the hosted auto-router is a third party.** The decision router sends routing requests to TypeSafe or OpenRouter. Use the chat router or a pinned model if that matters to you.
- **Caveat: the plugin host is experimental.** TypeScript extensions are feature-gated and limited to phase-1 capabilities.

*Sources: code at ecbf286, deepwiki-open wiki (12 pages), verified Q&A.*

## How codewhale-hq/Codewhale answers the Open-source coding agents questions

### How is the agent loop implemented? (answered)

Codewhale uses **a single, unified turn loop** — no planner/solver separation. The authoritative turn loop is `Engine::run_turn` in `crates/tui/src/core/engine/turn_loop.rs:1358`, and a guard test (`crates/core/tests/single_turn_loop.rs`) fails if a second loop is introduced.

**Turn structure:** Each call to `run_turn` iterates through four sequential phases, looped until a stop condition:
1. **prepare_model_step** (`preparation.rs`) — loads the tool catalog, refreshes MCP servers, drains sub-agent completions, processes pending steers (user mid-turn nudges), and applies soft-landing notices at ~80% of the step budget.
2. **run_model_step** (`model_step.rs`) — dispatches the model request (OpenAI Chat Completions, Anthropic Messages API, or OpenAI Responses API depending on provider), streams the response, decodes SSE events, and handles retries/stream resumes.
3. **continue_model_step** (`continuation.rs`) — handles output-limit truncation, RLM (reasoning language model) continuation, and prompts the model to continue from where it left off.
4. **run_tool_batch_phase** (`tool_batch.rs`) — plans tool executions (resolves tool schemas, applies Auto-Review gates, checks permission policy) and then executes them in parallel, collecting results into the session.

**Tool-call schema:** Tools are registered as `ToolSpec` implementations with JSON Schema inputs (`tools/spec/spec.rs`). The catalog for a turn is built from built-in "active native tools" (read, write, edit, bash, agent, workflow, todo_write) plus deferred/discovery tools accessed via `tool_search`. Tool calls follow the OpenAI/Anthropic function-calling wire format, with aliases for cross-harness parameter spellings (`old_string`→`search`, `file_path`→`path`, etc.).

**Stop conditions:** The loop exits when: step budget exhausted (`max_steps`), per-turn wall-clock timeout, user cancellation, tool-call budget exhausted, consecutive empty model responses, reasoning-only retry limit, no-progress detection from repeated permission denials, or the model stops emitting tool calls.

**Sub-agents:** Spawned via the `agent` tool (`tools/subagent/mod.rs:11506`). The model calls `agent` with an action (Start, Roster, Status, Message, Followup, etc.), and a background agent session runs with its own filtered toolset and workspace context. Foreground (turn-owned) sub-agents are tracked via `ForegroundChildRegistry` and cancelled when the turn ends. Background sub-agents complete asynchronously and deliver results via runtime messages. A fleet roster (`FleetRoster`) provides role-based model assignment for spawned agents.


Citations: [crates/tui/src/core/engine/turn_loop.rs:1358-1648](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop.rs#L1358-L1648) · [crates/tui/src/core/engine/turn_loop/preparation.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop/preparation.rs#L1-L80) · [crates/tui/src/core/engine/turn_loop.rs:1532-1591](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop.rs#L1532-L1591) · [crates/tui/src/tools/subagent/mod.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/subagent/mod.rs#L1-L80) · [crates/tui/src/core/engine/tool_catalog.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/tool_catalog.rs#L1-L80)

### How is repository context gathered and kept within the context window? (answered)

Codewhale gathers repository context through **explicit tool calls** rather than automated crawling or embedding-based retrieval. There are no embeddings or vector indexes — the model reads files it chooses.

**File reading:** The `read`/`ReadTool` and legacy `ReadFileTool` (`tools/file.rs:1092`) stream file contents with: small-file fast-path (whole file in one call when ≤4KB and ≤500 lines), line-range windows (`start_line`/`max_lines`), byte-size awareness with truncation footers naming the next offset, content hashing (SHA-256) reported in responses so the model can validate later edits. Binary files route to OCR (images) or PDF extraction.

**Search tools:** `search_content` (Grep) and `search_name` (FileSearch) provide pattern-based discovery (`tools/search.rs`, `tools/file_search.rs`). A deferred `tool_search` tool (`tool_catalog.rs:35`) with regex/BM25 matching allows discovery of tools and files by name or content. Web search is available via `web_search`.

**System prompt assembly:** The system prompt (`prompts/text.rs:44`, `BASE_PROMPT`) is composed from layers: constitution (binding core), personality overlay (currently CALM), approval-policy overlays, memory block (user memories from `runtime/src/native_memory.rs`), skills index, MCP server instructions, mode-specific instructions (Plan/Act/Operate), and project instruction files (CLAUDE.md/AGENTS.md). The session-pinned prefix is kept KV-cache stable — volatile facts (steers, completions, diagnostics) are appended as user-role messages, never spliced into the frozen prefix.

**Compaction for long sessions** (`compaction.rs:1`): When context pressure exceeds a model-specific token threshold (80% of context window), compaction triggers. The strategy is: first prune old tool results from the tail (cheapest, no LLM call), then if still over threshold, run an LLM summarization pass that compresses older user+assistant turns while keeping the **last user round verbatim** (the survival contract in `survival_contract.rs`, enforced by `last_round.rs`). Recent plain user messages are retained up to a configurable token budget (`retained_user_message_tokens`). The pre-compaction history is persisted as an artifact before any mutation. Context pressure warnings (warning at 70%, critical at 85%) are emitted mid-turn via `context_pressure_message`.

**Memory:** User memories (`tools/native_memory.rs`, `tools/remember.rs`) provide a persistent capture path — the model can ask to remember things, which are stored on disk and prepended to the system prompt on subsequent sessions.


Citations: [crates/tui/src/compaction.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/compaction.rs#L1-L80) · [crates/tui/src/compaction.rs:1340-1420](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/compaction.rs#L1340-L1420) · [crates/tui/src/compaction/last_round.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/compaction/last_round.rs#L1-L60) · [crates/tui/src/prompts/text.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/prompts/text.rs#L1-L60) · [crates/tui/src/tools/file.rs:1092-1170](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L1092-L1170) · [crates/tui/src/core/engine/turn_loop/preparation.rs:95-108](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/turn_loop/preparation.rs#L95-L108)

### How are code edits applied? (answered)

Codewhale provides three file-editing mechanisms with gradual complexity, plus git-based undo:

**1. `write` (`WriteFileTool`, `file.rs:1755`):** Full file replacement. Accepts `path`, `content`, and optional `expected_hash`. Creates parent directories if needed, preserves line-ending style (CRLF/LF) on overwrite, and writes atomically via `run_blocking_write_atomic`. The optional content-hash guard (`verify_expected_hash`, `file.rs:89-108`) refuses the write if the file changed since it was last read, preventing stale-state overwrites.

**2. `edit` (`EditFileTool`, `file.rs:2581`):** Search-and-replace with fuzzy matching. Accepts `path`, `search`, `replace`, and optional `expected_hash`. Matching is multi-layered: exact match first, then indentation-tolerant fuzzy match, then typographic-punctuation (smart quotes, em-dashes) normalization. Returns detailed error messages when search/replace are identical (a common model mistake). After the edit a Rust code formatter (`normalize_edit`, `rust_format.rs`) is applied, and LSP diagnostics are injected into the result via `lsp_diagnostics_for_paths`.

**3. `apply_patch` (`ApplyPatchTool`, `apply_patch.rs:1`):** Unified diff patching supporting multi-hunk patches with fuzzy matching (`MAX_FUZZ=50`, default `fuzz=3`). Hunks can be relocated when their stated line numbers are stale (requires ≥4 anchor lines for safety via `MIN_ANCHOR_LINES`). Preserves CRLF line endings and trailing-newline state from the base file.

**Cross-harness parameter compatibility:** All file tools translate known spellings from other coding harnesses (`old_string`/`new_string`, `file_path`, `filePath`, `offset`/`limit` for read windows) onto the canonical names via `apply_param_aliases` (`file.rs:228-255`). This prevents rejection loops where the model's training data names parameters differently. Unknown parameters are rejected with explicit errors rather than silently dropped.

**Validation pipeline:** All mutations go through `guard_edit` (`syntax_check.rs`) for syntax validation, `normalize_edit` for code formatting, and the content-hash guard. Auto-Review gates can flag edits for user approval based on file path patterns.

**Git/undo:** Workspace snapshots are taken into a side git repository (`turn.rs:8-16`) at pre-turn, pre-tool, and post-tool boundaries. The `/restore N` command and `revert_turn` tool consume these snapshots. The `git_tool` provides direct git operations (commit, push, branch creation).


Citations: [crates/tui/src/tools/file.rs:1755-1835](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L1755-L1835) · [crates/tui/src/tools/file.rs:2581-2670](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L2581-L2670) · [crates/tui/src/tools/file.rs:89-108](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L89-L108) · [crates/tui/src/tools/apply_patch.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/apply_patch.rs#L1-L80) · [crates/tui/src/tools/file.rs:228-255](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/file.rs#L228-L255) · [crates/tui/src/core/turn.rs:1-20](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/turn.rs#L1-L20)

### How are shell commands and file writes kept safe? (answered)

Execution safety in Codewhale is implemented at multiple layers: approval modes, command safety classification, OS sandboxing, and filesystem guards.

**Approval modes** (`execpolicy/src/approval_mode.rs:7`): Four modes — `Suggest` (ask for non-safe tools, the default), `Auto` (auto-review risky calls), `Bypass` (no approvals/"YOLO mode"), and `Never` (block all tools requiring approval). The user cycles through Suggest → Auto → Bypass with Shift+Tab. `ApprovalMode::from_config_value` accepts many CLI aliases ("yolo", "bypass-permissions", "full-access", etc.).

**Command safety** (`execpolicy/src/command_safety.rs:1`): A `COMMAND_ARITY` dictionary maps ~40+ command prefixes (`git status`, `npm run`, `cargo check`, `docker compose`, etc.) to their canonical forms. `Ruleset` objects (`execpolicy/src/lib.rs:29`) at three priority layers (BuiltinDefault → Agent → User) define `trusted_prefixes` (auto-allowed), `denied_prefixes` (always blocked), and `ask_rules` (typed tool-invocation patterns requiring approval). Permission actions are Allow/Ask/Deny. The command-safety system classifies commands into `SafetyLevel`s, and the `ExecPolicyContext` uses these to determine whether a shell invocation requires approval.

**OS sandbox** (`sandbox/mod.rs:1`):
- **macOS**: Seatbelt (`sandbox-exec`) when runtime probe succeeds — uses the native macOS sandbox.
- **Linux**: Bubblewrap (`bwrap`) only when user opts-in (`prefer_bwrap: true`) and `/usr/bin/bwrap` is executable. Creates a new mount namespace with `--unshare-all`, read-only bind of `/`, tmpfs `/tmp`, `/dev`, `/proc`, and writable roots per policy. `--share-net` is added for network-enabled policies. The extension host always uses bwrap when available without the opt-in.
- **Windows**: Process-tree containment via Job Object (planned, not yet sandboxed).
- `SandboxPolicy` (`sandbox/policy.rs:22`) has four levels: `DangerFullAccess` (no restrictions), `ReadOnly` (read-only everywhere), `ExternalSandbox` (already sandboxed, skip double-wrapping), `WorkspaceWrite` (read-only + writable workspace + optional extra roots + optional network).

The sandbox is **never silently downgraded** — if bwrap isn't available on Linux, the tool reports no OS sandbox rather than pretending.

**Filesystem guards:** A read deny-list (`tools/file.rs:458`) protects credential files, config paths, and sensitive system files. Paths are checked on both raw spelling (before symlink resolution) and canonical form (after). The content-hash guard prevents stale-state overwrites.

**Network restrictions:** A session-scoped `NetworkPolicyDecider` controls per-domain access. The `/network allow <host>` command persists session-scoped approvals. Plan mode forces a read-only sandbox with no network access.


Citations: [crates/execpolicy/src/approval_mode.rs:1-64](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/execpolicy/src/approval_mode.rs#L1-L64) · [crates/execpolicy/src/lib.rs:1-100](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/execpolicy/src/lib.rs#L1-L100) · [crates/execpolicy/src/command_safety.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/execpolicy/src/command_safety.rs#L1-L80) · [crates/tui/src/sandbox/mod.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/sandbox/mod.rs#L1-L80) · [crates/tui/src/sandbox/bwrap.rs:1-50](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/sandbox/bwrap.rs#L1-L50) · [crates/tui/src/sandbox/policy.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/sandbox/policy.rs#L1-L60)

### Which models are supported and how are they called? (answered)

Codewhale is **provider-neutral** and supports 46+ provider kinds through a unified `LlmClient` async trait.

**Supported providers** (`config/src/provider_kind.rs`): DeepSeek, OpenAI, Anthropic (native Messages API), Google (Gemini), Mistral AI, Ollama, Ollama Cloud, Hugging Face, Together, Groq, Fireworks, Novita, OpenRouter, Grok (xAI), Meta (Muse Spark), and many more. Custom providers use the `Custom` kind for arbitrary OpenAI-compatible endpoints. Dual-wire variants (`DeepSeekAnthropic`, `MinimaxAnthropic`, `ModelstudioTokenPlanAnthropic`) select the Anthropic Messages wire protocol instead of OpenAI Chat Completions.

**Wire protocols:** Three client implementations in `crates/tui/src/client/`:
- **`chat.rs`** — OpenAI Chat Completions (the production path for most providers): streaming, request building, SSE parsing.
- **`anthropic.rs`** — Native Anthropic Messages API: handles adaptive thinking, `cache_control` breakpoints, signed-thinking signatures, usage normalization.
- **`responses.rs`** — OpenAI Responses API (for newer OpenAI endpoints).
The `LlmClient` trait (`llm_client/mod.rs:54`) defines `create_message`, `create_message_stream`, and `provider_name`/`model()` with retry logic (`with_retry`), exponential backoff, jitter, and error classification.

**Model selection:**
- **Auto model detection** (`auto_model.rs`): For DeepSeek, a rule-based classifier scores the user's prompt to choose between "deepseek-v4-flash" (simple tasks) and "deepseek-v4-pro" (complex).
- **Model catalog** (`catalog.rs:1`): Layered resolution — bundled Models.dev snapshot < live Models.dev < corrections < signed cloud facts < live provider `/v1/models` < user config overrides.
- **Fleet roster**: Role-based model assignment for sub-agents via `FleetRoster` and `FleetRole`, where each role can pin a specific provider/model.
- **Route resolver** in `config/src/route/` compiles catalog rows into executable routes with resolved `base_url`, wire protocol, and limits.

**Per-model prompt tuning:** System prompt shape adapts per mode (Plan/Act/Operate). Provider-specific request shaping includes: `max_completion_tokens` vs `max_tokens` selection (`chat.rs:46`), reasoning effort parameters (`chat.rs:67`), token limits per route, and provider-specific tool-choice modes (`strict_tool_mode` for DeepSeek beta).

**Cost tracking:** `Usage` struct with input/output/cache tokens, `CostScopeToken` for billing attribution, and route-level `RouteLimits` with per-provider pricing.

> **Editor's note.** Correction: `crates/config/src/auto_model.rs` is a legacy DeepSeek-only scorer kept for API compatibility and is not used for `model = "auto"`; auto routing is done in `crates/tui/src/model_routing.rs`, which picks a big/cheap pair for the active provider via a chat router or a TypeSafe/OpenRouter decision router.

Citations: [crates/config/src/provider_kind.rs:1-303](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/provider_kind.rs#L1-L303) · [crates/tui/src/llm_client/mod.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/llm_client/mod.rs#L1-L80) · [crates/tui/src/client/chat.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/client/chat.rs#L1-L80) · [crates/tui/src/client/anthropic.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/client/anthropic.rs#L1-L80) · [crates/config/src/auto_model.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/auto_model.rs#L1-L60) · [crates/config/src/catalog.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/config/src/catalog.rs#L1-L60)

### How can it be extended and customised? (answered)

Codewhale is extensible through multiple mechanisms, ranked by integration depth:

**MCP (Model Context Protocol)** — `crates/tui/src/mcp/`: Multiple transport backends — stdio (`stdio.rs`, newline-framed JSON-RPC over child process), SSE (`sse.rs` for server-sent events), Streamable HTTP (`streamable_http.rs`), and raw HTTP (`http.rs`). OAuth support (`oauth.rs`) for authenticating MCP servers. Server configuration via `McpServerConfig`/`McpPool`, with per-turn catalog refresh. The MCP Registry (`tools/mcp_registry.rs`) provides a `registry_sync` tool for deferred discovery — the model queries for a capability and gets scored matches from a local snapshot.

**Extension Host** (`extension_host/mod.rs:1`) — A Bun/Node process running TypeScript plugins. Each plugin can contribute: **tools** (become `ToolSpec` entries behind the same permission gate as built-in tools), **slash commands** (user-invocable `/commands`), and **additive prompt sections** (injected into the system prompt). Plugin lifecycle: activation via `extension_host:reconcile_in_background`, defers tool registration to next catalog rebuild. `core/call` gate allows extension tools to invoke core tools through the same planning/approval pipeline. The extension host runs OS-sandboxed (Seatbelt on macOS, bubblewrap on Linux) with no direct network and restricted filesystem access.

**Hooks** (`hooks/config.rs:25`) — Executable subprocess hooks fired on lifecycle events: `SessionStart`, `SessionEnd`, `MessageSubmit`, `ToolCallBefore` (can deny with exit code 2), `ToolCallAfter`, `ModeChange`, `OnError`, `TurnEnd`, `SubagentSpawn`, `SubagentComplete`, `ShellEnv` (injects env vars), `SessionIdle`. Hooks are defined in `.codewhale/hooks.toml`, require a two-step approval process (review + approve with SHA-256 digest), and symlink paths are rejected.

**Instruction files** — AGENTS.md (project-scoped per directory), CLAUDE.md (repository instructions), `.codewhale/hooks.toml`, and configurable `instructions` sources. The `load_skill` tool dynamically loads skill prompts.

**Skills** — Discoverable from local filesystem directories (`skills_dir`), loaded on demand. The default active native tools include `load_skill` so models can call skills mid-session.

**Headless/SDK** — The engine is a library (`EngineConfig` with all state) embeddable in other processes. CLI `codewhale exec` runs headless tasks from scripts/CI. npm and Cargo packages wrap the same binary. The `codewhale-protocol` crate defines shared request types for programmatic use. `AppServer` (`crates/app-server`) provides an ACP (Agent Communication Protocol) server for remote agent orchestration.

**Workflow tool** — The model-visible `workflow` tool (`tools/workflow/`) allows the model to define and run multi-step workflows with sub-agents, similar to the human-facing `Workflow` orchestration tool.

> **Editor's note.** Correction: the TypeScript extension host is experimental and only active behind `[features] extension_host` (phase 1); it should not be presented as a stable plugin mechanism.

Citations: [crates/tui/src/extension_host/mod.rs:1-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/extension_host/mod.rs#L1-L80) · [crates/tui/src/hooks/config.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/hooks/config.rs#L1-L60) · [crates/tui/src/hooks/executor.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/hooks/executor.rs#L1-L60) · [crates/tui/src/mcp/stdio.rs:1-60](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/mcp/stdio.rs#L1-L60) · [crates/tui/src/core/engine/tool_catalog.rs:52-80](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/core/engine/tool_catalog.rs#L52-L80) · [crates/tui/src/tools/workflow/mod.rs:7223-7240](https://github.com/codewhale-hq/Codewhale/blob/ecbf2869ab9110486a9e1d6d75481d1d7adf8acf/crates/tui/src/tools/workflow/mod.rs#L7223-L7240)
