LLMs Technical Reviews

codewhale-hq/Codewhale

Rust terminal coding agent with ~50 provider kinds, Plan/Act/Operate modes, approval postures, OS sandboxing and sub-agents.

GitHub ↗★ 41kRustMITcommit ecbf286 · 2026-10-05homepage ↗

Overview

Codewhale is a terminal coding agent written in Rust. It works with many model providers. One codewhale binary gives you an interactive TUI, a headless exec mode for scripts and CI, an MCP server, an ACP/app server with an HTTP/SSE runtime API, and a small web client. The same engine runs behind all of them. The npm package and the crates.io package install that same binary.

The project is large. The workspace has 28 crates plus two test crates, and almost all of the code is in crates/tui, which holds well over a million lines of Rust: the engine, tools, providers, sandbox, MCP, hooks, sub-agents and the UI. The smaller crates hold configuration (config), command policy (execpolicy), protocol and state types, hooks, memory and workflow support. The code is heavily commented with issue numbers and “known limitations” notes, which makes it easier to audit than its size suggests.

Codewhale is in the Claude Code / Codex family, and the resemblance is deliberate: it accepts other harnesses’ tool parameter spellings, reads CLAUDE.md and AGENTS.md, and copies the deferred tool_search pattern. On top of that it adds a wide provider matrix (about 50 provider kinds), Plan/Act/Operate modes, git-backed restore points, multi-agent “fleets” and a decision-model auto-router.

Architecture

flowchart LR
  CLI["codewhale CLI"] --> TUI["TUI (ratatui)"]
  CLI --> EXEC["exec / serve / app-server"]
  TUI --> ENG["Engine event loop"]
  EXEC --> ENG
  ENG --> LOOP["run_turn: 4 phases"]
  LOOP --> ROUTE["Model routing + LlmClient"]
  ROUTE --> PROV["Chat / Messages / Responses wires"]
  LOOP --> GATE["Plan tools: hooks, approval, Auto-Review"]
  GATE --> TOOLS["Native tools + tool_search"]
  TOOLS --> SBX["Seatbelt / bwrap sandbox"]
  TOOLS --> MCP["MCP pool"]
  TOOLS --> SUB["agent tool: sub-agents"]
  LOOP --> COMP["Compaction"]
  LOOP --> SNAP["Side-git snapshots"]
Component Path Role
CLI entry crates/cli/ Argument parsing and subcommands (run, exec, serve, mcp-server, app-server, doctor, …)
Engine crates/tui/src/core/engine.rs Long-lived event loop; receives Ops, owns the session, MCP boot and sub-agent completions
Turn loop crates/tui/src/core/engine/turn_loop.rs, turn_loop/ run_turn and its prepare / model / continue / tool-batch phases
Tool catalog crates/tui/src/core/engine/tool_catalog.rs Eager native tools, deferred tools, tool_search
Tools crates/tui/src/tools/ File read/write/edit, apply_patch, shell, sub-agents, workflow, memory and more
Policy crates/execpolicy/ Approval modes and command prefix rules
Sandbox crates/tui/src/sandbox/ Seatbelt on macOS, opt-in bubblewrap on Linux, sandbox policies
Providers crates/config/src/provider_kind.rs, crates/tui/src/client/, llm_client/ Provider kinds, wire clients, retries
Routing crates/tui/src/model_routing.rs model = "auto" big/cheap routing
Compaction crates/tui/src/compaction.rs, compaction/ Context pruning and summarisation with a survival contract
Hooks and extensions crates/tui/src/hooks/, extension_host/ Subprocess hooks; experimental TypeScript plugin host

How a request flows

  1. Start. main resets SIGPIPE and calls codewhale_cli::run_cli(). The interactive and headless paths share the parser and the engine entry point (main.rs). spawn_engine builds an Engine and runs it as a supervised task (engine.rs).
  2. Receive. Engine::run_owned boots MCP, starts a stall watchdog, and loops over inputs: sub-agent completions, MCP updates, shell wake-ups and operations. A user message arrives as Op::SendMessage (engine.rs).
  3. Set up the turn. run_turn starts the wall-clock budget, revalidates tools activated earlier in the conversation, checks whether the tool surface changed against the pinned prefix fingerprint (for prompt caching), and creates the retry budgets for streams, empty replies and reasoning-only replies (turn_loop.rs).
  4. Loop. Each iteration runs four phases, and any of them can retry, break or return (turn_loop.rs). prepare_model_step builds the request (catalog, steers, sub-agent results). run_model_step dispatches and streams it (model_step.rs). continue_model_step handles truncated output, and run_tool_batch_phase runs any tool calls.
  5. Tools. run_tool_batch_phase connects the MCP pool if needed, then calls plan_tool_calls (hooks, approval policy, Auto-Review, call budget, deferred-tool activation), execute_planned_tools and process_tool_results. Steers that arrived during the batch are added as user messages (tool_batch.rs).
  6. Finish. When the model stops calling tools, or a budget or cancel ends the loop, run_turn reports Completed, Failed or Interrupted. Turn-owned sub-agents keep running in the background, and the session gets a note that they will report back (turn_loop.rs).

Key components

Tool catalog

Only a few tools go on the wire eagerly: read, write, edit, bash, agent, workflow, todo_write, the goal tools and load_skill. Everything else, including MCP tools, is found through tool_search and activated for the rest of the conversation (tool_catalog.rs). This keeps the prompt prefix small and cache-stable. The comments justify each eager entry in bytes.

File editing

edit tries an exact match first, in a CRLF-normalised view. If that fails, it tries an indentation-tolerant match, then a typographic-punctuation match (smart quotes, dashes, NBSP). It refuses ambiguous matches and shows the start of the search text when nothing matches (file.rs). Writes and edits accept an optional expected_hash and refuse to write if the file changed since it was read (file.rs). Aliases such as old_string and file_path are mapped to the canonical names, and conflicting duplicates are rejected (file.rs). After a write, a syntax guard and a formatter pass run, and LSP diagnostics are attached to the result. apply_patch handles unified diffs with a default fuzz of 3 (maximum 50) and needs at least four anchor lines before it will move a hunk (apply_patch.rs).

Modes, approvals and sandbox

There are two separate controls. The mode (Plan, Agent shown as Act, Operate) sets what kind of work is allowed. The mode cycle is defined in app_mode.rs. Plan has no file-mutation authority, and the write/edit fallbacks tell the model to switch modes first (tool_catalog.rs). The approval posture (Suggest/Ask by default, Auto review, Bypass, Never) sets when to ask, and Shift+Tab cycles through the first three (approval_mode.rs). Containment depends on the platform: Seatbelt on macOS when the probe succeeds, bubblewrap on Linux only if you opt in and /usr/bin/bwrap exists, and nothing on Windows yet (sandbox/mod.rs). The code says plainly when no OS sandbox is active, and that honesty is to its credit.

Providers and auto-routing

ProviderKind has roughly 50 variants: DeepSeek (the default), OpenAI, Anthropic, OpenRouter, Ollama, vLLM, SGLang, many Chinese clouds, and dual-wire variants that speak Anthropic Messages to non-Anthropic backends (provider_kind.rs). Three wire clients cover Chat Completions, Anthropic Messages and OpenAI Responses. With model = "auto", model_routing.rs chooses between a big and a cheap model for the active provider (model_routing.rs). It decides either with a chat router model or with a “decision” router that asks TypeSafe’s hosted Jev model, directly or through OpenRouter (system_one.rs). The older keyword scorer in auto_model.rs is kept only for compatibility and is not used for provider-neutral auto routing (auto_model.rs).

Context and restore points

There is no embedding index. The model gathers context with read and search tools. Compaction is on by default. It prunes old tool output first, then summarises older turns while keeping the last user round word for word (the “survival contract”). Real thresholds come from each model’s window, with an 800K-token fallback (compaction.rs). Before a turn, around each tool call and after the turn, the workspace is snapshotted into a side git repository. /restore N and the revert_turn tool use those snapshots (turn.rs).

Extending it

  • Hooks. Subprocesses run on events such as SessionStart, MessageSubmit, ToolCallBefore/After, ModeChange, SubagentSpawn, ShellEnv (injects env vars) and SessionIdle (hooks/config.rs). Exit code 2 blocks the action (executor.rs). Project hooks run only in a trusted workspace, after /hooks review and /hooks approve <digest> for those exact bytes (hooks/config.rs).
  • MCP. Supports stdio, SSE, streamable HTTP and OAuth. Tools from connected servers join the deferred catalog. codewhale mcp-server exposes Codewhale itself over MCP.
  • Skills. SKILL.md folders are listed in the prompt and loaded with load_skill.
  • Extension host (experimental). A Bun or Node process runs TypeScript plugins behind [features] extension_host. Plugins can add tools, slash commands, admission hooks and prompt sections, all behind the core’s approval gate (extension_host/mod.rs).
  • Automation. codewhale exec runs headless. app-server serves the /v1 runtime API (sessions, threads, turns) and a mobile control page. Sub-agents and workflows can be scripted through the agent and workflow tools.

Running it

  • Install. Use the install script, npm install -g codewhale (a wrapper around the release binaries) or cargo install codewhale-cli --locked. Building from source needs Rust 1.89 or later.
  • Configure. Add a provider key or OAuth login, or point a provider at a local Ollama/vLLM/SGLang endpoint. codewhale doctor checks the setup.
  • Optional. Install bwrap and opt in to get a Linux sandbox. Install git for restore points. Bun or Node is needed only for the extension host.

Strengths and caveats

  • Strength: careful edit tooling. Hash-guarded writes, alias normalisation, staged fuzzy matching and post-edit syntax and LSP checks deal directly with the usual ways models break edits.
  • Strength: provider breadth. About 50 provider kinds, three wire protocols and local engines, with a layered model catalog.
  • Strength: honest safety reporting. Modes, approval postures and sandboxing are separate controls, and the code says when a sandbox is not in effect.
  • Caveat: no sandbox by default on Linux or Windows. On Linux, approval is your only guard unless you opt in to bubblewrap.
  • Caveat: sheer size. More than a million lines in one crate and many feature flags. Understanding or forking it is a serious investment, and behaviour changes quickly between releases.
  • Caveat: the hosted auto-router is a third party. The decision router sends routing requests to TypeSafe or OpenRouter. Use the chat router or a pinned model if that matters to you.
  • Caveat: the plugin host is experimental. TypeScript extensions are feature-gated and limited to phase-1 capabilities.

Sources: code at ecbf286, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the Open-source coding agents questions

Each answer was drafted by a code-reading agent at commit ecbf286. Its citations were checked mechanically. Compare with the other open-source coding agents →

How is the agent loop implemented?

answered

Codewhale uses a single, unified turn loop — no planner/solver separation. The authoritative turn loop is Engine::run_turn in crates/tui/src/core/engine/turn_loop.rs:1358, and a guard test (crates/core/tests/single_turn_loop.rs) fails if a second loop is introduced.

Turn structure: Each call to run_turn iterates through four sequential phases, looped until a stop condition:

  1. prepare_model_step (preparation.rs) — loads the tool catalog, refreshes MCP servers, drains sub-agent completions, processes pending steers (user mid-turn nudges), and applies soft-landing notices at ~80% of the step budget.
  2. run_model_step (model_step.rs) — dispatches the model request (OpenAI Chat Completions, Anthropic Messages API, or OpenAI Responses API depending on provider), streams the response, decodes SSE events, and handles retries/stream resumes.
  3. continue_model_step (continuation.rs) — handles output-limit truncation, RLM (reasoning language model) continuation, and prompts the model to continue from where it left off.
  4. run_tool_batch_phase (tool_batch.rs) — plans tool executions (resolves tool schemas, applies Auto-Review gates, checks permission policy) and then executes them in parallel, collecting results into the session.

Tool-call schema: Tools are registered as ToolSpec implementations with JSON Schema inputs (tools/spec/spec.rs). The catalog for a turn is built from built-in "active native tools" (read, write, edit, bash, agent, workflow, todo_write) plus deferred/discovery tools accessed via tool_search. Tool calls follow the OpenAI/Anthropic function-calling wire format, with aliases for cross-harness parameter spellings (old_string→search, file_path→path, etc.).

Stop conditions: The loop exits when: step budget exhausted (max_steps), per-turn wall-clock timeout, user cancellation, tool-call budget exhausted, consecutive empty model responses, reasoning-only retry limit, no-progress detection from repeated permission denials, or the model stops emitting tool calls.

Sub-agents: Spawned via the agent tool (tools/subagent/mod.rs:11506). The model calls agent with an action (Start, Roster, Status, Message, Followup, etc.), and a background agent session runs with its own filtered toolset and workspace context. Foreground (turn-owned) sub-agents are tracked via ForegroundChildRegistry and cancelled when the turn ends. Background sub-agents complete asynchronously and deliver results via runtime messages. A fleet roster (FleetRoster) provides role-based model assignment for spawned agents.

How is repository context gathered and kept within the context window?

answered

Codewhale gathers repository context through explicit tool calls rather than automated crawling or embedding-based retrieval. There are no embeddings or vector indexes — the model reads files it chooses.

File reading: The read/ReadTool and legacy ReadFileTool (tools/file.rs:1092) stream file contents with: small-file fast-path (whole file in one call when ≤4KB and ≤500 lines), line-range windows (start_line/max_lines), byte-size awareness with truncation footers naming the next offset, content hashing (SHA-256) reported in responses so the model can validate later edits. Binary files route to OCR (images) or PDF extraction.

Search tools: search_content (Grep) and search_name (FileSearch) provide pattern-based discovery (tools/search.rs, tools/file_search.rs). A deferred tool_search tool (tool_catalog.rs:35) with regex/BM25 matching allows discovery of tools and files by name or content. Web search is available via web_search.

System prompt assembly: The system prompt (prompts/text.rs:44, BASE_PROMPT) is composed from layers: constitution (binding core), personality overlay (currently CALM), approval-policy overlays, memory block (user memories from runtime/src/native_memory.rs), skills index, MCP server instructions, mode-specific instructions (Plan/Act/Operate), and project instruction files (CLAUDE.md/AGENTS.md). The session-pinned prefix is kept KV-cache stable — volatile facts (steers, completions, diagnostics) are appended as user-role messages, never spliced into the frozen prefix.

Compaction for long sessions (compaction.rs:1): When context pressure exceeds a model-specific token threshold (80% of context window), compaction triggers. The strategy is: first prune old tool results from the tail (cheapest, no LLM call), then if still over threshold, run an LLM summarization pass that compresses older user+assistant turns while keeping the last user round verbatim (the survival contract in survival_contract.rs, enforced by last_round.rs). Recent plain user messages are retained up to a configurable token budget (retained_user_message_tokens). The pre-compaction history is persisted as an artifact before any mutation. Context pressure warnings (warning at 70%, critical at 85%) are emitted mid-turn via context_pressure_message.

Memory: User memories (tools/native_memory.rs, tools/remember.rs) provide a persistent capture path — the model can ask to remember things, which are stored on disk and prepended to the system prompt on subsequent sessions.

How are code edits applied?

answered

Codewhale provides three file-editing mechanisms with gradual complexity, plus git-based undo:

1. write (WriteFileTool, file.rs:1755): Full file replacement. Accepts path, content, and optional expected_hash. Creates parent directories if needed, preserves line-ending style (CRLF/LF) on overwrite, and writes atomically via run_blocking_write_atomic. The optional content-hash guard (verify_expected_hash, file.rs:89-108) refuses the write if the file changed since it was last read, preventing stale-state overwrites.

2. edit (EditFileTool, file.rs:2581): Search-and-replace with fuzzy matching. Accepts path, search, replace, and optional expected_hash. Matching is multi-layered: exact match first, then indentation-tolerant fuzzy match, then typographic-punctuation (smart quotes, em-dashes) normalization. Returns detailed error messages when search/replace are identical (a common model mistake). After the edit a Rust code formatter (normalize_edit, rust_format.rs) is applied, and LSP diagnostics are injected into the result via lsp_diagnostics_for_paths.

3. apply_patch (ApplyPatchTool, apply_patch.rs:1): Unified diff patching supporting multi-hunk patches with fuzzy matching (MAX_FUZZ=50, default fuzz=3). Hunks can be relocated when their stated line numbers are stale (requires ≥4 anchor lines for safety via MIN_ANCHOR_LINES). Preserves CRLF line endings and trailing-newline state from the base file.

Cross-harness parameter compatibility: All file tools translate known spellings from other coding harnesses (old_string/new_string, file_path, filePath, offset/limit for read windows) onto the canonical names via apply_param_aliases (file.rs:228-255). This prevents rejection loops where the model's training data names parameters differently. Unknown parameters are rejected with explicit errors rather than silently dropped.

Validation pipeline: All mutations go through guard_edit (syntax_check.rs) for syntax validation, normalize_edit for code formatting, and the content-hash guard. Auto-Review gates can flag edits for user approval based on file path patterns.

Git/undo: Workspace snapshots are taken into a side git repository (turn.rs:8-16) at pre-turn, pre-tool, and post-tool boundaries. The /restore N command and revert_turn tool consume these snapshots. The git_tool provides direct git operations (commit, push, branch creation).

How are shell commands and file writes kept safe?

answered

Execution safety in Codewhale is implemented at multiple layers: approval modes, command safety classification, OS sandboxing, and filesystem guards.

Approval modes (execpolicy/src/approval_mode.rs:7): Four modes — Suggest (ask for non-safe tools, the default), Auto (auto-review risky calls), Bypass (no approvals/"YOLO mode"), and Never (block all tools requiring approval). The user cycles through Suggest → Auto → Bypass with Shift+Tab. ApprovalMode::from_config_value accepts many CLI aliases ("yolo", "bypass-permissions", "full-access", etc.).

Command safety (execpolicy/src/command_safety.rs:1): A COMMAND_ARITY dictionary maps ~40+ command prefixes (git status, npm run, cargo check, docker compose, etc.) to their canonical forms. Ruleset objects (execpolicy/src/lib.rs:29) at three priority layers (BuiltinDefault → Agent → User) define trusted_prefixes (auto-allowed), denied_prefixes (always blocked), and ask_rules (typed tool-invocation patterns requiring approval). Permission actions are Allow/Ask/Deny. The command-safety system classifies commands into SafetyLevels, and the ExecPolicyContext uses these to determine whether a shell invocation requires approval.

OS sandbox (sandbox/mod.rs:1):

  • macOS: Seatbelt (sandbox-exec) when runtime probe succeeds — uses the native macOS sandbox.
  • Linux: Bubblewrap (bwrap) only when user opts-in (prefer_bwrap: true) and /usr/bin/bwrap is executable. Creates a new mount namespace with --unshare-all, read-only bind of /, tmpfs /tmp, /dev, /proc, and writable roots per policy. --share-net is added for network-enabled policies. The extension host always uses bwrap when available without the opt-in.
  • Windows: Process-tree containment via Job Object (planned, not yet sandboxed).
  • SandboxPolicy (sandbox/policy.rs:22) has four levels: DangerFullAccess (no restrictions), ReadOnly (read-only everywhere), ExternalSandbox (already sandboxed, skip double-wrapping), WorkspaceWrite (read-only + writable workspace + optional extra roots + optional network).

The sandbox is never silently downgraded — if bwrap isn't available on Linux, the tool reports no OS sandbox rather than pretending.

Filesystem guards: A read deny-list (tools/file.rs:458) protects credential files, config paths, and sensitive system files. Paths are checked on both raw spelling (before symlink resolution) and canonical form (after). The content-hash guard prevents stale-state overwrites.

Network restrictions: A session-scoped NetworkPolicyDecider controls per-domain access. The /network allow <host> command persists session-scoped approvals. Plan mode forces a read-only sandbox with no network access.

Which models are supported and how are they called?

answered

Codewhale is provider-neutral and supports 46+ provider kinds through a unified LlmClient async trait.

Supported providers (config/src/provider_kind.rs): DeepSeek, OpenAI, Anthropic (native Messages API), Google (Gemini), Mistral AI, Ollama, Ollama Cloud, Hugging Face, Together, Groq, Fireworks, Novita, OpenRouter, Grok (xAI), Meta (Muse Spark), and many more. Custom providers use the Custom kind for arbitrary OpenAI-compatible endpoints. Dual-wire variants (DeepSeekAnthropic, MinimaxAnthropic, ModelstudioTokenPlanAnthropic) select the Anthropic Messages wire protocol instead of OpenAI Chat Completions.

Wire protocols: Three client implementations in crates/tui/src/client/:

  • chat.rs — OpenAI Chat Completions (the production path for most providers): streaming, request building, SSE parsing.
  • anthropic.rs — Native Anthropic Messages API: handles adaptive thinking, cache_control breakpoints, signed-thinking signatures, usage normalization.
  • responses.rs — OpenAI Responses API (for newer OpenAI endpoints). The LlmClient trait (llm_client/mod.rs:54) defines create_message, create_message_stream, and provider_name/model() with retry logic (with_retry), exponential backoff, jitter, and error classification.

Model selection:

  • Auto model detection (auto_model.rs): For DeepSeek, a rule-based classifier scores the user's prompt to choose between "deepseek-v4-flash" (simple tasks) and "deepseek-v4-pro" (complex).
  • Model catalog (catalog.rs:1): Layered resolution — bundled Models.dev snapshot < live Models.dev < corrections < signed cloud facts < live provider /v1/models < user config overrides.
  • Fleet roster: Role-based model assignment for sub-agents via FleetRoster and FleetRole, where each role can pin a specific provider/model.
  • Route resolver in config/src/route/ compiles catalog rows into executable routes with resolved base_url, wire protocol, and limits.

Per-model prompt tuning: System prompt shape adapts per mode (Plan/Act/Operate). Provider-specific request shaping includes: max_completion_tokens vs max_tokens selection (chat.rs:46), reasoning effort parameters (chat.rs:67), token limits per route, and provider-specific tool-choice modes (strict_tool_mode for DeepSeek beta).

Cost tracking: Usage struct with input/output/cache tokens, CostScopeToken for billing attribution, and route-level RouteLimits with per-provider pricing.

Editor's note. Correction: crates/config/src/auto_model.rs is a legacy DeepSeek-only scorer kept for API compatibility and is not used for model = "auto"; auto routing is done in crates/tui/src/model_routing.rs, which picks a big/cheap pair for the active provider via a chat router or a TypeSafe/OpenRouter decision router.

How can it be extended and customised?

answered

Codewhale is extensible through multiple mechanisms, ranked by integration depth:

MCP (Model Context Protocol) — crates/tui/src/mcp/: Multiple transport backends — stdio (stdio.rs, newline-framed JSON-RPC over child process), SSE (sse.rs for server-sent events), Streamable HTTP (streamable_http.rs), and raw HTTP (http.rs). OAuth support (oauth.rs) for authenticating MCP servers. Server configuration via McpServerConfig/McpPool, with per-turn catalog refresh. The MCP Registry (tools/mcp_registry.rs) provides a registry_sync tool for deferred discovery — the model queries for a capability and gets scored matches from a local snapshot.

Extension Host (extension_host/mod.rs:1) — A Bun/Node process running TypeScript plugins. Each plugin can contribute: tools (become ToolSpec entries behind the same permission gate as built-in tools), slash commands (user-invocable /commands), and additive prompt sections (injected into the system prompt). Plugin lifecycle: activation via extension_host:reconcile_in_background, defers tool registration to next catalog rebuild. core/call gate allows extension tools to invoke core tools through the same planning/approval pipeline. The extension host runs OS-sandboxed (Seatbelt on macOS, bubblewrap on Linux) with no direct network and restricted filesystem access.

Hooks (hooks/config.rs:25) — Executable subprocess hooks fired on lifecycle events: SessionStart, SessionEnd, MessageSubmit, ToolCallBefore (can deny with exit code 2), ToolCallAfter, ModeChange, OnError, TurnEnd, SubagentSpawn, SubagentComplete, ShellEnv (injects env vars), SessionIdle. Hooks are defined in .codewhale/hooks.toml, require a two-step approval process (review + approve with SHA-256 digest), and symlink paths are rejected.

Instruction files — AGENTS.md (project-scoped per directory), CLAUDE.md (repository instructions), .codewhale/hooks.toml, and configurable instructions sources. The load_skill tool dynamically loads skill prompts.

Skills — Discoverable from local filesystem directories (skills_dir), loaded on demand. The default active native tools include load_skill so models can call skills mid-session.

Headless/SDK — The engine is a library (EngineConfig with all state) embeddable in other processes. CLI codewhale exec runs headless tasks from scripts/CI. npm and Cargo packages wrap the same binary. The codewhale-protocol crate defines shared request types for programmatic use. AppServer (crates/app-server) provides an ACP (Agent Communication Protocol) server for remote agent orchestration.

Workflow tool — The model-visible workflow tool (tools/workflow/) allows the model to define and run multi-step workflows with sub-agents, similar to the human-facing Workflow orchestration tool.

Editor's note. Correction: the TypeScript extension host is experimental and only active behind [features] extension_host (phase 1); it should not be presented as a stable plugin mechanism.