LLMs Technical Reviews
Home / Open-source coding agents / openinterpreter

openinterpreter/openinterpreter

Rust fork of OpenAI Codex that re-creates other agents' prompts, tools and wire formats ("harnesses") to get more out of open models.

GitHub ↗★ 69kRustApache-2.0commit 2767e5f · 2026-10-02homepage ↗

Overview

The current Open Interpreter has nothing in common with the original Python “let the LLM run code” project except the name. At the pinned commit the repository is a fork of OpenAI’s Codex CLI. The Rust workspace in codex-rs/ still uses Codex crate, protocol and config names. FORK_BRANDING.md explicitly tells forks not to rename codex globally and to change only the product identity in codex-rs/product-info. The installed commands are interpreter and i, and the README says the binary can stand in for codex under the Codex SDK.

The fork’s own idea is harness emulation. The authors’ bet is that a cheap model works best inside the harness it was trained or tuned with. So Open Interpreter re-implements the system prompts, tool schemas, tool names and request shapes of other agents on top of the Codex engine: Claude Code, Kimi Code, Kimi CLI, Z.ai’s ZCode, Qwen Code, DeepSeek-TUI, OpenCode, Pi, SWE-agent, mini-swe-agent, Terminus 2, little-coder and a minimal one. /harness switches between them, and a provider/model heuristic picks a default. Everything underneath is inherited from Codex: the session and turn loop, approvals, the OS sandboxes, MCP, hooks, rollouts and the app server.

Architecture

flowchart LR
  U["interpreter TUI / exec / ACP / SDK"] --> SES["Session + RegularTask"]
  SES --> TURN["run_turn loop"]
  TURN --> SAMP["run_sampling_request"]
  SAMP --> MC["ModelClientSession.stream"]
  MC --> ROUTE["resolve_stream_transport_route"]
  ROUTE --> RESP["Responses API (native)"]
  ROUTE --> CHAT["Chat harness request builders"]
  ROUTE --> MSG["Messages harness (claude-code, zcode)"]
  SAMP -->|"tool calls"| ROUTER["ToolRouter"]
  ROUTER --> NATIVE["Codex tools: unified_exec, apply_patch, MCP"]
  ROUTER --> ALIAS["Harness aliases: Read, Edit, Bash, Grep..."]
  NATIVE --> ORCH["Orchestrator: approval, sandbox, retry"]
  ALIAS --> ORCH
  ORCH --> SBX["Seatbelt / bwrap+Landlock / Windows token"]
Component Path Role
CLI codex-rs/cli/src/main.rs Subcommands: TUI (default), exec, review, mcp, mcp-server, acp, app-server, plugin
Task and turn loop codex-rs/core/src/tasks/regular.rs, core/src/session/turn.rs Outer task loop, per-turn sampling/tool loop, compaction triggers
Model client codex-rs/core/src/client.rs Picks a transport route per wire API and harness, then streams
Harness selection codex-rs/tools/src/harness.rs, core/src/harness/routing.rs Harness enum and the (wire API, harness) → transport table
Harness implementations codex-rs/core/src/harness/*.rs Per-harness system prompts, tool JSON and request builders
Harness tool aliases codex-rs/core/src/tools/handlers/harness_aliases.rs Claude-style Read/Edit/Bash/Grep/Task*, DeepSeek-TUI and ZCode tools
Provider info codex-rs/model-provider-info/src/lib.rs Providers, WireApi, default harness per provider/model
Tool orchestration codex-rs/core/src/tools/orchestrator.rs Approval → sandbox selection → attempt → escalated retry
Sandboxing codex-rs/sandboxing/, linux-sandbox/, windows-sandbox-rs/ Seatbelt profiles, bubblewrap + Landlock, Windows restricted tokens
Patching codex-rs/apply-patch/ Codex’s *** Begin Patch format, parser and applier

How a request flows

  1. Task. A user message starts a RegularTask. Its run() calls run_turn in a loop for as long as new input is queued, and stops on a terminal error (regular.rs).
  2. Turn. run_turn first runs pre-sampling compaction if needed (turn.rs). It then loops: drain steering input, run user-prompt hooks, and sample (turn.rs).
  3. Pick a transport. ModelClientSession::stream asks resolve_stream_transport_route(wire_api, harness). A native harness uses the Responses API (WebSocket with HTTP fallback). A Chat harness such as qwen-code or kimi-code goes to its own request builder. claude-code and zcode on an Anthropic-style endpoint use the Messages route (client.rs, routing.rs).
  4. Shape the request. For a Chat harness, build_chat_harness_request dispatches to the harness module. The Qwen Code harness, for example, sends the Qwen Code system prompt, a fake “startup context” user message with date, OS and a folder listing, a canned “Got it” assistant reply, and the Kimi-style tool list, with max_tokens: 8000 (qwen_code.rs, request.rs). The chat SSE stream is converted back into Codex ResponseEvents, so the rest of the engine does not know which harness is active.
  5. Dispatch tools. As output items finish, handle_output_item_done queues tool futures and sets needs_follow_up (turn.rs). Native calls go to Codex handlers. Harness tool names such as Read, Edit, Bash or exec_shell go to HarnessAliasHandler (harness_aliases.rs).
  6. Execute safely. Shell work goes through the orchestrator: ask for approval if the policy needs it, pick a sandbox, run, and retry with escalated permissions on denial without asking again (orchestrator.rs). An alias such as Claude-style Bash is translated into a unified_exec payload and takes the same route (harness_aliases.rs).
  7. Loop or stop. If any tool ran, the turn samples again. When the model answers without tool calls and no input is pending, the turn ends. Stop hooks and context-window compaction can step in on the way out.

Key components

Harness routing

The Harness enum has fifteen named variants plus Other(String) (harness.rs). The routing table makes the wire API part of the choice. Most emulated harnesses accept only wire_api = "chat". claude-code works over Responses, Chat or Messages. zcode uses Messages. Invalid pairs fail with an explicit error instead of degrading silently. When nothing is configured, default_harness_for_provider_model picks one: zcode for the Z.ai Anthropic endpoint, claude-code for Anthropic-style providers, kimi-code for Kimi/Moonshot, qwen-code for Qwen/DashScope, and claude-code-bare for DeepSeek, with a comment that this “gets the best results out of DeepSeek models” (lib.rs).

Harness tools

Each emulated harness exposes its own tool names. HarnessAliasHandler maps several dozen of them onto Codex machinery (harness_aliases.rs). The Claude-style Edit reproduces Claude Code’s semantics: exact old_string match, an error on multiple matches unless replace_all is set, and (in the bare and ZCode profiles) “File has not been read yet” when there was no prior Read. It writes the file in-process after a filesystem-policy check and emits Codex file-change events for the UI (harness_aliases.rs). In the native harness, edits use Codex’s apply_patch envelope instead.

Context

Project instructions come from AGENTS.md files gathered from the project root (found by .git markers) down to the working directory (agents_md.rs). Compaction runs before a turn, in the middle of a turn when the window overflows, and after a turn. It uses either a local summarisation prompt or the provider’s remote compaction. The rebuilt history keeps up to 20,000 tokens of recent user messages next to the summary (compact.rs). Kimi CLI brings its own compaction prompt. Across sessions, the inherited Codex memories pipeline extracts raw memories from past rollouts and consolidates them into a MEMORY.md handbook plus a short summary that is injected on later runs (memories README).

Sandboxing

get_platform_sandbox returns Seatbelt on macOS, the Linux sandbox (bubblewrap plus Landlock/seccomp helpers) on Linux, and a restricted token on Windows when enabled (manager.rs). Policies come from permission profiles with separate filesystem and network rules. These are Codex’s sandboxes, unchanged in design.

Extending it

  • MCP. Configure servers with interpreter mcp. The agent can itself be served over MCP (mcp-server) or over ACP (interpreter acp) for editors (main.rs).
  • Hooks. Twelve events, among them PreToolUse, PermissionRequest, PostToolUse, Pre/PostCompact, SessionStart/End, SubagentStart/Stop, Stop and Interrupt (hooks lib.rs).
  • Plugins and skills. Inherited Codex plugin manager and marketplace commands, plus skills directories.
  • A new harness. Add a Harness variant, a route in routing.rs, a request builder with its prompt and tool JSON in core/src/harness/, and any missing tool aliases.
  • SDKs. sdk/python and sdk/typescript, compatible with the Codex exec protocol.

Running it

  • Install with the project’s shell or PowerShell installer, then run interpreter (or i) in a repository.
  • Configure providers and models in config.toml. Set wire_api (responses, chat or messages) per provider, and optionally a harness, or switch with /harness during a session.
  • interpreter exec runs headless. interpreter review runs a non-interactive code review.
  • Building from source uses Cargo or Bazel. The workspace is large (well over a hundred crates).

Strengths and caveats

  • Strength: a mature engine underneath. The approvals, sandboxes, rollouts, compaction and app-server protocol are Codex’s, and are far more developed than in most open-model agents.
  • Strength: harness choice as a tuning knob. Being able to run a model inside the prompt and tool set it was tuned for is a real lever for open models, and the provider-based defaults encode the authors’ measurements.
  • Strength: compatibility. It works with Codex SDK clients and with ACP editors.
  • Caveat: emulation drifts. Each harness is a hand-written copy of someone else’s moving target (claude_code.rs alone is about 4,900 lines). Fidelity depends on how often the copies are updated.
  • Caveat: two edit paths. Native mode edits through apply_patch. The emulated harnesses write files in-process through alias handlers. Behaviour, error messages and read-before-write rules differ by harness.
  • Caveat: wire constraints. Most emulated harnesses need a Chat Completions endpoint and refuse messages, so not every provider/harness pair is possible.
  • Caveat: inherited naming. Config, crates and protocol still say codex, which confuses anyone reading the source or the logs.

Sources: code at 2767e5f, deepwiki-open wiki (16 pages), verified Q&A.

How it answers the Open-source coding agents questions

Each answer was drafted by a code-reading agent at commit 2767e5f. Its citations were checked mechanically. Compare with the other open-source coding agents →

How is the agent loop implemented?

answered

The agent loop is two-layer: an outer turn loop and an inner sampling loop. RegularTask (core/src/tasks/regular.rs:31-126) is the SessionTask implementation. Its run() emits TurnStarted, runs turn lifecycle, then enters a loop calling run_turn() while sess.input_queue.has_pending_input() returns true (line 120). Inside run_turn() (core/src/session/turn.rs:163-817), the inner loop builds a StepContext with a ToolRouter and calls run_sampling_request() (line 1592). The Prompt struct (line 1558) contains cloned history, model_visible_specs(), base_instructions, and output_schema. The model streams through ModelClientSession.stream() (line 2491). There is no separate planner -- plan mode (ModeKind::Plan) is a tool-mode flag on the ToolRouter enabling PlanDelta streaming within the same request. Sub-agents use multi-agent collaboration (core/src/tools/handlers/multi_agents.rs:1-78): spawn_agent/send_message create child threads via LocalAgentControl (core/src/agent/control.rs). AgentRegistry (core/src/agent/registry.rs:26-28) tracks agents and enforces depth/session limits. Stop conditions: terminal errors (line 117), no pending input (line 120), turn-stop hooks (line 646-691), context-window compaction (line 602).

How is repository context gathered and kept within the context window?

answered

Context is managed through compaction, AGENTS.md discovery, and per-turn assembly. The StepContext captures world state, environment snapshot, tool router, MCP binding, and model info. History is cloned from the session and filtered by model input modalities (turn.rs:513-519). Pre-turn compaction via run_pre_sampling_compact() (turn.rs:183-221) calls run_auto_compact() (compact.rs:1446-1504) which delegates to local summarization or remote compaction v2 depending on the provider's RemoteCompactionSupport. Mid-turn compaction triggers on context window exceeded (turn.rs:614-641); post-turn compaction can also run (line 706-735). The summary is capped at 20,000 tokens (compact.rs:62). Repository context comes from AGENTS.md files (core/src/agents_md.rs:1-18): walking from project root (found via .git markers) to cwd, concatenating into user instructions. Fallback: AGENTS.override.md. File-search tools (core/src/tools/handlers/tool_search.rs:28-33) provide BM25 search over tools. There are no vector embeddings or repo maps -- context is prompt concatenation and compaction.

Editor's note. Correction: the 20,000-token constant (COMPACT_USER_MESSAGE_MAX_TOKENS, compact.rs L62) caps the recent user messages retained next to the summary in the rebuilt history, not the summary length.

How are code edits applied?

answered

Code edits use the apply_patch system. The model calls apply_patch (freeform format with inline patches in shell commands). ApplyPatchHandler (core/src/tools/handlers/apply_patch.rs:79-81) routes patch commands. The core library (apply-patch/src/lib.rs) implements parse_patch() and StreamingPatchParser for unified diffs, plus maybe_parse_apply_patch() for extracting patches from shell commands using tree-sitter Bash grammar (apply-patch/src/invocation.rs:1-100). ApplyPatchFileChange supports Add, Delete, and Update (unified-diff hunks with optional move_path). These map to protocol FileChange::Add/Delete/Update (core/src/apply_patch.rs:67-84). Patch streaming sends PatchApplyUpdatedEvent deltas as hunks parse (handler lines 121-158). Validation: assess_patch_safety() checks permission profiles and sandbox routes. The ToolOrchestrator (core/src/tools/orchestrator.rs:1-60) sequences approval to sandbox selection to attempt to retry. Edits run inside the environment filesystem with FileSystemSandboxContext enforcement. There is no lint/test validation step or git integration; safety relies on sandboxing and approvals.

Editor's note. Correction: apply_patch is the native harness's edit path only; emulated harnesses (claude-code, zcode, deepseek-tui, …) use alias tools such as a Claude-style Edit that does exact-string replacement in-process (core/src/tools/handlers/harness_aliases.rs handle_edit).

How are shell commands and file writes kept safe?

answered

Safety uses approval policy plus platform sandboxing plus permission profiles. Approvals module (core/src/tools/approvals.rs:1-60) routes decisions through GuardianReviewContext, with ApprovalStore caching per-session (core/src/tools/sandboxing.rs:41-63). ToolOrchestrator (core/src/tools/orchestrator.rs:47-60) drives approval to sandbox selection to attempt to retry. Platform sandboxing is OS-specific. Linux (sandboxing/src/bwrap.rs:1-50) uses bubblewrap. macOS (seatbelt.rs:1-60) uses /usr/bin/sandbox-exec with SBPL policies for base/network. Windows uses WindowsRestrictedToken or WindowsMxc. SandboxManager (manager.rs:48-62) resolves get_platform_sandbox() by OS. Landlock (landlock.rs) provides Linux landlock enforcement. Permission profiles (FileSystemSandboxPolicy, NetworkSandboxPolicy) control read/write/exec. policy_transforms.rs computes effective policies. Shell/file commands go through unified_exec (core/src/tools/handlers/unified_exec.rs) with configurable timeouts and sandbox-permissions overrides. Allow/deny use ExecPolicyAmendment and ApprovalPolicy. Guardian module adds per-reviewer budget. No checkpoint/restore; sandbox is per-command.

Which models are supported and how are they called?

answered

Models use a provider abstraction: ModelProvider (model-provider/src/provider.rs:42-73) with ProviderCapabilities (namespace_tools, image_generation, web_search, remote_compaction). Catalog: models-manager/models.json (1663 lines) with slugs, context_window, max_context_window, reasoning levels (low/medium/high/xhigh/max), tool_mode, input_modalities, comp_hash. ModelsManager (models-manager/src/manager.rs) provides OpenAiModelsManager and StaticModelsManager backed by ModelsCache. Provider resolution via create_model_provider dispatches to Amazon Bedrock, OpenAI/Anthropic, or combined auth (model-provider/src/provider.rs:39). bundled_provider_catalog (model-provider-info/src/bundled_provider_catalog.rs) serves offline data. Model calls go through ModelClient (core/src/client.rs); turn-scoped ModelClientSession caches WebSocket connections. Per-turn settings (model, reasoning_effort, service_tier) pass to client_session.stream(). Tool calling: Prompt.tools = step_context.tool_router.model_visible_specs() (turn.rs:1566). Cost tracking via SessionTelemetry and AnalyticsEventsClient.

Editor's note. Correction: the answer misses the harness layer: resolve_stream_transport_route maps (wire_api = responses|chat|messages, harness) to a transport and per-harness request builder (core/src/harness/routing.rs), and default_harness_for_provider_model picks claude-code, kimi-code, qwen-code, zcode or claude-code-bare from the provider/model.

How can it be extended and customised?

answered

MCP (Model Context Protocol) is the primary extension mechanism: McpHandler (core/src/tools/handlers/mcp.rs:51-60) wraps MCP tools into ToolSpec for the ToolRouter. codex-mcp module (codex-rs/codex-mcp/src/) manages connections, bindings, catalogs, and discovery. mcp-server/ provides MCP server runtime with tool config and exec approval. Plugins: PluginsManager with build_plugin_injections() (turn.rs:1099), PluginDeclaration (hooks/src/declarations.rs). Hooks: hooks/src/lib.rs:23-36 supports 12 events (PreToolUse, PostToolUse, PreCompact, PostCompact, SessionStart, SessionEnd, UserPromptSubmit, SubagentStart, SubagentStop, Stop, Interrupt, PermissionRequest). Hooks run as subprocesses and can block or modify state. AGENTS.md files (core/src/agents_md.rs:42-49): per-project instructions searched from project root to cwd. Default: AGENTS.md, override: AGENTS.override.md. SDK/headless mode via Python (sdk/python/) and TypeScript (sdk/typescript/) SDKs. Multiple SessionTask types (Regular, Review, UserShell). ExtensionData and TurnInputContributor allow runtime extension injection.