LLMs Technical Reviews

openai/codex

OpenAI's Rust terminal coding agent: a Responses-API turn loop with apply_patch edits and OS-level sandboxes for every shell command.

GitHub ↗★ 128kRustApache-2.0commit ccde2fc · 2026-10-06

Overview

Codex is OpenAI’s terminal coding agent. At this commit it is almost entirely a Rust workspace (codex-rs/, more than 110 top-level crates). The npm package in codex-cli/ is a thin launcher for the native binary. You type a request in the TUI or pass it to codex exec. The agent then streams a model response, runs the shell commands and patches the model asks for, and loops until the model stops asking for tools.

Two design choices set Codex apart from most agents in this category. First, it speaks only one wire protocol. WireApi has a single variant, the OpenAI Responses API, and the built-in providers are OpenAI, Amazon Bedrock, Ollama and LM Studio (model-provider-info/src/lib.rs, L659-L691). Second, it puts its safety budget into OS sandboxing rather than prompts: shell commands run under Seatbelt on macOS, bubblewrap plus seccomp on Linux, and a restricted token on Windows.

The tool set is small. The model gets exec_command (a PTY-backed shell) and write_stdin, apply_patch, update_plan, view_image, MCP tools, and the multi-agent tools. There is no dedicated read, grep or list tool. The model reads the repository through the shell.

Architecture

flowchart LR
  TUI["TUI / codex exec"] --> ASC["app-server client (in-process)"]
  IDE["IDE / desktop app"] --> AS["app-server JSON-RPC"]
  ASC --> AS
  AS --> CORE["codex-core Session"]
  CORE --> TASK["RegularTask -> run_turn"]
  TASK --> MC["ModelClientSession (Responses API, WebSocket or SSE)"]
  TASK --> TR["ToolCallRuntime / ToolRouter"]
  TR --> ORCH["ToolOrchestrator: approve -> sandbox -> run"]
  ORCH --> SBX["Seatbelt / bwrap+seccomp / Windows token"]
  TR --> AP["apply-patch crate"]
  TR --> MCP["codex-mcp servers"]
  TR --> SUB["spawn_agent / send_message / wait_agent"]
  CORE --> HIST["history, compaction, rollout files"]
Component Path Role
CLI entry codex-rs/cli/src/main.rs Subcommands: exec, review, mcp, plugin, app-server, sandbox, execpolicy, apply, resume
TUI codex-rs/tui/ Interactive terminal UI; a client of the app-server protocol
Headless codex-rs/exec/ codex exec: prints the final message, or JSONL events with --json
App server codex-rs/app-server/, app-server-protocol/ JSON-RPC surface used by the TUI, exec, IDE extensions and the desktop app
Core codex-rs/core/src/session/ Session state, run_turn, sampling, compaction triggers
Tools codex-rs/core/src/tools/ Router, registry, parallel runtime, approval orchestrator, handlers
Patch engine codex-rs/apply-patch/ Parser and applier for the *** Begin Patch format
Sandboxing codex-rs/sandboxing/, linux-sandbox/, windows-sandbox-rs/ Per-platform process isolation
Exec policy codex-rs/execpolicy/, core/src/exec_policy.rs .rules files that allow, prompt or forbid command prefixes
Providers codex-rs/model-provider*/, models-manager/ Provider catalog, auth, model metadata

How a request flows

  1. Task start. A user turn becomes a RegularTask. Its run method calls run_turn in a loop (tasks/regular.rs).
  2. Pre-turn work. run_turn checks pending Guardian input, drains async hook results, and runs a pre-sampling compaction before it records the new input (session/turn.rs).
  3. Step loop. Each pass through the loop drains steering input the user typed while the model was working, runs hooks on it, and captures one StepContext. Context, the advertised tools and the tool calls all share this one view (turn.rs). The prompt input is the cloned history filtered for the model’s input modalities (turn.rs).
  4. Sampling. run_sampling_request builds a ToolCallRuntime and the prompt, then retries try_run_sampling_request within the provider’s stream retry budget (turn.rs). That function opens the stream through ModelClientSession::stream with the step’s model, reasoning effort and service tier (turn.rs).
  5. Tool dispatch. For every OutputItemDone, handle_output_item_done asks ToolRouter::build_tool_call whether the item is a tool call. If it is, the item is persisted at once and a tool future is queued, and the turn is marked needs_follow_up (stream_events_utils.rs). ToolCallRuntime::handle_tool_call runs it, in parallel when the tool supports that (tools/parallel.rs).
  6. Approval and sandbox. Shell and patch runtimes go through ToolOrchestrator: approval, then sandbox selection, then the attempt, then a retry with an escalated sandbox if the sandbox denied it (orchestrator.rs).
  7. Continue or stop. After sampling, the loop checks pending input and the token status. If a follow-up is needed and the context limit was hit, it runs run_auto_compact and continues. If no follow-up is needed, it runs Stop hooks, which can block the stop and force another pass (turn.rs).

Key components

Transport to the model

ModelClient lives for the session. A ModelClientSession lives for one turn. It caches a Responses WebSocket connection, prewarms it with a generate=false request so the next call can reuse previous_response_id, and keeps a sticky-routing token. On failure it falls back to the normal stream retry path (client.rs). Third-party providers are added through model_providers in config.toml, but they must also speak Responses. The comment in built_in_model_providers says plainly that OpenAI does not want to bundle third-party providers.

apply_patch

Edits use a custom patch grammar rather than search/replace: *** Begin Patch, then Add File, Delete File or Update File hunks with optional Move to, @@ context lines and +/-/space lines (apply-patch/src/parser.rs). The parser is deliberately lenient about whitespace around markers. A streaming parser lets the UI show the diff while the model is still writing it, and TurnDiffTracker collects the turn’s net diff.

Sandboxing

get_platform_sandbox picks MacosSeatbelt, LinuxSeccomp or WindowsRestrictedToken (the last only when enabled). Tools declare a SandboxablePreference of Auto, Require or Forbid (sandboxing/src/manager.rs). On Linux, bubblewrap enforces the filesystem view. The in-process helpers add no_new_privs and seccomp for network isolation, and Landlock remains only as a legacy fallback (linux-sandbox/src/landlock.rs).

Approvals and exec policy

AskForApproval has four values, of which three can be configured: on-request (the default; the model decides when to ask), granular (per-category switches, where false means auto-reject) and never. The fourth, untrusted (prompt unless an exec-policy rule allows the command), is retired as a setting — config rejects approval_policy = "untrusted" (config/mod.rs#L228-L231). Enum: (protocol.rs). Exec-policy rules live in rules/*.rules files with a default.rules (exec_policy.rs). codex-shell-command adds a dangerous-command classifier, for example for forced rm. A separate Guardian subsystem can have a model review an action before it is approved.

Context and instructions

AGENTS.md discovery walks up from the working directory to a project root marker (.git by default). It then concatenates every AGENTS.md from the root down to the working directory and does not go past the root (agents_md.rs). AGENTS.override.md takes precedence locally. Long sessions are kept in budget by auto-compaction, both inline (the model summarises the history) and remote, through the provider’s compaction endpoint.

Multi-agent

The v2 multi-agent handlers expose spawn_agent, send_message and wait_agent. Child agents are separate threads managed by the session’s agent control, not a separate planner. The main loop is still a single model-driven loop.

Extending it

  • MCP. codex mcp manages servers. The codex-mcp crate connects to them, and their tools reach the model through the MCP handler, along with resource tools.
  • Hooks. Events are PreToolUse, PermissionRequest, PostToolUse, PreCompact, PostCompact, SessionStart, SessionEnd, UserPromptSubmit, SubagentStart, SubagentStop, Stop and Interrupt (protocol.rs). A hook can be a command, an MCP tool or a prompt, configured in config.toml or a hooks.json.
  • Skills and plugins. Bundled sample skills include skill-creator, skill-installer, imagegen, openai-docs and review-agent. codex plugin and a marketplace command manage plugins.
  • Embedding. codex app-server exposes the full JSON-RPC protocol. codex exec --json gives a JSONL event stream for CI. The sdk/ folder wraps it for other languages.

Running it

  • Install. Use the install script, a GitHub release binary, or the npm/Homebrew packages. Sign in with a ChatGPT account or set an API key. --oss switches to a local Ollama or LM Studio server, which must expose a Responses-compatible endpoint.
  • Use. Run codex for the TUI, codex exec "..." for one-shot non-interactive runs, codex review for code review, and codex resume to continue a saved session.
  • Build from source. Cargo works inside codex-rs/. CI and releases use Bazel (MODULE.bazel, BUILD.bazel).

Strengths and caveats

  • Strength: real OS sandboxing. Few agents in this category isolate shell commands at the kernel level on all three desktop platforms. Codex does, and it retries with an escalated sandbox instead of failing silently.
  • Strength: one protocol, many front ends. The TUI, exec, the IDE extensions and the desktop app all go through the same app-server protocol, so behaviour is consistent.
  • Strength: mature turn engine. Steering input mid-turn, Stop hooks that can veto a stop, WebSocket prewarm and incremental requests all show production polish.
  • Caveat: Responses API only. No Chat Completions or Anthropic wire. Non-OpenAI models need a Responses-compatible endpoint or a proxy.
  • Caveat: large surface. More than 110 crates, plus feature flags for code mode, Guardian, realtime voice and cloud tasks. Reading or forking the core is a serious investment.
  • Caveat: the shell is the reader. Without dedicated read or search tools, context gathering depends on the model writing good rg and sed commands. Each one then passes through the approval and sandbox path.

Sources: code at ccde2fc, verified Q&A.

How it answers the Open-source coding agents questions

Each answer was drafted by a code-reading agent at commit ccde2fc. Its citations were checked mechanically. Compare with the other open-source coding agents →

How is the agent loop implemented?

answered

Codex CLI uses a single-loop architecture without a separate planner; the model itself plans and acts within each turn. The core loop lives in run_turn() in core/src/session/turn.rs:164 which accepts TurnInput (user messages, function-call outputs, inter-agent communications) and runs a loop { ... } (line 427) that repeats until the model produces a final assistant message with no pending input. Inside the loop, the session builds a prompt from its cloned history (line 527), invokes the OpenAI Responses API via run_sampling_request() (line 535), streams the response, and executes any tool calls the model makes through the ToolCallRuntime (line 1633). After each sampling round the loop checks for pending user input (line 431), token-limit thresholds for auto-compaction (line 612), and hook stop/block requests (line 663–719). Tool calls are dispatched through try_run_sampling_request() (line 1679) which invokes the ToolRouter — a registry that maps tool names to CoreToolRuntime implementations in core/src/tools/registry.rs. Sub-agents are supported via a multi-agent V2 system in core/src/tools/handlers/multi_agents_v2.rs:1 where tools like spawn_agent, send_message, and wait let the model create and communicate with child threads, coordinated by the AgentControl trait in core/src/agent/api.rs:41. The turn's tool-call schema is defined by the Responses API's native function-calling format, with tool definitions served by the ToolRouter and filtered per-step based on permissions and feature flags.

How is repository context gathered and kept within the context window?

answered

Repository context is built through a layered World State system in core/src/context/world_state/mod.rs:64. Each turn constructs a WorldState by merging context fragments from many sources: base instructions, model instructions, AGENTS.md instructions (loaded in core/src/agents_md.rs:58 by walking up from the project root to find files named AGENTS.md), environment configurations, plugin/skill instructions, and permission profiles. These fragments are placed as Prefix, Standalone, or Mergeable items (line 70–75) into the model's input. File search uses a fuzzy-file-search tool and a ripgrep-based content search tool (exposed through the unified exec tool surface). For context-window management, Codex CLI implements auto-compaction in core/src/compact.rs:1. When the token count exceeds the model's limit — tracked via ContextWindowTokenStatus in core/src/session/context_window.rs:8 — run_auto_compact() (called from turn.rs at lines 625, 734, 785) sends the conversation history to the model with a summarization prompt (SUMMARIZATION_PROMPT from prompts). The two main compaction strategies are inline auto-compact (line 625) and remote auto-compact v2 (imported from compact_remote_v2). Session history is stored via the thread store and cloned for each model request, with compaction producing CompactedHistoryMetadata that replaces verbose history. The WorldState also tracks per-section context diffs and token budgets, emitting context-window guidance fragments when approaching limits.

How are code edits applied?

answered

Codex CLI applies code edits through a dedicated apply-patch crate at codex-rs/apply-patch/src/lib.rs:1. The model sends edits as structured function calls containing patch chunks, which are parsed by the parser module into UpdateFileChunk objects (exported from line 27). The core logic in file_update.rs:25 reads the target file, computes replacements against the original content using compute_replacements() (line 60), and produces the new file contents. The patch format supports change-context matching via seek_sequence::seek_sequence() (line 72) which locates context lines before applying edits, allowing robust search/replace operations. The ApplyPatchFileUpdate struct (exported from line 33) tracks what changed. Edits go through the approval pipeline defined in core/src/tools/approvals.rs:67 where tool calls like ExecCommand and ApplyPatchFileUpdate require user approval based on the configured AskForApproval policy. Validation is handled through the Guardian review system: the tool orchestrator in core/src/tools/orchestrator.rs:1 runs approval → sandbox-selection → execution with retry escalation. There is no built-in linter integration; however, edits are recorded as file diffs and tracked via TurnDiffTracker for display. Undo is supported through Git — Codex discovers the project's Git repository and can diff against HEAD or the remote (the GitDiffToRemoteParams schema suggests git-aware diffing). The approval system also caches decisions per session via ApprovalStore (line 42 in sandboxing.rs) so repeated identical edits don't reprompt.

How are shell commands and file writes kept safe?

answered

Execution safety is enforced through multi-layer sandboxing with approval gates. On Linux, commands run inside bubblewrap (bwrap) containers for filesystem isolation, combined with seccomp filters for network restriction (implemented in codex-rs/linux-sandbox/src/landlock.rs:1). The bwrap launcher in codex-rs/linux-sandbox/src/linux_run_main.rs:1 creates a sandboxed process with controlled filesystem mounts, network namespaces, and process isolation. On macOS, the seatbelt sandbox is used (codex-rs/sandboxing/src/lib.rs:9). On Windows, restricted-token sandboxes apply. There are three sandboxable preferences: Auto, Require, and Forbid (codex-rs/sandboxing/src/manager.rs:42). The exec policy system in core/src/exec_policy.rs:1 evaluates commands against allow/deny rule files (*.rules files, line 55) loaded from the project, with a default.rules file installed by default. Dangerous command detection in codex-shell-command identifies known dangerous patterns (e.g., rm -rf /, > /dev/sda). The approval system in core/src/tools/approvals.rs:54 uses AskForApproval policies — Always, Auto, Never, or Granular — configured per model and environment. Before executing any shell command, the tool orchestrator (core/src/tools/orchestrator.rs:55) runs: approval decision → sandbox selection → execution → retry with escalated sandboxing on denial. Network restrictions enforce NetworkSandboxPolicy from PermissionProfile, which can block, allow-list, or monitor network access. Hook scripts (configured via settings.json hooks) can intercept tool calls before and after execution, adding a programmable policy layer. The Guardian auto-review system (core/src/guardian/) can independently review actions for approval.

Editor's note. Correction: the approval policies are untrusted, on-request (default), granular and never (codex-rs/protocol/src/protocol.rs L961-L984), not Always/Auto; and hooks are configured in config TOML or hooks.json, not settings.json.

Which models are supported and how are they called?

answered

Codex CLI connects primarily to OpenAI's Responses API as the default provider, with model catalog fetched from api.openai.com and cached in models_cache.json (managed by codex-rs/models-manager/src/manager.rs:34). The ModelProvider trait in codex-rs/model-provider/src/provider.rs:108 defines the abstraction for model backends, with implementations for: OpenAI's production API (the default, via CoreAuthProvider in lib.rs:27), Amazon Bedrock (amazon_bedrock in lib.rs:17), LM Studio (codex-rs/lmstudio/), and Ollama (codex-rs/ollama/). The ModelsEndpointClient trait (manager.rs:42) handles remote catalog fetching, while static and OpenAI models managers serve the bundled catalog (bundled_models_response() in lib.rs:13 from a bundled models.json). Model selection uses a pinned model per session, with the option to switch mid-conversation. The ModelInfo struct carries per-model metadata: context window size, reasoning-effort defaults, input/output modalities, auto-compact token limits, and ToolMode (standard, code-mode-only, or code-mode — core/src/tools/mod.rs:75). Models are called through the OpenAI Responses API streaming endpoints via the ModelClientSession in core/src/client.rs:1, which manages WebSocket connections for sticky routing, retries with exponential backoff (core/src/responses_retry.rs:33), and transport fallback. Per-model prompts are rendered via render_model_instructions() in prompts/src/model_instructions.rs:8, which uses the model's instruction template from its metadata. Cost tracking uses TokenUsageInfo collected per turn and aggregated in analytics events. Reasonable efforts (ReasoningEffort) controls are configurable per-turn.

How can it be extended and customised?

answered

Codex CLI is extensible through multiple mechanisms. MCP (Model Context Protocol) is the primary extension system: the codex-mcp crate in codex-rs/codex-mcp/src/lib.rs:1 implements MCP server management, tool catalog caching, and resource access. MCP servers are declared in config via McpServerConfig entries and discovered through the configured_mcp_servers() function, with their tools surfaced to the model through McpHandler (core/src/tools/handlers/mcp.rs). Skills (a plugin system) let users install task-focused packages — sample skills in codex-rs/skills/src/assets/samples/ include skill-installer, skill-creator, and imagegen — each with scripts and configurations. AGENTS.md files (core/src/agents_md.rs:42) provide per-project instructions: Codex discovers AGENTS.md (and AGENTS.override.md) by walking up from the project root to the current directory, concatenating their contents as system instructions. A fallback project_doc_fallback_filenames config option supports additional filenames. Hooks (core/src/hook_runtime.rs:1) support lifecycle callbacks at session start, user prompt submit, pre/post tool use, compact, and stop events. Hooks can inject context, block/allow actions, and request user permission. Plugins extend tool availability: core/src/plugins/ manages plugin discovery and installation, with tools like RequestPluginInstallHandler for runtime plugin requests. The config system (TOML-based with layered config in core/src/config/) allows full customization of model selection, permissions, approval policies, sandboxing levels, MCP server configuration, and environment definitions. Headless/SDK mode is supported via the exec crate which exposes Codex as an app-server with JSON-RPC (ClientRequest/ServerNotification in exec/src/lib.rs) — consumable by IDE extensions and automated tooling. The codex-cli/bin/package.json enables npm distribution as @openai/codex.

Editor's note. Correction: the JSON-RPC server for IDEs and other clients is the app-server crate (codex app-server); the exec crate is the non-interactive codex exec CLI, which itself talks to an in-process app-server.