openai/codex
OpenAI's Rust terminal coding agent: a Responses-API turn loop with apply_patch edits and OS-level sandboxes for every shell command.
Overview
Codex is OpenAI’s terminal coding agent. At this commit it is almost entirely a Rust workspace (codex-rs/, more than 110 top-level crates). The npm package in codex-cli/ is a thin launcher for the native binary. You type a request in the TUI or pass it to codex exec. The agent then streams a model response, runs the shell commands and patches the model asks for, and loops until the model stops asking for tools.
Two design choices set Codex apart from most agents in this category. First, it speaks only one wire protocol. WireApi has a single variant, the OpenAI Responses API, and the built-in providers are OpenAI, Amazon Bedrock, Ollama and LM Studio (model-provider-info/src/lib.rs, L659-L691). Second, it puts its safety budget into OS sandboxing rather than prompts: shell commands run under Seatbelt on macOS, bubblewrap plus seccomp on Linux, and a restricted token on Windows.
The tool set is small. The model gets exec_command (a PTY-backed shell) and write_stdin, apply_patch, update_plan, view_image, MCP tools, and the multi-agent tools. There is no dedicated read, grep or list tool. The model reads the repository through the shell.
Architecture
flowchart LR
TUI["TUI / codex exec"] --> ASC["app-server client (in-process)"]
IDE["IDE / desktop app"] --> AS["app-server JSON-RPC"]
ASC --> AS
AS --> CORE["codex-core Session"]
CORE --> TASK["RegularTask -> run_turn"]
TASK --> MC["ModelClientSession (Responses API, WebSocket or SSE)"]
TASK --> TR["ToolCallRuntime / ToolRouter"]
TR --> ORCH["ToolOrchestrator: approve -> sandbox -> run"]
ORCH --> SBX["Seatbelt / bwrap+seccomp / Windows token"]
TR --> AP["apply-patch crate"]
TR --> MCP["codex-mcp servers"]
TR --> SUB["spawn_agent / send_message / wait_agent"]
CORE --> HIST["history, compaction, rollout files"]
| Component | Path | Role |
|---|---|---|
| CLI entry | codex-rs/cli/src/main.rs |
Subcommands: exec, review, mcp, plugin, app-server, sandbox, execpolicy, apply, resume |
| TUI | codex-rs/tui/ |
Interactive terminal UI; a client of the app-server protocol |
| Headless | codex-rs/exec/ |
codex exec: prints the final message, or JSONL events with --json |
| App server | codex-rs/app-server/, app-server-protocol/ |
JSON-RPC surface used by the TUI, exec, IDE extensions and the desktop app |
| Core | codex-rs/core/src/session/ |
Session state, run_turn, sampling, compaction triggers |
| Tools | codex-rs/core/src/tools/ |
Router, registry, parallel runtime, approval orchestrator, handlers |
| Patch engine | codex-rs/apply-patch/ |
Parser and applier for the *** Begin Patch format |
| Sandboxing | codex-rs/sandboxing/, linux-sandbox/, windows-sandbox-rs/ |
Per-platform process isolation |
| Exec policy | codex-rs/execpolicy/, core/src/exec_policy.rs |
.rules files that allow, prompt or forbid command prefixes |
| Providers | codex-rs/model-provider*/, models-manager/ |
Provider catalog, auth, model metadata |
How a request flows
- Task start. A user turn becomes a
RegularTask. Itsrunmethod callsrun_turnin a loop (tasks/regular.rs). - Pre-turn work.
run_turnchecks pending Guardian input, drains async hook results, and runs a pre-sampling compaction before it records the new input (session/turn.rs). - Step loop. Each pass through the
loopdrains steering input the user typed while the model was working, runs hooks on it, and captures oneStepContext. Context, the advertised tools and the tool calls all share this one view (turn.rs). The prompt input is the cloned history filtered for the model’s input modalities (turn.rs). - Sampling.
run_sampling_requestbuilds aToolCallRuntimeand the prompt, then retriestry_run_sampling_requestwithin the provider’s stream retry budget (turn.rs). That function opens the stream throughModelClientSession::streamwith the step’s model, reasoning effort and service tier (turn.rs). - Tool dispatch. For every
OutputItemDone,handle_output_item_doneasksToolRouter::build_tool_callwhether the item is a tool call. If it is, the item is persisted at once and a tool future is queued, and the turn is markedneeds_follow_up(stream_events_utils.rs).ToolCallRuntime::handle_tool_callruns it, in parallel when the tool supports that (tools/parallel.rs). - Approval and sandbox. Shell and patch runtimes go through
ToolOrchestrator: approval, then sandbox selection, then the attempt, then a retry with an escalated sandbox if the sandbox denied it (orchestrator.rs). - Continue or stop. After sampling, the loop checks pending input and the token status. If a follow-up is needed and the context limit was hit, it runs
run_auto_compactand continues. If no follow-up is needed, it runs Stop hooks, which can block the stop and force another pass (turn.rs).
Key components
Transport to the model
ModelClient lives for the session. A ModelClientSession lives for one turn. It caches a Responses WebSocket connection, prewarms it with a generate=false request so the next call can reuse previous_response_id, and keeps a sticky-routing token. On failure it falls back to the normal stream retry path (client.rs). Third-party providers are added through model_providers in config.toml, but they must also speak Responses. The comment in built_in_model_providers says plainly that OpenAI does not want to bundle third-party providers.
apply_patch
Edits use a custom patch grammar rather than search/replace: *** Begin Patch, then Add File, Delete File or Update File hunks with optional Move to, @@ context lines and +/-/space lines (apply-patch/src/parser.rs). The parser is deliberately lenient about whitespace around markers. A streaming parser lets the UI show the diff while the model is still writing it, and TurnDiffTracker collects the turn’s net diff.
Sandboxing
get_platform_sandbox picks MacosSeatbelt, LinuxSeccomp or WindowsRestrictedToken (the last only when enabled). Tools declare a SandboxablePreference of Auto, Require or Forbid (sandboxing/src/manager.rs). On Linux, bubblewrap enforces the filesystem view. The in-process helpers add no_new_privs and seccomp for network isolation, and Landlock remains only as a legacy fallback (linux-sandbox/src/landlock.rs).
Approvals and exec policy
AskForApproval has four values, of which three can be configured: on-request (the default; the model decides when to ask), granular (per-category switches, where false means auto-reject) and never. The fourth, untrusted (prompt unless an exec-policy rule allows the command), is retired as a setting — config rejects approval_policy = "untrusted" (config/mod.rs#L228-L231). Enum: (protocol.rs). Exec-policy rules live in rules/*.rules files with a default.rules (exec_policy.rs). codex-shell-command adds a dangerous-command classifier, for example for forced rm. A separate Guardian subsystem can have a model review an action before it is approved.
Context and instructions
AGENTS.md discovery walks up from the working directory to a project root marker (.git by default). It then concatenates every AGENTS.md from the root down to the working directory and does not go past the root (agents_md.rs). AGENTS.override.md takes precedence locally. Long sessions are kept in budget by auto-compaction, both inline (the model summarises the history) and remote, through the provider’s compaction endpoint.
Multi-agent
The v2 multi-agent handlers expose spawn_agent, send_message and wait_agent. Child agents are separate threads managed by the session’s agent control, not a separate planner. The main loop is still a single model-driven loop.
Extending it
- MCP.
codex mcpmanages servers. Thecodex-mcpcrate connects to them, and their tools reach the model through the MCP handler, along with resource tools. - Hooks. Events are
PreToolUse,PermissionRequest,PostToolUse,PreCompact,PostCompact,SessionStart,SessionEnd,UserPromptSubmit,SubagentStart,SubagentStop,StopandInterrupt(protocol.rs). A hook can be a command, an MCP tool or a prompt, configured inconfig.tomlor ahooks.json. - Skills and plugins. Bundled sample skills include
skill-creator,skill-installer,imagegen,openai-docsandreview-agent.codex pluginand a marketplace command manage plugins. - Embedding.
codex app-serverexposes the full JSON-RPC protocol.codex exec --jsongives a JSONL event stream for CI. Thesdk/folder wraps it for other languages.
Running it
- Install. Use the install script, a GitHub release binary, or the npm/Homebrew packages. Sign in with a ChatGPT account or set an API key.
--ossswitches to a local Ollama or LM Studio server, which must expose a Responses-compatible endpoint. - Use. Run
codexfor the TUI,codex exec "..."for one-shot non-interactive runs,codex reviewfor code review, andcodex resumeto continue a saved session. - Build from source. Cargo works inside
codex-rs/. CI and releases use Bazel (MODULE.bazel,BUILD.bazel).
Strengths and caveats
- Strength: real OS sandboxing. Few agents in this category isolate shell commands at the kernel level on all three desktop platforms. Codex does, and it retries with an escalated sandbox instead of failing silently.
- Strength: one protocol, many front ends. The TUI,
exec, the IDE extensions and the desktop app all go through the same app-server protocol, so behaviour is consistent. - Strength: mature turn engine. Steering input mid-turn, Stop hooks that can veto a stop, WebSocket prewarm and incremental requests all show production polish.
- Caveat: Responses API only. No Chat Completions or Anthropic wire. Non-OpenAI models need a Responses-compatible endpoint or a proxy.
- Caveat: large surface. More than 110 crates, plus feature flags for code mode, Guardian, realtime voice and cloud tasks. Reading or forking the core is a serious investment.
- Caveat: the shell is the reader. Without dedicated read or search tools, context gathering depends on the model writing good
rgandsedcommands. Each one then passes through the approval and sandbox path.
Sources: code at ccde2fc, verified Q&A.
How it answers the Open-source coding agents questions
Each answer was drafted by a code-reading agent at commit ccde2fc. Its citations were checked mechanically. Compare with the other open-source coding agents →
How is the agent loop implemented?
answeredCodex CLI uses a single-loop architecture without a separate planner; the model itself plans and acts within each turn. The core loop lives in run_turn() in core/src/session/turn.rs:164 which accepts TurnInput (user messages, function-call outputs, inter-agent communications) and runs a loop { ... } (line 427) that repeats until the model produces a final assistant message with no pending input. Inside the loop, the session builds a prompt from its cloned history (line 527), invokes the OpenAI Responses API via run_sampling_request() (line 535), streams the response, and executes any tool calls the model makes through the ToolCallRuntime (line 1633). After each sampling round the loop checks for pending user input (line 431), token-limit thresholds for auto-compaction (line 612), and hook stop/block requests (line 663–719). Tool calls are dispatched through try_run_sampling_request() (line 1679) which invokes the ToolRouter — a registry that maps tool names to CoreToolRuntime implementations in core/src/tools/registry.rs. Sub-agents are supported via a multi-agent V2 system in core/src/tools/handlers/multi_agents_v2.rs:1 where tools like spawn_agent, send_message, and wait let the model create and communicate with child threads, coordinated by the AgentControl trait in core/src/agent/api.rs:41. The turn's tool-call schema is defined by the Responses API's native function-calling format, with tool definitions served by the ToolRouter and filtered per-step based on permissions and feature flags.
How is repository context gathered and kept within the context window?
answeredRepository context is built through a layered World State system in core/src/context/world_state/mod.rs:64. Each turn constructs a WorldState by merging context fragments from many sources: base instructions, model instructions, AGENTS.md instructions (loaded in core/src/agents_md.rs:58 by walking up from the project root to find files named AGENTS.md), environment configurations, plugin/skill instructions, and permission profiles. These fragments are placed as Prefix, Standalone, or Mergeable items (line 70–75) into the model's input. File search uses a fuzzy-file-search tool and a ripgrep-based content search tool (exposed through the unified exec tool surface). For context-window management, Codex CLI implements auto-compaction in core/src/compact.rs:1. When the token count exceeds the model's limit — tracked via ContextWindowTokenStatus in core/src/session/context_window.rs:8 — run_auto_compact() (called from turn.rs at lines 625, 734, 785) sends the conversation history to the model with a summarization prompt (SUMMARIZATION_PROMPT from prompts). The two main compaction strategies are inline auto-compact (line 625) and remote auto-compact v2 (imported from compact_remote_v2). Session history is stored via the thread store and cloned for each model request, with compaction producing CompactedHistoryMetadata that replaces verbose history. The WorldState also tracks per-section context diffs and token budgets, emitting context-window guidance fragments when approaching limits.
How are code edits applied?
answeredCodex CLI applies code edits through a dedicated apply-patch crate at codex-rs/apply-patch/src/lib.rs:1. The model sends edits as structured function calls containing patch chunks, which are parsed by the parser module into UpdateFileChunk objects (exported from line 27). The core logic in file_update.rs:25 reads the target file, computes replacements against the original content using compute_replacements() (line 60), and produces the new file contents. The patch format supports change-context matching via seek_sequence::seek_sequence() (line 72) which locates context lines before applying edits, allowing robust search/replace operations. The ApplyPatchFileUpdate struct (exported from line 33) tracks what changed. Edits go through the approval pipeline defined in core/src/tools/approvals.rs:67 where tool calls like ExecCommand and ApplyPatchFileUpdate require user approval based on the configured AskForApproval policy. Validation is handled through the Guardian review system: the tool orchestrator in core/src/tools/orchestrator.rs:1 runs approval → sandbox-selection → execution with retry escalation. There is no built-in linter integration; however, edits are recorded as file diffs and tracked via TurnDiffTracker for display. Undo is supported through Git — Codex discovers the project's Git repository and can diff against HEAD or the remote (the GitDiffToRemoteParams schema suggests git-aware diffing). The approval system also caches decisions per session via ApprovalStore (line 42 in sandboxing.rs) so repeated identical edits don't reprompt.
How are shell commands and file writes kept safe?
answeredExecution safety is enforced through multi-layer sandboxing with approval gates. On Linux, commands run inside bubblewrap (bwrap) containers for filesystem isolation, combined with seccomp filters for network restriction (implemented in codex-rs/linux-sandbox/src/landlock.rs:1). The bwrap launcher in codex-rs/linux-sandbox/src/linux_run_main.rs:1 creates a sandboxed process with controlled filesystem mounts, network namespaces, and process isolation. On macOS, the seatbelt sandbox is used (codex-rs/sandboxing/src/lib.rs:9). On Windows, restricted-token sandboxes apply. There are three sandboxable preferences: Auto, Require, and Forbid (codex-rs/sandboxing/src/manager.rs:42). The exec policy system in core/src/exec_policy.rs:1 evaluates commands against allow/deny rule files (*.rules files, line 55) loaded from the project, with a default.rules file installed by default. Dangerous command detection in codex-shell-command identifies known dangerous patterns (e.g., rm -rf /, > /dev/sda). The approval system in core/src/tools/approvals.rs:54 uses AskForApproval policies — Always, Auto, Never, or Granular — configured per model and environment. Before executing any shell command, the tool orchestrator (core/src/tools/orchestrator.rs:55) runs: approval decision → sandbox selection → execution → retry with escalated sandboxing on denial. Network restrictions enforce NetworkSandboxPolicy from PermissionProfile, which can block, allow-list, or monitor network access. Hook scripts (configured via settings.json hooks) can intercept tool calls before and after execution, adding a programmable policy layer. The Guardian auto-review system (core/src/guardian/) can independently review actions for approval.
untrusted, on-request (default), granular and never (codex-rs/protocol/src/protocol.rs L961-L984), not Always/Auto; and hooks are configured in config TOML or hooks.json, not settings.json.Which models are supported and how are they called?
answeredCodex CLI connects primarily to OpenAI's Responses API as the default provider, with model catalog fetched from api.openai.com and cached in models_cache.json (managed by codex-rs/models-manager/src/manager.rs:34). The ModelProvider trait in codex-rs/model-provider/src/provider.rs:108 defines the abstraction for model backends, with implementations for: OpenAI's production API (the default, via CoreAuthProvider in lib.rs:27), Amazon Bedrock (amazon_bedrock in lib.rs:17), LM Studio (codex-rs/lmstudio/), and Ollama (codex-rs/ollama/). The ModelsEndpointClient trait (manager.rs:42) handles remote catalog fetching, while static and OpenAI models managers serve the bundled catalog (bundled_models_response() in lib.rs:13 from a bundled models.json). Model selection uses a pinned model per session, with the option to switch mid-conversation. The ModelInfo struct carries per-model metadata: context window size, reasoning-effort defaults, input/output modalities, auto-compact token limits, and ToolMode (standard, code-mode-only, or code-mode — core/src/tools/mod.rs:75). Models are called through the OpenAI Responses API streaming endpoints via the ModelClientSession in core/src/client.rs:1, which manages WebSocket connections for sticky routing, retries with exponential backoff (core/src/responses_retry.rs:33), and transport fallback. Per-model prompts are rendered via render_model_instructions() in prompts/src/model_instructions.rs:8, which uses the model's instruction template from its metadata. Cost tracking uses TokenUsageInfo collected per turn and aggregated in analytics events. Reasonable efforts (ReasoningEffort) controls are configurable per-turn.
How can it be extended and customised?
answeredCodex CLI is extensible through multiple mechanisms. MCP (Model Context Protocol) is the primary extension system: the codex-mcp crate in codex-rs/codex-mcp/src/lib.rs:1 implements MCP server management, tool catalog caching, and resource access. MCP servers are declared in config via McpServerConfig entries and discovered through the configured_mcp_servers() function, with their tools surfaced to the model through McpHandler (core/src/tools/handlers/mcp.rs). Skills (a plugin system) let users install task-focused packages — sample skills in codex-rs/skills/src/assets/samples/ include skill-installer, skill-creator, and imagegen — each with scripts and configurations. AGENTS.md files (core/src/agents_md.rs:42) provide per-project instructions: Codex discovers AGENTS.md (and AGENTS.override.md) by walking up from the project root to the current directory, concatenating their contents as system instructions. A fallback project_doc_fallback_filenames config option supports additional filenames. Hooks (core/src/hook_runtime.rs:1) support lifecycle callbacks at session start, user prompt submit, pre/post tool use, compact, and stop events. Hooks can inject context, block/allow actions, and request user permission. Plugins extend tool availability: core/src/plugins/ manages plugin discovery and installation, with tools like RequestPluginInstallHandler for runtime plugin requests. The config system (TOML-based with layered config in core/src/config/) allows full customization of model selection, permissions, approval policies, sandboxing levels, MCP server configuration, and environment definitions. Headless/SDK mode is supported via the exec crate which exposes Codex as an app-server with JSON-RPC (ClientRequest/ServerNotification in exec/src/lib.rs) — consumable by IDE extensions and automated tooling. The codex-cli/bin/package.json enables npm distribution as @openai/codex.
app-server crate (codex app-server); the exec crate is the non-interactive codex exec CLI, which itself talks to an in-process app-server.