LLMs Technical Reviews

How are shell commands and file writes kept safe?

Approval modes; sandboxing (containers, seatbelt, landlock); allow/deny lists; network restrictions; checkpoints.

Verdict

Codex, and Open Interpreter, which inherits its engine, put every shell command in an OS sandbox on macOS, Linux and Windows. DeepSeek-Reasonix also jails bash and fails closed. The app builders get their safety from where the code runs.

OS sandbox as the default path.

  • Codex: approval, then sandbox selection (Seatbelt, bubblewrap plus seccomp, or a restricted token), then a retry with an escalated sandbox. The default approval policy is on-request.
  • DeepSeek-Reasonix: Seatbelt or bubblewrap, and it refuses to run bash where there is no backend (Windows). However, an Ask rule resolves to allow when no approver is attached.

Sandbox available, but conditional.

  • Qwen Code: bwrap or Landlock for shell commands, and Seatbelt, Docker or Podman for the whole CLI. auto mode puts a regex filter and a fail-closed LLM classifier in front of risky calls.
  • Codewhale: Seatbelt on macOS. Bubblewrap on Linux is opt-in, and Windows has no sandbox.

Approval only, with no OS sandbox.

  • Aider: shell prompts need an explicit yes even with --yes-always.
  • Cline: the CLI auto-approves every tool by default.
  • OpenCode: the default ruleset is "*": "allow". Only repeated identical calls, external directories and .env reads ask.
  • PI-Desktop: a Rust host process enforces risk tiers and concurrency budgets, but bash runs as you, and auto mode allows everything.

Isolated by architecture (app builders).

  • bolt.diy: code runs in a WebContainer in the browser tab, and MCP tools run only after an explicit approval. /api/chat has no rate limiting.
  • Open Lovable: an E2B or Vercel cloud sandbox with no approvals at all. Creating a sandbox kills every other one.
  • VibeSDK: the default Think engine has no shell (workspaceBash = false), and each app gets its own Durable Objects. preDeploySafetyGate checks React render-loop hazards, not secrets, and only on the legacy engines.

Pick: Codex for kernel-level isolation on all three desktop platforms. Pick: DeepSeek-Reasonix or Qwen Code for open-model agents with a real shell jail. Pick: bolt.diy if generated code must never leave the user’s browser.

More projects in this category are being researched.

Per-project answers

anomalyco/opencode

answered

Execution Safety

OpenCode uses a ruleset-based permission system with wildcard pattern matching. The Permission service (permission/index.ts) evaluates each tool call against a stack of rulesets. Rules have action in {"allow", "deny", "ask"}. For each tool execution, evaluate() (line 28-38) scans rulesets in reverse order and returns the first matching rule (defaulting to "ask"). If the action is "allow", execution proceeds silently. If "deny", a DeniedError is thrown. If "ask", a pending permission request is created and the user must reply (ask(), line 67-107). Approvals can be "once" or "always" (stored in the session's approved list for future auto-allow). Denials can include user feedback as a CorrectedError (line 125).

Permission rulesets are defined per-agent in agent/agent.ts. The build agent (default) allows most tools but "ask" for doom_loop, external directories, and .env file reads (lines 119-155). The plan agent denies all edit/write tools except plan files (lines 156-181). The explore agent denies everything except read-only tools (grep, glob, list, bash, read, webfetch, websearch) (lines 197-213). Custom agents and user-configured permissions merge on top of these defaults (lines 267-294).

Shell command safety (tool/shell.ts): The shell tool detects dangerous command types — it has sets for CWD (cd, pushd etc.), FILES (rm, cp, mv, chmod, chown, etc.), and potentially destructive commands. Output is truncated via the Truncate service. Timeouts are configurable (default 300s).

External directory checks: The assertExternalDirectoryEffect function (tool/external-directory.ts) prevents writing outside the project worktree unless explicitly allowed by the permission ruleset.

Doom loop detection in processor.ts (lines 353-381) checks if the same tool has been called with the same input 3+ times consecutively, and requires explicit user permission to continue.

Sandboxing/containers: There is no built-in container sandbox or Landlock/Seccomp integration visible in the codebase. Execution safety relies on the permission system operating at the tool-call level, not OS-level isolation. The Question tool provides structured input from the user. The codebase does not implement network egress filtering or allow/deny lists for network access beyond what the permission system covers.

Checkpoints via Snapshot (snapshot/index.ts): Before each LLM step, a git-based snapshot is taken (track()). On step completion, the diff is captured (patch()). The revert system allows rolling back snapshots.

Editor's note. Correction: the shell tool has no dangerous-command detector; its CWD/FILES sets only find paths outside the project for external_directory prompts, and each tree-sitter-parsed sub-command becomes a bash permission pattern. The default ruleset is "*": "allow", so shell and edits run unprompted unless configured.

openai/codex

answered

Execution safety is enforced through multi-layer sandboxing with approval gates. On Linux, commands run inside bubblewrap (bwrap) containers for filesystem isolation, combined with seccomp filters for network restriction (implemented in codex-rs/linux-sandbox/src/landlock.rs:1). The bwrap launcher in codex-rs/linux-sandbox/src/linux_run_main.rs:1 creates a sandboxed process with controlled filesystem mounts, network namespaces, and process isolation. On macOS, the seatbelt sandbox is used (codex-rs/sandboxing/src/lib.rs:9). On Windows, restricted-token sandboxes apply. There are three sandboxable preferences: Auto, Require, and Forbid (codex-rs/sandboxing/src/manager.rs:42). The exec policy system in core/src/exec_policy.rs:1 evaluates commands against allow/deny rule files (*.rules files, line 55) loaded from the project, with a default.rules file installed by default. Dangerous command detection in codex-shell-command identifies known dangerous patterns (e.g., rm -rf /, > /dev/sda). The approval system in core/src/tools/approvals.rs:54 uses AskForApproval policies — Always, Auto, Never, or Granular — configured per model and environment. Before executing any shell command, the tool orchestrator (core/src/tools/orchestrator.rs:55) runs: approval decision → sandbox selection → execution → retry with escalated sandboxing on denial. Network restrictions enforce NetworkSandboxPolicy from PermissionProfile, which can block, allow-list, or monitor network access. Hook scripts (configured via settings.json hooks) can intercept tool calls before and after execution, adding a programmable policy layer. The Guardian auto-review system (core/src/guardian/) can independently review actions for approval.

Editor's note. Correction: the approval policies are untrusted, on-request (default), granular and never (codex-rs/protocol/src/protocol.rs L961-L984), not Always/Auto; and hooks are configured in config TOML or hooks.json, not settings.json.

cline/cline

answered

Cline layers multiple safety mechanisms for shell commands and file writes. Tool-level approval is mediated by requestToolApproval callback on RuntimeBuilderInput; every tool call can be presented to the user for approval or rejection (runtime-builder.ts:87-89). Tool policies define per-tool enablement: the yolo preset auto-approves all tools, while the default preset requires user approval (presets.ts:133-143). The plan-mode command guard (command-guard-extension.ts) registers a beforeTool hook that blocks file-editing shell commands (rm, mv, cp, sed -i, git commit, redirects, etc.) from the run_commands tool when in plan mode, without prompting the user — the model receives the error directly (command-guard.ts:23-100). MCP server sandboxing validates server registrations using Zod schemas for transport type (stdio, SSE, streamable HTTP), restricts environment variable handling, and applies timeout policies (mcp/config-loader.ts:62-90, mcp/policies.ts). The RunCommandExecutionController tracks active shell processes and supports detaching them from the agent loop so long-running background commands persist across turns (run-command-execution-controller.ts). Checkpoints (git-based workspace snapshots) provide rollback with commit/rollback transactions (checkpoint-restore.ts). Network restrictions are not hard-coded — they depend on MCP server configurations. The filterToolsByPolicies function (runtime-builder.ts:85-99) lets tools be selectively disabled by name or via wildcard policies.

Editor's note. Correction: the default policy preset returns no policies, and the SDK treats unlisted tools as auto-approved. The CLI defaults to autoApprove: true for all tools; the VS Code extension routes read, edit, command, web and MCP tools to its approval callback, whose default settings auto-approve everything except commands. MCP config validation is not a sandbox.

openinterpreter/openinterpreter

answered

Safety uses approval policy plus platform sandboxing plus permission profiles. Approvals module (core/src/tools/approvals.rs:1-60) routes decisions through GuardianReviewContext, with ApprovalStore caching per-session (core/src/tools/sandboxing.rs:41-63). ToolOrchestrator (core/src/tools/orchestrator.rs:47-60) drives approval to sandbox selection to attempt to retry. Platform sandboxing is OS-specific. Linux (sandboxing/src/bwrap.rs:1-50) uses bubblewrap. macOS (seatbelt.rs:1-60) uses /usr/bin/sandbox-exec with SBPL policies for base/network. Windows uses WindowsRestrictedToken or WindowsMxc. SandboxManager (manager.rs:48-62) resolves get_platform_sandbox() by OS. Landlock (landlock.rs) provides Linux landlock enforcement. Permission profiles (FileSystemSandboxPolicy, NetworkSandboxPolicy) control read/write/exec. policy_transforms.rs computes effective policies. Shell/file commands go through unified_exec (core/src/tools/handlers/unified_exec.rs) with configurable timeouts and sandbox-permissions overrides. Allow/deny use ExecPolicyAmendment and ApprovalPolicy. Guardian module adds per-reviewer budget. No checkpoint/restore; sandbox is per-command.

Aider-AI/aider

answered

User approval for every shell command. Shell commands are suggested by the LLM inside ```bash blocks. They are extracted in EditBlockCoder.get_edits() (editblock_coder.py:33) or via run_shell_commands() (base_coder.py:2434-2486). Each command goes through handle_shell_commands() which calls self.io.confirm_ask("Run shell command?", explicit_yes_required=True) (base_coder.py:2456) — explicit_yes_required=True means the user must type 'y', not just press Enter. The ConfirmGroup mechanism allows "Skip all"/"Allow all" within a batch. The allow_never flag lets users permanently reject a specific command pattern.

--yes-always mode. The flag --yes-always / config yes-always: true (aider/args.py:760-764) bypasses all confirmation prompts. In confirm_ask() (aider/io.py:866-867): when self.yes is True, the response is "y" unless explicit_yes_required is set. This is the only automatic bypass — every shell command execution and file write by default requires explicit user confirmation per operation.

No sandboxing. There is no container isolation, seatbelt, Landlock, or any OS-level sandbox. The run_cmd() function (aider/run_cmd.py:11-23) uses pexpect.spawn (on Unix TTYs) or subprocess.Popen with shell=True. The linter (aider/linter.py) also calls run_cmd_subprocess directly. The sole protection is user approval prompts.

Network restrictions. None enforced by aider. The LLM calls go through litellm (external APIs), and web scraping uses Playwright. There are no allow/deny lists for network access.

File-write safety. Writes happen through self.io.write_text() only after the LLM has generated them and the user has seen and implicitly accepted them by not cancelling. The dry_run flag (base_coder.py:415) skips all writes. The git workflow (dirty_commit before edits, auto_commit after) provides an undo safety net. The aiderignore file (.aiderignore, parsed in repo.py) can exclude files from being read or edited.

Checkpoints via git. Each aider edit session auto-commits before making changes (dirty_commit, base_coder.py:2411), and commits after applying edits (auto_commit, base_coder.py:2375). The /undo command (commands.py:553) reverts the last aider commit using git revert. The aider_commit_hashes set tracks which commits are aider-generated.

codewhale-hq/Codewhale

answered

Execution safety in Codewhale is implemented at multiple layers: approval modes, command safety classification, OS sandboxing, and filesystem guards.

Approval modes (execpolicy/src/approval_mode.rs:7): Four modes — Suggest (ask for non-safe tools, the default), Auto (auto-review risky calls), Bypass (no approvals/"YOLO mode"), and Never (block all tools requiring approval). The user cycles through Suggest → Auto → Bypass with Shift+Tab. ApprovalMode::from_config_value accepts many CLI aliases ("yolo", "bypass-permissions", "full-access", etc.).

Command safety (execpolicy/src/command_safety.rs:1): A COMMAND_ARITY dictionary maps ~40+ command prefixes (git status, npm run, cargo check, docker compose, etc.) to their canonical forms. Ruleset objects (execpolicy/src/lib.rs:29) at three priority layers (BuiltinDefault → Agent → User) define trusted_prefixes (auto-allowed), denied_prefixes (always blocked), and ask_rules (typed tool-invocation patterns requiring approval). Permission actions are Allow/Ask/Deny. The command-safety system classifies commands into SafetyLevels, and the ExecPolicyContext uses these to determine whether a shell invocation requires approval.

OS sandbox (sandbox/mod.rs:1):

  • macOS: Seatbelt (sandbox-exec) when runtime probe succeeds — uses the native macOS sandbox.
  • Linux: Bubblewrap (bwrap) only when user opts-in (prefer_bwrap: true) and /usr/bin/bwrap is executable. Creates a new mount namespace with --unshare-all, read-only bind of /, tmpfs /tmp, /dev, /proc, and writable roots per policy. --share-net is added for network-enabled policies. The extension host always uses bwrap when available without the opt-in.
  • Windows: Process-tree containment via Job Object (planned, not yet sandboxed).
  • SandboxPolicy (sandbox/policy.rs:22) has four levels: DangerFullAccess (no restrictions), ReadOnly (read-only everywhere), ExternalSandbox (already sandboxed, skip double-wrapping), WorkspaceWrite (read-only + writable workspace + optional extra roots + optional network).

The sandbox is never silently downgraded — if bwrap isn't available on Linux, the tool reports no OS sandbox rather than pretending.

Filesystem guards: A read deny-list (tools/file.rs:458) protects credential files, config paths, and sensitive system files. Paths are checked on both raw spelling (before symlink resolution) and canonical form (after). The content-hash guard prevents stale-state overwrites.

Network restrictions: A session-scoped NetworkPolicyDecider controls per-domain access. The /network allow <host> command persists session-scoped approvals. Plan mode forces a read-only sandbox with no network access.

esengine/DeepSeek-Reasonix

answered

Shell commands and file writes have three independent layers of protection. Permission policy (internal/safety/permission/permission.go:1): each tool call is evaluated against a Policy of rules (allow/ask/deny). Rules are configured as [permission] entries in reasonix.toml (e.g. allow = ["Bash:git*"], deny = ["Bash:rm -rf /"]). A Gate wraps the policy with an interactive Approver that lets the user allow/deny/always-allow on each call. The default is Ask (prompt user). Read-only mode denies all writers with a clear refusal code (RefusalReadOnly). OS sandbox (internal/safety/sandbox/sandbox.go:1): on Linux, bash commands run under bubblewrap; on macOS, under sandbox-exec/Seatbelt. The Spec declares write roots (workspace + configured extras), forbid-read roots, a Network boolean, HostAuthorities for SSH/Docker/Podman, an Egress proxy route for external HTTP, and MinimalWrites for MCP processes. On Windows, OS-level bash sandboxing is not available and enforced sandbox fails closed (UnavailableMessage). Path confinement (internal/tools/builtin/confine.go:29): ConfineBash binds the bash tool to the sandbox spec; BindSessionTemp attaches session-private temp dirs. The write tools (write_file, edit_file, multi_edit, move_file, delete operations) apply confineWrite to ensure the target path falls within allowed write roots and outside protected Reasonix data directories (SessionDataGuard). Network access for tool calls is further governed by the egress policy (internal/safety/egress/policy.go) which routes external HTTP through a proxy. The sandbox enforces three dimensions independently: Integrity (write roots), Confidentiality (forbid-read roots), and Authority (host services). Checkpoints (session snapshots) are stored separately in the session store for rewind, not used for sandbox isolation.

firecrawl/open-lovable

answered

All code execution and file writes happen inside sandboxed cloud environments, never on the host machine. Two sandbox providers are available: Vercel Sandbox (@vercel/sandbox) and E2B (@e2b/code-interpreter), selected via the SANDBOX_PROVIDER env var (lib/sandbox/factory.ts:6-21). The abstract SandboxProvider class (lib/sandbox/types.ts:35-65) defines the interface. The SandboxManager (lib/sandbox/sandbox-manager.ts) manages lifecycle: creating sandboxes, reconnecting to existing ones (E2B only), and cleanup after 1 hour of inactivity. There is no user approval prompt for any command — the system writes files and runs commands against the sandbox without confirmation. There are no command allow/deny lists, no network restriction configuration, and no checkpoint/rollback mechanism. For package installation, npm install is run directly in the sandbox (vercel-provider.ts:266-323, e2b-provider.ts:185-229). File paths are normalized by prepending src/ when they don't match config-file patterns about config files (apply-ai-code-stream/route.ts:529,603-612). The build-validator.ts checks whether the sandbox serves working content after edits, but it's a read-only health check with no recovery beyond retrying. The safety model is entirely outsourced to the sandbox provider's own isolation — the host Next.js app has no sandboxing (no Landlock, no seccomp, no containerization) and global variables hold all mutable state.

Editor's note. Correction: build-validator.ts is dead code (validateBuild has no callers), so there is no post-apply health check at all.

QwenLM/qwen-code

answered

Safety operates at multiple layers. Approval modes (packages/core/src/config/approval-mode.ts) offer a progression: plan, default, auto-edit, auto, and yolo — controlling whether the agent must ask before running commands or editing files.

Permission system (packages/core/src/permissions/permission-manager.ts, types at packages/core/src/permissions/types.ts): rules match tool calls by tool name, optionally with a specifier (shell command glob, file path pattern, domain). Decisions can be allow, ask, deny, or default (falling through to approval mode). Rules can be scoped to system, user, workspace, or session levels.

Sandbox execution (packages/core/src/sandbox/sandbox-execution.ts): shell commands run through runtime-shell.ts which checks sandbox policy. Two backends — bwrap (Bubblewrap, bwrap-execution.ts) and Landlock (landlock-execution.ts) — enforce filesystem policy (read-only or workspace-write) and network policy (open or closed). The policy also specifies maskedPaths for additional restriction. Docker/Podman sandbox is available for additional isolation via execute-sandbox.ts.

Destructive commands (packages/core/src/permissions/destructive-commands.ts) use CommitAttributionService to track staged changes. Runtime file versioning (sandbox/runtime-file.ts) snapshots files before edits, enabling diff tracking. There are no container or VM-level checkpoints.

Editor's note. Correction: execute-sandbox.ts only dispatches to bwrap or Landlock; Seatbelt/Docker/Podman whole-CLI sandboxing is configured in packages/cli/src/config/sandboxConfig.ts. destructive-commands.ts is a deterministic regex pre-filter for AUTO mode that runs before a fail-closed two-stage LLM classifier (permissions/classifier.ts); it does not use CommitAttributionService.

stackblitz-labs/bolt.diy

answered

Code execution runs inside WebContainer, an in-browser sandboxed Node.js runtime that cannot access the host system, files, or network beyond what the browser permits. The WebContainer.boot() call at app/lib/webcontainer/index.ts line 26 uses coep: 'credentialless' for cross-origin embedding policy and manages execution through virtual filesystem and shell APIs. Shell commands are validated before execution in #validateShellCommand (action-runner.ts:577-670) — it intercepts rm commands to add -f flags when files don't exist, checks that cd targets exist (auto-creating them with mkdir -p), and verifies source files before cp/mv. Locked files are enforced at the prompt level: the system prompt lists locked files and instructs the LLM not to modify them (stream-text.ts lines 202-223). On the security side, app/lib/security.ts provides rate limiting per-IP (10 req/min for LLM, 100 for other API endpoints), security headers (CSP, X-Frame-Options, HSTS), and API key format validation. MCP tool execution requires user approval via TOOL_EXECUTION_APPROVAL: in mcpService.ts line 402, the tool's execute function only runs when the result equals 'Yes, approved.', and rejections return TOOL_EXECUTION_DENIED. The maxSteps parameter (default 5, from MCP store) limits the number of LLM reasoning steps. Network requests from within WebContainer are restricted by browser policies — the system prompt warns against native binaries and restricts package choices to JS-only implementations.

Editor's note. Correction: the security.ts rate limits are not applied to the LLM endpoints. withSecurity wraps only the GitHub/GitLab/Netlify/Vercel/Supabase routes; /api/chat and /api/llmcall have no rate limiting.

vastsa/PI-Desktop

answered

Shell commands and file writes go through a multi-layer safety system.

Permission modes — The PermissionManager in crates/host-core/src/permissions.rs evaluates every tool call against an effective permission mode. Tools are classified by risk: Read/Glob/Grep = Low, Write/Edit/Bash = High, plugins = Medium (or declared), MCP = Medium (permissions.rs:120-137). The permission mode can be auto (auto-allow low-risk, prompt for high-risk), accept-edits (auto-allow Write/Edit only), or stricter modes that require user approval. Plan/Goal modes have an explicit allowlist (plan_mode_allows(), permissions.rs:144-148) that only admits Read, Glob, Grep, Bash, and BrowserPreview — never Write, Edit, plugins, or Task.

Tool budgets — crates/host-core/src/tool_budget.rs enforces concurrency caps: max 16 in-flight tools globally, 4 shell, 8 reads, 2 mutations, 4 plugins, 1 mutation per session, and a queue depth of 64. This prevents resource exhaustion and serializes dangerous operations.

Bash execution — Commands run through a resolved shell (POSIX, PowerShell, cmd) selected by the user (crates/host-core/src/tools/shell.rs). The CommandShellOption travels with the session. The runtime enforces a timeout clamp (1–21,600 seconds, runtime.ts:1242-1250) and detects patch commands (isPatchCommand(), runtime.ts:1123-1134) to redirect the model to use Edit/Write instead. External-path Bash calls require explicit permission (runtime.ts:3097-3098).

Network policy — Host-core mirrors the networkPolicy settings for LAN/enforcement decisions (network_policy.rs:1-51), supporting both strict mode (no insecure LAN endpoints) and relaxed mode.

Checkpoint and audit — The compaction system creates checkpoints before each model request, providing rollback points. The transcript is append-only. File mutations are serialized per-path. Plugin permissions are sandboxed via capabilities derived from the manifest.

Editor's note. Correction: auto permission mode allows every tool (including outside-workspace paths) without a prompt; ask auto-allows only low-risk tools and accept-edits additionally allows Write/Edit (crates/host-core/src/permissions.rs L231-L256).

cloudflare/vibesdk

answered

Execution safety is implemented at multiple layers:

Tool boundaries — ThinkAgent: Bash execution is explicitly disabled (workspaceBash = false in ThinkAgent). The Think loop exposes only file workspace tools (read, write, edit, list, find, grep, delete), git/commit tools, deploy tools, browser-log inspection, and a structured question-asking tool. No shell or arbitrary command execution is possible through the Think agent.

Sandbox for Agentic/Phasic behaviors: The older agent paths use @cloudflare/sandbox (v0.5.6) for isolating command execution. The SandboxDockerfile builds a Docker image based on cloudflare/sandbox:0.5.6 with git, curl, and a process-monitoring system (container/ directory with cli-tools.ts, process-monitor.ts, storage.ts). Commands are executed inside the container via the BaseSandboxService in worker/services/sandbox/BaseSandboxService.ts which manages sandbox instances, file uploads, command execution, static analysis, and deployment results.

Process monitoring: The container includes a ProcessMonitor class in container/process-monitor.ts and a CLI tool (container/cli-tools.ts) that validates instance IDs with a strict regex (/^[a-zA-Z0-9][a-zA-Z0-9_-]*$/), caps ID length at 64 characters, and stores logs and errors in a managed SQLite storage layer within the container.

Preview authorization: Preview URLs use signed, branch-scoped tokens. getBrowserPreviewURL() in ThinkCodingBehavior invokes signSpacePreviewToken() to create a JWT-like token embedding spaceName, branch, userId, and a previewVersion. The token bootstraps an HttpOnly preview cookie and prevents unauthorized access to preview deployments.

Network restrictions: The sandbox container image does not appear to restrict network access, but the Think agent has no bash or curl tools to make arbitrary network requests.

Secrets: User API keys and tokens are handled through Cloudflare services (AI Gateway bindings, encrypted token blobs). The preDeploySafetyGate in worker/agents/utils/preDeploySafetyGate.ts checks generated code for hardcoded secrets before deployment.

Durable Object isolation: Each app maps to its own ThinkAgent DO, SpaceDO workspace, Artifacts repository, and generated App Facet — full resource isolation per project. VibeSDK does not use Landlock, seccomp, or other OS-level sandboxing beyond the Docker container and Cloudflare's DO isolation.

Editor's note. Correction: there is no pre-deploy secret scan; preDeploySafetyGate checks React render-loop hazards. The container sandbox applies only to Phasic/Agentic apps — the default Think path has no shell and previews through SpaceDO Dynamic Workers.

← How are code edits applied? · Which models are supported and how are they called? →