QwenLM/qwen-code
Terminal coding agent forked from Gemini CLI, with a multi-provider model layer, bwrap/Landlock shell sandboxing and an LLM auto-approver.
Overview
Qwen Code is the Qwen team’s terminal coding agent. It started as a fork of Google’s Gemini CLI, and the lineage is still visible: many files keep the Copyright 2025 Google LLC header, the internal message format is @google/genai’s Content/Part, and LlmClient is still exported under the alias GeminiClient (client.ts). Since the fork it has grown a lot. At the pinned commit (v0.25.0) the pnpm monorepo has about twenty packages: the qwen CLI, the shared core, SDKs for TypeScript, Python and Java, a VS Code companion, a Zed extension, a web shell, a desktop app and a set of chat “channels” (Telegram, DingTalk, Feishu, WeCom, GitHub, GitLab, email and others).
The core is a classic tool-calling loop. The model returns text and function calls. The front end runs the calls through a permission pipeline and sends the results back. Almost every feature sits on top of that loop: around fifty built-in tools, sub-agents and “teams”, cron jobs, goals, auto-memory, skills, hooks, MCP and extensions. Many of these copy Claude Code’s design closely. The hook event names, the skills directory layout, the [Old tool result content cleared] microcompaction placeholder and the 13,000-token autocompact buffer are all documented in comments as matching Claude Code (chatCompressionService.ts).
Its real differentiators are the model layer, which speaks Qwen OAuth, OpenAI Chat Completions, OpenAI Responses, Anthropic and Gemini natively, and a safety stack that combines rule-based permissions, an LLM classifier for “auto” mode and an OS sandbox for shell commands.
Architecture
flowchart LR
U["User / IDE / SDK / channel"] --> CLI["qwen CLI (llm.tsx main)"]
CLI --> TUI["Interactive UI"]
CLI --> NI["runNonInteractive"]
CLI --> ACP["ACP agent"]
TUI --> LC["LlmClient.sendMessageStream"]
NI --> LC
ACP --> LC
LC --> CHAT["LlmChat + compression"]
CHAT --> CG["ContentGenerator (OpenAI / Anthropic / Gemini / Qwen)"]
LC -->|"ToolCallRequest events"| SCH["Tool scheduler"]
SCH --> PERM["Permissions + AUTO classifier"]
PERM --> TOOLS["Built-in tools / MCP / agents"]
TOOLS --> SBX["Shell sandbox (bwrap / Landlock)"]
TOOLS -->|"function responses"| LC
| Component | Path | Role |
|---|---|---|
| CLI entry | packages/cli/index.ts, packages/cli/src/llm.tsx |
Parses flags, picks interactive UI, non-interactive run or ACP mode |
| Non-interactive driver | packages/cli/src/nonInteractiveCli.ts |
Headless loop: stream a turn, execute tool calls, feed results back |
| Agent client | packages/core/src/core/client.ts |
LlmClient.sendMessageStream: one model turn plus loop guards, hooks, memory, IDE context |
| Turn / chat | packages/core/src/core/turn.ts, llm-chat.ts |
Streams model output into typed events; owns history and auto-compaction |
| Content generators | packages/core/src/core/contentGenerator.ts and *ContentGenerator/ |
Per-protocol adapters behind one ContentGenerator interface |
| Providers | packages/core/src/providers/ |
Presets (Alibaba plans, DeepSeek, Grok, MiniMax, Z.ai, Moonshot, OpenRouter, …) |
| Tools | packages/core/src/tools/ |
Edit, write, read, grep, glob, shell, web, agent, cron, worktree, MCP bridge |
| Permissions | packages/core/src/permissions/ |
Rules, AUTO-mode classifier, destructive-command pre-filter |
| Sandbox | packages/core/src/sandbox/, packages/cli/src/config/sandboxConfig.ts |
bwrap/Landlock for shell commands; Seatbelt/Docker/Podman for the whole CLI |
| Memory and compaction | packages/core/src/memory/, services/microcompaction/, services/chatCompressionService.ts |
Auto-memory, old-output clearing, summarisation |
| Extensions | packages/core/src/extension/, skills/, hooks/ |
Installable bundles of MCP servers, skills, agents, hooks, commands |
How a request flows
Take a headless run such as qwen -p "fix the failing test":
- Entry.
main()inllm.tsxbuilds the config and branches. With the ACP flag it importsrunAcpAgent, in a TTY it callsstartInteractiveUI, and otherwise it callsrunNonInteractive(llm.tsx, L1463-L1468, L1650-L1662). - Send a turn.
runNonInteractiveenters awhile (true)loop and callsllmClient.sendMessageStream(parts, signal, promptId, { type: sendType, … })(nonInteractiveCli.ts). TheSendMessageTypetells the client whether this is a user query, a tool result, a retry, a cron prompt, a teammate message or a goal continuation (client.ts). - Guard the turn.
sendMessageStreamcaps recursion atMAX_TURNS = 100(client.ts), injects IDE context and recalled memories, and streams events from aTurn. On everyToolCallRequestit runs the always-on loop checks (identical consecutive calls, shell stagnation, per-turn call cap) before the optional heuristicLoopDetectionService(client.ts). - Execute tools. The client does not run tools itself. It yields
ToolCallRequestevents, and the front end runs them. Headless mode callsexecuteToolCallfor each request (nonInteractiveCli.ts). The interactive UI does the same throughuseReactToolSchedulerandCoreToolScheduler. Each call goes through permission rules, the approval mode, the AUTO classifier when it applies, and hooks. - Edit. For a file change the model calls
editwithold_string/new_string.checkPriorReadrejects the edit unless the file was read in this session and has not changed on disk since.applyReplacementthen does a literal replacement (edit.ts). - Feed back. The function responses become the next message (
currentMessages = [{ role: 'user', parts: toolResponseParts }]) and the loop sends again asToolResult(nonInteractiveCli.ts). - Finish or continue. A turn with no tool calls ends the run, but first a
checkNextSpeakerside query asks whether the model meant to keep going. If so, the client recurses with a synthetic “Please continue.” (client.ts).
Key components
Model layer
createContentGenerator switches on AuthType and lazily imports one of five adapters: OpenAI chat, OpenAI Responses, Qwen OAuth, Anthropic, or Gemini/Vertex (contentGenerator.ts). Every adapter translates to and from the @google/genai shapes that the rest of the code uses. Thirteen provider presets supply base URLs, env keys and per-model defaults (all-providers.ts). For OpenAI-compatible providers a model can set wireApi: "responses" to switch to the Responses API (modelRegistry.ts). Local servers such as Ollama or vLLM go through the custom OpenAI-compatible provider.
Tool surface
ToolNames lists the built-ins: edit, write_file, read_file, grep_search, glob, run_shell_command, web_fetch/web_search, agent, skill, the plan-mode, cron, task and team tools, worktree tools, notebook_edit, lsp, and tool_search/tool_call for deferred MCP tools (tool-names.ts). Edits are always string replacement or whole-file writes. Nothing lints or runs tests automatically after an edit.
Context management
Three layers keep the window in check. Project instructions come from QWEN.md-style context files and .qwen/rules/. Microcompaction replaces old outputs of read, shell, grep, glob, web and edit tools with [Old tool result content cleared] (microcompact.ts). Full compression runs a side query that produces a <state_snapshot> summary when usage passes a threshold that defaults to 85% of the window (chatCompressionService.ts). Auto-memory under ~/.qwen/memories/ and a per-project root runs background extract, “dream” and recall cycles (paths.ts). There is no embedding index. Code is found with grep, glob and the optional LSP tool.
Safety
Approval modes are plan, default, auto-edit, auto and yolo (approval-mode.ts). In auto, a two-stage LLM classifier decides each risky call. A 32-token fast pass handles the common case, a reasoning pass reviews blocks, and any failure blocks the call (fail-closed) (classifier.ts). A regex layer runs before the classifier and blocks git reset --hard, git clean -f, a non-session --amend and terraform destroy unless the prompt clearly asks for them (destructive-commands.ts). Shell commands can additionally run under bwrap or Landlock with a read-only/workspace-write filesystem policy and an open/closed network policy (sandbox-execution.ts). Separately, the whole CLI can be relaunched under macOS Seatbelt, Docker or Podman (sandboxConfig.ts).
Extending it
- MCP. Stdio, SSE and streamable-HTTP servers, with OAuth. Large tool sets can be deferred behind
tool_search/tool_call. - Extensions. One installable unit (local path, git, GitHub release or npm) can ship MCP servers, LSP servers, context files, commands, skills, sub-agents, workflows, hooks and channel adapters (extensionManager.ts).
- Skills.
SKILL.mddirectories in.qwen/skills/or~/.qwen/skills/, plus extension-provided and bundled skills (skills/index.ts). Bundled skills includebrowser-use, which drives the user’s Chrome through the Qwen Browser SDK (packages/browser-use), andcomputer-use, which is backed by the vendored Cua driver (packages/cua-driver) (browser-use SKILL.md). - Hooks. Claude-Code-style events (
PreToolUse,PostToolUse,UserPromptSubmit,SessionStart,Stop,SubagentStart,PreCompact, …) (hooks/types.ts). - Embedding. ACP mode for editors (Zed, the VS Code companion) and the TypeScript, Python and Java SDKs for programs that drive the agent.
Running it
- Install. The workspace root requires Node.js 22 or newer. The CLI publishes a
qwenbinary (cli package.json). - Auth. Qwen OAuth, or an API key for any configured provider (
OPENAI_API_KEY,ANTHROPIC_API_KEY,GEMINI_API_KEY, or provider-specific keys), selected through settings. - Modes. Interactive TUI by default,
-pfor headless runs, ACP for editors, and aservedaemon for the web and desktop shells. - Sandbox. Shell sandboxing needs
bwrapor Landlock on Linux. Whole-process sandboxing needssandbox-execon macOS, or Docker or Podman with the published image.
Strengths and caveats
- Strength: broad, native provider support. Anthropic, Gemini, OpenAI Chat and Responses are first-class adapters, not one OpenAI-compatible shim. That matters for thinking blocks and caching.
- Strength: layered, fail-closed safety. Deterministic rules come first, then the LLM classifier, then an OS sandbox for shell commands. A classifier outage blocks the call instead of allowing it.
- Strength: stale-write protection. Edits and writes are refused unless the file was read in this session and is unchanged on disk, with a recheck after reading to close the race window.
- Caveat: size and churn.
client.tsalone is close to 6,000 lines, and features (goals, teams, cron, channels, omni media tools) keep piling onto one loop. Behaviour is hard to predict from the docs alone. - Caveat: tool execution lives in the front ends. The headless driver and the React scheduler each implement the tool loop. Embedders have to go through the SDK or ACP rather than one core
run()function. - Caveat: Gemini-shaped internals. Every provider is translated through
@google/genaitypes, so non-Gemini features depend on the fidelity of each adapter. - Caveat: no built-in verification. There is no automatic lint or test step after edits. Checking the work is left to the model’s prompts and to hooks.
Sources: code at ac81c07, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the Open-source coding agents questions
Each answer was drafted by a code-reading agent at commit ac81c07. Its citations were checked mechanically. Compare with the other open-source coding agents →
How is the agent loop implemented?
answeredThe agent loop centres on LlmClient.sendMessageStream() in packages/core/src/core/client.ts, an async generator that orchestrates one complete model interaction turn at a time. The file caps turns at MAX_TURNS = 100 (line 212) and supports multiple message types — UserQuery, ToolResult, Retry, Cron, Notification, Teammate, and Goal (lines 218–237). Each turn: the model generates text plus tool calls, the CoreToolScheduler (core/coreToolScheduler.ts) executes them, and results are fed back as function responses. The loop ends on model text (no tool calls), turn limit, session token limit, loop detection (LoopDetectionService), user interrupt, or stop hooks. There is no separate Planner loop — plan mode is gated by enterPlanMode/exitPlanMode tools that the model itself calls.
Sub-agents use two modes: forked agent (packages/core/src/agents/forkedAgent.ts) — a single-turn cache-sharing call with or without tools, used for auto-memory tasks — and headless/multi-turn via AgentCore (packages/core/src/agents/runtime/agent-core.ts), which owns runReasoningLoop(), tool scheduling, stats, and event emission. AgentInteractive (packages/core/src/agents/runtime/agent-interactive.ts) composes AgentCore with persistent state (message queue, three-level cancellation). Workflows orchestrate many sub-agents in parallel via workflow-orchestrator.ts. Tool-call schemas follow @google/genai's FunctionDeclaration format throughout.
How is repository context gathered and kept within the context window?
answeredRepository context is gathered through several mechanisms. getDirectoryContextString() in packages/core/src/core/environmentContext.ts (lines 53–83) reads folder structures via getFolderStructure() for all workspace directories and injects them as a system prompt section. The getEnvironmentContext() function (lines 92–100) assembles the current date, OS platform, and directory tree into the startup context. IDE context blocks are generated per‑turn by LlmClient.sendMessageStream() and wrapped as <system-reminder> tags (lines 4429–4443 of client.ts).
Auto-Memory is the primary persistent-memory system. MemoryManager (packages/core/src/memory/manager.ts) runs background extract/dream/recall/forget cycles, building a tree of markdown memory files under .claude/memory/. The auto-memory prompt is injected into the system prompt before each turn. Microcompaction (packages/core/src/services/microcompaction/microcompact.ts) replaces old tool outputs (Read, Grep, Shell, WebFetch, etc.) with placeholder text ([Old tool result content cleared]), keeping the conversation within context limits. Chat compression (packages/core/src/services/chatCompressionService.ts) runs a separate model side-query to summarize the conversation when token thresholds are breached, then replaces the compacted history with a summary. File content is read on-demand via dedicated tools (Read, Grep, Glob) — there is no automatic full‑repo embedding or vector index.
~/.qwen/memories/ (user scope) and a per-project auto-memory root, not .claude/memory/ (packages/core/src/memory/paths.ts).How are code edits applied?
answeredCode edits are applied through two tools. Edit (packages/core/src/tools/edit.ts) uses a search-and-replace format: the model provides file_path, old_string, and new_string. The core function applyReplacement() (lines 65–85) uses safeLiteralReplace() to perform the text substitution, with a replace_all option for multiple occurrences and support for new‑file creation (empty old_string + isNewFile flag). WriteFile (packages/core/src/tools/write-file.ts) writes an entire file from scratch with writeRuntimeFile(). Both tools require the file to have been read first in the same turn — this is enforced by checkPriorRead from packages/core/src/tools/priorReadEnforcement.ts.
Before writing, captureRuntimeFileVersion() takes a snapshot of the current file content for diff/undo tracking. After writing, createPatchSmart() in packages/core/src/tools/diffOptions.ts computes the diff, getDiffStat() summarizes it, and the CommitAttributionService tracks session edits. There is no automatic lint, test, or retry loop after edits — the model is expected to verify its own changes by running tools like shell or code review. The edit format is always search-and-replace, never unified diff or whole-file patch. Undo is handled through git integration (session commit tracking and CommitAttributionService), not an explicit undo command.
How are shell commands and file writes kept safe?
answeredSafety operates at multiple layers. Approval modes (packages/core/src/config/approval-mode.ts) offer a progression: plan, default, auto-edit, auto, and yolo — controlling whether the agent must ask before running commands or editing files.
Permission system (packages/core/src/permissions/permission-manager.ts, types at packages/core/src/permissions/types.ts): rules match tool calls by tool name, optionally with a specifier (shell command glob, file path pattern, domain). Decisions can be allow, ask, deny, or default (falling through to approval mode). Rules can be scoped to system, user, workspace, or session levels.
Sandbox execution (packages/core/src/sandbox/sandbox-execution.ts): shell commands run through runtime-shell.ts which checks sandbox policy. Two backends — bwrap (Bubblewrap, bwrap-execution.ts) and Landlock (landlock-execution.ts) — enforce filesystem policy (read-only or workspace-write) and network policy (open or closed). The policy also specifies maskedPaths for additional restriction. Docker/Podman sandbox is available for additional isolation via execute-sandbox.ts.
Destructive commands (packages/core/src/permissions/destructive-commands.ts) use CommitAttributionService to track staged changes. Runtime file versioning (sandbox/runtime-file.ts) snapshots files before edits, enabling diff tracking. There are no container or VM-level checkpoints.
execute-sandbox.ts only dispatches to bwrap or Landlock; Seatbelt/Docker/Podman whole-CLI sandboxing is configured in packages/cli/src/config/sandboxConfig.ts. destructive-commands.ts is a deterministic regex pre-filter for AUTO mode that runs before a fail-closed two-stage LLM classifier (permissions/classifier.ts); it does not use CommitAttributionService.Which models are supported and how are they called?
answeredModels are accessed through a multi-protocol architecture. AuthType (from core/contentGenerator.ts) defines supported protocols: openai, openai-responses, anthropic, gemini, vertex-ai, and qwen-oauth. Constants in packages/core/src/models/constants.ts (lines 74–105) map each protocol to environment variables (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY).
There are 13 provider presets registered in packages/core/src/providers/all-providers.ts (lines 55–69): Alibaba (coding plan, token plan, standard), DeepSeek, Grok, Minimax, Z.ai, Moonshot, IdeaLab, ModelScope, OpenRouter, Requesty, and a generic Custom provider. Each provider defines envKey, baseUrl, modelNamePrefix, and optional per-model generation config (thinking, context window, modalities) — built by provider-config.ts.
Content generators translate the protocol into actual API calls: OpenAIContentGenerator (core/openaiContentGenerator/) for all OpenAI-compatible backends, AnthropicContentGenerator (core/anthropicContentGenerator/), and GeminiChat/geminiContentGenerator for Gemini. The model registry (models/modelRegistry.ts) resolves a provider ID + model config into the correct AuthType via resolveModelProtocol(), supporting wireApi routing between chat-completions and responses (lines 69–92). Per-provider tuning lives in openaiContentGenerator/provider/ (DeepSeek, Fireworks, Grok, Mistral, etc.). Cost tracking is in telemetry/loggers.ts. Local models (Ollama/vLLM) connect via the Custom provider.
How can it be extended and customised?
answeredExtensibility has multiple dimensions. MCP (Model Context Protocol): a full client implementation (packages/core/src/tools/mcp-client.ts, mcp-client-manager.ts) connects to stdio, SSE, and streamable HTTP MCP servers. MCP tools are discovered at runtime and registered into the tool surface as mcp__<server>__<tool> entries (with tool call bridging via tool_search/tool_call).
Plugin system: ExtensionManager (packages/core/src/extension/extensionManager.ts) loads extensions from git repos, npm packages, GitHub releases, and archive URLs. Extensions can contribute MCP servers, skills, subagents, hooks, workflows, context files, settings, channel integrations, and commands. The agent-plugins-v1 subsystem provides manifest-based plugin discovery.
Skills and Hooks: SkillManager (packages/core/src/skills/skill-manager.ts) loads .claude/skills/ markdown-defined skills. The HookSystem (packages/core/src/hooks/hookSystem.ts) orchestrates lifecycle hooks (pre/post tool use, session start/end, stop failure, etc.) with function, command, and HTTP hook runners (lines 60–100).
Headless/SDK use: ACP (Agent Communication Protocol) enables host-controlled agent sessions; SDK packages exist for TypeScript, Python, and Java. Workspace agents (agents/workspace-agents/) support A2A (agent-to-agent) contracts for multi-agent collaboration. Custom tools register via ToolRegistry. Per-project rules files (CLAUDE.md, AGENTS.md) are discovered by rulesDiscovery.ts.
.qwen/skills/ and ~/.qwen/skills/ (plus extension and bundled skills), not .claude/skills/.