aaif-goose/goose
Rust agent with a CLI and Electron desktop app, built on MCP extensions, many LLM providers and summary-based context compaction.
Overview
goose is a general-purpose AI agent written in Rust. It is aimed at developers who want one agent for coding and for wider automation. The same core library powers an interactive terminal CLI, an Electron desktop app, a headless goose run mode and an Agent Client Protocol (ACP) server. The project describes itself as “beyond code suggestions”: the agent reads files, edits them, runs shell commands and calls any tool exposed over the Model Context Protocol (MCP).
The design choice that shapes everything is that goose has almost no built-in tools of its own. All tools come from extensions. A few are “platform” extensions compiled into the binary. The developer extension provides write, edit, shell, tree and read_image. Others are ordinary MCP servers started over stdio or Streamable HTTP. The agent loop only knows how to talk to a provider, check permissions, and dispatch tool calls to the extension that owns them.
At the pinned commit the codebase is in the middle of a rewrite of that loop. The default path is still a long streaming loop in Agent::reply_internal. A new, ordered state-machine pipeline (goose-agent crate plus state_machine/ops_*.rs) runs only when GOOSE_STATE_MACHINE=1 is set, or when an ACP client asks for it. Both paths share the providers, extensions, permission inspectors and compaction code.
Architecture
flowchart LR
CLI["goose CLI session"] --> AG["Agent::reply"]
DESK["Electron desktop"] --> SRV["goose serve (ACP HTTP/WS)"]
SRV --> AG
AG --> LOOP["reply_internal loop"]
AG -.-> SM["StateMachine (opt-in)"]
LOOP --> PROV["Provider trait"]
LOOP --> INSP["ToolInspectionManager"]
LOOP --> CTX["context_mgmt: compaction"]
INSP --> EXT["ExtensionManager"]
EXT --> PLAT["Platform extensions (developer...)"]
EXT --> MCP["MCP servers: stdio / HTTP"]
PROV --> LLM["LLM APIs / local llama.cpp"]
LOOP --> SESS["SessionManager (persisted)"]
| Component | Path | Role |
|---|---|---|
| Agent | crates/goose/src/agents/agent.rs |
Agent::reply, the legacy loop, inspector setup, state-machine assembly |
| State machine | crates/goose-agent/src/machine.rs, crates/goose/src/agents/state_machine/ |
Ordered Operation pipeline (opt-in) |
| Extension manager | crates/goose/src/agents/extension_manager/ |
Loads extensions, prefixes tool names, dispatches calls |
| Platform extensions | crates/goose/src/agents/platform_extensions/ |
In-process tools: developer, todo, summon, orchestrator, code_execution, analyze… |
| Providers | crates/goose-providers/, crates/goose/src/providers/ |
Provider implementations, declarative OpenAI-compatible definitions, toolshim |
| Provider types | crates/goose-provider-types/src/base.rs |
Provider trait, GooseMode |
| Security | crates/goose/src/security/, crates/goose/src/permission/ |
Security, egress, adversary, permission and repetition inspectors |
| Context management | crates/goose/src/context_mgmt/, crates/goose-context-management/ |
Threshold check and LLM summarisation |
| Hooks | crates/goose/src/hooks/mod.rs |
Plugin lifecycle hooks (hooks.json) |
| CLI | crates/goose-cli/ |
goose binary: session, run, recipe, acp, serve, configure |
| ACP server | crates/goose/src/acp/ |
Agent Client Protocol over stdio or HTTP/WebSocket |
| Desktop | ui/desktop/ |
Electron app that spawns goose serve and talks ACP |
How a request flows
Take a prompt typed into the CLI, on the default (non-state-machine) path:
- Entry. The CLI session calls
agent.reply(message, session_config, state_machine::enabled(), cancel)(session/mod.rs).reply_implhandles elicitation replies, and then either hands off to the state machine or continues on the legacy path (agent.rs). The desktop app goes the same way throughgoose serve. Its ACP handler reads anunrolledAgentLoopflag from request metadata (acp/server.rs). - Prepare.
reply_internalcollects tools from all enabled extensions, builds the system prompt, appends project instructions, and resumes the provider-side session if there is one (agent.rs). - Turn budget. Each iteration increments
turns_takenand stops with a fixed message once it passesGOOSE_MAX_TURNS(default 1000) (agent.rs). - Stream.
stream_response_from_providersends system prompt, messages and tools to the provider and streams back the reply (agent.rs).categorize_tool_requestssplits tool requests from text and fixes tool names that the model mangled (agent.rs). - Inspect and approve. All registered inspectors run over the batch of tool requests. The permission inspector then sorts them into approved, needs-approval and denied (agent.rs). Approved calls start at once, and calls that need approval wait for the user.
- Dispatch. Each call goes to
ExtensionManager::dispatch_tool_call, which finds the owning extension through a per-session lease (extension_manager/mod.rs). Results are appended as tool responses and the loop iterates. - Overflow. If the provider reports that the context is too long, the loop calls
compact_messagesand retries. It gives up after two compaction attempts. - Finish. A reply with no tool calls ends the turn, subject to recipe retry checks and stop hooks.
Key components
Two loops, one agent
The state machine is the clearer of the two designs. StateMachine::step walks an ordered list of operations. The first one that returns Applied produces effects, which are persisted, and run repeats until no step applies or one yields to the client (machine.rs). Before an inference step, every operation can contribute tools, prompt parts and “moim” parts. create_state_machine shows the order: steer, max turns, ! shell, compaction, tool-pair compaction, tool approval, doctor, project, skills, foreground subagents, recipe, tool execution, unknown tool, retry, stop hook and exit-on-error (agent.rs). The gate is an environment variable that defaults to off (state_machine/mod.rs).
Extensions and tool naming
ExtensionConfig has stdio, builtin, platform and streamable_http variants, among others (extension.rs). Tools are exposed as <extension>__<tool>. Platform extensions run in-process with direct access to the agent. The developer extension’s tool set is pinned by a test to write, edit, shell, tree and read_image (developer/mod.rs).
Editing
There is no diff or patch format. edit is a plain string replacement that requires exactly one match. On zero matches it returns a “did you mean” hint and a file preview, and on several matches it lists the first two locations (edit.rs). Git, tests and linting go through shell.
Safety layers
create_tool_inspection_manager registers five inspectors in priority order: a pattern-based security scanner, an egress inspector that pulls URLs, SSH remotes and S3 targets out of commands, an LLM “adversary” reviewer, the permission inspector and a repetition detector (agent.rs). The adversary reviewer is opt-in and runs only when ~/.config/goose/adversary.md exists (adversary_inspector.rs). GooseMode decides how much is asked: Auto (the default, approve everything), Approve, SmartApprove and Chat (no tools). There is no filesystem checkpointing and no OS sandbox in the core. A Container handle can route execution into Docker.
Context compaction
check_if_compaction_needed compares the session’s token count (or an estimate) with the model’s context limit. The threshold is GOOSE_AUTO_COMPACT_THRESHOLD, default 0.8 (context_mgmt/mod.rs). compact_messages keeps the latest plain user message, asks the model for a summary, and then flips visibility flags. Old messages stay visible to the user but are hidden from the agent (context_mgmt/mod.rs). An optional tool-pair summariser shrinks old tool call/result pairs. There is no embedding index.
Providers
The Provider trait has stream, complete, resume and model-info methods (base.rs). Native implementations cover Anthropic, OpenAI, Google, Ollama, OpenRouter, Bedrock, Vertex, Databricks, Snowflake, local llama.cpp and more. 48 declarative JSON definitions add OpenAI-compatible vendors. Several “ACP providers” (Claude, Codex, Copilot, Amp, pi) drive another agent as if it were a model. For models without native tool calling, the toolshim sends the text to an interpreter model, which extracts the tool calls.
Extending it
- MCP servers. Add any stdio or HTTP MCP server as an extension in config, in the desktop UI or with
--with-extension. - Hooks. Plugins ship
hooks/hooks.jsonwithcommandactions for events such asPreToolUse,PostToolUse,BeforeShellExecution,AfterFileEditandStop(hooks/mod.rs). - Recipes. YAML or JSON files that bundle instructions, extensions, parameters, retry checks and sub-recipes. They run with
goose run --recipeand can be scheduled. - Skills. Filesystem skills loaded on demand through a
load_skilltool. - Embedding. Use the
goosecrate directly (Agent::replyreturns a stream ofAgentEvent), drive it over ACP, or use thegoose-sdkbindings. - Providers. Implement the
Providertrait, or add a declarative definition for an OpenAI-compatible endpoint.
Running it
- CLI. Install the
goosebinary, rungoose configureto pick a provider and key, thengoose session(interactive) orgoose run -t "..."(headless). - Desktop. The Electron app in
ui/desktopspawnsgoose servelocally and talks ACP over HTTP/WebSocket. - Agent server.
goose acpserves ACP on stdio, which editors can use as an external agent. - Required. An API key or local model. Building from source needs the pinned Rust toolchain, plus Node/pnpm for the desktop app.
Strengths and caveats
- Strength: MCP-native. Every tool, built-in or third-party, goes through the same extension and permission path, so adding capabilities needs no core changes.
- Strength: broad provider support, including local llama.cpp, a toolshim for models without tool calling, and the unusual option of wrapping other agents as providers.
- Strength: layered tool inspection. A pattern scanner, an egress check, an optional LLM reviewer and per-tool permissions all run before execution.
- Caveat: permissive by default.
GooseMode::Autoapproves tool calls without asking. Nothing rolls back file changes or sandboxes the shell unless you set it up. - Caveat: two loops in flight. The legacy loop in
agent.rs(a 6,900-line file) and the opt-in state machine overlap. Behaviour can differ depending on a flag, and the code is in transition. - Caveat: minimal editor. Exact single-match string replacement, with no fuzzy fallback and no multi-edit, means more retries on large refactors than diff-based agents need.
Sources: code at 540df77, verified Q&A.
How it answers the Open-source coding agents questions
Each answer was drafted by a code-reading agent at commit 540df77. Its citations were checked mechanically. Compare with the other open-source coding agents →
How is the agent loop implemented?
answeredGoose has two agent-loop implementations: the legacy loop in agent.rs and the new state machine in state_machine/. The reply() method dispatches between them via a boolean flag; the state machine is gated by GOOSE_STATE_MACHINE=1 (mod.rs:94-98).
The state machine (goose-agent/src/machine.rs:51) is an ordered pipeline of Steps, each wrapping either an Operation or an Inference. Its run() loop loads the session, then calls step() — which walks the steps in order calling each operation's run() — and then apply() to persist effects. The first operation that returns Applied yields the turn; NotApplicable falls through to the next step. The loop repeats until no step applies or yield_to_client is set (machine.rs:167-182).
goose defines ~20 concrete operations for the state machine (create_state_machine at agent.rs:1694): EntryHookOperation, SteerOperation (queued user guidance delivered between turns), MaxTurnsOperation (default 1000-turn budget, agent.rs:88), BangShellOperation (the ! shortcut), CompactionOperation (auto-compact when context exceeds threshold), ToolApprovalOperation (permission routing), ToolExecutionOperation (dispatch tool calls to extensions/MCP servers), RetryOperation (retry with goal/grind), ExitOnErrorOperation, RecipeOperation, ForegroundSubagentOperation, StopHookOperation, plus prompt-part contributors like ProjectOperation and SkillOperation.
Tool calling flows through ToolExecutionOperation (ops_toolcalling.rs): it collects pending tool requests from the conversation, runs PreToolUse hooks (blocking — can deny execution), categorises the tool (Shell/Read/Write/Other), emits BeforeShellExecution/BeforeReadFile extended hooks per category, dispatches to the ExtensionManager, and emits PostToolUse/PostToolUseFailure on completion. Tool names are prefixed by their owning extension with __ (e.g. developer__shell) and the manager canonicalises name mangling from providers.
Stop conditions are: max turns reached (MaxTurnsOperation), a hard error (ExitOnErrorOperation), explicit yield-to-client from any operation (notably RetryOperation after goal satisfaction), and the turn always ending after the LLM responds without tool calls.
Sub-agents are implemented by ForegroundSubagentOperation (ops_foreground_subagent.rs): it spawns a child Agent with its own Session and recipe, runs it to completion, and captures the final output text. The subagent reuses the same provider and extension manager context. There is also a subagent_execution_tool module for background multi-task sub-agent execution.
How is repository context gathered and kept within the context window?
answeredContext management in goose centres on conversation compaction via LLM summarisation, not vector embeddings or RAG. The context_mgmt module (context_mgmt/mod.rs) provides two key functions:
check_if_compaction_needed() (mod.rs:224-274) compares current context usage against a configurable threshold (GOOSE_AUTO_COMPACT_THRESHOLD, defaulting to ~75% of the model's context limit from DEFAULT_COMPACTION_THRESHOLD in goose-context-management). Total tokens come from session metadata or are estimated on the fly by a token counter.
compact_messages() (mod.rs:70-207) runs when the threshold is breached. It preserves the most recent user text-only message, sends the rest to the LLM for summarisation via do_compact(), then replaces the original messages with:
- Original messages marked agent-invisible (kept for user viewing)
- A summary message (agent-only)
- A continuation prompt depending on the mode (tool-loop or conversation)
The CompactionOperation (ops_compaction.rs) triggers this from the state machine and also emits a "goose is compacting..." status line.
Tool-pair compaction (ToolPairCompactionOperation, gated by GOOSE_TOOL_PAIR_SUMMARIZATION) summarises batches of tool-call/result pairs inline when they exceed a threshold, reducing token waste from long tool results.
For repository context, the ProjectOperation (ops_project.rs:1-39) loads project-level instructions from sources::read_project() into the system prompt. SkillOperation (ops_skills.rs) loads skill files from disk into inference. The hooks system (hooks/mod.rs) provides BeforeReadFile events that fire before the agent reads any file.
There is no embeddings database, no semantic search, and no vector store. Context is managed purely through: (a) incremental token counting, (b) LLM-based summarisation when the model's context window approaches capacity, (c) message visibility flags that hide old messages from the agent while keeping them in the conversation for user display, and (d) project/skill instructions injected directly into the system prompt.
How are code edits applied?
answeredGoose does not have a built-in file editor. All code editing is performed through MCP tool extensions — the primary editing tool is text_editor provided by the built-in developer platform extension (platform_extensions/developer/). The tool categorisation in both agent.rs:129-137 and ops_toolcalling.rs:37-51 recognises edit tool names by matching tokens like write, edit, patch, write_file, edit_file after stripping any extension prefix.
Editing tools are exposed through the standard MCP lifecycle: the ExtensionManager (extension_manager/mod.rs:285) loads extensions via stdio or Streamable HTTP, each publishing its tool schemas. When the LLM requests a tool call, ToolExecutionOperation routes it through the pre-tool hook chain (which includes AdversaryInspector — LLM-based safety review), tool approval, and execution against the owning extension.
There is no automatic lint, test, or retry cycle for edits. The RetryOperation (ops_retry.rs) retries the entire LLM turn when the model indicates an error, but it does not apply per-edit validation. Hook events AfterFileEdit and AfterShellExecution fire post-tool and can be subscribed to by hook scripts for custom validation.
Git integration is available only through the shell tool (e.g. developer__shell running git commands). There is no automatic git undo, diff, or commit mechanism — the agent must explicitly use shell commands for version control. The StopHookOperation (ops_stop_hook.rs) can block turn-ending if policy hooks detect uncommitted work, acting as a soft safeguard.
In summary: edits are applied by MCP extension tools (primarily text_editor), pass through the security/permission pipeline, fire lifecycle hooks, and rely on the agent's own explicit shell-based git usage for version control. There is no diff format or patch-apply mechanism in the core agent loop.
write and edit (exact single-match string replacement), alongside shell, tree and read_image; there is no text_editor tool at this commit.How are shell commands and file writes kept safe?
answeredGoose implements a defence-in-depth security architecture across several layers:
Permission system: Tools are categorised into Shell, Read, Write, and Other (agent.rs:129-137). The PermissionManager (permission/mod.rs) tracks decisions per tool name at levels: AlwaysAllow, AskOnce, DenyOnce, AlwaysDeny. ToolApprovalOperation (ops_tool_approval.rs:26-41) resolves pending approvals by consulting session-stored permissions and marking tools as executable or not before the LLM sees them. In Chat mode (GooseMode::Chat), approval is skipped — all tools are run with explicit user confirmation per tool call.
Adversary Inspector (adversary_inspector.rs:56-61): an LLM-based inspector that reviews tool calls against configurable rules defined in ~/.config/goose/adversary.md. Default rules block exfiltration, destructive commands, malware, and privilege escalation. The inspector returns Block or Allow decisions that feed into the ToolInspectionManager pipeline.
Egress Inspector (egress_inspector.rs:49-80): scans shell commands for URLs (HTTP, FTP), git SSH remotes, and S3 URIs, extracting domains and destinations to detect data exfiltration attempts.
Shell/Read hooks (ops_toolcalling.rs:179-207): before executing shell commands or reading files, specific hook events fire (BeforeShellExecution, BeforeReadFile) with the command/path as a matcher. Hook scripts can block the operation by returning a denial.
Docker sandboxing: a lightweight Container struct (container.rs:2-5) wraps a Docker container ID. The extension system supports running MCP extensions inside containers, but the sandboxing is delegated to the container runtime rather than Goose enforcing it.
Checkpoints: not implemented in this codebase. There is no filesystem snapshot or rollback mechanism. Safety is purely through pre-execution permission/inspection gates.
Network restrictions: not implemented in the core. Hook scripts or container-level network policies would be needed to restrict outbound access.
Which models are supported and how are they called?
answeredGoose supports 15+ model providers through a two-tier architecture. The base Provider trait is defined in goose-providers/src/base.rs with key methods stream(), complete(), resume(), and fetch_model_info(). Providers support tool calling via the standard MCP tool schema and streaming responses.
Core providers (in goose-providers/src/): Anthropic (anthropic.rs — default model claude-sonnet-4-5), OpenAI (openai.rs — default gpt-4o), Google (google.rs), Ollama (ollama.rs), OpenRouter (openrouter.rs), Snowflake (snowflake.rs), Azure Foundry (azure_foundry.rs), Databricks (databricks.rs, databricks_v2.rs), plus local inference (local_inference.rs via llama-cpp-2).
Additional providers (in crates/goose/src/providers/): Bedrock (bedrock.rs), Google Vertex AI (gcpvertexai.rs), GitHub Copilot (githubcopilot.rs), ChatGPT Codex (chatgpt_codex.rs), xAI (xai.rs, xai_oauth.rs), Gemini CLI/OAuth (gemini_cli.rs, gemini_oauth.rs), LiteLLM (litellm.rs), NVIDIA (nanogpt.rs), HuggingFace (huggingface.rs), SageMaker TGI (sagemaker_tgi.rs), Kimi Code (kimicode.rs), Tetrate (tetrate.rs), Cursor Agent (cursor_agent.rs), and several ACP-based providers (claude_acp.rs, codex_acp.rs, copilot_acp.rs, amp_acp.rs, pi_acp.rs).
Declarative providers (goose-providers/src/declarative.rs): a data-driven system supporting 40+ third-party providers via config files in declarative/definitions/, including DeepSeek, Groq, Together, Fireworks, Perplexity, Mistral, and many more. Each is a YAML/JSON descriptor that maps API endpoints to the OpenAICompatible format.
Tool-calling format: providers use the MCP tool schema (rmcp::model::Tool) for tool definitions. The GooseInferenceProvider wrapper (ops_llm.rs:100-160) enriches tool error messages, canonicalises tool names, and assembles the advertised_tools note. Providers that don't natively support tool calling transparently fall back through text-format deserialisation in reply_parts.rs.
Thinking/effort support: normalize_legacy_provider_thinking_effort() in agent.rs:100-119 maps ThinkingEffortSupport across providers. Per-provider thinking effort is configurable via the model config.
Cost tracking: goose-providers/src/conversation/token_usage/ provides ProviderUsage, Usage, and CostSource types with cost calculation per provider format. Usage is persisted to sessions and displayed via /status (StatusOperation).
How can it be extended and customised?
answeredGoose is designed for extension through several mechanisms:
MCP (Model Context Protocol): The primary extension system. Extensions are loaded by ExtensionManager (extension_manager/mod.rs:285) via stdio subprocess or Streamable HTTP. Each extension publishes tools, resources, and prompts. Tool names are prefixed with the extension key plus __ (e.g. filesystem__read). The manager supports dynamic enable/disable per session, caching of extension states, and recovery from mangled tool names (recover_mangled_tool_name). First-class extensions are built-in Rust extensions (is_first_class_extension), while third-party extensions from the MCP registry load via the standard MCP protocol.
Platform extensions (in platform_extensions/): built-in Rust extensions shipped with Goose, including developer (shell, text editor, websearch), summarize, orchestrator, chatrecall, todo, scheduler, tom, code_execution, analyze, and summon. These are first-class extensions integrated at compile time.
Hook system (hooks/mod.rs): JSON-configurable lifecycle hooks that subscribe to events (PreToolUse, PostToolUse, SessionStart, SessionEnd, BeforeReadFile, AfterFileEdit, BeforeShellExecution, AfterShellExecution, Stop, etc.) and run shell commands. Hook scripts receive the event context on stdin. Defined per plugin in <plugin-root>/hooks/hooks.json.
Skills (SkillOperation, ops_skills.rs): filesystem-based skills that an agent can load at runtime, providing additional guidance and tools. Loadable via /skill slash command. The load_skill tool allows the LLM to request skill loading.
Recipes (RecipeOperation, ops_recipe.rs): declarative recipe files (JSON) that configure the agent session: system prompt, extensions, retry/on-failure commands, success checks, max turns, and goal/grind parameters. Recipes can be run via the CLI (goose run --recipe) or loaded from session config.
SDK crate (goose-sdk): provides the goose crate as a library. The Agent struct is instantiable directly, allowing headless integration. Agent::reply() returns a stream of AgentEvents that callers consume. Built-in extensions use an MCP server pattern (mcp_server_runner.rs in goose-mcp) to embed goose inside other MCP hosts.
CLI and ACP (goose-cli/src/): the CLI exposes the full agent lifecycle. The ACP protocol (ui/goose-acp/) enables desktop ↔ agent communication via a WebSocket/HTTP protocol with streaming events.
Custom providers: new providers implement the Provider trait (providers/base.rs) with stream()/complete()/resume() methods. Declarative providers need only a config file mapping to an OpenAI-compatible endpoint.