cloudflare/vibesdk
Self-hostable prompt-to-app platform on Cloudflare Workers; a Durable Object agent writes, deploys and previews whole web apps.
Overview
VibeSDK is Cloudflare’s open-source prompt-to-app builder: a hosted web product in the Lovable / bolt.new mould where a user types “build me a habit tracker”, watches an agent write the files, and gets a live preview URL and a one-click deploy. It is not a terminal or IDE coding agent. It never touches a developer’s local checkout. The “repository” it edits is a per-app workspace that lives in a Durable Object, and the output is a running Cloudflare Worker, not a pull request. The intended user is a team that wants to run its own vibe-coding service on its Cloudflare account.
At the pinned commit, the default engine for the app project type is the Think behavior. It is a single model-and-tool loop built on @cloudflare/think, running in its own ThinkAgent Durable Object. Its file tools operate on a companion SpaceDO, which holds git-backed files and serves previews by loading the built app as a Dynamic Worker. Two older engines remain in the tree: Phasic (blueprint → phases → implement → review, with a container sandbox) and Agentic (recursive tool calling over the same sandbox). New apps from the home page get Think. Presentations, workflows and “general” projects still get Agentic.
The whole platform is itself a Worker: a React 19 + Vite front end, a Hono API, D1 for users and apps, and roughly a dozen Cloudflare bindings (Durable Objects, Containers, Worker Loaders, Browser Rendering, Artifacts, Workers for Platforms, AI Gateway).
Architecture
flowchart LR
UI["React app (chat + editor + preview)"] -->|"POST /api/agent"| API["Hono API: CodingAgentController"]
UI <-->|"WebSocket"| CGA["CodeGeneratorAgent DO"]
API --> CGA
CGA --> TB["ThinkCodingBehavior"]
TB -->|"configureVibe / chat()"| TA["ThinkAgent DO (@cloudflare/think loop)"]
TA --> GW["AI Gateway (OpenAI-compatible)"]
TA -->|"read/write/edit/grep, commit, deploy_space"| SP["SpaceDO (files + git)"]
TA --> BR["Browser Rendering: console logs"]
SP --> AR["Artifacts (git history)"]
SP -->|"LOADER.get()"| DW["Dynamic Worker preview + App facet"]
TB -->|"publish"| WFP["Workers for Platforms namespace"]
| Component | Path | Role |
|---|---|---|
| API routes | worker/api/routes/codegenRoutes.ts |
POST /api/agent creates an app; /api/agent/:id/ws attaches the editor |
| Agent controller | worker/api/controllers/agent/controller.ts |
Rate limits, resolves behavior and project type, starts the agent DO, streams NDJSON |
| Coding agent DO | worker/agents/core/codingAgent.ts |
One DO per app; picks the Think, Phasic or Agentic behavior |
| Think behavior | worker/agents/core/behaviors/think.ts |
Host side: configures ThinkAgent, drains user prompts, maps stream chunks to WebSocket events |
| ThinkAgent | worker/agents/think/ThinkAgent.ts |
The model/tool loop: model wiring, tools, step budget, context trimming |
| Think tools | worker/agents/think/*-tool.ts |
deploy_space, commit, set_title, ask_questions, get_browser_console_logs |
| SpaceDO | space/src/space/durable-object.ts |
Workspace file system, git, deploy, rollback, preview serving, app DB inspector |
| Deploy engine | space/src/space/deploy-engine.ts |
Bundles a branch with @cloudflare/worker-bundler and records the deployment |
| Legacy engines | worker/agents/core/behaviors/{phasic,agentic}.ts |
Blueprint/phase pipeline and recursive tool loop backed by a container sandbox |
| SDK | sdk/src/ |
VibeClient / BuildSession for driving builds over HTTP + WebSocket |
How a request flows
- Create. The home page always sends
behaviorType=thinkfor app projects (home.tsx).POST /api/agentreachesstartCodeGeneration, which checks prompt length and the app-creation rate limit, then callsresolveBehaviorType(controller.ts). For Think it skips template selection entirely, writes the WebSocket URL to the NDJSON stream and callsinitializeon the agent DO (controller.ts). - Pick an engine.
CodeGeneratorAgentinstantiatesThinkCodingBehavior,PhasicCodingBehaviororAgenticCodingBehaviorfrom props, persisted state, or the feature registry (codingAgent.ts, features/types.ts). - Configure the loop.
ThinkCodingBehavior.initializenames the project, thenconfigureThinkAgentresolves AI Gateway credentials, builds the per-app system prompt and pushes both into theThinkAgentDO withconfigureVibe(think.ts).seedEmptySpacecommits a marker file so the SpaceDO has amainbranch. - Run a turn. A
GENERATE_ALLWebSocket message leads tobuild(), which joins queued prompts and callsrunPrompt. That callsstub.chat(text, forwarder)on the ThinkAgent and translatestext-delta,reasoning-deltaandtool-*chunks into editor events (think.ts). - Model ↔ tools. Inside the ThinkAgent,
beforeTurntrims history,beforeStepmeters credits and forces a tool-free wrap-up on step 25, andgetToolsexposes SpaceDO-backedread/write/edit/list/find/grep/deleteplus the custom tools. Bash is off (ThinkAgent.ts). - Deploy and check. The system prompt tells the model to call
deploy_space, thenget_browser_console_logs, and to repeat until the build and the console are clean (think.ts).deploy_spacecommits the tree and callsSpaceDO.deploy(deploy-tool.ts), which bundles the branch and stores a deployment row (deploy-engine.ts). - Preview. The host sees the
deploy_spaceoutput, signs a branch-scoped preview token and broadcastsDEPLOYMENT_COMPLETED(think.ts). Requests to/space/:name/preview/:branch/hitservePreview. It serves static assets, then loads the app as a Dynamic Worker withLOADER.get()and forwards to itsAppclass running as a Durable Object facet (durable-object.ts).
Key components
ThinkAgent: the loop
ThinkAgent extends Think and overrides the hooks it needs (ThinkAgent.ts). maxSteps is 25 and workspaceBash is false. getModel builds an @ai-sdk/openai chat provider against the AI Gateway /compat endpoint. Its custom fetch patches Gemini’s streamed tool-call deltas and round-trips Gemini “thought signatures”, which the AI SDK would otherwise drop (ThinkAgent.ts). That much shim code is a sign of how tightly the loop is tuned to one model. THINK_MODEL_ID is hard-coded to google-ai-studio/gemini-3.6-flash (model-config.ts), and configureThinkAgent uses it regardless of the user’s per-action model settings.
Context handling
There is no embedding index. The model explores the workspace with find/grep/read. Before each turn, selectThinkContextMessages keeps the first message, the last five text-bearing messages, and the complete tool-call/result pairs for ask_questions, activate_skill and set_title. It strips every other tool call out of older assistant turns (context-selector.ts). The base prompt is picked per model family from prompts/*.txt (prompts.ts).
SpaceDO: files, git, previews
SpaceDO is published as a separate package (space/). Each app gets one instance, keyed by the agent id. It offers file RPCs (readFile, editFile, grep, patch), git operations (gitCommit, gitLog, gitDiff), and deploy, rollbackToCommit and servePreview. With ENABLE_ARTIFACTS on, history is pushed to Cloudflare Artifacts after each successful deploy (durable-object.ts). Rollback restores the target commit’s bytes as a new commit, so history is never rewritten (durable-object.ts). The generated app must export class App extends DurableObject. SpaceDO wraps it with inspector methods, which drive the read-only DB tab.
Verification in a real browser
get_browser_console_logs opens the preview in Cloudflare Browser Rendering (or a local sidecar in development). It returns console output, page errors and failed requests, and it accepts an optional interact_script to click through the page first (browser-logs-tool.ts). Together with build errors from deploy_space, this is the agent’s only feedback signal. There is no test runner and no type-checker in the Think path.
Legacy Phasic / Agentic engines
These engines use worker/agents/inferutils/core.ts with chat-completions tool calling, structured Zod outputs, and search/replace or unified-diff parsers in output-formats/diff-formats/. They run code in a @cloudflare/sandbox container image (SandboxDockerfile). Before deploy they run preDeploySafetyGate, a Babel AST scan for React render-loop hazards such as module-level JSX, unstable selectors, and setState during render.
Extending it
- Think tools. Add a
tool()factory underworker/agents/think/and register it inThinkAgent.getTools()(ThinkAgent.ts). New SpaceDO calls need typings inspace-workspace-ops.ts. - Skills.
worker/agents/think/skills/*/SKILL.md(four today: app file structure and three frontend-design variants) are injected into the system prompt by Think’s skill loader. - Prompts. Per-family base prompts live in
worker/agents/think/prompts/. The VibeSDK-specific deploy-and-verify workflow is appended inbuildSystemPrompt. - MCP.
MCPManagercan connect to SSE MCP servers for the legacy engines, but theMCP_SERVERSlist is empty at this SHA. - Headless.
sdk/exposesVibeClient.build(prompt), which returns aBuildSessionwithwaitForPreviewDeployed()and similar helpers, so you can drive builds from scripts or tests.
Running it
VibeSDK is a deployment, not a library you npm install into a project. Use the “Deploy to Cloudflare” button, or run bun run setup and then bun run dev for local work. Production needs a Workers Paid plan with Workers for Platforms for publishing generated apps, a D1 database, R2/KV, Containers (for the legacy sandbox), Browser Rendering, Worker Loaders, AI Gateway with at least one provider key, and a custom domain with wildcard DNS for preview subdomains. All of these are declared in wrangler.jsonc. Feature toggles such as ENABLE_ARTIFACTS are set in the dashboard, not in the repo.
Strengths and caveats
- Strength: preview without a VM. Generated apps are bundled in-process and served as Dynamic Workers inside the SpaceDO. No container needs to boot for the default Think path, so preview latency is close to a Worker deploy.
- Strength: real per-app isolation. Each app has its own agent DO, ThinkAgent DO, SpaceDO, git history and App facet storage. Previews need signed, branch-scoped tokens.
- Strength: honest restore points. The model decides when to
commit, and rollback re-commits old content, so history is never rewritten. - Caveat: Cloudflare-only by design. Nearly every subsystem is a Cloudflare binding. You cannot run this elsewhere, and generated apps must target Workers + Durable Objects.
- Caveat: one model for the main path. Think is hard-wired to a Gemini Flash model through AI Gateway, with Gemini-specific stream patches. The rich per-action model configuration UI only affects the legacy engines.
- Caveat: two architectures in one repo. Think, Phasic and Agentic share
BaseCodingBehavior(about 2,000 lines) but have different tools, sandboxes and edit formats. Reading the code, it is easy to mistake legacy paths for live ones. - Caveat: stale scripts.
package.jsonstill definescliandtuiscripts, but thecli/directory they point to does not exist at this commit.
Sources: code at 9da158d, deepwiki-open wiki (11 pages), verified Q&A.
How it answers the Open-source coding agents questions
Each answer was drafted by a code-reading agent at commit 9da158d. Its citations were checked mechanically. Compare with the other open-source coding agents →
How is the agent loop implemented?
answeredVibeSDK has three distinct agent-loop architectures, selected per app at creation: Think (model-and-tool loop via @cloudflare/think), Agentic (recursive inference with tool calling), and Phasic (deterministic phase orchestration).
The Think loop is the primary/canonical loop, implemented in ThinkAgent (extends @cloudflare/think) at worker/agents/think/ThinkAgent.ts. It is a single model-and-tool loop, not a planner+executor split. Think.getTools() returns a ToolSet of SpaceDO-backed workspace tools (read, write, edit, list, find, grep, delete) from @cloudflare/think/tools/workspace, plus VibeSDK-specific tools: commit, deploy_space, set_title, ask_questions, get_browser_console_logs. Bash is disabled (workspaceBash = false). Each turn is driven by ThinkAgent.chat(), called from ThinkCodingBehavior.runPrompt() which submits user text to the DO and translates streamed UIMessageChunk frames into WebSocket events. Tools use the OpenAI-compatible function-calling schema.
The Agentic loop (worker/agents/core/behaviors/agentic.ts) uses the infer() function in worker/agents/inferutils/core.ts — a generic inference engine that assembles tool-call requests, streams responses, accumulates tool-call deltas from chat-completions chunks, executes tools via executeToolCallsWithDependencies(), and recurses until a depth limit or CompletionDetector signal fires. Stop conditions include: max tool-calling depth per action key, a max message limit (64000 in constants), and completion signals from specific tools. The Agentic loop also includes conversation compactification triggers every 9 tool calls.
The Phasic loop (worker/agents/core/behaviors/phasic.ts) is a deterministic orchestrator that runs blueprint generation, phase generation, phase implementation, and review cycles in a fixed sequence, then iterates with the user.
Think has no sub-agents in the traditional sense; it issues tool calls to the SpaceDO workspace and other services. The infer() function does support recursive tool calling (model calls tool, tool result feeds back to model), which is its form of sub-loop.
Turn structure (Think): beforeTurn() selects context messages, beforeStep() enforces rate limits and injects a max-steps prompt on the final step. getModel() wraps an OpenAI-compatible provider over AI Gateway with custom SSE patching for Google Gemini compatibility.
How is repository context gathered and kept within the context window?
answeredRepository context is gathered and managed through several mechanisms:
Codebase reading tools: In the Think loop, workspace tools include read, list, find, and grep from @cloudflare/think/tools/workspace which operate on the SpaceDO workspace. The ThinkAgent has no separate glob tool beyond the generic workspace tools, but the system prompt instructs the model to use search tools extensively. The Gemini system prompt (worker/agents/think/prompts/gemini.txt) explicitly tells the agent: "Use grep and glob search tools extensively (in parallel if independent) to understand file structures, existing code patterns, and conventions. Use read to understand context and validate any assumptions."
Context selection: ThinkAgent.beforeTurn() calls selectThinkContextMessages() in worker/agents/think/context-selector.ts to trim the message history before each turn. It preserves all tool-result pairs for tool names matching a protected list (deploy_space, commit, ask_questions), keeps a minimum of RECENT_MESSAGE_COUNT (5) recent messages, and uses a diff-based strategy to preserve complete tool-call/tool-result pairs around named tools rather than naive truncation.
Conversation compactification: The older Agentic/Phasic behaviors use worker/agents/utils/conversationCompactifier.ts which triggers summarization at 40 conversation turns or 100k estimated tokens. It preserves the last 10 messages uncompacted, calls an LLM to summarize older history into a concise context block wrapped in <system_context> tags, and replaces compacted messages with the summary.
File-reading strategy: The Think model is instructed to read files to understand context before making changes. The workspace operations layer reads and writes through SpaceDO RPC calls — not local filesystem — so file I/O goes through Durable Object stubs. The codebaseContext.ts utility provides a CodebaseContextBuilder for the Agentic behavior that can scan directories, collect file contents, and build structured context blocks for the LLM.
VibeSDK does not use embeddings or vector search for repository context. It relies entirely on LLM-driven search (grep/find within workspace tools) and conversation history management.
How are code edits applied?
answeredCode edits are applied through two parallel systems depending on the agent behavior:
ThinkAgent (primary): Edits happen via SpaceDO-backed workspace tools registered in ThinkAgent.getTools(). These use @cloudflare/think/tools/workspace tool creators (createReadTool, createWriteTool, createEditTool, createDeleteTool) backed by SpaceWorkspaceOps which route through Durable Object RPC to SpaceDO. The model can write (whole file), edit (search/replace on existing files), or delete. Edits land directly in SpaceDO's git-backed workspace. Think has no separate diff format validation — the edit tool uses Think's built-in search/replace with exact matching. The model is instructed to keep changes minimal and to deploy after editing.
Agentic/Phasic behaviors: These use a more elaborate system. The agent emits structured outputs matching Zod schemas in worker/agents/schemas.ts — FileOutputSchema for whole-file writes, with a format field supporting full_content or unified_diff. Two diff parsers exist in worker/agents/output-formats/diff-formats/: search-replace.ts (a <<<<<<< SEARCH / ======= / >>>>>>> REPLACE format with multiple matching strategies: exact, whitespace-insensitive, indentation-preserving, fuzzy) and udiff.ts (unified diff format with indentation normalization). Both parse the model's text output into structured edits.
Validation: The Agentic/Phasic flow uses the preDeploySafetyGate utility (worker/agents/utils/preDeploySafetyGate.ts) which runs before deployment to check for hardcoded secrets/API keys in generated code. Static analysis is performed during build. After deploy, browser console logs are inspected. The Think flow deploys via deploy_space and checks browser logs via get_browser_console_logs, forming an automatic repair loop.
Undo/git: All edits go through SpaceDO, which is backed by Cloudflare Artifacts for durable git history. Each commit tool call creates a restore point. rollbackToCommit() in ThinkCodingBehavior restores a prior commit and redeploys without rewriting history. There is no per-edit undo — users roll back to a named commit.
How are shell commands and file writes kept safe?
answeredExecution safety is implemented at multiple layers:
Tool boundaries — ThinkAgent: Bash execution is explicitly disabled (workspaceBash = false in ThinkAgent). The Think loop exposes only file workspace tools (read, write, edit, list, find, grep, delete), git/commit tools, deploy tools, browser-log inspection, and a structured question-asking tool. No shell or arbitrary command execution is possible through the Think agent.
Sandbox for Agentic/Phasic behaviors: The older agent paths use @cloudflare/sandbox (v0.5.6) for isolating command execution. The SandboxDockerfile builds a Docker image based on cloudflare/sandbox:0.5.6 with git, curl, and a process-monitoring system (container/ directory with cli-tools.ts, process-monitor.ts, storage.ts). Commands are executed inside the container via the BaseSandboxService in worker/services/sandbox/BaseSandboxService.ts which manages sandbox instances, file uploads, command execution, static analysis, and deployment results.
Process monitoring: The container includes a ProcessMonitor class in container/process-monitor.ts and a CLI tool (container/cli-tools.ts) that validates instance IDs with a strict regex (/^[a-zA-Z0-9][a-zA-Z0-9_-]*$/), caps ID length at 64 characters, and stores logs and errors in a managed SQLite storage layer within the container.
Preview authorization: Preview URLs use signed, branch-scoped tokens. getBrowserPreviewURL() in ThinkCodingBehavior invokes signSpacePreviewToken() to create a JWT-like token embedding spaceName, branch, userId, and a previewVersion. The token bootstraps an HttpOnly preview cookie and prevents unauthorized access to preview deployments.
Network restrictions: The sandbox container image does not appear to restrict network access, but the Think agent has no bash or curl tools to make arbitrary network requests.
Secrets: User API keys and tokens are handled through Cloudflare services (AI Gateway bindings, encrypted token blobs). The preDeploySafetyGate in worker/agents/utils/preDeploySafetyGate.ts checks generated code for hardcoded secrets before deployment.
Durable Object isolation: Each app maps to its own ThinkAgent DO, SpaceDO workspace, Artifacts repository, and generated App Facet — full resource isolation per project. VibeSDK does not use Landlock, seccomp, or other OS-level sandboxing beyond the Docker container and Cloudflare's DO isolation.
Which models are supported and how are they called?
answeredVibeSDK supports models across multiple providers through Cloudflare AI Gateway routing. The model master list is defined in worker/agents/inferutils/config.types.ts with over 20+ models classified by size (Lite, Regular, Large), each with a provider, creditCost, and contextSize.
Supported providers: Google AI Studio (Gemini 2.5 Flash/Pro, Gemini 3 Flash/Pro Preview), Anthropic (Claude 3.7 Sonnet, Claude 4/4.5 Sonnet/Opus/Haiku), OpenAI (GPT-5, GPT-5.1, GPT-5.2, GPT-5 Mini), Grok (Grok 4 Fast, Grok 4.1 Fast, Grok Code Fast 1), Vertex AI (GPT-OSS 120B, Kimi K2 Thinking, Qwen 3 Coder 480B). Several models (Cerebras, OpenAI OSS/Codex variants) are commented out in the master list.
ThinkAgent model: Configured in worker/agents/think/model-config.ts with THINK_MODEL_ID = 'google-ai-studio/gemini-3.6-flash'. The default config in worker/agents/inferutils/config.ts selects a Gemini-only set; when PLATFORM_MODEL_PROVIDERS env var is set, it uses a multi-provider config with Gemini 3 Pro Preview for blueprints, Grok 4.1 Fast for conversational response and deep debugging, and OpenAI 5 Mini as fallback.
How models are called: All model calls go through AI Gateway. The getConfigurationForModel() function in worker/agents/inferutils/core.ts resolves a model's baseURL, apiKey, and headers by checking: (1) runtime overrides (SDK BYOK keys), (2) user's encrypted gateway token, (3) user's custom gateway override, (4) environment variable per provider (${PROVIDER}_API_KEY), (5) platform AI Gateway binding. The OpenAI SDK (openai package) is used as the inference client with client.chat.completions.create(). The ThinkAgent uses @ai-sdk/openai via createOpenAI() with a custom fetch wrapper for Gemini compatibility patching.
Per-model prompt tuning: The Think agent has tuned system prompts per model family in worker/agents/think/prompts/: gemini.txt, gpt.txt, anthropic.txt, trinity.txt, kimi.txt, codex.txt, beast.txt (for GPT-4/o1/o3), and default.txt. These differ in tone, detail, and instructions; e.g., the Gemini prompt emphasizes security and tool-use patterns, while the Anthropic prompt is more concise and direct.
Cost tracking: Each model has a creditCost relative to GPT-5 Mini baseline (1.0 credit = $0.25/1M input). Rate limiting is enforced via RateLimitService.enforceLLMCallsRateLimit() which checks configurable budgets per user. Multi-provider configs have per-action model assignments with fallback chains (e.g., blueprint: Gemini 3 Pro Preview → Gemini 2.5 Flash).
No local model support: All models are called remotely through AI Gateway. There is no Ollama, llama.cpp, or local model infrastructure.
How can it be extended and customised?
answeredVibeSDK can be extended and customized through several mechanisms:
Skills (Think agent): The Think agent supports a skill catalog via worker/agents/think/skills.ts. Skills are SKILL.md files in worker/agents/think/skills/ with frontmatter (name, description, compatibility, allowedTools, metadata). Currently four skills: app-file-structure, frontend-design, frontend-design-landing-page, frontend-design-saas. Each skill is loaded at build time via ?raw import, parsed with agents/skills frontmatter parser, and exposed as a SkillSource via fromManifest(). Think automatically injects the skill catalog into the system prompt and registers a skill-loading tool.
Custom tools (Think): Adding a Think tool requires: (1) creating the tool under worker/agents/think/, (2) adding SpaceDO RPC typing to space-workspace-ops.ts, (3) registering it in ThinkAgent.getTools(), and (4) updating the relevant prompt or skill. Existing custom tools include ask-questions-tool.ts, browser-logs-tool.ts, commit-tool.ts, deploy-tool.ts, set-title-tool.ts.
Toolkit system (Agentic): The Agentic/Phasic behaviors use a different tool registration in worker/agents/tools/toolkit/ (e.g., exec-commands.ts, generate-files.ts, read-files.ts, deploy-preview.ts) with a buildTools()/buildDebugTools() pattern in worker/agents/tools/ registered via customTools.ts.
MCP (Model Context Protocol): VibeSDK includes an MCPManager class in worker/agents/tools/mcpManager.ts that uses @modelcontextprotocol/sdk/client with SSEClientTransport to connect to MCP servers. It connects via SSE, lists tools, and integrates them into the agent's toolset. The MCP_SERVERS array in mcpManager.ts is currently empty (just a commented-out cloudflare-docs example), but the full connection infrastructure is in place.
Rules/instruction files: The project has CLAUDE.md (project instructions for Claude Code) and AGENTS.md (detailed guidance for AI agents covering tooling, verification, frontend UI, boundaries, change paths, and constraints). These are checked into the repository.
SDK/headless use: The sdk/ directory is an independent Bun package providing a programmatic client. VibeClient (sdk/src/client.ts) offers HTTP methods, AgenticClient and PhasicClient extend it with preset behavior types, and BuildSession (sdk/src/session.ts) manages a live agent session with WebSocket streaming, file watching, and phase tracking. Applications can install the SDK as a dependency and script agent interactions.
CLI/TUI: bun run cli and bun run tui in package.json point to cli/index.ts and cli/tui.tsx for command-line and terminal-UI interaction modes.
Configuration: Model configs, agent action configs, per-action constraints (allowedModels, enabled), rate limit settings, and user-configurable settings are all adjustable without code changes through environment variables and AI Gateway configuration.