LLMs Technical Reviews

stackblitz-labs/bolt.diy

Open-source bolt.new: a Remix prompt-to-app builder that streams LLM file and shell actions into an in-browser WebContainer.

GitHub ↗★ 20kTypeScriptMITcommit db479e8 · 2026-10-04homepage ↗

Overview

bolt.diy is the community open-source version of StackBlitz’s bolt.new: a prompt-to-app builder that runs in the browser. You describe an app. The model streams back a <boltArtifact> of file writes and shell commands. Your browser tab executes them inside a WebContainer, StackBlitz’s WebAssembly Node.js runtime, and shows the dev server in a preview pane next to an editor and terminal. You can then push the result to GitHub or GitLab or deploy it to Netlify or Vercel. It is not a terminal or IDE coding agent working on your local repository. The project lives in the tab’s virtual file system, with chats persisted in IndexedDB.

That split is the key architectural fact. The server (a Remix app on Cloudflare Pages, Electron or Docker) only builds prompts and streams LLM text. All execution happens client-side. useMessageParser parses the stream as it arrives, and an ActionRunner writes files with webcontainer.fs.writeFile and runs commands in a jsh shell. Nothing runs on a server VM, so the sandbox is the browser itself.

The agent is a single streamed completion per user turn, not a tool-calling loop over the workspace. The optional “context optimization” mode adds two cheap pre-calls, a chat summary and a file selection. MCP tools can be exposed with maxSteps, but the model edits files only through the artifact text format.

Architecture

flowchart LR
  subgraph Browser
    CH["Chat.client.tsx (useChat)"] --> MP["EnhancedStreamingMessageParser"]
    MP --> WB["workbenchStore"]
    WB --> AR["ActionRunner"]
    AR --> WC["WebContainer (fs + jsh shell)"]
    WC --> PV["Preview iframe"]
    CH --> IDB["IndexedDB chat history"]
  end
  CH -->|"POST /api/chat"| API["api.chat.ts chatAction"]
  API --> SUM["createSummary"]
  API --> SEL["selectContext"]
  API --> ST["streamText (stream-text.ts)"]
  ST --> LLM["LLMManager providers"]
  API --> MCP["MCPService"]
Component Path Role
Chat UI app/components/chat/Chat.client.tsx useChat from @ai-sdk/react; sends files, mode, prompt id and MCP step limit
Chat endpoint app/routes/api.chat.ts Summary, context selection, main streamText, continuation on length
Prompt + model call app/lib/.server/llm/stream-text.ts Picks provider and model, injects context buffer and locked files, sets token params
Context helpers app/lib/.server/llm/create-summary.ts, select-context.ts LLM-written chat summary and an <updateContextBuffer> file picker
Stream parser app/lib/runtime/message-parser.ts, enhanced-message-parser.ts Incremental <boltArtifact>/<boltAction> parser with a code-block fallback
Action runner app/lib/runtime/action-runner.ts Serial queue executing file, shell, start, build, supabase actions
WebContainer boot app/lib/webcontainer/index.ts WebContainer.boot, preview error forwarding, inspector script
Providers app/lib/modules/llm/ BaseProvider, 22 provider classes, LLMManager registry
MCP app/lib/services/mcpService.ts stdio, SSE and streamable-HTTP MCP clients with approve/reject gating
Prompts app/lib/common/prompts/, prompt-library.ts default, original and optimized system prompts

How a request flows

  1. Send. useChat posts messages, the current files map, promptId, contextOptimization, chatMode, Supabase state and maxLLMSteps to /api/chat (Chat.client.tsx). The chosen model and provider travel inside the user message text as [Model: …] and [Provider: …] prefixes (constants.ts). API keys come from cookies (api.chat.ts).
  2. Resolve MCP results. processToolInvocations runs any tool the user approved in the previous turn, or records a denial (mcpService.ts).
  3. Optimize context (optional). If files exist and context optimization is on, createSummary writes a structured chat summary and selectContext asks the model which files belong in the context buffer. Both are reported to the UI as progress annotations (api.chat.ts).
  4. Stream. streamText looks up the model, builds the system prompt from PromptLibrary, appends the context buffer and summary (and then keeps only the last three messages), lists locked files, and calls the AI SDK. Reasoning models get maxCompletionTokens and lose sampling parameters (stream-text.ts). In discuss mode a different, non-building prompt is used.
  5. Continue. If the model stops with finishReason === 'length', onFinish appends a continue prompt and streams again, at most MAX_RESPONSE_SEGMENTS = 2 times (api.chat.ts, constants.ts).
  6. Parse in the browser. For each chunk, EnhancedStreamingMessageParser fires onArtifactOpen, onActionOpen, onActionStream and onActionClose. File actions are queued as soon as they open and stream into place. Shell and start actions are queued when they close (useMessageParser.ts).
  7. Execute. ActionRunner.runAction chains every action onto one promise, so they run strictly in order (action-runner.ts). Files are written with mkdir + writeFile. Shell commands go through BoltShell after a pre-check. start runs the dev server without blocking (action-runner.ts).
  8. Preview. WebContainer’s server-ready and port events populate the preview list. Uncaught exceptions inside the preview are forwarded back as an action alert, which the user can send to the model with one click (webcontainer/index.ts).

Key components

Artifact protocol

The system prompt describes the WebContainer: no native binaries, stdlib-only Python without pip, no C compiler, no git. It insists on full-file rewrites because “WebContainer CANNOT execute diff or patch editing” (prompts.ts). The model replies with one <boltArtifact> containing ordered <boltAction type="file|shell|start|build|supabase"> children. StreamingMessageParser.parse is a resumable character scanner that keeps per-message state, so a half-received tag is picked up on the next chunk (message-parser.ts). The enhanced subclass wraps bare markdown code blocks into artifacts when a model ignores the format.

Context optimization

Without it, every request carries the full conversation. With it, the server spends two extra LLM calls: a templated summary, then an <includeFile>/<excludeFile> selection. The selection prompt asks for at most five files, but that limit is only an instruction to the model (select-context.ts). The selected files are re-serialised as boltAction entries in the system prompt. There is no embedding search.

Providers

There are 22 provider classes, among them OpenAI, Anthropic, Google, OpenRouter, Bedrock, Mistral, DeepSeek, Ollama and LM Studio (registry.ts). Each extends BaseProvider, which resolves keys and base URLs from cookies, UI settings or env and can fetch dynamic model lists. At this SHA, however, LLMManager registers only OpenRouter, Anthropic, OpenAI and Google. The rest are skipped and never reach the model picker (manager.ts).

MCP with approval

MCPService connects to servers configured in the UI over stdio, SSE or streamable HTTP and builds an AI SDK ToolSet. api.chat.ts passes toolsWithoutExecute to the model, so a tool call becomes a UI prompt. The tool runs on the next request only if the user’s result is exactly TOOL_EXECUTION_APPROVAL.APPROVE. The default maxLLMSteps is 5 (mcp.ts). MCP is the only place where bolt.diy has a real tool loop, and it runs on the server, not in the WebContainer.

Extending it

  • Providers. Add a BaseProvider subclass under app/lib/modules/llm/providers/, export it from registry.ts, and add its name to ENABLED_PROVIDERS in manager.ts. Without that last step it stays invisible.
  • Prompts. Register a template in PromptLibrary. Users switch between default, original and optimized in settings.
  • Tools. Configure MCP servers in the settings UI (stored in localStorage and pushed to /api/mcp-update-config).
  • Integrations. Deploy and VCS hooks live in app/components/deploy/ (GitHub, GitLab, Netlify, Vercel), and Supabase migrations and queries are a supabase action type.

Running it

Run pnpm install and then pnpm run dev (Remix + Vite; Node ≥ 18.18). For production, pnpm run build and then pnpm run start, which runs wrangler pages dev. You can also deploy to Cloudflare Pages, use the provided Dockerfile, or build the Electron desktop app (electron:build:*). Provider keys go in .env.local or the settings UI. The browser needs cross-origin isolation for WebContainers, which the app configures. The README notes that WebContainer itself needs a commercial StackBlitz licence for for-profit production use.

Strengths and caveats

  • Strength: zero server-side execution. Generated code, npm install and the dev server all run in the user’s tab, so a self-hosted instance needs no sandbox infrastructure and no per-user VMs.
  • Strength: streaming feels live. File actions write into the editor while the model is still typing, and preview errors loop back with one click.
  • Strength: broad integration surface. Git import/export, Netlify/Vercel deploy, Supabase, MCP, Electron and a choice of prompts.
  • Caveat: WebContainer limits. JS/WASM only: Python is limited to the standard library, there is no pip and no git binary, and native modules and native-binary databases do not work. Browser support for WebContainers is narrower than for ordinary web apps.
  • Caveat: full-file rewrites. Every edit re-emits whole files, which is costly on large files and runs into output-token ceilings. The continuation hack allows only two extra segments.
  • Caveat: “any LLM” is aspirational at this commit. Only four providers are registered. Ollama, LM Studio and the other providers need a code change despite the README’s setup sections.
  • Caveat: thin server hardening. /api/chat has no rate limiting. withSecurity wraps only the GitHub/GitLab/Netlify/Vercel/Supabase routes. Locked files are protected by a prompt instruction, not by the runner.

Sources: code at db479e8, verified Q&A.

How it answers the Open-source coding agents questions

Each answer was drafted by a code-reading agent at commit db479e8. Its citations were checked mechanically. Compare with the other open-source coding agents →

How is the agent loop implemented?

answered

There is no Planner sub-agent — Bolt uses a single-turn streaming LLM loop orchestrated from the server. The entry point is app/routes/api.chat.ts, the server-side action handler for /api/chat. The client (Chat.client.tsx) calls useChat from the @ai-sdk/react SDK which sends a POST to that route. On the server, chatAction creates a SwitchableStream and runs a createDataStream block with an execute callback (line 93). Inside that callback, it first processes MCP tool invocations with mcpService.processToolInvocations (line 103), then optionally runs createSummary (line 122) and selectContext (line 164) — two separate LLM calls that create a chat summary and choose relevant files. The main generation is streamText (line 311) which calls the Vercel AI SDK's streamText with model, system prompt, and messages. Tool calls are exposed to the LLM via mcpService.toolsWithoutExecute set as options.tools (line 214), with toolChoice: 'auto' and maxSteps configurable per request. When the LLM response hits finishReason === 'length', the loop constructs a new user message with CONTINUE_PROMPT and calls streamText again, up to MAX_RESPONSE_SEGMENTS = 2 times (lines 251-298). The client uses @ai-sdk/react's useChat which handles streaming state, abort via stop(), and renders tool-call annotations in the UI.

How is repository context gathered and kept within the context window?

answered

Context gathering is a multi-phase server-side process triggered on every request when contextOptimization is enabled. First, createSummary (app/lib/.server/llm/create-summary.ts) sends all processed messages to the LLM with a structured prompt requesting a summary in a defined template format (project overview, decisions, implementation status, etc.). The summary is stored as a chatSummary annotation on the data stream. Next, selectContext (app/lib/.server/llm/select-context.ts) retrieves all file paths from the project (filtered by an ignore-list similar to .gitignore via the ignore package), reads the current context buffer from annotations, and asks the LLM which files are relevant by outputting <updateContextBuffer> XML with <includeFile> / <excludeFile> tags. Only 5 files are selected at a time (line 170). The file content is read from the in-memory FileMap (populated by FilesStore from the WebContainer filesystem) and formatted as boltAction type="file" entries inside a boltArtifact wrapper via createFilesContext in app/lib/.server/llm/utils.ts. For long sessions, messageSliceId is set to messages.length - 3 (api.chat.ts line 106-107), keeping only the last 3 messages, while the summary carries forward the earlier history. The summary and code context are injected into the system prompt before the main LLM call. The system prompt builder getSystemPrompt includes environment constraints, database instructions, artifact formatting rules, and design guidance — all composed into a single system message.

How are code edits applied?

answered

Code edits use a structured XML artifact format parsed in real-time during streaming. The LLM emits <boltArtifact> containing <boltAction type="file"> tags with filePath attributes and full file content — Bolt does NOT use diffs or search/replace due to WebContainer limitations (the system prompt explicitly states this at prompts.ts line 36). The StreamingMessageParser (app/lib/runtime/message-parser.ts) incrementally parses the stream: when it sees <boltArtifact> it emits an onArtifactOpen callback (line 289), then for each <boltAction type="file"> it extracts the filePath and accumulates content until </boltAction> closes. On close, it calls onActionClose (line 165) which triggers the ActionRunner. The ActionRunner (app/lib/runtime/action-runner.ts) executes each action sequentially via #executeAction (line 150). For file actions, #runFileAction (line 310) uses WebContainer's filesystem: it resolves the relative path, creates parent folders via mkdir -p, and writes the full file content with webcontainer.fs.writeFile. The EnhancedStreamingMessageParser (app/lib/runtime/enhanced-message-parser.ts) adds a fallback: if the LLM output contains no boltArtifact tags at all, it scans markdown code blocks and wraps them into artifacts automatically using file-path heuristics. There is no git integration for undo — file history is saved separately via saveFileHistory (action-runner.ts line 359) which stores previous versions as JSON under .history/. Shell commands (type="shell") are executed via the terminal emulator, and type="start" launches dev servers. Validation includes pre-flight shell command checks (#validateShellCommand at line 577) that verify paths exist before rm/cd/cp/mv operations.

Editor's note. Correction: saveFileHistory/getFileHistory have no callers, so no .history/ snapshots are written; there is no undo beyond what the chat UI keeps.

How are shell commands and file writes kept safe?

answered

Code execution runs inside WebContainer, an in-browser sandboxed Node.js runtime that cannot access the host system, files, or network beyond what the browser permits. The WebContainer.boot() call at app/lib/webcontainer/index.ts line 26 uses coep: 'credentialless' for cross-origin embedding policy and manages execution through virtual filesystem and shell APIs. Shell commands are validated before execution in #validateShellCommand (action-runner.ts:577-670) — it intercepts rm commands to add -f flags when files don't exist, checks that cd targets exist (auto-creating them with mkdir -p), and verifies source files before cp/mv. Locked files are enforced at the prompt level: the system prompt lists locked files and instructs the LLM not to modify them (stream-text.ts lines 202-223). On the security side, app/lib/security.ts provides rate limiting per-IP (10 req/min for LLM, 100 for other API endpoints), security headers (CSP, X-Frame-Options, HSTS), and API key format validation. MCP tool execution requires user approval via TOOL_EXECUTION_APPROVAL: in mcpService.ts line 402, the tool's execute function only runs when the result equals 'Yes, approved.', and rejections return TOOL_EXECUTION_DENIED. The maxSteps parameter (default 5, from MCP store) limits the number of LLM reasoning steps. Network requests from within WebContainer are restricted by browser policies — the system prompt warns against native binaries and restricts package choices to JS-only implementations.

Editor's note. Correction: the security.ts rate limits are not applied to the LLM endpoints. withSecurity wraps only the GitHub/GitLab/Netlify/Vercel/Supabase routes; /api/chat and /api/llmcall have no rate limiting.

Which models are supported and how are they called?

answered

Bolt supports 20+ providers registered in app/lib/modules/llm/registry.ts. Only 4 are enabled by default via ENABLED_PROVIDERS in the manager (OpenRouter, Anthropic, OpenAI, Google), but all 22 registered providers (Anthropic, OpenAI, Google, Groq, Cohere, DeepSeek, HuggingFace, Mistral, Ollama, OpenRouter, LMStudio, AmazonBedrock, Perplexity, Together, xAI, Fireworks, Hyperbolic, Cerebras, GitHub, Moonshot, OpenAILike, Z-AI) can be enabled. Each provider extends BaseProvider (app/lib/modules/llm/base-provider.ts) which provides common infrastructure: getting base URL and API key from environment variables, cookie settings, or server env; Docker URL rewriting; and model caching with cache-key generation. The getModelInstance method creates an SDK model instance per provider — for example, AnthropicProvider uses createAnthropic() from @ai-sdk/anthropic, OllamaProvider uses createOllama() from ollama-ai-provider. Models can be static (hardcoded in the provider class, e.g., Anthropic's three fallback models) or dynamic (fetched at runtime via getDynamicModels — Ollama fetches from /api/tags, Anthropic from api.anthropic.com/v1/models). The LLMManager singleton (app/lib/modules/llm/manager.ts) orchestrates provider registration and model list updates. It caches dynamic models based on a cache key derived from API keys and settings. The Vercel AI SDK serves as the unified interface: streamText in the SDK handles streaming, tool calling, token counting, and error handling. Reasoning models (o1, o3, GPT-5) get special handling — they use maxCompletionTokens instead of maxTokens, and have unsupported parameters (temperature, topP, etc.) filtered out (stream-text.ts lines 240-280). Per-model prompt tuning is handled by the PromptLibrary which supports swapping system prompts (default, original, optimized). No cost tracking is implemented in this codebase.

Editor's note. Correction: at this commit LLMManager registers only OpenRouter, Anthropic, OpenAI and Google (ENABLED_PROVIDERS in manager.ts). The other 18 provider classes, including Ollama and LM Studio, are skipped and need a source edit to become usable.

How can it be extended and customised?

answered

The primary extension mechanism is MCP (Model Context Protocol) via MCPService (app/lib/services/mcpService.ts). Users configure MCP servers in the UI (stored in localStorage), supporting three transport types: stdio (local subprocess), sse (Server-Sent Events), and streamable-http. The MCPService validates configs with Zod schemas, creates clients via experimental_createMCPClient from the AI SDK, registers tools into a shared ToolSet, and handles tool call execution with user approval. Tools are exposed to the LLM through the streaming options as tools: mcpService.toolsWithoutExecute (with execute functions stripped), and actual execution happens via processToolInvocations (mcpService.ts:377). New LLM providers can be added by extending BaseProvider in app/lib/modules/llm/providers/ and registering in registry.ts — the LLMManager dynamically discovers them. The system prompt is customizable via PromptLibrary (app/lib/common/prompt-library.ts) which provides a plugin-like registry of prompt templates — users can select between default, original, or optimized prompts. The contextOptimization feature is toggleable per-chat. For rules/instructions files, Bolt does not support AGENTS.md or similar — the system prompt is hardcoded. The app runs as both a web app (Remix/Cloudflare) and an Electron desktop app (electron/ directory). The headless/SDK path is through Cloudflare Workers actions: the api.chat.ts route accepts POST requests with messages, files, and settings as JSON. WebContainer provides the sandboxed execution environment. Settings and provider configurations are stored in cookies (API keys) and localStorage (MCP config, provider settings). No plugin system or custom-tool registration API beyond MCP is provided.