LLMs Technical Reviews

CopilotKit/OpenBot

Self-hosted platform where AG-UI agents get their own browser computer, with every action policy-checked and audited first.

GitHub ↗★ 6.1kTypeScriptMITcommit f4bc60b · 2026-10-05homepage ↗

Overview

OpenBot is a self-hosted platform for running AI “coworkers” (Bots) that each get their own computer: a Chromium browser, a /workspace file area and a shell, all in a container. A person talks to a Bot in a web, desktop (Tauri) or mobile (Expo) client. The Bot can browse, read and write files, call third-party tools through MCP or Composio, and hand work to another Bot.

The core idea is that the server owns every side effect. Bots never reach their computer directly. Each browser or file action goes through one gateway, which resolves the element on the server, checks a CEL policy, writes an audit row and only then acts. A second layer, the approvals service, decides whether to act, ask the person, hand the action off to them, or refuse.

OpenBot is built on CopilotKit. The server runs @copilotkit/runtime/v2 in Intelligence mode, so threads, memory and “learning” depend on a CopilotKit Intelligence project (managed, or self-hosted). Bots are AG-UI agents. Built-in ones run in-process as BuiltInAgent. Remote ones are any AG-UI endpoint (LangGraph, Mastra, Pydantic AI, CrewAI, ADK and others), reached as HttpAgent.

Architecture

flowchart LR
  UI["app (React + Vite)"] --> API["server (Hono on Bun)"]
  Desk["desktop / mobile"] --> API
  API --> RT["CopilotRuntime"]
  RT --> INT["CopilotKit Intelligence"]
  RT --> BI["Built-in Bots"]
  RT --> RB["Remote AG-UI Bots"]
  BI --> GW["Computer gateway"]
  RB --> GW
  GW --> POL["CEL policy"]
  GW --> APR["Approvals service"]
  GW --> AUD["Audit (Postgres)"]
  GW --> SUP["supervisor"]
  SUP --> COMP["agent-computer (Chromium)"]
  BI --> PLG["Plugins: MCP, Composio"]
  WRK["worker (routines)"] --> API
Component Path Role
Web app app/ React UI with TanStack Router, built with Vite
API server server/src/app.ts Hono app: auth, admin, policy, audit, plugins, channels; mounts the CopilotKit handler
Runtime wiring server/src/copilot.ts Builds Bots, attaches middleware, mounts CopilotRuntime
Computer gateway server/src/computer/gateway.ts The only path from a Bot to its computer
Action policy server/src/computer/policy.ts CEL allow/deny rules, dry-run or enforce
Approvals server/src/approvals/ Allow / ask / hand off / deny per action, with optional LLM auto-review
Supervisor supervisor/, server/src/computer/supervisor.ts One computer container per Bot
Bot computer agent-computer/ Chromium, workspace files and shell behind a token
Plugins server/src/plugins/ MCP, Composio, Google Drive REST, OAuth, tool selection
Memory server/src/memory/ Personal memories in Postgres, injected by middleware
Worker worker/ Fires due routines by handing runs to the server
Bot adapters agent-*/ 14 example AG-UI Bots (frameworks plus a proof-of-concept agent-bot), and agent-computer

How a request flows

  1. The browser sends an AG-UI run to /api/copilotkit. The handler is mounted at the root with its own base path (app.ts).
  2. mountCopilotRuntime creates the CopilotRuntime with an Intelligence client and an identifyUser callback that scopes threads and memory to the signed-in person (copilot.ts).
  3. The runtime asks createRequestAgents for the Bots this person may see. The resolver builds them with buildAgents, which loads the granted tools, standing instructions and learned skills, then attaches PersonalMemoryMiddleware to each Bot (copilot.ts).
  4. The model picks a computer tool (for example computer_click). The call reaches ComputerGateway. Its private govern function loads the stored snapshot, resolves the ref to a real element, and builds a PolicyContext with the tool, page host, element, file, command and initiator (gateway.ts).
  5. evaluateActionPolicy runs deny rules first, then allow rules. No match means refused (policy.ts). The gateway writes the audit row whatever the outcome, and throws ActionRefusedError if the decision does not forward.
  6. If the policy allows, the gateway calls the approvalGate with an effect (write or read) and a scope. The approvals service may replay an earlier answer, pause for the person, or let it run (gateway.ts).
  7. The action is sent to the Bot’s computer, found through the supervisor. A failure after an allowed decision gets its own audit row.

Key components

Computer gateway

The header comment states the contract: resolve the ref from the server’s own snapshot, ask the policy, write the row, then act (gateway.ts). Resolving on the server matters: a model cannot evade “never click Submit” by labelling the click “Continue”. Every policy field is bound even when empty, because cel-js throws on unknown identifiers and a throwing deny rule counts as a match.

Policy and approvals

The CEL policy has mode (dry-run or enforce), deny and allow lists (policy.ts). Dry-run records refusals but lets work continue, so an operator can test rules on real traffic. Approvals sit on top. decideAction checks built-in safety requirements first; regexes for credentials, security settings and payments always hand the action to the person and cannot be configured (approvals/policy.ts). Then team and personal rules apply (allow, pre_approved, ask, hand_off), then an optional LLM auto-review (approvals/policy.ts).

Computers and the supervisor

With COMPUTER_SUPERVISOR_URL set, the server asks the supervisor for each Bot’s computer address. Without it, every Bot shares the single AGENT_COMPUTER_URL computer (supervisor.ts). The server never talks to Docker itself.

Memory and learning

PersonalMemoryMiddleware is an AG-UI middleware that loads a person’s memories and injects them before each run (memory/tools.ts). Memories live in Postgres. Thread history and learned skills live in CopilotKit Intelligence, not in OpenBot’s own database.

Models

Only three provider IDs exist: openai, anthropic and google (model-providers.ts). Defaults per provider and per Bot directory come from model-providers.json. BOT_PROVIDER and BOT_MODEL override them in runtimeModelForEnvironment (copilot.ts). A local model is possible only through an OpenAI-compatible OPENAI_BASE_URL.

Extending it

  • Bring an agent. Register any AG-UI endpoint as a coworker from /agents, or declare it in a tenant package’s agents.yaml as built-in (a system prompt) or remote-ag-ui (an endpoint). TENANT_PACKAGE_DIR points at the package; examples/fintech is the default.
  • Add tools. server/src/plugins/transport.ts routes each catalogue entry to MCP, Composio, Google Drive REST or built-in routines behind one listTools/callTool shape (transport.ts).
  • Write rules. Add CEL deny/allow expressions to the action policy, and team or personal approval rules. Granted plugin tools are checked by the same policy, with mcp.server, mcp.tool and mcp.effect bound, so one rule set covers browser and tool actions.
  • Add a provider. One row in model-providers.json plus one entry in PROVIDER_IDS in both the TypeScript and Python loaders.

Running it

Requirements: Docker, Bun 1.3+, a CopilotKit Intelligence project, and a model key. Copy .env.example to .env, run bun install, then bun scripts/setup-learning.ts to provision the Intelligence key and Learning container. bash scripts/start.sh starts Postgres (pgvector image), migrations, the computer containers, the API on 3001, the routine worker and the app on 3010. For production, one image (ghcr.io/copilotkit/openbot) carries the app, API and browser, with optional embedded Postgres. A Helm chart lives in charts/openbot. OPENBOT_SINGLE_USER=true in the example env skips OAuth sign-in setup.

Strengths and caveats

  • Strength: governance is structural. No code path acts on a computer without a policy decision and an audit row written first. Refs resolve on the server, and rule errors fail closed.
  • Strength: framework-neutral. Any AG-UI server can be a Bot, and the repo ships example adapters for many frameworks.
  • Strength: dry-run policy mode lets operators tune rules against real traffic before enforcing them.
  • Caveat: hard dependency on CopilotKit Intelligence. Config refuses to boot without it. Threads, learning and history are not in your Postgres alone.
  • Caveat: three providers only. No first-class OpenRouter or local-model support beyond base-URL overrides.
  • Caveat: heavy stack. Postgres, per-Bot Chromium containers, a supervisor and optional SPIRE make it more than a single-process assistant.
  • Caveat: large, dense modules. server/src/plugins/store.ts alone is about 7,000 lines and server/src/index.ts over 3,000. The comments are thorough, but the surface to learn is big.

Sources: code at f4bc60b, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (61 pages), verified Q&A.

How it answers the Open-source personal assistants questions

Each answer was drafted by a code-reading agent at commit f4bc60b. Its citations were checked mechanically. Compare with the other open-source personal assistants →

How is the assistant architected?

answered

Agent loop and runtime. The runtime is built on @copilotkit/runtime/v2 in Intelligence mode (server/src/copilot.ts:52-62). The server mounts a CopilotRuntime via Hono that drives agents over the AG-UI protocol. Built-in agents (the 13 shipped coworkers) run as BuiltInAgent instances; external agents are HttpAgent instances — anything that speaks AG-UI (LangGraph, Mastra, CrewAI, Pydantic AI, Google ADK, or hand-written servers) works without adapters. The runtime is always in Intelligence mode because CopilotKit Intelligence is required for durable threads, memory, and learning.

Frontend/backend split. The browser app (app/) is a React application built with TanStack Router (app/src/router.tsx:1-8). It communicates with the server via TanStack Query (React Query). The server (server/) is a Hono app running on Bun (server/src/app.ts). Both are colocated in the monorepo and share the shared/ package for types. The frontend renders the AG-UI stream and manages sign-in; the server handles all agent orchestration, policy, memory, and data.

Main packages. The monorepo at /work/project/package.json declares workspaces: app, server, worker. Thirteen agent-* directories each implement one agent framework adapter (agent-mastra, agent-langgraph, agent-crewai, agent-pydantic-ai, agent-adk, agent-ag2, agent-agno, agent-langroid, agent-llamaindex, agent-strands, agent-claude-sdk, agent-bot, agent-computer). The desktop/ and mobile/ directories hold Tauri desktop and mobile builds. supervisor/ manages per-Bot Docker containers. shared/ contains model provider specs, bot prompts, and utilities.

Request flow. A user types in the browser → the AG-UI stream arrives at the server → the runtime resolves the target agent via ActorAgentResolver (server/src/agents/agent-resolver.ts) → the agent's prompt is assembled with standing role, memory and learning injected as system messages → the agent runs against its provider (OpenAI/Anthropic/Google) → every tool call the agent makes passes through the gateway (server/src/computer/gateway.ts:1-15), which: (1) resolves snapshot refs on the server, never from what the model claimed; (2) evaluates the action policy (server/src/computer/policy.ts); (3) passes through the approvals gate for human-in-the-loop; (4) records every action in the audit trail — then and only then acts.

Editor's note. Correction: the server runs Hono on Bun, not Elysia (see server/package.json, server/src/app.ts). Desktop is Tauri, but mobile is an Expo app, not Tauri. There are 15 agent-* packages, including agent-microsoft and agent-langgraph-agui.

How are integrations (email, calendar, chat, docs) implemented?

answered

Supported services. Integrations are accessed through three mechanisms: (1) Composio (server/src/plugins/composio.ts) — a broker for 200+ SaaS apps (Gmail, Slack, GitHub, Notion, etc.). Each action is exposed as a tool the agent can call; the deployment holds one Composio API key and connections are per-person. (2) MCP (server/src/plugins/mcp.ts:1-7) — the Model Context Protocol for connecting arbitrary external servers. Each MCP server is a catalogue entry with a URL, optional token, and tool list. (3) Built-in routines (server/src/plugins/builtin-routines.ts) — server-side scheduled tasks like daily summaries. A dedicated Google Drive REST adapter (server/src/plugins/google-drive-rest.ts) exists because Google's Drive MCP server is gated behind a developer preview — it implements the same listTools/callTool interface as MCP, making transports swappable per catalogue entry.

OAuth flow. The OAuth module (server/src/plugins/oauth.ts) implements the PKCE authorization code flow. A sealed, encrypted state token carries the connect origin ("settings" or "admin"), PKCE verifier, and target user; the callback is the single path /api/plugins/oauth/callback. The state is encrypted (not just signed) so proxy logs holding the callback URL cannot redeem the code. Dynamic MCP client registration is handled; the state expires after 10 minutes.

Token storage. Credentials are encrypted at rest using AES-GCM (server/src/credentials.ts:7-13) with a keyEncryptionKey from deployment config. Stored in the credentials table with kind, provider, keyId, metadata, and encryptedValue. Keys are write-only — there is no way to read a stored credential back.

Sync vs on-demand. Tool listing is on-demand: the system calls the vendor to list available tools at grant time via refreshTools. Composio actions are paginated with a limit of 1000 per page and followed to the end via cursor. MCP servers get a 15-second listing timeout. When a Bot calls a tool, it goes through the approval gate, the transport (MCP/Composio/Drive/built-in), and the result (capped at 20,000 characters) is returned to the model.

How is memory and user context stored and retrieved?

answered

Storage — PostgreSQL. Personal memories live in the personalMemories table in PostgreSQL, accessed through Drizzle ORM (server/src/db/schema/memory). The memory store (server/src/memory/store.ts:16-80) provides CRUD operations: list (up to 500, newest-first), create, formMemory (Bot-observed facts awaiting human review), update, remove. Deduplication uses a SHA-256 hash of lowercased content (deleted memories stay forgotten).

Memory types. Personal memories are free-text records with optional sourceApp and sourceLink. Bot-formed memories (formMemory) are observed facts the person can confirm or dismiss. A MemorySource records where imported data came from (a connected app). Memory import from connected apps is handled by normalizeConnectorRecords (server/src/memory/ingestion.ts:16-80) which parses JSON responses from external services (searching items, files, or results keys), normalizes records to text, and caps at 50 records or 256KB.

Injection into prompts. The PersonalMemoryMiddleware (server/src/memory/tools.ts:29-70) is an AG-UI Middleware that injects the person's memories as both a system message and a context entry on every agent run. The system message id is "openbot:personal-memory" — the middleware also deduplicates stale ones. Memory content is loaded asynchronously and injected before the run begins.

Learning (skills). ``CopilotKit Intelligence's Learning system (server/src/learning/runtime.ts:1-35) provides learned skills — reusable capabilities published from observed patterns. The runtime exposes copilotkit_load_skill and copilotkit_read_skill_file tools to agents, backed by CopilotKit Intelligence. A RemoteLearnedSkillsMiddleware injects learned skill context into agent runs. Learning targets are per-agent or default, configured through admin settings (server/src/learning/settings.ts).

How are actions on the user's behalf gated?

answered

The gateway — mandatory choke point. Every action a Bot takes on its computer goes through ComputerGateway (server/src/computer/gateway.ts:1-15). The gateway has three jobs: (1) resolve the element ref from the server-fetched snapshot (never from what the model claims); (2) evaluate the action policy; (3) write an audit row and only then act. "An action that was not recorded did not happen, because there is no path that acts without writing the row first" is a stated invariant.

Action policy — CEL expressions. The policy (server/src/computer/policy.ts) uses cel-js to evaluate allow/deny rules against a PolicyContext that includes tool name, bot ID, page URL/host, actor ID, element role/name/type, and keyboard key. Deny beats allow. The policy supports dry-run mode (records decisions without blocking) for safe rule authoring. The three policy layers are: (1) CEL expressions (allow/deny lists), (2) approve/ask/hand-off decisions, (3) learning-based classification of action effects.

Approval gate — human-in-the-loop. The approvals service (server/src/approvals/service.ts) evaluates every action against a four-tier model: built-in safety requirements → team rules → personal rules → auto-review → default behavior. Outcomes are allow, ask, hand_off, or deny (server/src/approvals/types.ts:31-52). Safety requirements (server/src/approvals/policy.ts:83-108) automatically flag password changes, security settings, and payments for hand-off using regex patterns — these are non-configurable.

Auto-review model. Risky actions that are not caught by safety rules go through an LLM-based auto-review (server/src/approvals/policy.ts:210-218). The reviewer model gets the action context and user's request, then outputs a structured verdict (proceed | needs_approval | hand_off) with safety classification. It fails closed: errors become needs_approval. Actions marked as pre_approved skip approval if the person explicitly requested them.

Audit trail. Every policy decision, approval request, and action is recorded in the auditEvents table (server/src/audit.ts). Audit event types include everything from configuration.changed and credential.created to detailed agent action records. Sensitive fields (tokens, passwords, credentials, prompts, tool arguments/results) are identified by key name and stripped from audit logs (server/src/audit.ts:17-46).

How are LLM providers selected and configured?

answered

Supported providers. Three providers are hardcoded in TypeScript and Python: openai, anthropic, google (shared/model-providers.ts:42). Each has a spec row in shared/model-providers.json (shared/model-providers.json:1-38) defining its API key environment variable, base URL override variable, and default model. The default models are gpt-5.5 (OpenAI), claude-sonnet-4-5 (Anthropic), and gemini-2.5-flash (Google). Each agent-framework adapter chooses which SDK it can drive — adding Google's spec doesn't make agent-mastra (which loads no Google module) able to call Gemini.

Configuration surface. Provider and model are selected through environment variables (server/src/copilot.ts:205-220): BOT_PROVIDER and BOT_MODEL override the defaults from the spec file. An empty provider string defaults to OpenAI (matching what the desktop writes when switching back). Base URLs can be overridden per provider (OPENAI_BASE_URL, ANTHROPIC_BASE_URL, GOOGLE_GENERATIVE_AI_BASE_URL) — the Anthropic variant automatically versions to /v1 if missing. Per-Bot defaults are defined in the bots section of model-providers.json, keyed by agent directory name (e.g., agent-bot defaults to OpenAI/gpt-5.5, agent-langgraph to OpenAI/gpt-5.5).

Tool-calling and structured output. The built-in agents use the AG-UI client's AbstractAgent interface for tool calling, with middleware for memory injection and learning. The approval auto-review uses structured JSON schema output (server/src/approvals/policy.ts:188-194) via Zod: { verdict, preApproved, reason, safety }. The Composio integration passes action schemas (as published by the vendor) directly to the model as function definitions. MCP tools are listed with their JSON Schema parameters and returned as model-callable tools.

Local model support. There is no built-in local model support beyond the base URL override mechanism. Any provider that speaks an OpenAI-compatible API can be pointed at by setting OPENAI_BASE_URL to a local endpoint (e.g., Ollama, vLLM). The Google provider similarly supports GOOGLE_GENERATIVE_AI_BASE_URL. No quantized, local-first, or on-device model is shipped.

Editor's note. Addition: gpt-5.5 is the default in the JSON spec, but when the provider is switched through environment variables, runtimeModelForEnvironment falls back to gpt-5.6-terra.

How is it deployed and self-hosted?

answered

Runtime dependencies. The required infrastructure is: PostgreSQL (specifically pgvector/pgvector:pg17 from docker-compose), CopilotKit Intelligence (managed or self-hosted), and a model provider API key. Docker Compose (docker-compose.yml) brings up postgres, migrate (Drizzle migrations), agent-computer (a Playwright/Chromium container per Bot), and optionally SPIRE for workload identity. A single all-in-one Dockerfile (Dockerfile:1-60) packages the full stack — app, API, and Chromium browser — for serverless deployment. The supervisor (server/src/computer/supervisor.ts:1-76) manages per-Bot computer containers via Docker on a laptop; absent a supervisor, all Bots share one computer (AGENT_COMPUTER_URL).

Quick start path. The .env.example sets OPENBOT_SINGLE_USER=true so a fresh clone works without any OAuth client registration. The startup script (scripts/start.sh) handles .env generation, docker compose up for postgres/computer, bun run dev for the API and frontend. A CopilotKit Intelligence project is required — the setup-learning.ts script provisions it. The example package ships 13 coworkers under examples/ (General Assistant, Knowledge, Risk Analyst, and 10 fintech-specific agents).

One-click Docker. The Dockerfile at root builds a single container (openbot-server) with Bun, Playwright/Chromium, and all workspace code. It runs the API server + static file serving + a shared browser. The SERVER_IMAGE and COMPUTER_IMAGE env vars in docker-compose let you publish named release images via IMAGE_PULL_POLICY=missing. docker-compose.yml is the primary deployment vehicle — it separates postgres onto an isolated data network so Bot containers cannot reach the database directly.

Required external accounts. A model provider key (OpenAI, Anthropic, or Google), a CopilotKit Intelligence project (free tier available, can be self-hosted via INTELLIGENCE_API_URL), and optionally: Composio API key for 200+ SaaS integrations, Google/Microsoft/Okta OAuth credentials for sign-in, and Google Drive API credentials for the Drive connector.