# LLMs Technical Reviews — full text > Source-level reviews of open-source LLM tooling: personal assistants, browser control, API connectors, AI scraping and open-source DeepWiki generators. API: https://llms-technical-reviews.com/api/v1/ (OpenAPI https://llms-technical-reviews.com/openapi.json; public read-only, no auth / no API key / no OAuth). MCP: https://llms-technical-reviews.com/mcp. # Category: Open-source personal assistants > Self-hostable AI assistants that act on your mail, calendar, docs and chat on your behalf. Personal assistants here are self-hosted agent platforms that act for you: they browse, read and send mail, edit documents and call third-party APIs, usually from a computer or sandbox of their own. Because they act on real accounts, the code that matters most is not the chat UI but what sits between the model and the side effect. When choosing, check how actions are gated (per-action policy, approval rules, audit trail), where memory and thread history are stored and whether that store is yours, how integrations are reached (MCP, Composio-style catalogues, or custom clients) and where tokens are kept, which model providers and local models are supported, and which external services a deployment cannot run without. ## Projects - [CopilotKit/OpenBot](https://llms-technical-reviews.com/p/openbot/) — Self-hosted platform where AG-UI agents get their own browser computer, with every action policy-checked and audited first. (★6111, TypeScript) - [elie222/rakazo](https://llms-technical-reviews.com/p/rakazo/) — Self-hosted persistent AI bots on a Pi agent runtime, with Postgres job queue, pluggable sandboxes and rule-based tool approval. (★3365, TypeScript) In the research queue (not yet published): yc-software/qm. ## Comparison questions - [How is the assistant architected?](https://llms-technical-reviews.com/personal-assistants/q/architecture/) — Agent loop and runtime; frontend/backend split; main packages; how a user request flows to an action. - [How are integrations (email, calendar, chat, docs) implemented?](https://llms-technical-reviews.com/personal-assistants/q/integrations/) — Which services; API clients vs MCP; OAuth flow; token storage; sync vs on-demand fetch. - [How is memory and user context stored and retrieved?](https://llms-technical-reviews.com/personal-assistants/q/memory/) — Storage (DB, vector store, files); what is remembered; how it is injected into prompts; summarization. - [How are actions on the user's behalf gated?](https://llms-technical-reviews.com/personal-assistants/q/action-safety/) — Approval / human-in-the-loop flows; permission scopes; dry-run or draft modes; audit trail. - [How are LLM providers selected and configured?](https://llms-technical-reviews.com/personal-assistants/q/models/) — Supported providers; config surface; tool-calling / structured-output usage; local-model support. - [How is it deployed and self-hosted?](https://llms-technical-reviews.com/personal-assistants/q/deploy/) — Runtime dependencies (DB, queues, browsers); Docker/one-click paths; required external accounts. Full comparison: https://llms-technical-reviews.com/compare/personal-assistants/ ## Verdict: How is the assistant architected? Both published assistants split a web API from the agent loop, but they put the loop in different places. [OpenBot](/p/openbot/) runs the loop inside its Hono server through CopilotKit's `CopilotRuntime` in Intelligence mode. Built-in Bots run in-process as `BuiltInAgent`; any other AG-UI server (LangGraph, Mastra, Pydantic AI and more) plugs in as an `HttpAgent`. Every browser and file action then passes through one `ComputerGateway`, which resolves the target, checks policy and writes an audit row before acting. Threads and learning live in CopilotKit Intelligence, an external service. [Rakazo](/p/rakazo/) keeps the API thin. `threads.send` writes a queued run and enqueues a Graphile job; a separate worker leases the run and drives the Pi agent runtime from a large `executor.ts`. All state, including the job queue and realtime fanout, sits in Postgres, and every backend (runtime, sandbox, connectors, memory) is an `adapter-kit` interface. Choose OpenBot if you already build agents in several frameworks and want one governed front door for them. Choose Rakazo if you want one self-contained stack where runs are durable background jobs that survive restarts and can wait for approval or a takeover. More projects in this category are being researched. ## Verdict: How are integrations (email, calendar, chat, docs) implemented? Both projects lean on managed catalogues plus MCP instead of hand-written Gmail or Calendar clients. [OpenBot](/p/openbot/) routes every catalogue entry through one transport seam with MCP's own `listTools`/`callTool` shape. Behind it sit MCP servers, Composio (one deployment key, per-person connections), a Google Drive REST adapter (because Drive's MCP server is gated), and built-in routines. OAuth uses PKCE with an encrypted state token, and credentials are AES-GCM encrypted at rest and write-only. Tools are listed at grant time; calls go through the same policy and approvals path as computer actions. [Rakazo](/p/rakazo/) builds a connector stack in the worker: user-installed connectors (remote MCP, Treg, OpenAPI documents), Composio, Pipedream Connect, and MCP (stdio only with `MCP_STDIO_ENABLED` and an allowlist). Connector secrets go into an encrypted store. Tools are discovered when a run starts, with no background sync. Rakazo also connects chat platforms (Slack, Telegram, WhatsApp, Sendblue) as places to talk to bots. OpenBot suits an admin-curated catalogue with uniform governance. Rakazo suits users who want to bring their own APIs, including any OpenAPI spec, without operator work. More projects in this category are being researched. ## Verdict: How is memory and user context stored and retrieved? The two projects store memory very differently. [OpenBot](/p/openbot/) keeps two kinds. Personal memories are free-text rows in Postgres, deduplicated by a content hash. Bots can propose observed facts that the person confirms or dismisses, and connected apps can be imported. `PersonalMemoryMiddleware`, an AG-UI middleware, injects them as a system message on every run. Conversation threads and learned skills are held by CopilotKit Intelligence, so a large part of the context is outside your database. [Rakazo](/p/rakazo/) keeps everything in its own Postgres. Messages are stored per thread and old history is compacted into an LLM summary that is injected into later runs. A `MarkdownMemoryStore` holds versioned Markdown documents per user, bot and scope, written and read through `remember`/`recall_memory` tools. Optional semantic providers (Serenity, Supermemory) add vector recall of up to five items, but only once a thread has been compacted. Memory text is passed through secret redaction before it reaches the prompt. Pick OpenBot if a human-reviewed list of facts is what you want and you accept a hosted thread store. Pick Rakazo for self-contained, portable memory and long threads that need summarisation. More projects in this category are being researched. ## Verdict: How are actions on the user's behalf gated? This is where the two projects differ most. [OpenBot](/p/openbot/) makes governance structural. A `ComputerGateway` is the only path to a Bot's computer: it resolves the clicked element from the server's own snapshot, evaluates CEL `deny`/`allow` rules (deny wins, errors deny, `dry-run` mode available), writes an audit row, and only then acts. An approvals service adds non-configurable safety hand-offs for credentials, security settings and payments, then team and personal rules (`allow`, `pre_approved`, `ask`, `hand_off`), then an optional LLM auto-review that fails closed. [Rakazo](/p/rakazo/) gates at the tool level. `toolRequiresApproval` exempts computer, browser, shell and file tools, always asks for destructive builtins, and judges connector tools by name prefix (unknown names ask). User rules by tool, connector or category override the default; Auto Review can only escalate to `ask`. Webhook-triggered runs are stricter. Each effect is recorded with an idempotency key. OpenBot fits teams that need an audit trail and fine-grained control over what a browser agent clicks. Rakazo fits personal use where the sandbox is the boundary and approval is reserved for outward-facing tool calls. More projects in this category are being researched. ## Verdict: How are LLM providers selected and configured? [OpenBot](/p/openbot/) supports exactly three provider IDs: `openai`, `anthropic` and `google`, defined in `shared/model-providers.json` and checked by both TypeScript and Python loaders. `BOT_PROVIDER` and `BOT_MODEL` override the defaults, and per-provider base URL variables allow an OpenAI-compatible endpoint, which is the only route to a local model. Remote AG-UI Bots choose their own models, so the platform-level setting mostly affects built-in Bots. The approval auto-review uses structured JSON output. [Rakazo](/p/rakazo/) delegates to the Pi framework, which brings a large built-in catalogue. The deployment default is OpenRouter (`PI_DEFAULT_PROVIDER`, `PI_DEFAULT_MODEL`), with Anthropic as an alternative deployment key. Users connect their own credentials in the UI, by API key or by OAuth for ChatGPT, Claude, GitHub Copilot and xAI subscriptions. Bots can pin a model and thinking level, and local Ollama-style or OpenAI-compatible endpoints are first-class. Choose Rakazo if users should bring their own models or subscriptions, or you want local models. OpenBot's narrow list is enough when an admin sets one provider for the deployment and agent frameworks handle the rest. More projects in this category are being researched. ## Verdict: How is it deployed and self-hosted? Both ship Docker paths, but their required services differ. [OpenBot](/p/openbot/) needs Postgres (the compose file uses the pgvector image), a model key and a CopilotKit Intelligence project, managed or self-hosted; the server refuses to boot without Intelligence. `scripts/start.sh` brings up Postgres, migrations, Bot computers, the API, the routine worker and the app. Production can use a single published image carrying app, API and Chromium, with optional embedded Postgres, or the Helm chart in `charts/openbot`. A supervisor gives each Bot its own computer container; without it, all Bots share one. [Rakazo](/p/rakazo/) needs only Postgres 16, Docker and a model key. `install-images.sh` writes `.env` with random secrets and starts `postgres`, `supervisor`, `api`, `worker` and `web` from published `edge` images. There is no Redis: jobs and realtime run on Postgres. Computers can be local Docker or hosted E2B, Daytona, CreateOS or Box. `infra/compose` adds a Caddyfile, host hardening and egress restriction scripts. Rakazo is the simpler fully self-hosted install with no required third-party service. OpenBot fits organisations that already use CopilotKit or want Kubernetes and SSO-oriented deployment. More projects in this category are being researched. # Project: CopilotKit/OpenBot > Self-hosted platform where AG-UI agents get their own browser computer, with every action policy-checked and audited first. — https://llms-technical-reviews.com/p/openbot/ (commit f4bc60b) ### Overview OpenBot is a self-hosted platform for running AI "coworkers" (Bots) that each get their own computer: a Chromium browser, a `/workspace` file area and a shell, all in a container. A person talks to a Bot in a web, desktop (Tauri) or mobile (Expo) client. The Bot can browse, read and write files, call third-party tools through MCP or Composio, and hand work to another Bot. The core idea is that the server owns every side effect. Bots never reach their computer directly. Each browser or file action goes through one gateway, which resolves the element on the server, checks a CEL policy, writes an audit row and only then acts. A second layer, the approvals service, decides whether to act, ask the person, hand the action off to them, or refuse. OpenBot is built on CopilotKit. The server runs `@copilotkit/runtime/v2` in Intelligence mode, so threads, memory and "learning" depend on a CopilotKit Intelligence project (managed, or self-hosted). Bots are AG-UI agents. Built-in ones run in-process as `BuiltInAgent`. Remote ones are any AG-UI endpoint (LangGraph, Mastra, Pydantic AI, CrewAI, ADK and others), reached as `HttpAgent`. ### Architecture ```mermaid flowchart LR UI["app (React + Vite)"] --> API["server (Hono on Bun)"] Desk["desktop / mobile"] --> API API --> RT["CopilotRuntime"] RT --> INT["CopilotKit Intelligence"] RT --> BI["Built-in Bots"] RT --> RB["Remote AG-UI Bots"] BI --> GW["Computer gateway"] RB --> GW GW --> POL["CEL policy"] GW --> APR["Approvals service"] GW --> AUD["Audit (Postgres)"] GW --> SUP["supervisor"] SUP --> COMP["agent-computer (Chromium)"] BI --> PLG["Plugins: MCP, Composio"] WRK["worker (routines)"] --> API ``` | Component | Path | Role | | --- | --- | --- | | Web app | `app/` | React UI with TanStack Router, built with Vite | | API server | `server/src/app.ts` | Hono app: auth, admin, policy, audit, plugins, channels; mounts the CopilotKit handler | | Runtime wiring | `server/src/copilot.ts` | Builds Bots, attaches middleware, mounts `CopilotRuntime` | | Computer gateway | `server/src/computer/gateway.ts` | The only path from a Bot to its computer | | Action policy | `server/src/computer/policy.ts` | CEL allow/deny rules, `dry-run` or `enforce` | | Approvals | `server/src/approvals/` | Allow / ask / hand off / deny per action, with optional LLM auto-review | | Supervisor | `supervisor/`, `server/src/computer/supervisor.ts` | One computer container per Bot | | Bot computer | `agent-computer/` | Chromium, workspace files and shell behind a token | | Plugins | `server/src/plugins/` | MCP, Composio, Google Drive REST, OAuth, tool selection | | Memory | `server/src/memory/` | Personal memories in Postgres, injected by middleware | | Worker | `worker/` | Fires due routines by handing runs to the server | | Bot adapters | `agent-*/` | 14 example AG-UI Bots (frameworks plus a proof-of-concept `agent-bot`), and `agent-computer` | ### How a request flows 1. The browser sends an AG-UI run to `/api/copilotkit`. The handler is mounted at the root with its own base path ([app.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/app.ts#L1235-L1243)). 2. `mountCopilotRuntime` creates the `CopilotRuntime` with an Intelligence client and an `identifyUser` callback that scopes threads and memory to the signed-in person ([copilot.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/copilot.ts#L2038-L2165)). 3. The runtime asks `createRequestAgents` for the Bots this person may see. The resolver builds them with `buildAgents`, which loads the granted tools, standing instructions and learned skills, then attaches `PersonalMemoryMiddleware` to each Bot ([copilot.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/copilot.ts#L479-L590)). 4. The model picks a computer tool (for example `computer_click`). The call reaches `ComputerGateway`. Its private `govern` function loads the stored snapshot, resolves the ref to a real element, and builds a `PolicyContext` with the tool, page host, element, file, command and initiator ([gateway.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/gateway.ts#L505-L602)). 5. `evaluateActionPolicy` runs deny rules first, then allow rules. No match means refused ([policy.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/policy.ts#L337-L380)). The gateway writes the audit row whatever the outcome, and throws `ActionRefusedError` if the decision does not forward. 6. If the policy allows, the gateway calls the `approvalGate` with an effect (`write` or `read`) and a scope. The approvals service may replay an earlier answer, pause for the person, or let it run ([gateway.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/gateway.ts#L657-L683)). 7. The action is sent to the Bot's computer, found through the supervisor. A failure after an allowed decision gets its own audit row. ### Key components ### Computer gateway The header comment states the contract: resolve the ref from the server's own snapshot, ask the policy, write the row, then act ([gateway.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/gateway.ts#L1-L19)). Resolving on the server matters: a model cannot evade "never click Submit" by labelling the click "Continue". Every policy field is bound even when empty, because cel-js throws on unknown identifiers and a throwing deny rule counts as a match. ### Policy and approvals The CEL policy has `mode` (`dry-run` or `enforce`), `deny` and `allow` lists ([policy.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/policy.ts#L1-L35)). Dry-run records refusals but lets work continue, so an operator can test rules on real traffic. Approvals sit on top. `decideAction` checks built-in safety requirements first; regexes for credentials, security settings and payments always hand the action to the person and cannot be configured ([approvals/policy.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/approvals/policy.ts#L75-L108)). Then team and personal rules apply (`allow`, `pre_approved`, `ask`, `hand_off`), then an optional LLM auto-review ([approvals/policy.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/approvals/policy.ts#L279-L330)). ### Computers and the supervisor With `COMPUTER_SUPERVISOR_URL` set, the server asks the supervisor for each Bot's computer address. Without it, every Bot shares the single `AGENT_COMPUTER_URL` computer ([supervisor.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/computer/supervisor.ts#L1-L14)). The server never talks to Docker itself. ### Memory and learning `PersonalMemoryMiddleware` is an AG-UI middleware that loads a person's memories and injects them before each run ([memory/tools.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/memory/tools.ts#L29-L40)). Memories live in Postgres. Thread history and learned skills live in CopilotKit Intelligence, not in OpenBot's own database. ### Models Only three provider IDs exist: `openai`, `anthropic` and `google` ([model-providers.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/shared/model-providers.ts#L33-L42)). Defaults per provider and per Bot directory come from [model-providers.json](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/shared/model-providers.json#L1-L38). `BOT_PROVIDER` and `BOT_MODEL` override them in `runtimeModelForEnvironment` ([copilot.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/copilot.ts#L206-L238)). A local model is possible only through an OpenAI-compatible `OPENAI_BASE_URL`. ### Extending it - **Bring an agent.** Register any AG-UI endpoint as a coworker from `/agents`, or declare it in a tenant package's `agents.yaml` as `built-in` (a system prompt) or `remote-ag-ui` (an endpoint). `TENANT_PACKAGE_DIR` points at the package; `examples/fintech` is the default. - **Add tools.** `server/src/plugins/transport.ts` routes each catalogue entry to MCP, Composio, Google Drive REST or built-in routines behind one `listTools`/`callTool` shape ([transport.ts](https://github.com/CopilotKit/OpenBot/blob/f4bc60bf9b12c65eb3d7640173f1b3432f6b2a60/server/src/plugins/transport.ts#L1-L20)). - **Write rules.** Add CEL `deny`/`allow` expressions to the action policy, and team or personal approval rules. Granted plugin tools are checked by the same policy, with `mcp.server`, `mcp.tool` and `mcp.effect` bound, so one rule set covers browser and tool actions. - **Add a provider.** One row in `model-providers.json` plus one entry in `PROVIDER_IDS` in both the TypeScript and Python loaders. ### Running it Requirements: Docker, Bun 1.3+, a CopilotKit Intelligence project, and a model key. Copy `.env.example` to `.env`, run `bun install`, then `bun scripts/setup-learning.ts` to provision the Intelligence key and Learning container. `bash scripts/start.sh` starts Postgres (pgvector image), migrations, the computer containers, the API on 3001, the routine worker and the app on 3010. For production, one image (`ghcr.io/copilotkit/openbot`) carries the app, API and browser, with optional embedded Postgres. A Helm chart lives in `charts/openbot`. `OPENBOT_SINGLE_USER=true` in the example env skips OAuth sign-in setup. ### Strengths and caveats - **Strength: governance is structural.** No code path acts on a computer without a policy decision and an audit row written first. Refs resolve on the server, and rule errors fail closed. - **Strength: framework-neutral.** Any AG-UI server can be a Bot, and the repo ships example adapters for many frameworks. - **Strength: dry-run policy mode** lets operators tune rules against real traffic before enforcing them. - **Caveat: hard dependency on CopilotKit Intelligence.** Config refuses to boot without it. Threads, learning and history are not in your Postgres alone. - **Caveat: three providers only.** No first-class OpenRouter or local-model support beyond base-URL overrides. - **Caveat: heavy stack.** Postgres, per-Bot Chromium containers, a supervisor and optional SPIRE make it more than a single-process assistant. - **Caveat: large, dense modules.** `server/src/plugins/store.ts` alone is about 7,000 lines and `server/src/index.ts` over 3,000. The comments are thorough, but the surface to learn is big. *Sources: code at f4bc60b, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (61 pages), verified Q&A.* # Project: elie222/rakazo > Self-hosted persistent AI bots on a Pi agent runtime, with Postgres job queue, pluggable sandboxes and rule-based tool approval. — https://llms-technical-reviews.com/p/rakazo/ (commit 4cdf3e2) ### Overview Rakazo is a self-hosted platform for persistent AI "bots": each bot has its own threads, memory, routines and a computer (browser, terminal, files, optional graphical desktop). People reach bots from a React web app, an Electron desktop app, an Expo mobile app, or chat platforms (Slack, Telegram, WhatsApp, Sendblue). It is a TypeScript monorepo: `apps/` holds deployable surfaces and `packages/` holds domain logic and adapters. The design is adapter-first. `packages/adapter-kit` defines interfaces for the agent runtime, sandbox, connectors, memory, jobs, realtime and more, and `packages/adapters` implements them. The default agent runtime is Pi (`@earendil-works/pi-agent-core` and `pi-ai`), so model choice is broad: OpenRouter by default, plus Anthropic, OpenAI, Google, ChatGPT/Claude subscription OAuth, OpenAI-compatible and local endpoints. Work is asynchronous. The API only writes a queued run and enqueues a job. A separate Graphile Worker process leases the run in Postgres and drives the model loop, so runs survive restarts and can pause for approval or a human takeover. ### Architecture ```mermaid flowchart LR WEB["web / desktop / mobile"] --> API["api (Hono + oRPC)"] CHAT["Slack / Telegram / WhatsApp"] --> API API --> PG["Postgres (Prisma)"] API --> Q["Graphile jobs"] Q --> WK["worker"] WK --> EX["Run executor"] EX --> PI["PiAgentRuntime"] PI --> LLM["Model providers"] EX --> GATE["Approval gate"] EX --> CON["Connectors: Composio, Pipedream, MCP, OpenAPI"] EX --> MEM["Memory: Markdown + semantic"] EX --> SBX["Sandbox provider"] SBX --> SUP["supervisor"] SUP --> COMP["computer container"] EX --> PG PG --> RTF["Realtime fanout"] RTF --> WEB ``` | Component | Path | Role | | --- | --- | --- | | API server | `apps/api/src/app.ts`, `apps/api/src/router.ts` | Better Auth sessions, oRPC procedures, webhooks, voice routes | | Worker | `apps/worker/src/index.ts` | Wires adapters and runs Graphile Worker job handlers | | Run executor | `packages/adapters/src/executor.ts` | Leases a run, builds context and tools, gates and records each tool call | | Pi runtime | `packages/adapters/src/pi-runtime.ts` | Streams the model and executes tool calls | | Approval rules | `packages/core/src/action-approval.ts` | Which tools need approval; rule resolution; auto-review plan | | Adapter interfaces | `packages/adapter-kit/src/interfaces.ts` | Contracts for runtime, sandbox, connectors, memory, jobs | | Memory | `packages/memory/src/index.ts` | Markdown documents with revisions in Postgres | | Sandboxes | `packages/adapters/src/*-sandbox.ts`, `infra/sandboxes/` | Docker (via supervisor), E2B, Daytona, CreateOS, Box, desktop | | Clients | `apps/web`, `apps/desktop`, `apps/mobile` | React + Vite, Electron, Expo | | Deploy | `infra/compose/` | Compose files, installer, Caddy and hardening scripts | ### How a request flows 1. The client calls the `threads.send` oRPC procedure. The `/rpc/*` middleware resolves the Better Auth session and the space membership first ([app.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/apps/api/src/app.ts#L553-L567)). 2. `threads.send` refuses if no model is connected, then calls `sendThreadMessage` ([router.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/apps/api/src/router.ts#L1924-L1933)). A `run` row is created with status `queued` and `runContinueJob(run.id)` is enqueued ([router.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/apps/api/src/router.ts#L680-L693)). 3. The worker's `run.continue` handler calls `executor.continueRun` ([background-job-handlers.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/background-job-handlers.ts#L55-L69)). 4. `continueRun` takes a fenced lease on the run with a conditional `updateMany`, so only one worker owns it ([executor.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/executor.ts#L3166-L3196)). It then resolves the model, loads memory and scratchpad context, recalls semantic memories and discovers connector tools. 5. It calls `deps.runtime.run(...)` with the prompt and instructions; memory and scratchpad context pass through `redactSecrets` first ([executor.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/executor.ts#L6134-L6160)). 6. For each tool call, the executor checks `toolRequiresApproval`, user rules and webhook-trigger rules, then `planActionGate` returns `ask`, `allow` or `judge` ([executor.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/executor.ts#L4088-L4145)). An `ask` writes an approval block and the run waits for the person. 7. Effects are recorded with idempotency keys (`recordEffect`), message blocks are persisted, and Postgres realtime fanout pushes updates to clients. If chat platforms are configured, bot replies are mirrored to them. ### Key components ### Worker wiring `main()` in the worker is the clearest map of the system. It picks `PiAgentRuntime` or `ScriptedAgentRuntime` from `AGENT_RUNTIME`, builds the sandbox from `SANDBOX_PROVIDER`, the MCP connector with stdio disabled unless `MCP_STDIO_ENABLED=true`, Composio and Pipedream when keys exist, and a connector stack of installed connectors, managed catalogues and MCP ([worker/index.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/apps/worker/src/index.ts#L62-L150)). ### Approval gate Tools fall into sets: approval-exempt (computer, browser, file, shell, `remember`, scheduling), always-approve builtins (`delete_bot`, `forget_memory`, `cloud_agent_launch` and others), and `create_space`, which always asks ([action-approval.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/core/src/action-approval.ts#L1-L37)). Connector tools are judged by name: a declared write or a mutating or compound name needs approval, and anything not clearly read-only does too ([action-approval.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/core/src/action-approval.ts#L80-L125)). Auto Review can only escalate to `ask`, and judge errors fail closed for consequential tools ([action-approval.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/core/src/action-approval.ts#L215-L245)). ### Memory `MarkdownMemoryStore` stores documents and revisions per space, user, scope and bot, with plain search ([memory/index.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/memory/src/index.ts#L16-L40)). Optional semantic providers (Serenity, Supermemory) implement `SemanticMemoryProvider` with `recall`, `save` and `purgeHistory` ([interfaces.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapter-kit/src/interfaces.ts#L202-L232)). Old thread history is compacted into an LLM summary. Semantic recall runs only after a thread has been compacted, and returns at most five memories. ### Models `resolveDeploymentModel` sets the deployment default: `PI_DEFAULT_PROVIDER` (default `openrouter`) with `OPENROUTER_API_KEY` or `ANTHROPIC_API_KEY`, and `PI_DEFAULT_MODEL` ([deployment-model.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/deployment-model.ts#L1-L27)). Users add their own credentials (API key or OAuth) in the UI, and bots can pin a model. ### Sandboxes `resolveSandboxProvider` defaults to `docker` and falls back to `none` when a remote provider has no API key, or when production Docker has no supervisor token ([sandbox-provider-env.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapters/src/sandbox-provider-env.ts#L4-L22)). ### Extending it - **New adapter.** Implement an interface from `adapter-kit`, such as `AgentRuntime` ([interfaces.ts](https://github.com/elie222/rakazo/blob/4cdf3e23315c2633b6b3fde8977f197b986f06fa/packages/adapter-kit/src/interfaces.ts#L239-L246)) or `ConnectorProvider` (`discoverTools`), and wire it in the worker and API. - **Tools without code.** Users add a remote MCP server, a Treg endpoint or an OpenAPI JSON document from the Integrations screen; operators enable Composio or Pipedream with env keys. - **Skills.** Bots load skills through `skill-tools.ts`; `builtin-skills.ts` ships defaults. - **Approval rules.** Per-tool, per-connector or per-category (`email`, `purchase`) rules decide always-allow versus require-approval. ### Running it The fastest path is `infra/compose/install-images.sh`: it downloads the Compose files, writes `.env` with random secrets, pulls the `edge` images and starts `postgres` (16), `supervisor`, `api`, `worker` and `web` on port 5173. No Redis is used; jobs and realtime both run on Postgres. For development you need Node 22.22.2+/24/26+, pnpm 9 and Docker, then `pnpm install`, `pnpm db:generate`, `pnpm db:migrate`, `pnpm sandbox:build` and `pnpm dev`. Set `SANDBOX_PROVIDER` to `e2b`, `daytona`, `createos` or `box` to use a hosted computer. Production overlays, a Caddyfile, `harden-host.sh` and `restrict-computer-egress.sh` live in `infra/compose/`. ### Strengths and caveats - **Strength: durable runs.** Fenced leases, a job reconciler and idempotent effect records mean a crashed worker does not double-send. - **Strength: model and sandbox choice.** Pi gives a wide provider catalogue including local and subscription OAuth; six sandbox backends share one interface. - **Strength: few hard vendor dependencies.** Only Postgres and a model key are required. - **Caveat: computer actions are not gated in normal runs.** `shell`, `browser_act`, `computer_act` and `write_file` are approval-exempt (webhook-triggered runs are stricter). A shell guard blocks commands that would break a bot's desktop, but otherwise safety on the computer relies on the sandbox, not per-action approval. - **Caveat: name-based connector gating.** Read-only detection uses tool-name prefixes; an oddly named write tool is caught only because unknown names default to approval. - **Caveat: very large core file.** `executor.ts` is over 8,000 lines and `router.ts` over 6,000, which makes changes to the run loop costly to review. *Sources: code at 4cdf3e2, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (17 pages), verified Q&A.* # Category: Browser & computer control > Frameworks that let an LLM drive a browser (or a whole desktop/phone) to complete tasks. These projects let a language model operate a real browser, or in some cases a whole desktop or phone. The model gets a view of the screen, picks an action, and the framework carries it out with clicks, typing and navigation. Designs range from full autonomous agents with their own step loop, to SDKs that add a single AI-driven `act` or `extract` call to an ordinary script. When choosing, check five things. How is the page shown to the model: DOM text, accessibility tree, screenshots, or a mix? That decides token cost and which models work. Who owns the loop: the framework, or your code or agent harness? What happens when an action fails? Look for retries, self-healing and caching of known-good actions. Which model providers are supported? And are stealth, proxies and CAPTCHA handling in the open-source code, or only in a paid cloud browser? ## Projects - [browser-use/browser-use](https://llms-technical-reviews.com/p/browser-use/) — Async Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP. (★117260, Python) - [browserbase/stagehand](https://llms-technical-reviews.com/p/stagehand/) — Browser automation SDK (TS, Python, Go) whose act/observe/extract run inside a Chrome extension that drives pages over CDP. (★25547, TypeScript) In the research queue (not yet published): browser-use/browser-harness, browser-use/workflow-use, browser-use/jev-ultrafast, Skyvern-AI/skyvern, hyperbrowserai/HyperAgent, trycua/cua, lahfir/agent-desktop, awlevin/typesafe-computer-use, omdsh-dev/dsh-browser, jkudish/jev-browser, droidrun/mobile-jev, ekzhang/openjev-sglang, savka777/jev-use, razaanstha/ulka. ## Comparison questions - [How is the page represented to the model?](https://llms-technical-reviews.com/browser-control/q/page-perception/) — DOM serialization, accessibility tree, screenshots, set-of-marks/element indexes; size limits and pruning. - [How are actions executed and how are elements targeted?](https://llms-technical-reviews.com/browser-control/q/action-execution/) — CDP / Playwright / OS-level input; selectors vs indexes vs coordinates; typing, scrolling, file upload, tabs. - [How is the agent loop / planning implemented?](https://llms-technical-reviews.com/browser-control/q/agent-loop/) — Step loop; planner vs executor; tool-call schema; stop conditions; memory between steps. - [How are failures, retries and self-healing handled?](https://llms-technical-reviews.com/browser-control/q/reliability/) — Error classes caught; retries; replanning; caching of successful actions or workflows; timeouts. - [Which models are supported and how are they called?](https://llms-technical-reviews.com/browser-control/q/models/) — Providers; vision requirement; structured output / tool calling; small or specialised models. - [How are browser sessions, profiles, auth and anti-bot handled?](https://llms-technical-reviews.com/browser-control/q/sessions/) — Local vs remote/cloud browsers; persistent profiles and cookies; stealth; proxies; CAPTCHA handling. Full comparison: https://llms-technical-reviews.com/compare/browser-control/ ## Verdict: How is the page represented to the model? Both projects send the model a text rendering of the page built from CDP, not raw HTML. They differ in how much they add on top. [browser-use](/p/browser-use/) merges `DOMSnapshot.captureSnapshot`, the full DOM and the accessibility tree. Its `DOMTreeSerializer` drops elements hidden by paint order and redundant wrappers, then prints interactive nodes as `[index]`. The index is the CDP `backend_node_id` where possible, so it stays stable across steps. New elements are starred. A screenshot is taken every step and sent when `use_vision` is on. Serialized text is capped at 40,000 characters. [Stagehand](/p/stagehand/) builds a "hybrid snapshot": an accessibility outline merged with DOM data, with iframes stitched in and shadow DOM pierced. Each line carries a `frameOrdinal-backendNodeId` id that maps to an XPath. Callers can scope it to a locator or exclude parts. There are no screenshots in `act` or `observe`. `extract` can attach a plain viewport PNG on request, with no set-of-marks overlay. Pick browser-use when pages are visual or ambiguous and a vision model is worth the extra tokens on every step. Pick Stagehand for cheaper, text-only prompts scoped to part of a page, especially for extraction. More projects in this category are being researched. ## Verdict: How are actions executed and how are elements targeted? Both projects keep the model away from raw coordinates by default. The model names an element id, and the library performs the input with CDP. [browser-use](/p/browser-use/) exposes a fixed action catalogue: `click`, `input_text`, `scroll`, `send_keys`, `upload_file`, `switch_tab`, `navigate`, `extract` and more. Elements are targeted by the numeric index from its DOM dump. Clicks go through an event bus to a watchdog that checks occlusion and sends `Input.dispatchMouseEvent`, with a JS-click fallback. For some Claude, Gemini 3 Pro and Browser Use models, coordinate clicking is also enabled. One model reply may contain up to five actions. The queue stops when the URL or tab changes, or after an action marked `terminates_sequence`. [Stagehand](/p/stagehand/) asks for exactly one `{elementId, method, arguments}` per inference. The method comes from an enum (click, fill, type, press, scroll, selectOption, hover, drag and similar). The id becomes an XPath, and its CDP-based `Locator` resolves it, scrolls it into view, reads the box model and clicks at the centre. A `twoStep` flag covers custom dropdowns. File inputs are filled with in-page `File` objects. Use browser-use when the model should run whole tasks with many action types and tabs. Use Stagehand when you want individual, auditable actions inside your own script. More projects in this category are being researched. ## Verdict: How is the agent loop / planning implemented? This is where the two designs differ most. [browser-use](/p/browser-use/) is an agent. `Agent.run` loops through `step()`: capture state, call the LLM for an `AgentOutput` (thinking, evaluation of the last goal, memory, next goal, optional plan update, action list), execute, then post-process. Planning is optional. It keeps a plan list with statuses, and nudges the model to replan after repeated failures or to make a plan after five steps without one. Memory is the message history, compacted by an LLM summary on long runs, plus a free-text `memory` field. The run stops on `done`, at `max_steps` (where the schema is forced to `done`), after `max_failures` consecutive failures, or on a user stop. [Stagehand](/p/stagehand/) has no loop at this commit, and its SDK has no `agent()` method. `act`, `observe` and `extract` each make one or two structured LLM calls and keep no memory between calls. Multi-step behaviour is left to your own code, or to an agent harness that uses Stagehand's integration tools (`run`, `snapshot`, `screenshot`) over MCP or natively. Choose browser-use for "give it a goal and let it go" tasks. Choose Stagehand if you already have an agent framework or prefer scripted control flow with AI only at the fuzzy steps. More projects in this category are being researched. ## Verdict: How are failures, retries and self-healing handled? Both projects bound each operation with timeouts. They differ in where recovery happens. [browser-use](/p/browser-use/) recovers inside the loop. Steps time out after 180 s and actions after 180 s by default. LLM calls get a per-model timeout. An empty action list gets one clarification retry. Rate-limit, auth, 5xx and truncation errors can switch the run once to a `fallback_llm`. Dropped CDP WebSockets get three reconnect attempts. An `ActionLoopDetector` watches repeated actions and unchanged pages and adds escalating hints to the prompt. After `max_failures` consecutive failed steps, one final recovery call is made. Nothing is cached between runs. `rerun_history` can replay a saved run and re-match elements. [Stagehand](/p/stagehand/) recovers per call. With `selfHeal`, a failed action triggers a fresh snapshot and a new inference. `TimeoutError` is always passed to the caller. Before each snapshot it waits for DOM and network quiet. Its main reliability tool is a server-side action cache keyed on the accessibility tree. A hit replays the stored actions without the LLM and falls back to inference if replay fails. The cache needs a Browserbase API key and session. Use browser-use for long unattended runs that must keep going. Use Stagehand's cache when the same flows repeat often on Browserbase. More projects in this category are being researched. ## Verdict: Which models are supported and how are they called? Both projects require structured output. They differ in how many providers they reach and how. [browser-use](/p/browser-use/) defines a small `BaseChatModel` protocol (`ainvoke(messages, output_format)`) with native adapters for about fifteen backends. These include OpenAI, Anthropic, Google, Azure, Bedrock, DeepSeek, Groq, Mistral, Ollama, OpenRouter, Cerebras, LiteLLM and its own hosted `ChatBrowserUse`, which is the default when no model is given. The agent always asks for a Pydantic `AgentOutput` whose action field is a discriminated union. Vision is on by default and switched off for DeepSeek and some Grok models. Timeouts, screenshot size and coordinate clicking are tuned by model name. [Stagehand](/p/stagehand/) calls models from inside its browser extension through the Vercel AI SDK. Five providers are supported (OpenAI, Anthropic, Google, Groq, Cerebras), with allow-listed model ids. Every call uses JSON Schema structured output generated from Zod. Without a provider key, calls go through the Browserbase Model Gateway. A client-side `generate` hook lets you route inference through any model you control. Vision is only used for optional `extract` screenshots. Pick browser-use for the widest provider choice, including local Ollama. Pick Stagehand when a short list of frontier providers, or your own `generate` function, is enough. More projects in this category are being researched. ## Verdict: How are browser sessions, profiles, auth and anti-bot handled? Both projects run local Chrome over CDP or a vendor cloud browser. Anti-bot features are mostly in the paid clouds. [browser-use](/p/browser-use/) configures everything in `BrowserProfile`. Settings include `user_data_dir` for persistent profiles, `storage_state` files for cookies, HTTP/SOCKS proxies and allowed or prohibited domain lists, which a security watchdog enforces. Local Chrome starts with a list of flags such as `--disable-blink-features=AutomationControlled`. `use_cloud=True` requests a Browser Use Cloud browser with proxies, country selection and fingerprinting. When that cloud's solver is running, the agent pauses on CAPTCHA events and reports the result to the model. [Stagehand](/p/stagehand/) offers two factories. `localBrowser.launch` starts Chrome with a temp or fixed `userDataDir`, a proxy flag and headless mode, then loads its extension. `browserbase.launch` creates a Browserbase session. Its options include advanced stealth, fingerprint settings, managed or external proxies, regions, CAPTCHA solving with selector hints, persistent contexts, recording and "verified" sessions. Cookies and extra headers are set through context methods. Any browser it drives must accept the Stagehand extension. For self-hosted runs, both offer only profiles and proxies. Choose by cloud: Browser Use Cloud with browser-use, or Browserbase with Stagehand. Note that Stagehand needs a browser that can load extensions. More projects in this category are being researched. # Project: browser-use/browser-use > Async Python agent loop that serializes pages into indexed DOM text plus screenshots and drives Chrome over raw CDP. — https://llms-technical-reviews.com/p/browser-use/ (commit 7be96ed) ### Overview browser-use is a Python library for an autonomous browser agent. You give an `Agent` a task in plain language and a chat model. It then runs a step loop. Each step reads the page as compact indexed DOM text plus a screenshot, asks the model for a structured `AgentOutput` (reasoning fields and a list of actions), and executes those actions in Chrome. The loop stops when the model calls `done`, the step budget runs out, or too many steps fail. The library drives Chromium directly over the Chrome DevTools Protocol through the `cdp-use` client. Playwright is not a package dependency at this commit; it is only invoked as a subprocess to install Chromium when no local browser is found. Browser work is organised around a `bubus` event bus: `BrowserSession` owns the CDP connection, and a set of "watchdogs" handle clicks, DOM capture, downloads, popups, security rules, storage state and CAPTCHA events. The library is aimed at Python developers who want a ready-made agent loop with many provider adapters, and who are fine with an LLM call on every step. The repo also ships more around the library: an MCP server, a `@sandbox` decorator for running code on Browser Use Cloud, a cloud browser client and history replay. Note that the `browser-use` console script now mostly delegates to the separate `browser-harness` package; the agent described here is the Python API. ### Architecture ```mermaid flowchart LR T["Task + LLM"] --> A["Agent.run / step"] A --> MM["MessageManager"] A --> LLM["BaseChatModel.ainvoke"] LLM --> OUT["AgentOutput (actions)"] OUT --> MA["Agent.multi_act"] MA --> TO["Tools.act -> Registry"] TO --> EB["bubus EventBus"] EB --> WD["Watchdogs"] WD --> CDP["cdp-use client"] A --> BS["BrowserSession state summary"] BS --> DOM["DomService + DOMTreeSerializer"] DOM --> CDP CDP --> CH["Chrome (local or cloud)"] ``` | Component | Path | Role | |---|---|---| | Agent | `browser_use/agent/service.py` | Step loop, retries, planning, loop detection, history, rerun | | Agent schemas | `browser_use/agent/views.py` | `AgentSettings`, `AgentOutput`, `ActionLoopDetector`, history models | | Message manager | `browser_use/agent/message_manager/` | Builds per-step prompts, compacts old history | | Tools and registry | `browser_use/tools/service.py`, `tools/registry/service.py` | Built-in actions, `@tools.action` decorator, dynamic action union | | Browser session | `browser_use/browser/session.py`, `profile.py` | CDP connection, targets, reconnection, launch and profile settings | | Watchdogs | `browser_use/browser/watchdogs/` | Event handlers for actions, DOM, screenshots, downloads, popups, security, CAPTCHA | | DOM pipeline | `browser_use/dom/service.py`, `dom/serializer/` | Merge DOMSnapshot + DOM + AX tree, filter, index, serialize | | LLM adapters | `browser_use/llm/` | `BaseChatModel` protocol and one package per provider | | Cloud and extras | `browser/cloud/`, `sandbox/`, `mcp/`, `skills/` | Cloud browsers, remote execution, MCP server, hosted skills | ### How a request flows 1. **Construct.** `Agent.__init__` falls back to `ChatBrowserUse` when no LLM is passed and turns on flash mode for it. It auto-sets screenshot size for Claude Sonnet and an LLM timeout per model family ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L225-L277)). It also enables coordinate clicking for some model names ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L326-L333)). 2. **Run.** `Agent.run(max_steps=500)` starts the session and then calls `_execute_step` until done. Each step runs under `asyncio.wait_for(step_timeout)`. 3. **Step.** `step()` first waits if a CAPTCHA solver is active. It then runs `_prepare_context`, `_get_next_action`, `_execute_actions` and `_post_process`, with one `_handle_step_error` for any exception ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1035-L1086)). 4. **Perceive.** `_prepare_context` calls `BrowserSession.get_browser_state_summary(include_screenshot=True)`. It filters actions by URL, builds the state message, and injects budget, replan, exploration and loop-detection nudges ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1087-L1160)). The summary request is a `BrowserStateRequestEvent` on the event bus ([session.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/browser/session.py#L1597-L1640)). 5. **Decide.** `get_model_output` calls `llm.ainvoke(messages, output_format=AgentOutput)`, trims to `max_actions_per_step`, and switches to `fallback_llm` on 401/402/429/5xx or truncated output ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1945-L1981), [L1983-L2025](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L1983-L2025)). 6. **Act.** `multi_act` runs the actions in order. It stops early on `done`, on an error, on an action marked `terminates_sequence`, or when the URL or focused tab changed ([service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/agent/service.py#L2730-L2848)). Each action goes through `Tools.act`, which calls `registry.execute_action` under a per-action timeout of 180 s by default ([tools/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L2178-L2240)). 7. **Execute in the browser.** For `click`, `_click_by_index` looks up the node in the selector map and dispatches `ClickElementEvent` ([tools/service.py](https://github.com/browser-use/browser-use/blob/7be96ed8bafa8dfe1eef228b59cf5c884b8b2431/browser_use/tools/service.py#L704-L734)). `DefaultActionWatchdog` refuses `