LLMs Technical Reviews

yc-software/qm

Self-hosted agent server for Slack and web that gives every person and channel its own memory, sandbox and keychain.

GitHub ↗★ 15kTypeScriptMITcommit 23af31b · 2026-10-05homepage ↗

Overview

QM (from YC Software) is an agent server for a whole company, not for one person. It runs as a headless TypeScript service. People talk to it from Slack or from its own web app. Each person, channel, team, group and the org itself is a scope (personal:, channel:, team:, group:, org: in src/types.ts). A scope owns its own memory notebook, files, keychain view, crons, published web apps and a durable sandbox computer. The design goal is that a personal assistant and a shared channel assistant can run on one deployment without leaking into each other.

QM does not ship its own agent loop. It wraps existing coding-agent harnesses, Pi (the default, via the vendored @earendil-works/pi-ai and pi-coding-agent packages), OpenCode, Codex and Claude Code. All four sit behind one Harness interface. Everything around the loop belongs to QM: the turn queue, the system prompt, memory capture, tool authorization, approvals, content screening, credentials and the sandbox. As a personal assistant, its standout feature is the built-in Inbox loop. It scans Gmail and Slack on a schedule and keeps a ledger of messages that wait for your reply, each with a draft in your voice.

This is a big codebase (about 2,400 files; src/core/orchestrator.ts alone is 4,666 lines). It is aimed at a startup that will run it in its own cloud account, with Postgres, and maintain a deployment repository for it. It is not a single-user desktop tool.

Architecture

flowchart LR
  SL["Slack plugin (Bolt)"] --> APP["app.turn()"]
  WEB["Web UI / admin (Lit)"] --> API["Fastify /v1/turns"]
  API --> APP
  APP --> Q["Run queue (Postgres runs)"]
  Q --> W["Worker: processRun"]
  W --> ORCH["Orchestrator.handleTurn"]
  ORCH --> MEM["Memory service + strategy"]
  ORCH --> TOOLS["Tool context + policy"]
  ORCH --> HR["Harness router"]
  HR --> H["Pi / OpenCode / Codex / Claude"]
  H --> LLM["Model catalog + providers"]
  TOOLS --> SBX["Per-scope sandbox"]
  ORCH --> TAPE["Session store / tape"]
Component Path Role
Entry point src/index.ts, src/wiring.ts Loads config, builds every service, runs migrations, starts the HTTP server, scheduler and Slack reconciler
HTTP API src/api/ Fastify routes (routes/turns.ts, admin, connectors, loops, crons, webhooks) and app-turn.ts, which admits and enqueues turns
Run queue and workers src/runs/ Postgres runs table with leases and heartbeats; worker.ts claims runs and calls the orchestrator
Orchestrator src/core/orchestrator.ts, src/core/orchestrator/ Builds one turn: identity, scope, system prompt, memory recall, tools, approvals, screening, then dispatches to a harness
Harnesses src/harness/ Harness contract, harness-router.ts, and adapters for Pi, OpenCode, Codex and Claude Code; agent-tools.ts defines the tool surface
Models src/model/ Catalog, overlays, custom providers and gateway models; default agent model id claude-opus-5
Memory src/memory/ Per-scope memory/MEMORY.md notebook with revisions, capture strategies, and optional external MCP providers
Policy and security src/policy/, src/security/ Command policy regexes, security posture, the LLM content-screening classifier
Sandboxes src/sandbox/ Sandbox interface with local (Docker), AWS microVM, Sprites, smolmachines, E2B, Modal, Porter, Agent37 and Superserve backends
Connectors and MCP src/connectors/, src/credentials/, src/mcp/ OAuth providers, encrypted keychain, admin-registered MCP servers
Loops and crons src/loops/, src/cron/ Background work: the Inbox loop, scheduled jobs, inbound webhooks
Slack src/slack/ In-process Bolt plugin that turns events into core turns and renders approvals
Surfaces plugins/ web-ui, admin, portal, auth, onboarding
CLI cli/ qm init, check, plan, up for deployment directories (Docker, Fly, AWS backends)

How a request flows

Take a Slack mention: “@qm summarise my unread email”.

  1. Startup. src/index.ts loads config, calls buildApp, runs the registered Postgres migrations, initialises sandbox resources, hydrates identity and custom providers, starts the run runtime and the HTTP server (src/index.ts).
  2. Ingress. The Bolt plugin handles app_mention and message events in src/slack/events.ts. callCore submits the turn with async: true (src/slack/turn-flow.ts). The Slack plugin runs in-process, so submitTurn is a direct call to app.turn with surface: "slack" (src/api/slack-core-client.ts). The web UI reaches the same function over POST /v1/turns (src/api/routes/turns.ts).
  3. Admission and enqueue. app.turn resolves the actor and refuses non-internal principals. It handles steering into an in-flight run, then calls runs.enqueue with the thread as the session id and an optional dedup key (src/api/app-turn.ts, src/api/app-turn.ts).
  4. Claim. A worker claims the oldest pending run with FOR UPDATE SKIP LOCKED. It skips any run whose session already has a running or earlier pending sibling, so turns in one thread run strictly in order (src/runs/postgres-run-store.ts). processRun heartbeats the lease and cancels the turn if the lease is lost (src/runs/worker.ts).
  5. Build the turn. Orchestrator.handleTurn checks access, takes a session turn lease, and resolves scope and sharing. It reads a memory snapshot (src/core/orchestrator.ts). It then assembles the system prompt from the mode frame, the resolved scope prompt, shared core instructions and the security-policy text (src/core/orchestrator.ts). Skills, cron, keychain and connected-app blocks are appended after that.
  6. Tools. createToolContext binds the sandbox, the credential handles and the command policy for this scope (src/core/orchestrator.ts). A tool call that matches a require_approval rule throws NeedsApproval. A deny rule throws CommandDenied (src/tools/primitives.ts).
  7. Recall and dispatch. Recalled memory goes into the turn’s environment note as a delta, so facts already visible in the history are not repeated (src/core/orchestrator.ts). Then deps.harness.turns.runTurn is called (src/core/orchestrator.ts). The router resolves the runtime choice, resets provider sessions if the thread switched harness, and wraps adapters that lack native goal enforcement (src/harness/harness-router.ts).
  8. Finish. Entries go to the session store as the turn runs. On a pause for approval, the run ends with pausedOnApproval, and Slack renders an approval card. Otherwise the memory strategy’s onTurnEnd is queued for the scope (src/core/orchestrator.ts), and the worker marks the run complete.

Key components

Harness contract

Harness has four parts: an adapter profile (control transport, tool transport, capabilities such as steer, native-tape and goal-enforcement), turns.runTurn, a set of optional model utilities (oneShot, judge, screenSecurity, compactHistory, generateTitle), and tool-name presentation (src/harness/harness.ts). The utilities matter, because memory extraction, titles and security screening are all one-shot calls through the active harness. QM therefore needs no separate LLM client for them. HARNESS accepts mock, pi, opencode, codex or claude (src/config.ts).

Tool surface

The tool set is fixed and small. It is defined in src/harness/agent-tools.ts: execute, sandbox, read, write, skill, memory, history, subagents, sessions, background, cron, webhook, share, publish, attach, finish_silently and a few others. Most real work goes through execute, which runs shell commands in the scope’s sandbox. Connectors and skills are reached through that sandbox, not through hundreds of typed tools. Admin-registered MCP servers are the exception. Their tools are namespaced <serverId>_<toolName>, and every call is audited (src/mcp/mcp-tool-service.ts).

Memory

Memory is a Markdown bullet notebook per scope, memory/MEMORY.md (src/memory/memory-service.ts). Recall is capped at the last 6,000 characters (src/memory/memory-service.ts). With Postgres, every write is a row in memory_revisions, so history and restore come for free (src/memory/postgres-memory-service.ts). Capture is pluggable: per-turn (default), scratch-promote or agent-only (src/memory/strategy.ts). The default does not extract after every message. It buffers a “burst” of turns per scope, conversation and actor. It flushes after 3 minutes of quiet or 10 turns, then makes one extraction call with a strict provenance prompt: only facts the user stated, directives kept verbatim, no secrets (src/memory/strategies/per-turn.ts, src/memory/strategies/per-turn.ts). Retrieval is not vector-based. The prompt gets the tail of the notebook, and query is a keyword match over bullets.

Security posture and command policy

An org picks dangerous, auto or strict. Narrower scopes can only tighten it (src/security/security-posture.ts). strict sends every tool call to human approval, and auto blocks private networks. Content screening (off/observe/enforce) is a separate deployment setting, and dangerous caps it at observe. Independently of posture, a regex command policy applies everywhere. The org floor asks for approval on recursive rm, force-push, DROP/TRUNCATE TABLE and curl | sh, and hard-denies mkfs and fork bombs. Scope rules are appended, and an org allowlist cannot be loosened to a denylist (src/policy/command-policy.ts).

Sandboxes

The Sandbox interface covers provisioning from workspace layers, run, file I/O, staging blobs, long-running processes with stdin and signals, and home snapshots (src/sandbox/sandbox.ts). SANDBOX_BACKEND defaults to local, which uses docker exec. Eight other backends are accepted (src/config.ts). Some hosted backends run with open egress unless you configure an egress proxy, and config prints a warning when that happens.

Inbox loop

The Inbox loop is a recurring agent task (every 15 minutes by default). It scans Gmail and Slack for conversations where the last message waits on the user, investigates context, writes a draft reply in the user’s voice, and posts items to a ledger API. If nothing changed, the turn ends silently (src/loops/inbox-loop.ts).

Extending it

  • Deployment directory. Org config, custom tools, skills, sandbox image and infrastructure live outside core, in a deployment repository that pins @yc-software/qm. The qm CLI validates and deploys it (check, plan, up).
  • Skills. Skills are scope-owned folders, seeded from skills-seed/ (Gmail voice profile, morning digest, Google Workspace, Linear, publish and others). They can be shared by grant, promoted org-wide by an admin, or imported as packs from git (src/skills/).
  • Models. Admins can register custom providers that speak OpenAI Chat Completions, OpenAI Responses or Anthropic Messages at any http(s) base URL. That includes self-hosted endpoints. These models are reachable only through harnesses that route via pi-ai (src/model/custom-providers.ts).
  • MCP servers. Admins (not end users, by design) register HTTP MCP servers with no auth, a bearer token or client-credentials auth.
  • Memory providers. src/memory/provider-router.ts can route a scope kind to an external MCP memory provider and keep the built-in notebook.
  • Harnesses. A new harness is a defineHarness(profile, implementation) call plus registration in wiring.

Running it

  • Requirements. Node 24.15+, at least one model credential (Anthropic, OpenAI, OpenRouter, a Codex or Claude subscription credential, or a custom provider) and, for anything real, Postgres. Without SESSION_STORE=postgres and DATABASE_URL, sessions live in process memory and disappear on restart.
  • Local development. npm start runs src/index.ts directly on Node, and npm run worker runs a separate worker. npm run dev-instance:web / :slack brings up a real-model, real-Postgres instance. The local sandbox backend needs Docker.
  • Production. qm init --org <slug> --target <fly-or-aws> scaffolds a deployment repository. Deploy providers are docker, aws, fly and porter, with Helm charts under deploy/helm. Slack is optional: socket mode or HTTP events, with bot and app tokens. Published web apps need a wildcard DNS record.

Strengths and caveats

  • Strength: real multi-tenancy. Scope-keyed memory, keychains, sandboxes and grants, plus an isolated default sharing posture, is the core of the design. Most personal-assistant projects never attempt this.
  • Strength: durable execution. Turns are Postgres runs with leases, heartbeats, per-session ordering and crash-loop parking. A dead worker does not lose a conversation.
  • Strength: layered safety. Posture, regex command policy, human approvals with session/always grants, an LLM injection screen over tool results, and audit sinks for egress and credential use all work together.
  • Strength: harness and sandbox portability. You can swap Pi, Codex, OpenCode or Claude Code, and nine sandbox backends, without changing the surrounding product.
  • Caveat: weight. It needs Postgres, a sandbox backend, OAuth app registrations and a deployment repository. The orchestrator is one 4,600-line closure, which makes it hard to read and to fork.
  • Caveat: simple memory. Memory is a flat bullet list with tail recall and keyword search. There are no embeddings, so large notebooks lose their oldest facts from the prompt.
  • Caveat: shell-first integrations. Most connectors are used by the agent through execute and skills in the sandbox. This is flexible, but it costs more turns and tokens than typed tools.
  • Caveat: contributions are text-only. The project accepts design notes in adrs/, not code PRs.

Sources: code at 23af31b, deepwiki-open wiki (20 pages), OpenDeepWiki wiki (35 pages), verified Q&A.

How it answers the Open-source personal assistants questions

Each answer was drafted by a code-reading agent at commit 23af31b. Its citations were checked mechanically. Compare with the other open-source personal assistants →

How is the assistant architected?

answered

QM is a headless core service with plugin surfaces, all in TypeScript on Node.js (≥24.15). The core (src/index.ts → src/wiring.ts) starts a Fastify HTTP server, attaches identity/directory services, the orchestrator, and optional Slack and portal plugins. There's an agent loop with durable execution against a central write-ahead log: the "session tape" (src/sessions/session-store.ts, the write-ahead log abstraction). A session offers up a write lease, a worker claims it from the run queue, and the orchestrator (src/core/orchestrator.ts) assembles system prompts, memory, context history, tools, and security screening before dispatching to a swappable harness (src/harness/harness-router.ts). Harnesses include Pi (src/harness/pi-harness.ts), OpenCode, Codex, and Claude — each implements runTurn(), calling into the model catalog (src/model/pi-models.ts, src/model/model-catalog.ts) which resolves model IDs and provider credentials. A user request arrives via Slack (Bolt-based events plugin at src/slack/) or the web UI (plugins/web-ui/, Vite+Lit). The API routes (src/api/routes/) validate auth, dispatch a turn, run the orchestrate→harness→model pipeline, and stream results back. Tool calls (src/tools/primitives.ts, src/harness/agent-tools.ts) are fulfilled against the scope's sandbox (src/sandbox/sandbox.ts) — an isolated compute environment backed by AWS, Fly, Docker, Modal, Sprites, etc. The model response and tool results are committed to the session tape. The whole system scales to swarms of agents for background work (crons, webhooks, sub-agents).

How are integrations (email, calendar, chat, docs) implemented?

answered

External services (email/calendar/chat/docs) are connected via OAuth flows routed through a central connectors/oauth system (src/connectors/oauth.ts). Each OAuth provider (Google, GitHub, Slack, X, etc.) defines a provider config with authUrl, tokenUrl, scopes, clientIdEnv, PKCE support, and custom exchange/refresh adapters. The PROVIDERS map in oauth.ts centralizes all connector definitions. OAuth tokens are stored encrypted per-scope in the keychain (src/credentials/keychain.ts), persisted in Postgres, with encrypted fields (encryptSecret/decryptSecret in connector-client-store.ts). Token refresh uses stored refresh tokens. There is no MCP-based integration layer for external services — direct API clients are used. Slack itself is unique: it's an in-process plugin (src/slack/index.ts) using Bolt for events and the web API client for messaging, started and supervised directly by core. Gmail push notifications use a separate createGmailPushClient (src/wiring.ts:27) for background mail syncing. Gmail sent mail handling is at src/connectors/gmail-sent.ts. The web UI (plugins/web-ui/src/connectors.ts, plugins/web-ui/src/connector-widget.ts) renders the connection management UI. A Composio bridge exists (src/api/routes/composio.ts) for third-party app integrations with on-demand consent flows. Connector status freshness is tracked via src/credentials/connector-status.ts. Sync is on-demand (per-turn keychain manifest injection) — there is no general background sync engine for connectors, though Gmail push and crons provide some background access.

Editor's note. Correction: QM does have an MCP integration layer: admins register HTTP MCP servers (none/bearer/client-credentials auth) and src/mcp/mcp-tool-service.ts exposes their tools to agents as audited <serverId>_<toolName> tools. Background sync also exists: the Inbox loop (src/loops/inbox-loop.ts) rescans Gmail and Slack every 15 minutes by default.

How is memory and user context stored and retrieved?

answered

Memory is per-scope (personal, channel, team, org) and stored as a sequence of Markdown bullet points in a write-ahead revision log in Postgres (src/memory/postgres-memory-service.ts). Each scope has a MEMORY.md file (path constant memory/MEMORY.md in src/memory/memory-service.ts:17) rendered as an append-only series of dated bullet facts. The revision table (memory_revisions) tracks scope_id, seq, op, body, author, and at timestamp. On each turn, the orchestrator calls memory.recall(scopeId), which appends the last 6,000 chars of memory into the system prompt (src/memory/notebook.ts:1 — RECALL_MAX_CHARS = 6_000). Memory extraction happens automatically after each turn via the per-turn strategy (src/memory/strategies/per-turn.ts), which uses a one-shot model call to extract durable facts from the last conversation exchange. Facts are deduplicated and appended. The system supports multiple memory backends via a provider router (src/memory/provider-router.ts), which can route by scope kind to different providers: Postgres built-in, MCP-based memory providers, or "Memorable" procedural memory (src/memory/memorable/). The config surface is at src/memory/provider-config.ts. Recall is trimmed via capTail to the character limit; there is no vector store for retrieval — it's a flat tail-of-history capture. Context boundaries (src/memory/context-boundary.ts) track which entries the current memory snapshot covers to avoid double-injection. Context compaction (src/harness/context-compaction.ts) and summarization (src/sessions/session-store.ts:39 with ContextSummaryPayload) handle prompt length.

Editor's note. Correction: the default per-turn strategy does not extract after every exchange; it buffers a burst and extracts after 180 s of quiet or 10 turns (per-turn.ts L14-L15, L88-L140). Recalled memory is injected as a delta into the turn's environment note, not appended to the system prompt (orchestrator.ts L3560-L3563).

How are actions on the user's behalf gated?

answered

Actions are gated at three layers: command policy, approvals, and security screening. Command policy (src/policy/command-policy.ts) defines an org-floor denylist of dangerous patterns (recursive delete, force push, destructive SQL, fork bombs, pipe-to-shell) with decisions of allow, deny, or require_approval. Policies are composed as org-floor + scope-specific rules, in either denylist or allowlist mode. The NeedsApproval / CommandDenied exceptions (src/tools/primitives.ts:118-149) are thrown by the authorization layer in createToolContext() and caught by the orchestrator. When a command requires approval, a PendingApprovalRecord is written to the approval store (src/core/approval-store.ts), which enqueues a delivery to the user via Slack. The orchestrator pauses the turn (pausedOnApproval) and awaits human resolution. Approval grants can be one-shot (session) or remembered (always). The security posture system (src/security/security-posture.ts) has three levels: dangerous (no tool approvals, no private network blocking), auto (no tool approvals but blocks private networks), strict (requires approval for every tool call). Security screening (SECURITY_SCREEN) can be off, observe (log only), or enforce (block), and uses an LLM classifier to detect prompt injection in tool results before they reach the agent. The orchestrator's security-screen.ts classifies each tool result and can taint entries or abort the turn. The audit trail includes src/admin/postgres-audit-log.ts, src/admin/egress-audit-sink.ts (outbound data), and src/admin/credential-usage-sink.ts. Background work (crons, webhooks) requires explicit grants before credentials can be used (SPEC.md line 28).

How are LLM providers selected and configured?

answered

LLM providers are configured via environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY, OPENROUTER_API_KEY) and the HARNESS config (pi, opencode, codex, claude). The model catalog (src/model/model-catalog.ts) aggregates from four sources: built-in models from @earendil-works/pi-ai (Anthropic/OpenAI base models), model overlays (src/model/model-overlay.ts — admin-customizable price/speed/label overrides), custom providers (src/model/custom-providers.ts), and gateway models (src/model/gateway-models.ts). Model selection (src/model/pi-models.ts) resolves model IDs with resolveModel(), which looks up the built-in catalog then falls through to overlays and custom providers. The default agent model is claude-opus-5 (DEFAULT_AGENT_MODEL_ID). Tool-calling and structured output are managed through the Pi harness (src/harness/pi-harness.ts) which uses pi-ai's ModelRuntime — standard tool definitions are created via createTurnTools() with Pi's ToolDefinition interface. Effort/thinking levels include auto, default, adaptive, low through max/ultracode. Fast mode is supported on all harnesses (harnessSupportsFastMode). There is no local-model support — the system requires API access to Anthropic, OpenAI, or OpenRouter. The admin UI has routes for model registry management (src/api/routes/admin/model-registry.ts, src/api/routes/admin/model-providers.ts, src/api/routes/admin/custom-providers.ts). A model verification system (src/model/model-verification.ts) and browser-model fallback (src/model/browser-model.ts) exist for edge cases.

Editor's note. Correction: local or self-hosted models are possible: admins can register custom providers at any http(s) base URL that speaks OpenAI Chat Completions, OpenAI Responses or Anthropic Messages (src/model/custom-providers.ts L1-L20). They are only reachable from harnesses that route through pi-ai.

How is it deployed and self-hosted?

answered

Deployment requires Postgres as the durable store (required for production; SESSION_STORE=postgres with DATABASE_URL). No queue server is needed — the system uses pg-boss-style advisory locks and its own session/run queue over Postgres (src/sessions/session-store.ts). The core runs on Node 24.15+ and is a single Fastify server. The one-click path uses Docker (deploy/core/Dockerfile) or Docker Compose via the CLI (cli/src/backends/docker.ts). Several deploy backends exist: Docker (src/deploy/docker-deploy-provider.ts), AWS (src/deploy/aws-deploy-provider.ts), Fly.io (src/deploy/fly-deploy-provider.ts), and Porter (src/deploy/porter-deploy-provider.ts). Each implements DeployProvider.apply() to stand up the full stack. The CLI (cli/bin/qm.ts) has qm sandbox, qm setup, qm init, and infra commands. Deploy profiles are managed via DEPLOY_PROVIDER env var. Required external accounts: at least one LLM API key (Anthropic/OpenAI/OpenRouter); optionally Slack (SLACK_BOT_TOKEN, SLACK_APP_TOKEN), email (Resend), and a container host (Docker daemon, Fly account, AWS, etc.) for app publishing. The web UI serves from plugins/web-ui/ via Vite and is bundled into the service or deployed separately (deploy/web-ui/Dockerfile). Admin UI is at plugins/admin/, auth broker at plugins/auth/, portal at plugins/portal/. No browser is required at runtime. There is no docker-compose.yml in the repo — the dev-instance script (scripts/dev-instance.sh) and CLI handle local orchestration. The sandbox backend can be any of ~10 providers (AWS, local, Sprites, Smolmachines, E2B, Modal, Porter, Agent37, Superserve).