agenvoy/Agenvoy
Single-binary Go daemon for a local personal agent, with LLM-routed model choice, sandboxed tools, cron skills and chat bots.
Overview
Agenvoy is a personal AI agent written in Go and shipped as one binary, agen. You run it on your own laptop or workstation. A background daemon owns everything: sessions, tools, schedules, chatbot connections and model routing. The terminal UI, the web dashboard, the Telegram and Discord bots and an OpenAI-compatible HTTP endpoint are all clients of that one daemon. State lives under ~/.config/agenvoy/ in SQLite and an embedded ToriiDB store, and API keys live in the OS keychain.
It is built for one technical user who wants a long-running assistant that can run shell commands, edit files, call APIs, run scheduled “skills” and answer from a phone chat app. It is not a multi-user server. The HTTP server binds to 127.0.0.1 only, and most management routes are also wrapped in a localhostOnly() check.
Two design choices stand out. First, model choice is automatic per request: a dispatcher LLM classifies the request as code, research, work, chat or fetch, and that label picks both the model tier and the reasoning effort. Second, the safety model is layered: a confirmation prompt for side-effecting tools, a sudo-password gate for paths outside the home directory or on a sensitive list, and a bubblewrap sandbox with no network by default for shell commands. The project is dual-licensed: AGPL-3.0, or a commercial license for network-served derivatives.
Architecture
flowchart LR
TUI["TUI (bubbletea)"] --> API["Gin HTTP API on 127.0.0.1"]
WEB["Web dashboard"] --> API
OAI["OpenAI-compatible clients"] --> API
BOTS["Telegram / Discord bots"] --> EXEC
CRON["Scheduler (cron skills)"] --> EXEC
API --> EXEC["exec.Start / Execute loop"]
EXEC --> SEL["Dispatcher: work type + model"]
EXEC --> LLM["go-llm-router providers"]
EXEC --> TOOLS["Tool registry"]
TOOLS --> SBX["bubblewrap sandbox"]
TOOLS --> MCPC["MCP client servers"]
TOOLS --> SUB["Subagents"]
EXEC --> MEM["SQLite + ToriiDB memory"]
STDIN["stdin pipe"] --> MCPS["MCP server mode"]
| Component | Path | Role |
|---|---|---|
| Entry point | cmd/app/main.go |
Chooses TUI, --daemon, stop, update, or MCP-server mode when stdin is a pipe |
| Daemon boot | cmd/app/daemon.go |
Opens storage, registers tools, starts MCP manager, scheduler, cron jobs, bots and the HTTP server |
| HTTP routes | internal/runtime/routes/ |
/v1/send, /v1/chat/completions, sessions, models, keys, tool confirmations |
| Agent loop | internal/agents/exec/ |
Prepare, Start, Execute, toolCall, model selection, compaction, summaries, subagents |
| Model config | internal/agents/keychain/, internal/session/config/ |
Provider credentials, work kinds, tiers and reasoning levels |
| Tools | internal/tools/ |
Built-in tools plus api_, script_ and ext_ tool groups |
| File boundary | internal/tools/file/boundary/ |
Work-dir confinement, sensitive and denied paths, per-session grants |
| MCP | internal/runtime/mcp/ |
MCP client (stdio/HTTP, OAuth) and a stdio MCP server that exposes Agenvoy’s tools |
| Storage | internal/runtime/store/, internal/runtime/torii/ |
SQLite history with FTS5; ToriiDB for tool cache, session history and error memory |
| Chatbots | internal/runtime/chatbot/ |
Telegram and Discord adapters |
| Web UI | page/, internal/runtime/webapp/ |
Embedded dashboard; vendor assets fetched once per version |
How a request flows
Take a message typed into the web dashboard:
- Boot (once).
Daemon()initialises the filesystem, ToriiDB, SQLite stores, the tool registry and subagents, starts the MCP manager in a goroutine, registers two system crons (summaries every 15 minutes, session cleanup every 30), connects Discord and Telegram, and servesroutes.New()on127.0.0.1(daemon.go). - Receive.
POST /v1/sendbinds the request, resolveswork_dir(default: home directory), creates atemp-orchat-session id and callsexec.Prepare, which auto-matches a skill from the message text. If the session is already running, the text is queued as a “steer” message instead of starting a second run (send.go, start.go). - Pick a model.
StartcallsResolveAgent, which runs the dispatcher.selectAgentDispatchersends the registered models and the request to a routing model with no reasoning and parses back a work type plus model names; the work type maps to a reasoning level (start.go, selectAgent.go). - Loop.
Executeruns up toMAX_TOOL_ITERATIONS(256) rounds. Each round injects pending steer messages, compacts history if the last input passed 80% of the model’s window, assembles messages withcompact.AssembleMessages, and streams the call throughstreamSend. A watchdog probes an unresponsive provider and switches to a fallback model if needed (execute.go). - Tools. If the reply has tool calls,
toolCallchecks each one. Restricted paths andtoolNeedsConfirmationdecide whether to ask the user throughruntime.Ask. Tools marked concurrent run in parallel in up to five slots, background tools are fired off, and the rest run in order. Results go back into the next round (toolCall.go). - Finish. A plain-text reply ends the loop. If the text contains the guardrail ban tag, it is replaced with a templated refusal (execute.go). Events stream back to the dashboard over SSE and are written to the session log.
Key components
Model routing
Models are registered as provider@model names. keychain.Config maps the provider prefix to a keychain secret or an OAuth flow: Claude, OpenAI, Gemini, Grok, DeepSeek, Mistral and others, plus a local claude-code subprocess provider (keychain.go). WorkKinds defines five work types. Each has a reasoning level (code is xhigh, chat is none) and a tier preference order, so chat goes to cheap B/C models first and code goes to S models first (tier.go). An optional beta dispatcher replaces the routing LLM with a hosted classifier at api.typesafe.ai, which receives the request text and recent turns (selectAgentBeta.go).
Safety layers
toolNeedsConfirmation always asks before reading sensitive files and before remove/restore modes. It skips the prompt for read-only tools, list/read/search modes, HTTP GETs, an allowlist of read-only shell commands, and rules the user saved with “remember” (toolCall.go). boundary.Restricted flags paths outside the home directory, and sensitive paths inside it, unless the session already has a grant. Those calls need a confirmation that also carries the system password (boundary.go), and sudo.Verify checks that password with sudo -S -v (password.go). If the approving channel cannot collect a password, the call is refused. run_command takes a bare binary name, blocks denied commands and sudo, turns rm into a move to trash, and runs the command through go_pkg_sandbox.Wrap with network denied unless the call sets network: true (runCommand.go).
Memory
There are three layers. Short-term, AssembleMessages builds system prompts, old histories, a summary message, the user input and tool history. Compaction triggers at 80% of the window, with a 128k fallback when the window is unknown (checkThreshold.go). Per session, the 15-minute cron folds new history chunks into a JSON summary dictionary through summary.Generate (generate.go). Across sessions, ToriiDB keeps four databases: tool cache, session history, tool-error memory and online markers (torii.go). Error memory search tries vector similarity first and falls back to a keyword scan (search.go). Embeddings use the OPENAI_API_KEY from the keychain, so without it the “semantic” part is keyword matching.
Tool registry
tools.init registers built-ins (files, run_command, fetch_page, search_web, schedules, ask_user) and three dynamic groups: api_ (JSON-described HTTP APIs), script_ (script tools in a folder) and ext_ (extension-installed tools) (register.go). Not every schema goes into every prompt. find_tools is always loaded and lets the model search the registry and pull in a tool’s schema when it needs it. MCP tools are registered under mcp__<server>__<tool> when a server’s tool list refreshes (mcp.go).
Subagents and schedules
The subagents tool runs a task in a separate session with its own model, reasoning level and excluded tools through exec.ExecWithSubagent, then saves the output as a report (invoke.go). The scheduler (go-scheduler) runs skills on cron specs or one-off timers through the RunSkill runner, so a scheduled job is just a skill run in a session.
Extending it
- Skills.
SKILL.mdfolders in~/.config/agenvoy/skills/or<workdir>/.config/agenvoy/skills/.Preparematches them from the message text, or the caller names one. The bundledskill-creatorandscheduler-skill-creatorskills let the agent write new ones. - Script and API tools. A folder per tool with a description document and a script goes in
tools/script/. The translator skips names starting with_or.and runs scripts through the same sandbox wrapper. A JSON API description goes intools/api/. Both also have per-work-directory variants (translator.go). The agent can author these itself withedit_toolandtest_tool. - MCP servers. Add stdio or HTTP servers to the MCP config. OAuth is handled for HTTP servers.
- Agenvoy as an MCP server. When
agenis started with stdin as a pipe, it serves its script, API and extension tools plus a set of bridged built-ins over stdio (main.go, expose.go). - HTTP clients.
/v1/chat/completionsand/v1/modelsspeak the OpenAI format, so other tools can use the daemon as a model endpoint (new.go).
Running it
- Install. The README’s install script, or
make build, which builds with thefts5tag and copies bundled skills into~/.config/agenvoy/skills/.system(makefile). - Start.
agenattaches the TUI and spawns the daemon if it is not running.agen stopstops it. An opt-in systemd user unit (Linux) or launchd agent (macOS) starts it at login. - Required. At least one provider key or OAuth login,
bubblewrapon Linux for the sandbox, and network access on first boot of each version, becausewebapp.SyncAssetdownloads the dashboard’s vendor files (asset.go). There is no Docker, Postgres or Redis. Bot tokens are needed only for Telegram or Discord.
Strengths and caveats
- Strength: real execution guardrails. Network-off sandboxing, rm-to-trash, a sudo-backed gate for paths outside the home directory and per-argument allowlists are more than most personal assistants ship.
- Strength: one daemon, many front doors. TUI, web, chat apps, cron, OpenAI-compatible HTTP and MCP all reach the same session store and tool registry.
- Strength: resilient loop. Health-probe watchdogs, model fallback, rate-limit cooldowns, compaction with fallbacks and resumable pending tasks are all in the main loop.
- Caveat: routing costs an extra call. Every new request with more than one registered model makes a dispatcher LLM round trip before work starts. The beta dispatcher sends request text to a third-party service.
- Caveat: memory is narrow. There is no general fact or user-profile store. Long-term memory is session summaries plus tool-error records, and vector search depends on an OpenAI key.
- Caveat: no first-party mail or calendar. Personal-data integrations come from MCP servers or tools you add.
- Caveat: large and fast-moving. About 54k lines of Go in one daemon package tree, with many interacting timeouts and retry counters. Expect behaviour to shift between releases.
Sources: code at d19ab5a, verified Q&A.
How it answers the Open-source personal assistants questions
Each answer was drafted by a code-reading agent at commit d19ab5a. Its citations were checked mechanically. Compare with the other open-source personal assistants →
How is the assistant architected?
answeredAgenvoy is a single Go binary that runs as a local daemon (cmd/app/daemon.go). The daemon boots, initialises filesystem/DB/ToriiDB, spawns a Gin HTTP server on 127.0.0.1:17989, and serves an embedded web dashboard (internal/runtime/webapp/) plus a TUI (internal/runtime/tui/). The agent execution loop lives in internal/agents/exec/execute.go::Execute(). A user request arrives via HTTP POST to /v1/send, which calls Execute() with an ExecuteMeta struct containing the agent config, content, session ID, tools, and skill. Inside Execute() there is a for range limit loop (max 256 iterations) that: assembles messages via compact.AssembleMessages(), streams the LLM call (via go-llm-router/core), handles tool calls via toolCall(), and feeds results back to the model. The LLM is wrapped as a provider.Agent from go-llm-router. Tools are executed through an Executor that dispatches by name to registered handler groups (tools/register.go registers built-in tools like run_command, read_files, edit_file, plus dynamic groups api_, script_, ext_ for API tools, script tools, and extension tools). The full pipeline: request → route handler → Execute() → streamSend() → LLM response → tool call → tool result → loop back to LLM → final text. The frontend has no separate server — the daemon serves the web UI from embedded assets (page/page.go) and all API endpoints are localhost-only except session chat and confirm endpoints, which CORS-restrict to local origins.
How are integrations (email, calendar, chat, docs) implemented?
answeredAgenvoy does not have dedicated API clients for email, calendar, or docs. Instead, external services are accessed through MCP servers — the built-in MCP client connects to arbitrary MCP servers over stdio or HTTP (internal/runtime/mcp/config.go). The MCP Client interface (internal/runtime/mcp/client.go) wraps the go-sdk to call ListTools and CallTool. Tools from connected MCP servers are registered at runtime under the mcp__<server>__<tool> prefix via toolRegister.Regist(). OAuth for HTTP MCP servers is handled by internal/runtime/mcp/oauth.go. Chatbot integrations (Telegram and Discord) are first-class: the daemon connects outbound to these platforms using bot tokens stored in the OS keychain (internal/runtime/chatbot/telegram/ and discord/), with full message loop, file attachment handling, and session management via session IDs prefixed tg- and dc-. There is no background sync daemon — data is fetched on demand when the LLM calls a tool. Sync vs on-demand: all external data flows through tool calls the LLM makes during a run — no polling or webhook ingestion exists beyond the Telegram/Discord long-poll loops. LLM providers themselves are integrated via go-llm-router which supports 14+ providers (internal/agents/keychain/keychain.go): API-key-based (OpenAI, Claude, Gemini, Grok, DeepSeek, Mistral, NVIDIA, OpenRouter, Cloudflare, Ollama Cloud), OAuth-based (GitHub Copilot, OpenAI Codex, Grok-OAuth), Claude Code as a subprocess, and a generic compat provider for any OpenAI-compatible endpoint. API keys are stored in the OS keychain. OAuth tokens are loaded and refreshed by go-llm-router's oauth helpers.
How is memory and user context stored and retrieved?
answeredAgenvoy uses a multi-layered memory system. Persistent storage uses two backends: SQLite (via go-sqlkit) for structured data — session messages (store/migrate.sql creates messages, session, file_history, action_history tables with FTS5 full-text search), and ToriiDB (an embedded vector database) for tool cache, session history, error memory, and online markers (torii/torii.go defines four DBs: DBToolCache, DBSessionHist, DBErrorMemory, DBOnline). Tool error memory (internal/agents/exec/memory/) stores structured Record entries with tool name, symptom, cause, action, outcome, and keywords. Records are persisted via torii.DB(DBErrorMemory).SetVector() with vector embeddings for semantic search. On tool failure, memory.Search() is called to find up to 16 related past errors by vector similarity or keyword scan. Session conversation memory uses summarization: every 15 minutes a system cron runs app.GenerateSummary (internal/agents/exec/summary/generate.go::Generate()), which chunks new session history records, sends them to the summary model with a prompt (configs.SummaryPrompt), and stores a JSON summary dictionary per session in ~/.config/agenvoy/sessions/<id>/summary.json. Context injection happens in systemPrompt.go — the getSystemPrompt() function builds the full system prompt by templating {{.AvailableSkills}}, {{.OfficialGuide}} (model-specific capability docs), {{.PermissionMode}}, {{.GuardrailRules}}, {{.BotPersona}}, and agent guide files (CLAUDE.md/AGENTS.md). The compact/ package handles prompt assembly: AssembleMessages() builds the message array from system prompts, summary message, old histories, user input, and tool histories. Context window management is handled by compact.CheckThreshold() with an 80% ratio trigger, and compact.ExtractOldHistories() moves older messages into a compacted summary when token limits approach.
How are actions on the user's behalf gated?
answeredActions are gated through a multi-level permission system in internal/agents/exec/toolCall.go. The toolNeedsConfirmation() function determines if a tool call requires user approval by checking: (1) whether reading sensitive paths (isSensitiveReadFile), (2) whether the tool is in destructive mode (remove/restore), (3) whether the tool is read-only (list/read/search), (4) whether the HTTP method is a GET, (5) whether the run_command is in the ReadOnlyCommand list, and (6) whether the tool is on the user's allowlist. If confirmation is needed, the runtime.Ask() function sends a runtime.Request{Kind: KindToolConfirm} to the TUI or web dashboard. The user can approve/deny, mark as "remember" (adds an allowlist rule), or "allow this turn". Restricted paths (outside work dir or on the sensitive path list from sensitive_path.json) trigger a password verification gate via boundary.Restricted(). The allowlist is session-level and tool-argument-aware: allow/tool/list.go loads work-dir specific allowlist entries. Sandboxing via bubblewrap (go_pkg_sandbox) wraps command execution: run_command.go and script_ tools execute in a sandbox that blocks sensitive paths and credentials. The configs/jsons/guardrail_rules.json embeds refusal rules, and configs/jsons/refusal_messages.json provides refusal response templates. File boundary enforcement (internal/tools/file/boundary/) includes the sensitive-path checker and the ExcludeList/DeniedPath lists. An audit trail is recorded: all tool calls are logged via sessionLog.Record() with full tool name, arguments, results, and errors. The action_history SQLite table stores every action per task hash with model, reasoning level, objective, tool results, and final reply, supporting replay and review.
boundary.Restricted), not to paths outside the work directory; remembered allowlist rules are stored per work directory and persist, not per session.How are LLM providers selected and configured?
answeredLLM provider selection and configuration is handled by go-llm-router (a Go package github.com/pardnchiu/go-llm-router). The provider.Agent interface provides Send() and SendStream() methods. The keychain/keychain.go resolves provider config from the OS keychain: it supports API-key-based providers (claude, openai, gemini, grok, deepseek, mistral, nvidia, openrouter, cloudflare, ollama-cloud), OAuth-based providers (copilot, codex, grok-oauth with token refresh), Claude Code as a subprocess (claude-code), and a generic compat provider for OpenAI-compatible endpoints. Models are registered in ~/.config/agenvoy/config.json as a flat list of model names like [email protected]. Tiered routing (internal/session/config/tier.go) classifies model names into tiers S/A/B/C based on heuristics (e.g., claude-fable → S, claude-sonnet → A, claude-haiku → B, gemma → C), with a fallback priority rank. LLM-driven agent selection (internal/agents/exec/selectAgent.go) calls a dispatcher model with a prompt listing available agents and the user request, asking it to return which model(s) to use and a work type (code/research/work/chat/fetch). The work type determines the reasoning level (xhigh/high/medium/low/none). If no dispatcher model is configured, the first registered model is used. An optional dispatcher_beta mode uses a faster Jev classifier for the routing decision. Structured output / tool calling is handled by go-llm-router's native support — the agent sends tool definitions to the LLM, and tool calls are extracted from the response choices at execute.go:706-763. Reasoning effort is specified per-model through the provider.Reasoning enum (none/low/medium/high/xhigh) from the go-llm-router, and reasoning tags like </think> are parsed as a fallback. Context window management uses model-specific thresholds from compact.CheckThreshold() (based on 80% of the model's context window). Local models are supported via the compat provider pointing at a local endpoint (e.g., Ollama at http://localhost:11434/v1) and the ollama-cloud provider (which is Ollama's cloud offering, not local — but the compat provider covers truly local endpoints).
dispatcher_beta classifier is a hosted third-party service (api.typesafe.ai, model jev-latest) that receives the request text and recent turns, not a local or faster in-process classifier.How is it deployed and self-hosted?
answeredAgenvoy is deployed as a single Go binary (make build produces /usr/local/bin/agen). It has zero infrastructure dependencies: no Docker, PostgreSQL, Redis, or Celery. The binary embeds the web UI assets (page/page.go → page/public/) and all configuration JSON files and prompt text files via //go:embed directives. Storage uses SQLite (via go-sqlkit) for structured data and ToriiDB (an embedded vector DB) for vector search, caching, and session history — all stored under ~/.config/agenvoy/. The daemon listens on 127.0.0.1:17989. Installation is a one-liner (curl -fsSL https://agenvoy.com/scripts/install.sh | bash). Updates use agen update which downloads and runs a shell script. Startup integration supports Linux systemd user services (internal/startup/linux.service with ExecStart="/usr/local/bin/agen" --daemon) and macOS launchd (internal/startup/darwin.plist). The TUI (cmd/app/tui.go) and web dashboard are two UIs for the same daemon. Required external accounts depend on the LLM providers used: users need API keys for their chosen providers (OpenAI, Anthropic, etc.) or OAuth tokens (Copilot, Codex, Grok). For Telegram/Discord chatbots, a bot token is needed. The make dev command serves the web UI from disk for development. Sandboxing requires bubblewrap to be installed on the host. There is a built-in config watcher (app/watchConfig.go) and session watcher (app/watchSession.go). The daemon lifecycle includes system cron jobs for summary generation (every 15 min) and session cleaning (every 30 min).
index.html and public/ are embedded; on the first boot of each version webapp.SyncAsset downloads the dashboard's vendor files from the URLs in vendor.json, so first start needs network access.