ShenSeanChen/waku-agent
Readable Python assistant harness: one Anthropic-shaped tool loop, gated SQLite/FTS5 memory, and built-in traces and evals.
Overview
Waku is a personal assistant you run on your own laptop, and it is written to be read. The core is about 19k lines of Python under waku/. The agent loop itself is one 209-line file. Each module opens with a docstring that explains the design choice (“THE LOOP — observe → reason → act → repeat. This file is the whole trick.”). Everything Waku keeps lives in ~/.waku/: one SQLite database (state.db), a SOUL.md persona file, skills, an outbox, a local calendar and JSONL traces.
It works in four areas. A synchronous tool-use loop. Memory with three parts (semantic facts, episodic summaries, procedural SKILL.md files), plus two small-model helpers: a retrieval gate and a batched consolidator. Providers behind two wire formats. Built-in LLM ops: traces, deterministic and judge evals, a release gate and per-turn cost receipts. You reach it from a terminal chat, a local browser dashboard on port 7777, voice, or Telegram, Discord and WhatsApp gateways.
Waku suits a single user who wants to understand and modify their assistant, and developers who want a small working reference for loop, memory and evals. It is not a multi-user product. The separate hosted/ directory (Elastic License 2.0, not shipped on PyPI) runs one container per tenant, but the agent inside each container is the same single-user Waku.
Architecture
flowchart LR
GW["Gateways: CLI, dashboard, voice, chat apps"] --> APP["Waku.respond (app.py)"]
APP --> TRI["Triage graph (optional)"]
TRI --> FULL["_run_full_turn"]
APP --> FULL
FULL --> SES["Session.build_system"]
SES --> GATE["Retrieval gate (small model)"]
GATE --> MEM["Facts + episodes (SQLite FTS5)"]
SES --> SK["Skill matcher"]
FULL --> LOOP["run_loop (agent.py)"]
LOOP --> LLM["Client: Anthropic or OpenAI-compat"]
LOOP --> REG["ToolRegistry"]
REG --> MCP["MCP bridge"]
APP --> CONS["Consolidation every N turns"]
APP --> TR["Tracer: JSONL + OTel"]
| Component | Path | Role |
|---|---|---|
| Wiring | waku/app.py |
Waku builds config, db, client, memory, tools, session and tracer; respond() runs one turn |
| Loop | waku/loop/agent.py |
run_loop: call the model, run tools, append results, repeat; two exit guardrails |
| Providers | waku/loop/models.py, waku/providers.toml |
Provider rows as data; get_client returns an Anthropic client or OpenAICompatClient |
| Working memory | waku/runtime/session.py |
System prompt from SOUL.md, clock, gated memory and matched skills; sliding history window |
| Long-term memory | waku/memory/ |
Memory facade, retrieval gate, consolidation, pluggable fact and episode stores, skill loader |
| Storage | waku/db.py |
state.db schema: facts, episodes (each with FTS5), chat_log, calendar_events |
| Tools | waku/tools/ |
Registry plus calendar, notes, outbox messages, search, memory admin, Apple, GitHub, MCP |
| Graph | waku/graph/ |
Wave-based node/edge engine; triage and gather workflows |
| Gateways | waku/gateway/ |
CLI, voice, Telegram, Discord, WhatsApp; runner.py serialises async gateways onto one thread |
| Ops | waku/ops/ |
Dashboard, tracing, receipts, pricing, turn evals, judge, arenas, release gate |
| Evals | evals/ |
deterministic/ (pytest, scripted model), judge/ (LLM-graded), hosted_docker/ |
How a request flows
Take “move my 3pm with Alex to tomorrow” typed in the dashboard.
- Entry. The dashboard’s chat route calls
agent.respond(message, observer=..., source="dashboard", stream=True). Telegram and Discord go throughGatewayRunner. It runsrespondon a single-worker executor, so turns are serialised and the SQLite connection stays on one thread (waku/gateway/runner.py). - Turn wrapper.
Waku.respondopens a trace turn. IfWAKU_GRAPH_WORKFLOWSis on, it first tries the triage graph, and any exception falls back to the plain path (waku/app.py). - Working memory.
_run_full_turncallsSession.build_system(waku/app.py). That function concatenatesSOUL.md, the local time with timezone, the active model, a lookup rule and the list of unavailable tools. Then it asks the memory facade for gated retrieval and matching skills (waku/runtime/session.py). Only the lasthistory_turns * 2rows (default 24) of chat history are sent. - Retrieval gate. A small-model call returns
{"retrieve", "query", "reason"}JSON. On any error or missing JSON it fails open to retrieval (waku/memory/retrieval_gate.py). If retrieval goes ahead, the top-k facts (default 4) and up to 3 episodes are searched (waku/memory/init.py). - Loop.
run_loopsendssystem,messagesandtools.schemas()toclient.messages.create(or.stream). It appends the assistant content, runs eachtool_usethroughToolRegistry.executeand appends thetool_resultblocks (waku/loop/agent.py). Tool exceptions come back to the model asError running ...strings (waku/tools/registry.py). - Exit. If the model asks for no tools, its text is the reply. After
max_iterations(default 15), one final call withtool_choice: noneand a “step limit” note asks for an answer from what was gathered (waku/loop/agent.py). - Persist.
respondstores the exchange with a compact[tools used: ...]note, so the model does not repeat actions next turn. It callsmaybe_consolidate, re-exportsMEMORY.md, then builds a cost receipt (waku/app.py).
Key components
The loop and its client contract
The loop speaks only the Anthropic Messages shape: content blocks, tool_use, tool_result. Every provider is reduced to that shape. Rows in providers.toml declare kind = "anthropic" or "openai", the key variable, default and small models, and a price estimate. That includes Anthropic, OpenAI, OpenRouter, Gemini, DeepSeek, MiniMax, Kimi, GLM, xAI, two OpenCode endpoints and the hosted free tier. get_client returns the Anthropic SDK client or OpenAICompatClient, which translates messages and tool calls to chat.completions and back (waku/loop/models.py). WAKU_BASE_URL overrides the endpoint, which is how you point it at a local OpenAI-compatible server. There is no structured-output mode, and tools return plain strings.
Memory
state.db holds facts and episodes. Each table has an FTS5 shadow table that triggers keep in sync (waku/db.py). The facade picks the fact store from WAKU_SEMANTIC_STORE (sqlite, supabase, mem0, zep, langmem) and the episode store from WAKU_EPISODIC_STORE (sqlite, notion) (waku/memory/init.py). Consolidation is batched. Once consolidate_every (default 6) unconsolidated exchanges exist, a small model distils them into facts and one dated episode. Code filters then drop facts about the agent’s own operating state (balances, rate limits, tool errors) and facts that only restate what the turn recalled (waku/memory/consolidation.py). The model can also edit memory directly through manage_memory, update_soul and create_skill.
Skills
Skills are SKILL.md folders from the bundled skills/ directory and ~/.waku/skills. Matching is deliberately simple: keyword overlap between the message and each skill’s name and description. At least two shared words are required, and at most two skills are injected (waku/memory/procedural/loader.py). waku skill install <url> adds community skills, and waku skill export copies them to Claude Code or Codex.
Graph workflows
waku/graph/engine.py runs a node/edge graph in waves. Parallel nodes write disjoint keys, plain-Python routers choose edges (“no LLM ever decides control flow”), and per-node max_visits plus a global max_steps bound execution (waku/graph/engine.py). The triage workflow sends trivial messages to a small-model quick reply and everything else to the full loop as a node. waku gather fetches GitHub, web, calendar and memory in parallel for a digest.
Ops
Every turn writes JSONL to ~/.waku/traces/<date>.jsonl. When OTEL_EXPORTER_OTLP_ENDPOINT is set, it also exports OpenTelemetry spans with GenAI semantic-convention names (waku/ops/tracing.py). python -m waku.ops.release_gate runs evals/deterministic (must pass 100%) and, with a key present, the judge evals, and exits 0 only on a pass (waku/ops/release_gate.py). The dashboard adds turn evals, model arenas and a SQL console.
Extending it
- Tools. Write a
Tool(name, description, input_schema, fn)that returns a string, and register it inbuild_registry(waku/tools/init.py). Long-running tools can setwants_notifyto stream progress. - MCP. Put servers in
~/.waku/mcp.json(needs the[mcp]extra).waku mcp loginhandles OAuth for remote servers. Their tools are registered next to the built-in ones. - Providers. Add a row to
waku/providers.tomlplus a logo. The picker, pricing and.env.exampleare generated from that table. - Memory backends. Implement the
FactStoreprotocol inwaku/memory/semantic/base.py. A conformance suite holds every backend to it. - Persona and skills. Edit
SOUL.md, or drop a folder with aSKILL.mdinto~/.waku/skills.
Running it
pip install waku-agent(the package metadata says Python 3.11+), setANTHROPIC_API_KEYor another provider key plusWAKU_PROVIDER, then runwakufor terminal chat orwaku dashboardfor the browser UI.- The dashboard binds to
127.0.0.1by default. It has no authentication, and its SQL console runs arbitrary SQL, so binding elsewhere prints a warning (waku/ops/dashboard.py). - Optional extras: voice, Telegram, Discord, WhatsApp (needs a public URL), MCP, tracing, and the alternative memory stores. Apple Calendar, Apple apps and GitHub tools are opt-in flags.
- Hosted mode (
hosted/, git checkout only): one VM with Docker Compose, a gateway, a container spawner, a metering proxy for the free tier, and one stock Waku container per tenant.
Strengths and caveats
- Strength: legibility. The wiring fits in one file, the loop in another, and decisions are explained in place. This is one of the easiest assistant codebases to read end to end.
- Strength: cheap, honest memory. The retrieval gate skips memory for small talk, consolidation is batched, and code filters (
is_ops_state,restates) stop the known ways memory gets polluted. Every fact is also a plain file you can inspect. - Strength: evals ship with the product. There are deterministic evals with a scripted model, judge evals, a release gate, per-turn receipts with cost, and OTel export.
- Strength: safe defaults for outbound actions.
send_messageonly writes a draft to the local outbox and never sends (messages.py). Real-calendar writes are opt-in flags. - Caveat: no approval gate. Once a tool is enabled, the model calls it freely. Apple Mail/Notes writers, calendar writes and MCP tools have no per-call confirmation.
- Caveat: simple retrieval. The default store uses FTS5 keyword search, and skill matching is word overlap. Recall depends on the gate writing good keywords.
- Caveat: synchronous, single-user core. One conversation thread per gateway, with blocking calls. Fine on a laptop, but not a server design.
- Caveat: Anthropic-shaped internals. Provider-specific features (structured outputs, OpenAI reasoning tools) are lost in translation through the compatibility client.
Sources: code at cd698a5, verified Q&A.
How it answers the Open-source personal assistants questions
Each answer was drafted by a code-reading agent at commit cd698a5. Its citations were checked mechanically. Compare with the other open-source personal assistants →
How is the assistant architected?
answeredWaku is wired together in waku/app.py (the Waku class). Entry points in waku/main.py dispatch to gateways: CLI, dashboard, Telegram, Discord, WhatsApp, voice. Each gateway feeds text into Waku.respond(), which builds working memory via Session.build_system() (waku/runtime/session.py): a SOUL.md persona file, current datetime, gated memory retrieval, and matching skill instructions. The core loop is run_loop() in waku/loop/agent.py: a synchronous while loop that calls the LLM with tool schemas from waku/tools/registry.py, executes tool calls, feeds results back as new messages, and repeats until the model stops asking for tools or hits max_iterations (default 15). Guardrail 2 produces a final tools-off call. Working memory is a sliding window of the last 24 rows (history_turns * 2). An optional graph layer (waku/graph/engine.py) wraps the loop with deterministic node-and-edge workflows; the triage workflow sends trivial messages through a quick small-model path and everything else through the full loop as a graph node. Frontend is waku/ops/dashboard.py, a stdlib HTTP server serving static files. The loop uses Anthropic Messages wire format natively and adapts OpenAI wire format via OpenAICompatClient in waku/loop/models.py.
How are integrations (email, calendar, chat, docs) implemented?
answeredIntegrations are declared in waku/integrations.py as a tuple of Integration dataclasses grouped into Channels (Telegram, Discord, WhatsApp), Calendar and Productivity (Google Calendar, Apple Calendar, Apple Tools), Memory and Storage (Notion, Mem0, Zep, LangMem, Supabase, TypeSafe Jev), and Search and Observability (Tavily, treg, OpenTelemetry). Each declares env fields and optionally a probe that validates connectivity on save. OAuth runs through waku/connect.py with google, treg, and waku-memory opening a browser window and caching the token. Google uses googleapiclient OAuth; treg and waku-memory use MCP OAuth via waku/tools/mcp_oauth.py (discover auth server, DCR, open browser, catch redirect on loopback port 41765, store via FileTokenStorage as JSON). Tools are used synchronously on-demand, no background sync. Calendar tools create events in Apple Calendar and Google Calendar as opt-in write targets. The search tool calls Tavily on demand.
How is memory and user context stored and retrieved?
answeredMemory is three pillars managed by waku/memory/init.py (Memory class). Semantic memory (facts) and episodic memory (dated events) live in state.db, a single SQLite file with FTS5 search (schema in waku/db.py). Procedural memory is SKILL.md files loaded via SkillLoader. The retrieval gate (waku/memory/retrieval_gate.py) is the hero pattern: a cheap model decides if memory is needed before touching any store. If yes, the semantic store is searched for top-k facts and episodic store for top 3 episodes. Semantic storage can be swapped from SQLite+FTS5 to Supabase pgvector, Mem0, Zep, or LangMem via WAKU_SEMANTIC_STORE. Episodic storage can switch to Notion. Consolidation (waku/memory/consolidation.py) is batched: after every N exchanges (default 6), a cheap model distills unconsolidated chat_log rows into semantic facts and an episodic summary. Operating state (balances, credits, errors) is filtered by is_ops_state(). Memory is injected into the system prompt in Session.build_system(). Read memory passes as recalled so same facts are not stored again. An optional Jev slot gate (waku/memory/slot_gate.py) scores each fact.
How are actions on the user's behalf gated?
answeredWaku does not have a general human-in-the-loop approval gate for tool execution. Safety is layered: ToolRegistry.execute() catches exceptions and returns error strings. The graph engine enforces per-node max_visits and global max_steps. The loop has two guardrails: stop when model does not ask for tools, and a final tools-off call at max_iterations. OAuth integrations require explicit browser sign-in that cannot be automated. The MCP client isolates external tools via subprocess or network boundaries. The retrieval gate fails open to retrieve. The slot gate defaults to keeping every fact when Jev is unavailable. Consolidation is_ops_state() prevents operating state from persisting. Tools are opt-in via Settings flags (apple_tools, apple_calendar, google_calendar, gh_tool, experimental default off). SOUL.md stores the assistant persona rules. There is no dry-run/draft mode or explicit permission-scope system per tool. Audit trail is trace JSONL files and chat_log, not a dedicated actions table.
send_message writes each message to ~/.waku/outbox for the user to review and send, and nothing is sent (waku/tools/messages.py L1-L22). Real calendar writes (Apple, Google) are also behind opt-in flags.How are LLM providers selected and configured?
answeredLLM providers are configured in waku/providers.toml and loaded as Provider dataclasses in waku/loop/models.py. PROVIDERS supports: anthropic, openai, gemini, deepseek, minimax, kimi, glm, openrouter, opencode_zen, opencode_go, and waku-platform (hosted free tier). Each has kind (anthropic or openai), key_env, default model ids, and optional endpoints. Anthropic-wire providers use the Anthropic SDK directly. OpenAI-wire providers go through OpenAICompatClient (waku/loop/models.py lines 393-513) which translates formats. User selects provider via WAKU_PROVIDER env var. Tool calling uses tool_use pattern: schemas from ToolRegistry.schemas() passed as tools= parameter. Structured output is not used; tools return strings. small_model is used for retrieval gate, consolidation, and triage. Model ids are resolved through models_for() with env overrides and cross-provider leakage prevention. The hosted free tier uses TurnTagged client with X-Waku-Turn headers for metering. Timeout defaults to 120 seconds. Local model support works through OpenAI-compatible endpoints.
How is it deployed and self-hosted?
answeredWaku has two deployment modes. Local: pip install waku-agent, then waku for CLI or waku dashboard for web UI on localhost:7777. Only dependency is Python 3.12+ with optional extras for Telegram, Discord, voice, MCP, Notion, Mem0, Zep, LangMem, Supabase, OpenTelemetry. SQLite is the sole storage engine. Hosted (hosted/README.md): single Ubuntu 24.04 VM with Docker Compose, XFS data disk (project quotas), Supabase for auth. Each tenant gets a Docker container with their own state and memory. The gateway handles login, session routing, tenant provisioning. The metering proxy sits in front of Anthropic for the free tier. The spawner manages Docker containers. Caddy handles TLS via DNS-01 wildcard certificates. Required external accounts: Anthropic API key, optionally Supabase (hosted auth), treg.to, Tavily, S3 storage (backups). The hosted code ships only from git checkout (not PyPI) under Elastic License 2.0. One-click deploy via install.sh on a prepared Ubuntu VM.