LLMs Technical Reviews

ShenSeanChen/waku-agent

Readable Python assistant harness: one Anthropic-shaped tool loop, gated SQLite/FTS5 memory, and built-in traces and evals.

GitHub ↗★ 1.9kPythonMITcommit cd698a5 · 2026-10-05homepage ↗

Overview

Waku is a personal assistant you run on your own laptop, and it is written to be read. The core is about 19k lines of Python under waku/. The agent loop itself is one 209-line file. Each module opens with a docstring that explains the design choice (“THE LOOP — observe → reason → act → repeat. This file is the whole trick.”). Everything Waku keeps lives in ~/.waku/: one SQLite database (state.db), a SOUL.md persona file, skills, an outbox, a local calendar and JSONL traces.

It works in four areas. A synchronous tool-use loop. Memory with three parts (semantic facts, episodic summaries, procedural SKILL.md files), plus two small-model helpers: a retrieval gate and a batched consolidator. Providers behind two wire formats. Built-in LLM ops: traces, deterministic and judge evals, a release gate and per-turn cost receipts. You reach it from a terminal chat, a local browser dashboard on port 7777, voice, or Telegram, Discord and WhatsApp gateways.

Waku suits a single user who wants to understand and modify their assistant, and developers who want a small working reference for loop, memory and evals. It is not a multi-user product. The separate hosted/ directory (Elastic License 2.0, not shipped on PyPI) runs one container per tenant, but the agent inside each container is the same single-user Waku.

Architecture

flowchart LR
  GW["Gateways: CLI, dashboard, voice, chat apps"] --> APP["Waku.respond (app.py)"]
  APP --> TRI["Triage graph (optional)"]
  TRI --> FULL["_run_full_turn"]
  APP --> FULL
  FULL --> SES["Session.build_system"]
  SES --> GATE["Retrieval gate (small model)"]
  GATE --> MEM["Facts + episodes (SQLite FTS5)"]
  SES --> SK["Skill matcher"]
  FULL --> LOOP["run_loop (agent.py)"]
  LOOP --> LLM["Client: Anthropic or OpenAI-compat"]
  LOOP --> REG["ToolRegistry"]
  REG --> MCP["MCP bridge"]
  APP --> CONS["Consolidation every N turns"]
  APP --> TR["Tracer: JSONL + OTel"]
Component Path Role
Wiring waku/app.py Waku builds config, db, client, memory, tools, session and tracer; respond() runs one turn
Loop waku/loop/agent.py run_loop: call the model, run tools, append results, repeat; two exit guardrails
Providers waku/loop/models.py, waku/providers.toml Provider rows as data; get_client returns an Anthropic client or OpenAICompatClient
Working memory waku/runtime/session.py System prompt from SOUL.md, clock, gated memory and matched skills; sliding history window
Long-term memory waku/memory/ Memory facade, retrieval gate, consolidation, pluggable fact and episode stores, skill loader
Storage waku/db.py state.db schema: facts, episodes (each with FTS5), chat_log, calendar_events
Tools waku/tools/ Registry plus calendar, notes, outbox messages, search, memory admin, Apple, GitHub, MCP
Graph waku/graph/ Wave-based node/edge engine; triage and gather workflows
Gateways waku/gateway/ CLI, voice, Telegram, Discord, WhatsApp; runner.py serialises async gateways onto one thread
Ops waku/ops/ Dashboard, tracing, receipts, pricing, turn evals, judge, arenas, release gate
Evals evals/ deterministic/ (pytest, scripted model), judge/ (LLM-graded), hosted_docker/

How a request flows

Take “move my 3pm with Alex to tomorrow” typed in the dashboard.

  1. Entry. The dashboard’s chat route calls agent.respond(message, observer=..., source="dashboard", stream=True). Telegram and Discord go through GatewayRunner. It runs respond on a single-worker executor, so turns are serialised and the SQLite connection stays on one thread (waku/gateway/runner.py).
  2. Turn wrapper. Waku.respond opens a trace turn. If WAKU_GRAPH_WORKFLOWS is on, it first tries the triage graph, and any exception falls back to the plain path (waku/app.py).
  3. Working memory. _run_full_turn calls Session.build_system (waku/app.py). That function concatenates SOUL.md, the local time with timezone, the active model, a lookup rule and the list of unavailable tools. Then it asks the memory facade for gated retrieval and matching skills (waku/runtime/session.py). Only the last history_turns * 2 rows (default 24) of chat history are sent.
  4. Retrieval gate. A small-model call returns {"retrieve", "query", "reason"} JSON. On any error or missing JSON it fails open to retrieval (waku/memory/retrieval_gate.py). If retrieval goes ahead, the top-k facts (default 4) and up to 3 episodes are searched (waku/memory/init.py).
  5. Loop. run_loop sends system, messages and tools.schemas() to client.messages.create (or .stream). It appends the assistant content, runs each tool_use through ToolRegistry.execute and appends the tool_result blocks (waku/loop/agent.py). Tool exceptions come back to the model as Error running ... strings (waku/tools/registry.py).
  6. Exit. If the model asks for no tools, its text is the reply. After max_iterations (default 15), one final call with tool_choice: none and a “step limit” note asks for an answer from what was gathered (waku/loop/agent.py).
  7. Persist. respond stores the exchange with a compact [tools used: ...] note, so the model does not repeat actions next turn. It calls maybe_consolidate, re-exports MEMORY.md, then builds a cost receipt (waku/app.py).

Key components

The loop and its client contract

The loop speaks only the Anthropic Messages shape: content blocks, tool_use, tool_result. Every provider is reduced to that shape. Rows in providers.toml declare kind = "anthropic" or "openai", the key variable, default and small models, and a price estimate. That includes Anthropic, OpenAI, OpenRouter, Gemini, DeepSeek, MiniMax, Kimi, GLM, xAI, two OpenCode endpoints and the hosted free tier. get_client returns the Anthropic SDK client or OpenAICompatClient, which translates messages and tool calls to chat.completions and back (waku/loop/models.py). WAKU_BASE_URL overrides the endpoint, which is how you point it at a local OpenAI-compatible server. There is no structured-output mode, and tools return plain strings.

Memory

state.db holds facts and episodes. Each table has an FTS5 shadow table that triggers keep in sync (waku/db.py). The facade picks the fact store from WAKU_SEMANTIC_STORE (sqlite, supabase, mem0, zep, langmem) and the episode store from WAKU_EPISODIC_STORE (sqlite, notion) (waku/memory/init.py). Consolidation is batched. Once consolidate_every (default 6) unconsolidated exchanges exist, a small model distils them into facts and one dated episode. Code filters then drop facts about the agent’s own operating state (balances, rate limits, tool errors) and facts that only restate what the turn recalled (waku/memory/consolidation.py). The model can also edit memory directly through manage_memory, update_soul and create_skill.

Skills

Skills are SKILL.md folders from the bundled skills/ directory and ~/.waku/skills. Matching is deliberately simple: keyword overlap between the message and each skill’s name and description. At least two shared words are required, and at most two skills are injected (waku/memory/procedural/loader.py). waku skill install <url> adds community skills, and waku skill export copies them to Claude Code or Codex.

Graph workflows

waku/graph/engine.py runs a node/edge graph in waves. Parallel nodes write disjoint keys, plain-Python routers choose edges (“no LLM ever decides control flow”), and per-node max_visits plus a global max_steps bound execution (waku/graph/engine.py). The triage workflow sends trivial messages to a small-model quick reply and everything else to the full loop as a node. waku gather fetches GitHub, web, calendar and memory in parallel for a digest.

Ops

Every turn writes JSONL to ~/.waku/traces/<date>.jsonl. When OTEL_EXPORTER_OTLP_ENDPOINT is set, it also exports OpenTelemetry spans with GenAI semantic-convention names (waku/ops/tracing.py). python -m waku.ops.release_gate runs evals/deterministic (must pass 100%) and, with a key present, the judge evals, and exits 0 only on a pass (waku/ops/release_gate.py). The dashboard adds turn evals, model arenas and a SQL console.

Extending it

  • Tools. Write a Tool(name, description, input_schema, fn) that returns a string, and register it in build_registry (waku/tools/init.py). Long-running tools can set wants_notify to stream progress.
  • MCP. Put servers in ~/.waku/mcp.json (needs the [mcp] extra). waku mcp login handles OAuth for remote servers. Their tools are registered next to the built-in ones.
  • Providers. Add a row to waku/providers.toml plus a logo. The picker, pricing and .env.example are generated from that table.
  • Memory backends. Implement the FactStore protocol in waku/memory/semantic/base.py. A conformance suite holds every backend to it.
  • Persona and skills. Edit SOUL.md, or drop a folder with a SKILL.md into ~/.waku/skills.

Running it

  • pip install waku-agent (the package metadata says Python 3.11+), set ANTHROPIC_API_KEY or another provider key plus WAKU_PROVIDER, then run waku for terminal chat or waku dashboard for the browser UI.
  • The dashboard binds to 127.0.0.1 by default. It has no authentication, and its SQL console runs arbitrary SQL, so binding elsewhere prints a warning (waku/ops/dashboard.py).
  • Optional extras: voice, Telegram, Discord, WhatsApp (needs a public URL), MCP, tracing, and the alternative memory stores. Apple Calendar, Apple apps and GitHub tools are opt-in flags.
  • Hosted mode (hosted/, git checkout only): one VM with Docker Compose, a gateway, a container spawner, a metering proxy for the free tier, and one stock Waku container per tenant.

Strengths and caveats

  • Strength: legibility. The wiring fits in one file, the loop in another, and decisions are explained in place. This is one of the easiest assistant codebases to read end to end.
  • Strength: cheap, honest memory. The retrieval gate skips memory for small talk, consolidation is batched, and code filters (is_ops_state, restates) stop the known ways memory gets polluted. Every fact is also a plain file you can inspect.
  • Strength: evals ship with the product. There are deterministic evals with a scripted model, judge evals, a release gate, per-turn receipts with cost, and OTel export.
  • Strength: safe defaults for outbound actions. send_message only writes a draft to the local outbox and never sends (messages.py). Real-calendar writes are opt-in flags.
  • Caveat: no approval gate. Once a tool is enabled, the model calls it freely. Apple Mail/Notes writers, calendar writes and MCP tools have no per-call confirmation.
  • Caveat: simple retrieval. The default store uses FTS5 keyword search, and skill matching is word overlap. Recall depends on the gate writing good keywords.
  • Caveat: synchronous, single-user core. One conversation thread per gateway, with blocking calls. Fine on a laptop, but not a server design.
  • Caveat: Anthropic-shaped internals. Provider-specific features (structured outputs, OpenAI reasoning tools) are lost in translation through the compatibility client.

Sources: code at cd698a5, verified Q&A.

How it answers the Open-source personal assistants questions

Each answer was drafted by a code-reading agent at commit cd698a5. Its citations were checked mechanically. Compare with the other open-source personal assistants →

How is the assistant architected?

answered

Waku is wired together in waku/app.py (the Waku class). Entry points in waku/main.py dispatch to gateways: CLI, dashboard, Telegram, Discord, WhatsApp, voice. Each gateway feeds text into Waku.respond(), which builds working memory via Session.build_system() (waku/runtime/session.py): a SOUL.md persona file, current datetime, gated memory retrieval, and matching skill instructions. The core loop is run_loop() in waku/loop/agent.py: a synchronous while loop that calls the LLM with tool schemas from waku/tools/registry.py, executes tool calls, feeds results back as new messages, and repeats until the model stops asking for tools or hits max_iterations (default 15). Guardrail 2 produces a final tools-off call. Working memory is a sliding window of the last 24 rows (history_turns * 2). An optional graph layer (waku/graph/engine.py) wraps the loop with deterministic node-and-edge workflows; the triage workflow sends trivial messages through a quick small-model path and everything else through the full loop as a graph node. Frontend is waku/ops/dashboard.py, a stdlib HTTP server serving static files. The loop uses Anthropic Messages wire format natively and adapts OpenAI wire format via OpenAICompatClient in waku/loop/models.py.

How are integrations (email, calendar, chat, docs) implemented?

answered

Integrations are declared in waku/integrations.py as a tuple of Integration dataclasses grouped into Channels (Telegram, Discord, WhatsApp), Calendar and Productivity (Google Calendar, Apple Calendar, Apple Tools), Memory and Storage (Notion, Mem0, Zep, LangMem, Supabase, TypeSafe Jev), and Search and Observability (Tavily, treg, OpenTelemetry). Each declares env fields and optionally a probe that validates connectivity on save. OAuth runs through waku/connect.py with google, treg, and waku-memory opening a browser window and caching the token. Google uses googleapiclient OAuth; treg and waku-memory use MCP OAuth via waku/tools/mcp_oauth.py (discover auth server, DCR, open browser, catch redirect on loopback port 41765, store via FileTokenStorage as JSON). Tools are used synchronously on-demand, no background sync. Calendar tools create events in Apple Calendar and Google Calendar as opt-in write targets. The search tool calls Tavily on demand.

How is memory and user context stored and retrieved?

answered

Memory is three pillars managed by waku/memory/init.py (Memory class). Semantic memory (facts) and episodic memory (dated events) live in state.db, a single SQLite file with FTS5 search (schema in waku/db.py). Procedural memory is SKILL.md files loaded via SkillLoader. The retrieval gate (waku/memory/retrieval_gate.py) is the hero pattern: a cheap model decides if memory is needed before touching any store. If yes, the semantic store is searched for top-k facts and episodic store for top 3 episodes. Semantic storage can be swapped from SQLite+FTS5 to Supabase pgvector, Mem0, Zep, or LangMem via WAKU_SEMANTIC_STORE. Episodic storage can switch to Notion. Consolidation (waku/memory/consolidation.py) is batched: after every N exchanges (default 6), a cheap model distills unconsolidated chat_log rows into semantic facts and an episodic summary. Operating state (balances, credits, errors) is filtered by is_ops_state(). Memory is injected into the system prompt in Session.build_system(). Read memory passes as recalled so same facts are not stored again. An optional Jev slot gate (waku/memory/slot_gate.py) scores each fact.

How are actions on the user's behalf gated?

answered

Waku does not have a general human-in-the-loop approval gate for tool execution. Safety is layered: ToolRegistry.execute() catches exceptions and returns error strings. The graph engine enforces per-node max_visits and global max_steps. The loop has two guardrails: stop when model does not ask for tools, and a final tools-off call at max_iterations. OAuth integrations require explicit browser sign-in that cannot be automated. The MCP client isolates external tools via subprocess or network boundaries. The retrieval gate fails open to retrieve. The slot gate defaults to keeping every fact when Jev is unavailable. Consolidation is_ops_state() prevents operating state from persisting. Tools are opt-in via Settings flags (apple_tools, apple_calendar, google_calendar, gh_tool, experimental default off). SOUL.md stores the assistant persona rules. There is no dry-run/draft mode or explicit permission-scope system per tool. Audit trail is trace JSONL files and chat_log, not a dedicated actions table.

Editor's note. Correction: outbound messaging is draft-only. send_message writes each message to ~/.waku/outbox for the user to review and send, and nothing is sent (waku/tools/messages.py L1-L22). Real calendar writes (Apple, Google) are also behind opt-in flags.

How are LLM providers selected and configured?

answered

LLM providers are configured in waku/providers.toml and loaded as Provider dataclasses in waku/loop/models.py. PROVIDERS supports: anthropic, openai, gemini, deepseek, minimax, kimi, glm, openrouter, opencode_zen, opencode_go, and waku-platform (hosted free tier). Each has kind (anthropic or openai), key_env, default model ids, and optional endpoints. Anthropic-wire providers use the Anthropic SDK directly. OpenAI-wire providers go through OpenAICompatClient (waku/loop/models.py lines 393-513) which translates formats. User selects provider via WAKU_PROVIDER env var. Tool calling uses tool_use pattern: schemas from ToolRegistry.schemas() passed as tools= parameter. Structured output is not used; tools return strings. small_model is used for retrieval gate, consolidation, and triage. Model ids are resolved through models_for() with env overrides and cross-provider leakage prevention. The hosted free tier uses TurnTagged client with X-Waku-Turn headers for metering. Timeout defaults to 120 seconds. Local model support works through OpenAI-compatible endpoints.

How is it deployed and self-hosted?

answered

Waku has two deployment modes. Local: pip install waku-agent, then waku for CLI or waku dashboard for web UI on localhost:7777. Only dependency is Python 3.12+ with optional extras for Telegram, Discord, voice, MCP, Notion, Mem0, Zep, LangMem, Supabase, OpenTelemetry. SQLite is the sole storage engine. Hosted (hosted/README.md): single Ubuntu 24.04 VM with Docker Compose, XFS data disk (project quotas), Supabase for auth. Each tenant gets a Docker container with their own state and memory. The gateway handles login, session routing, tenant provisioning. The metering proxy sits in front of Anthropic for the free tier. The spawner manages Docker containers. Caddy handles TLS via DNS-01 wildcard certificates. Required external accounts: Anthropic API key, optionally Supabase (hosted auth), treg.to, Tavily, S3 storage (backups). The hosted code ships only from git checkout (not PyPI) under Elastic License 2.0. One-click deploy via install.sh on a prepared Ubuntu VM.