leon-ai/leon
Self-hosted Node.js assistant: a tool-calling agent loop over JSON-declared toolkits, layered SQLite memory and optional remote-device tools.
Overview
Leon is a self-hosted personal assistant server written in TypeScript on Node.js 24. Clients (the bundled Vite app, the newer React web app, or your own) connect over Socket.IO and send “utterances”. Leon routes each one either to a tool-calling agent loop or to a deterministic skill. Leon has existed for years as an intent-classification assistant. At this commit it has been rebuilt around an LLM agent, and the old skill system survives as the “controlled” route.
Three routing modes exist in config.yml: agent (the default), controlled and smart. In agent mode every request goes to one long-running provider tool-calling loop. Tools come from a catalog of toolkits declared in JSON. They cover the shell, files, ripgrep, browser control, desktop control through Cua, Gmail, Notion, TickTick, media generation, speech, and more. Their schemas are loaded progressively, so the model sees only what it needs. Around the loop sit a layered SQLite memory, generated “context files” that describe the host machine, a private self-model, and a “pulse” that can start autonomous tasks on a timer.
It is aimed at one owner, or a few profiles, running Leon on their own computer. It assumes broad local access, because shell and desktop control are first-class tools. Thirteen LLM providers are supported, including local llama.cpp and SGLang servers.
Architecture
flowchart LR
CL["Clients (app, web-app, custom)"] --> SO["Socket.IO server"]
HT["Fastify HTTP API"] --> NLU
SO --> NLU["NLU: routing decision"]
NLU -->|"agent"| RE["ReActLLMDuty"]
NLU -->|"controlled"| SK["Skill router + action calling"]
RE --> LOOP["runAgentLoop"]
LOOP --> PROV["LLM provider (13 backends)"]
LOOP --> REG["ToolkitRegistry (progressive schemas)"]
LOOP --> EXE["ToolExecutor"]
EXE --> WK["Per-tool worker processes"]
EXE --> SAT["Satellite on another device"]
SK --> BR["Node.js / Python bridges"]
LOOP --> MEM["MemoryManager (SQLite + QMD)"]
PULSE["PulseManager"] --> RE
| Component | Path | Role |
|---|---|---|
| Boot | server/src/index.ts |
Starts the HTTP and Socket.IO servers, the optional Python TCP server (ASR/TTS), pulse and context managers |
| Socket server | server/src/core/socket-server.ts |
Client, satellite and widget events; utterance handling |
| NLU / router | server/src/core/nlp/nlu/nlu.ts |
Chooses the agent or controlled route, runs skill selection and slot filling |
| Agent duty | server/src/core/llm-manager/llm-duties/react-llm-duty.ts |
Builds the prompt, tool catalog and history, then calls the loop |
| Agent loop | server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts |
Iterates model calls, tool batches, recovery and completion review |
| Providers | server/src/core/llm-manager/llm-providers/ |
OpenAI, Anthropic, OpenRouter, Groq, DeepSeek, llama.cpp, SGLang and others |
| Tool manager | server/src/core/tool-manager/ |
Toolkit registry, executor and worker process pool |
| Toolkits | tools/<domain>/<tool>/tool.json |
38 tool declarations with function schemas, guidance and connection needs |
| Skills | skills/native/, skills/agent/ |
Deterministic actions (with locales) and SKILL.md guides for the agent |
| Memory | server/src/core/memory-manager/ |
SQLite schema, QMD retrieval, daily summaries |
| Connections | server/src/core/connections/ |
OAuth and API-key setup, encrypted per profile |
How a request flows
Take “find last week’s invoices in my Downloads folder and total them”, typed into the web app:
- Receive. The socket server binds
socket.on('utterance', ...)tohandleOwnerMessagefor that client (socket-server.ts). The profile credential on the connection decides which profile’s config, memory and tools are used. - Route.
NLU.processcallsgetRoutingDecision. Attachments force the agent route. If the route’s LLM target is not configured, Leon replies with a hint to run/model <provider> <model>(nlu.ts). In agent mode it callsrunReAct, which builds aReActLLMDutywith any forced agent skill or tool and executes it (nlu.ts). - Prepare. The duty assembles the system prompt (stable rules first, then volatile runtime state), memory recall, one-line context-file summaries and a small starting tool catalog. The catalog includes a toolkit loader, an agent-skill loader and
request_clarification. - Loop.
runAgentLoopruns up toagent_max_iterations(default 256) turns. Each turn calls the model with the transcript and current tools. It retries once on truncated or empty output, and once with a smaller context if the provider reports context pressure (agent-loop.ts). At most eight tool calls per turn are executed; the rest are deferred. - Execute tools. Here the model would load
operating_system_control, then callripgrepandfile. Each call is validated, checked for duplicate inputs, and run in a separate Node worker process spawned per toolkit and tool (tool-worker-manager.ts). Results, including errors, return as observations. Large outputs are written to artifact logs, and only a preview goes back to the model. - Review. When the model answers without tool calls after doing work,
reviewAgentCompletionasks the same provider for acomplete/continue/blockedverdict.continuesends the loop back to work. - Answer and remember. The final text streams back over Socket.IO.
MemoryManager.observeTurnstores the exchange as a daily event and a discussion note with a five-day TTL, refreshes the day summary, and prunes old discussion items (memory-manager.ts).
Key components
The agent loop
The loop keeps one transcript for the whole task. Checkpoints and compaction keep it within a fixed input budget: old tool exchanges and inactive toolkit schemas are dropped first. When the budget runs out, a final “finishing pass” without tools has to answer from the evidence already collected. The system prompt is long and specific. It ranks tools from dedicated APIs, through browser and OS tools, down to bounded shell commands. It has a <safety> block that forbids guessing values before side effects and forbids using automation to bypass CAPTCHAs or anti-bot controls (agent-loop.ts). The completion reviewer is told to treat tool output as evidence, not instructions. That is a light prompt-injection guard.
Toolkits and connections
A toolkit is a domain folder with a toolkit.json. Each tool has a tool.json that owns its function schemas, guidance, binaries and connection requirements, plus a Node.js or Python implementation built on the bridge SDK. “Installed”, “enabled” and “available” are separate states. A tool is available only when its required settings exist, and availability.tools.allowed / disabled in config narrow the set (config.sample.yml). When a call fails for lack of credentials, the loop adds a setup_connection tool limited to the blocked providers. OAuth uses PKCE. Secrets are stored with AES-256-GCM under a per-profile key.
Providers
LLM_PROVIDERS_MAP lists llama.cpp, SGLang, Groq, OpenRouter, Z.AI, DeepSeek, MiniMax, OpenAI, Anthropic, Moonshot, Cerebras, Hugging Face and Celeris (llm-provider.ts). llama.cpp and SGLang are the local providers (llm-routing.ts). You set llm.default as <provider>/<model>, optionally with separate workflow and agent targets. API keys live in the profile .env (config.sample.yml). Setup can also discover credentials already stored by Codex, Claude Code, OpenCode, Pi, Hermes, OpenClaw and T3 Code, and offer to reuse them.
Memory
The SQLite schema (WAL mode) stores memory_items with a scope (persistent, daily, discussion), a kind, importance, confidence, an expiry and a pin flag (schema.sql). Items are chunked into memory_chunks with an FTS5 index and optional embeddings (schema.sql). Recall goes through the QMD library, which does hybrid lexical and vector search. Results are then weighted by namespace: persistent memory 1.35, discussion 0.65 (memory-manager.ts). Weak first results trigger adaptive second passes. Separate from memory, an OWNER.md profile and machine-generated context files (activity, host system, network, inventory and others) ground questions about the environment.
Pulse and self-model
With runtime.pulse_enabled (on in the sample config), PulseManager ticks every 30 minutes. It can generate an autonomous “matter” from memory, context changes and the self-model, run at most one per tick, and back off after the owner declines similar matters. Cooldowns and suppression policies are stored per matter fingerprint (pulse-manager.ts). The private diary/self-model distills repeated habits into principles and injects a compact snapshot into the first agent request.
Extending it
- A tool. Add
tools/<domain>/<tool>/tool.jsonand an implementation insrc/nodejs/(extending the SDKTool) orsrc/python/(extendingBaseTool). Declare dependencies locally. Setup installs them with managed Node, Python, pnpm and uv. - An agent skill. Drop a
skills/agent/<name>/SKILL.mdwith discovery frontmatter. The agent loads it on demand and follows it inside the same loop. Coding, document authoring, live tutorial and media production ship this way. - A native skill. Use
skills/native/<skill>/withskill.json,locales/and action entry points, for deterministic flows in controlled mode. - A client. Speak the Socket.IO contract with a
<profile>:<token>credential. HTTP plugins can extend the API without patching core. - Remote devices. Run Leon Satellite on a laptop and map tools to it in config (for example
computer_use.cua: my-laptop). Those calls then execute on that device.
Running it
- Bare metal only. There is no Dockerfile. Install Node 24 (pinned through Volta) and pnpm, then run
pnpm install. Thepostinstallhook runs the setup pipeline (Python env, llama.cpp build tools, tool dependencies, profile prompts). Then runpnpm build && pnpm start, orpnpm dev:server(package.json). - Profiles. Each profile has its own config,
.env, encrypted connections, memory and logs, and one server can serve several profiles at once. - Model. Configure one with
/model <provider> <model>in the chat, or withllm.default. Voice (ASR, TTS, wake word) starts an extra Python TCP server.
Strengths and caveats
- Strength: serious agent runtime. Progressive tool loading, bounded recovery, completion review, artifact-backed large outputs and resumable checkpoints are more than most personal-assistant projects ship.
- Strength: tool hygiene. Tools declare schemas, settings and connections in one manifest, run in isolated worker processes, and can be pinned to a remote device.
- Strength: local-first. Everything is stored in per-profile SQLite and files, and local llama.cpp/SGLang are supported.
- Caveat: no action approval gate. Shell, file and desktop tools execute when the model calls them. Safety is a system prompt, allow/deny lists and connection prerequisites. There is no per-action confirmation.
- Caveat: autonomous by default. The pulse and the self-model are enabled in the sample config. Leon may act on its own every 30 minutes unless you turn that off.
- Caveat: big moving target. The agent loop alone is over 2,000 lines and the setup pipeline compiles native pieces. Expect setup friction and fast churn.
- Caveat: two paradigms. Native skills, NLU and slot filling remain beside the agent loop, so there are two ways to build features.
Sources: code at 2471443, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the Open-source personal assistants questions
Each answer was drafted by a code-reading agent at commit 2471443. Its citations were checked mechanically. Compare with the other open-source personal assistants →
How is the assistant architected?
answeredLeon's architecture centers on three runtime modes — agent (default), controlled, and smart — selected in config.sample.yml via routing.mode (server/src/../config.ts). In agent mode, the core loop lives in server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts (lines 1–54): a continuous, single-transcript provider tool-calling loop where the LLM calls tools, observations are fed back, and recovery/completion-review check progress. The loop is orchestrated by react-llm-duty.ts which imports from agent-constants.ts, tool-execution.ts, agent-plan.ts, and agent-history-manager.ts. The Brain (server/src/core/brain/brain.ts, lines 59–77) manages the answer queue, paraphrasing, TTS, and skill execution via handlers (dialog-action-skill-handler.ts, logic-action-skill-handler.ts). Tool schemas load progressively from the ToolkitRegistry (server/src/core/tool-manager/toolkit-registry.ts), and the ToolExecutor (server/src/core/tool-manager/tool-executor.ts) runs calls concurrently in isolated workers. The frontend/backend split uses client-agnostic Socket.IO (server/src/core/socket-server.ts) — text (and attached files) comes in as utterances, and the response streams back as tokens and final answers. Legacy [controlled mode] routes through NLU (natural language understanding) which classifies intent and triggers native skill actions (Skills -> Actions -> Tools -> Functions). Packages are organized as a pnpm workspace: server/ (core + Node.js bridge), app/ (terminal UI), web-app/ (React web UI), bridges/ (Node.js & Python bridges), skills/ (native + agent skills), tools/ (toolkits with JSON declarations). A user request flows: client emits an utterance via Socket.IO → socket-server.ts dispatches to the NLU or directly to the agent loop → based on routing mode, the LLM chooses/calls tools → results return as observations → final answer is queued through Brain.talk() and emitted back to the client.
How are integrations (email, calendar, chat, docs) implemented?
answeredLeon integrates external services via the Toolkit + Connection system rather than MCP. Services are declared in JSON toolkit manifests under /tools/ — organized by domain (e.g. communication/gmail, productivity_collaboration/notion, search_web/hosted, calendar_scheduling). Each tool has a tool.json (e.g. tools/communication/gmail/tool.json, lines 1–136) describing its functions, parameters, and connection requirements. Connections support two auth methods: OAuth and API key, both managed through server/src/core/connections/. The OAuth flow goes through oauth-manager.ts (lines 30–88): it generates PKCE challenge/verifier, redirects via browser to the provider's authorization URL, then exchanges the code for tokens at the token URL. Tokens are encrypted at rest in connection-store.ts using AES-256-GCM with a per-profile encryption key (LEON_CONNECTIONS_ENCRYPTION_KEY, line 13). The connection-tool.ts (lines 55–60) exposes a setup_connection tool the agent can call to submit credentials, and connection-catalog.ts retrieves credentials scoped to only declared fields. Each profile gets isolated encrypted connection storage. Lee also discovers Fellows (other AI coding tools like Claude Code, Codex, etc.) via server/src/core/llm-manager/fellows/ to reuse their existing API keys — fellows.json lists their config file locations and key names. Token refresh is handled transparently via refreshTokenIfNeeded using the supports_refresh field in the OAuth config (Gmail example has it enabled, line 33). The Satellite system (server/src/satellite.ts) allows running device-bound tools (e.g. Gmail, file system) from a remote Leon server by proxying tool calls through an encrypted Socket.IO tunnel. This is an on-demand architecture: tools declare their connection needs; the agent loads the toolkit, discovers it needs a connection, presents a setup widget, and only then uses the API.
How is memory and user context stored and retrieved?
answeredLeon's memory is layered into three scopes — persistent, daily, and discussion — plus context files as a separate grounding source. Core storage is SQLite via better-sqlite3 using the schema at server/src/core/memory-manager/sql/schema.sql: memory_items table stores records (lines 1–26) with fields like scope, kind (fact, preference, event, summary, etc.), content_md, importance (0–1), confidence, expires_at, and is_pinned. Chunks are stored in memory_chunks with FTS5 full-text search and optional embeddings. On top of SQLite, QMD (@tobilu/qmd package) provides retrieval via server/src/core/memory-manager/qmd-backend.ts (lines 37–55): content is mirrored into QMD collections, searched with hybrid (lexical + vector) retrieval, and reranked by namespace weight (memory_persistent gets 1.35, discussion gets 0.65). The flow in memory-manager.ts (lines 142–167): recall starts with QMD retrieval for the top-K chunks (default 12), then may run adaptive second passes if results are weak. Extracted durable facts use an LLM schema (EXTRACT_PERSISTENT_MEMORY_SCHEMA, lines 58–75). Memory is injected into prompts via renderRecallPrompt() which formats as "Memory Recall:" with facts and relevant chunks — this string is added to the agent's system prompt context (memory-manager.ts, lines 142–167). Turn observations feed daily and discussion memory automatically via MemoryManager.observeTurn(). Summarization happens through summarizer.ts for daily markdown summaries. Older short-term content is compacted and pruned based on TTLs (5 days for discussion, 30 days active retention, 180 days cold archive). The SelfModelManager (server/src/core/self-model-manager.ts, lines 1–60) distills repeated patterns into behavioral principles from conversation digests, injected as a compact snapshot. Context files (under server/src/core/context-manager/context-files/) are a separate system: files like ACTIVITY.md, HOST_SYSTEM.md, GPU_COMPUTE.md, and HOME.md are refreshed periodically and read via structured_knowledge tools to ground environment-aware answers.
How are actions on the user's behalf gated?
answeredActions on the user's behalf are gated through several layers. First, the connection system: before any external API tool (Gmail, Notion, etc.) can execute, it must present a declared connection with required settings (tools/communication/gmail/tool.json, lines 13–14: required_settings: ["access_token"]). The agent cannot call a tool until the user completes the OAuth or API-key setup via a browser or the setup_connection tool — and the runtime validates declared fields, never exposes raw secrets (connection-tool.ts, lines 55–60: "Returns only a safe status, never provider errors"). Tool availability scoping is explicit: config.sample.yml (lines 121–130) defines availability.tools.allowed and availability.tools.disabled lists — when allowed is non-empty, only those tools are available. Native skills use a parallel availability.skills block. In the agent loop, the system prompt (agent-loop.ts, lines 108–154) enforces a safety policy: "Verify required paths, identifiers, accepted values, and prerequisites before side effects", "Do not invent... tool-produced facts", and "Never use computer or browser automation to hide automation, spoof identity, bypass CAPTCHA". Duplicate call detection prevents repeated bad calls (agent-loop.ts, line 29: findDuplicateToolInputMatch). The completion_review system prompt (agent-loop.ts, lines 184–196) serves as an audit trail: before a final answer is accepted, the LLM self-reviews its evidence against the original request, enforcing completion standards and flagging blockers. Argument validation uses AJV schemas (server/src/ajv.ts) with human-readable errors. The tool executor (tool-executor.ts) validates and parses inputs before execution. The PulseManager (server/src/core/pulse-manager.ts, lines 60–65) tracks proactive actions with a suppression policy that learns from owner declines, so Leon does not repeatedly push unwanted proactive behaviors. Satellite architecture provides physical isolation: bound tool calls route to the user's local device, and tools become unavailable when Satellite disconnects, preventing remote execution without the device.
How are LLM providers selected and configured?
answeredLeon supports 13 LLM providers, mapped by LLM_PROVIDERS_MAP in llm-provider.ts (lines 36–50): OpenAI, Anthropic, DeepSeek, Groq, OpenRouter, ZAI, MiniMax, MoonshotAI, Cerebras, HuggingFace, LlamaCPP (local), SGLang (local), and Celeris. Each has a provider class under server/src/core/llm-manager/llm-providers/ (e.g. openai-llm-provider.ts extends AISDKRemoteLLMProvider, which is a thin OpenAI/AI SDK responses-API shim). Local model support comes via llamacpp-llm-provider.ts (node-llama-cpp binary) and sglang-llm-provider.ts (OpenAI-compatible server), both marked in llm-routing.ts (lines 26–29) as LOCAL_PROVIDERS. Configuration surfaces in config.sample.yml (lines 16–92): set llm.default to a <provider>/<model> string, with optional per-mode overrides under llm.workflow and llm.agent. API keys are read from environment variables (e.g. LEON_OPENAI_API_KEY). The provider config catalog is at llm-provider-account-configs.ts (lines 19–40). Tool-calling & structured outputs are central to the agent loop: react-llm-duty.ts builds OpenAI-compatible tool definitions from the toolkit registry, bundles them into the provider call, and the provider must support tool calling (LEON.md, line 37: "Agent models must support tool calling"). Structured output uses JSON Schema on completions via ajv validation (llm-provider.ts line 5, inference.ts lines 1–50). For autonomous sub-tasks, the runInference() helper (inference.ts, lines 26–50) accepts a jsonSchema parameter for structured results. The AI SDK wrapper (@ai-sdk/* packages in package.json lines 80–90: @ai-sdk/anthropic, @ai-sdk/openai, @ai-sdk/deepseek, etc.) provides a unified provider abstraction. Fellow discovery (server/src/core/llm-manager/fellows/) can auto-detect API keys from other AI tools (Claude Code, Codex, OpenCode, Pi, Hermes, T3 Code, OpenClaw) to reuse existing credentials without re-entering them.
How is it deployed and self-hosted?
answeredLeon is self-hosted and designed to run on your own hardware. The primary deploy path uses Node.js ≥24 (managed via Volta, per package.json line 81 and README.md line 82). There is no Dockerfile or docker-compose file in the repository — deployment is bare-metal via pnpm install followed by pnpm run build && pnpm start (or pnpm run dev:server for development). Runtime dependencies: SQLite through better-sqlite3 (the only database — no separate DB server required); optional Python runtime for bridges/python and tcp_server components; optional FFmpeg for media processing. The setup pipeline (scripts/setup/setup.js, lines 1–60) orchestrates: Node.js & pnpm validation, Python environment creation, CMake/Ninja/jq for local LLM compilation, node-llama-cpp setup, nltk data, PyTorch, QMD LLM indexing, skill dependencies, tool settings, and profile preferences. Profile isolation: each usage context gets its own profile directory (~/.leon/profiles/<profile-id>/) with isolated config, encrypted connections, memory, logs, and settings (ARCHITECTURE.md, lines 19–22). Multiple profiles can run concurrently. Satellite (server/src/satellite.ts, lines 1–60) is an optional companion process run via pnpm dev:satellite or pnpm start:satellite on a user device, connecting to a remote Leon server to proxy device-local tool execution. Required external accounts depend on the LLM provider: at minimum you need an API key for your chosen provider (Anthropic, OpenAI, etc.) stored via environment variable in the profile .env file (referenced from config.sample.yml as env: fields). For integrations (Gmail, Notion, TickTick), each needs its own OAuth application credentials. The startup sequence runs a pre-check (server/src/pre-check.ts) that validates the environment, then the main server starts with Fastify for HTTP and Socket.IO for real-time client communication. Built-in commands provide runtime configuration: /model <provider> <model name> to set the active model, /connection for tool connections.