# leon-ai/leon

> Self-hosted Node.js assistant: a tool-calling agent loop over JSON-declared toolkits, layered SQLite memory and optional remote-device tools.

- Category: [Open-source personal assistants](https://llms-technical-reviews.com/personal-assistants/)
- Repository: https://github.com/leon-ai/leon (reviewed at commit `24714435b41cb6de7abf7b4a963f2ca4d0bb05ce`, 2026-10-06)
- Stars: 17556 · Language: TypeScript · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/leon/

## Overview

Leon is a self-hosted personal assistant server written in TypeScript on Node.js 24. Clients (the bundled Vite app, the newer React web app, or your own) connect over Socket.IO and send "utterances". Leon routes each one either to a tool-calling agent loop or to a deterministic skill. Leon has existed for years as an intent-classification assistant. At this commit it has been rebuilt around an LLM agent, and the old skill system survives as the "controlled" route.

Three routing modes exist in `config.yml`: `agent` (the default), `controlled` and `smart`. In agent mode every request goes to one long-running provider tool-calling loop. Tools come from a catalog of toolkits declared in JSON. They cover the shell, files, ripgrep, browser control, desktop control through Cua, Gmail, Notion, TickTick, media generation, speech, and more. Their schemas are loaded progressively, so the model sees only what it needs. Around the loop sit a layered SQLite memory, generated "context files" that describe the host machine, a private self-model, and a "pulse" that can start autonomous tasks on a timer.

It is aimed at one owner, or a few profiles, running Leon on their own computer. It assumes broad local access, because shell and desktop control are first-class tools. Thirteen LLM providers are supported, including local llama.cpp and SGLang servers.

## Architecture

```mermaid
flowchart LR
  CL["Clients (app, web-app, custom)"] --> SO["Socket.IO server"]
  HT["Fastify HTTP API"] --> NLU
  SO --> NLU["NLU: routing decision"]
  NLU -->|"agent"| RE["ReActLLMDuty"]
  NLU -->|"controlled"| SK["Skill router + action calling"]
  RE --> LOOP["runAgentLoop"]
  LOOP --> PROV["LLM provider (13 backends)"]
  LOOP --> REG["ToolkitRegistry (progressive schemas)"]
  LOOP --> EXE["ToolExecutor"]
  EXE --> WK["Per-tool worker processes"]
  EXE --> SAT["Satellite on another device"]
  SK --> BR["Node.js / Python bridges"]
  LOOP --> MEM["MemoryManager (SQLite + QMD)"]
  PULSE["PulseManager"] --> RE
```

| Component | Path | Role |
|---|---|---|
| Boot | `server/src/index.ts` | Starts the HTTP and Socket.IO servers, the optional Python TCP server (ASR/TTS), pulse and context managers |
| Socket server | `server/src/core/socket-server.ts` | Client, satellite and widget events; `utterance` handling |
| NLU / router | `server/src/core/nlp/nlu/nlu.ts` | Chooses the agent or controlled route, runs skill selection and slot filling |
| Agent duty | `server/src/core/llm-manager/llm-duties/react-llm-duty.ts` | Builds the prompt, tool catalog and history, then calls the loop |
| Agent loop | `server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts` | Iterates model calls, tool batches, recovery and completion review |
| Providers | `server/src/core/llm-manager/llm-providers/` | OpenAI, Anthropic, OpenRouter, Groq, DeepSeek, llama.cpp, SGLang and others |
| Tool manager | `server/src/core/tool-manager/` | Toolkit registry, executor and worker process pool |
| Toolkits | `tools/<domain>/<tool>/tool.json` | 38 tool declarations with function schemas, guidance and connection needs |
| Skills | `skills/native/`, `skills/agent/` | Deterministic actions (with locales) and `SKILL.md` guides for the agent |
| Memory | `server/src/core/memory-manager/` | SQLite schema, QMD retrieval, daily summaries |
| Connections | `server/src/core/connections/` | OAuth and API-key setup, encrypted per profile |

## How a request flows

Take "find last week's invoices in my Downloads folder and total them", typed into the web app:

1. **Receive.** The socket server binds `socket.on('utterance', ...)` to `handleOwnerMessage` for that client ([socket-server.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/socket-server.ts#L1256-L1258)). The profile credential on the connection decides which profile's config, memory and tools are used.
2. **Route.** `NLU.process` calls `getRoutingDecision`. Attachments force the agent route. If the route's LLM target is not configured, Leon replies with a hint to run `/model <provider> <model>` ([nlu.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/nlp/nlu/nlu.ts#L1480-L1600)). In agent mode it calls `runReAct`, which builds a `ReActLLMDuty` with any forced agent skill or tool and executes it ([nlu.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/nlp/nlu/nlu.ts#L860-L890)).
3. **Prepare.** The duty assembles the system prompt (stable rules first, then volatile runtime state), memory recall, one-line context-file summaries and a small starting tool catalog. The catalog includes a toolkit loader, an agent-skill loader and `request_clarification`.
4. **Loop.** `runAgentLoop` runs up to `agent_max_iterations` (default 256) turns. Each turn calls the model with the transcript and current tools. It retries once on truncated or empty output, and once with a smaller context if the provider reports context pressure ([agent-loop.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts#L806-L1000)). At most eight tool calls per turn are executed; the rest are deferred.
5. **Execute tools.** Here the model would load `operating_system_control`, then call `ripgrep` and `file`. Each call is validated, checked for duplicate inputs, and run in a separate Node worker process spawned per toolkit and tool ([tool-worker-manager.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/tool-manager/tool-worker-manager.ts#L85-L125)). Results, including errors, return as observations. Large outputs are written to artifact logs, and only a preview goes back to the model.
6. **Review.** When the model answers without tool calls after doing work, `reviewAgentCompletion` asks the same provider for a `complete` / `continue` / `blocked` verdict. `continue` sends the loop back to work.
7. **Answer and remember.** The final text streams back over Socket.IO. `MemoryManager.observeTurn` stores the exchange as a daily event and a discussion note with a five-day TTL, refreshes the day summary, and prunes old discussion items ([memory-manager.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/memory-manager.ts#L1191-L1240)).

## Key components

### The agent loop

The loop keeps one transcript for the whole task. Checkpoints and compaction keep it within a fixed input budget: old tool exchanges and inactive toolkit schemas are dropped first. When the budget runs out, a final "finishing pass" without tools has to answer from the evidence already collected. The system prompt is long and specific. It ranks tools from dedicated APIs, through browser and OS tools, down to bounded shell commands. It has a `<safety>` block that forbids guessing values before side effects and forbids using automation to bypass CAPTCHAs or anti-bot controls ([agent-loop.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts#L108-L200)). The completion reviewer is told to treat tool output as evidence, not instructions. That is a light prompt-injection guard.

### Toolkits and connections

A toolkit is a domain folder with a `toolkit.json`. Each tool has a `tool.json` that owns its function schemas, guidance, binaries and connection requirements, plus a Node.js or Python implementation built on the bridge SDK. "Installed", "enabled" and "available" are separate states. A tool is available only when its required settings exist, and `availability.tools.allowed` / `disabled` in config narrow the set ([config.sample.yml](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L115-L135)). When a call fails for lack of credentials, the loop adds a `setup_connection` tool limited to the blocked providers. OAuth uses PKCE. Secrets are stored with AES-256-GCM under a per-profile key.

### Providers

`LLM_PROVIDERS_MAP` lists llama.cpp, SGLang, Groq, OpenRouter, Z.AI, DeepSeek, MiniMax, OpenAI, Anthropic, Moonshot, Cerebras, Hugging Face and Celeris ([llm-provider.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-provider.ts#L36-L50)). llama.cpp and SGLang are the local providers ([llm-routing.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-routing.ts#L26-L29)). You set `llm.default` as `<provider>/<model>`, optionally with separate `workflow` and `agent` targets. API keys live in the profile `.env` ([config.sample.yml](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L1-L60)). Setup can also discover credentials already stored by Codex, Claude Code, OpenCode, Pi, Hermes, OpenClaw and T3 Code, and offer to reuse them.

### Memory

The SQLite schema (WAL mode) stores `memory_items` with a scope (`persistent`, `daily`, `discussion`), a kind, importance, confidence, an expiry and a pin flag ([schema.sql](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/sql/schema.sql#L1-L30)). Items are chunked into `memory_chunks` with an FTS5 index and optional embeddings ([schema.sql](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/sql/schema.sql#L41-L66)). Recall goes through the QMD library, which does hybrid lexical and vector search. Results are then weighted by namespace: persistent memory 1.35, discussion 0.65 ([memory-manager.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/memory-manager.ts#L93-L108)). Weak first results trigger adaptive second passes. Separate from memory, an `OWNER.md` profile and machine-generated context files (activity, host system, network, inventory and others) ground questions about the environment.

### Pulse and self-model

With `runtime.pulse_enabled` (on in the sample config), `PulseManager` ticks every 30 minutes. It can generate an autonomous "matter" from memory, context changes and the self-model, run at most one per tick, and back off after the owner declines similar matters. Cooldowns and suppression policies are stored per matter fingerprint ([pulse-manager.ts](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/pulse-manager.ts#L40-L70)). The private diary/self-model distills repeated habits into principles and injects a compact snapshot into the first agent request.

## Extending it

- **A tool.** Add `tools/<domain>/<tool>/tool.json` and an implementation in `src/nodejs/` (extending the SDK `Tool`) or `src/python/` (extending `BaseTool`). Declare dependencies locally. Setup installs them with managed Node, Python, pnpm and uv.
- **An agent skill.** Drop a `skills/agent/<name>/SKILL.md` with discovery frontmatter. The agent loads it on demand and follows it inside the same loop. Coding, document authoring, live tutorial and media production ship this way.
- **A native skill.** Use `skills/native/<skill>/` with `skill.json`, `locales/` and action entry points, for deterministic flows in controlled mode.
- **A client.** Speak the Socket.IO contract with a `<profile>:<token>` credential. HTTP plugins can extend the API without patching core.
- **Remote devices.** Run Leon Satellite on a laptop and map tools to it in config (for example `computer_use.cua: my-laptop`). Those calls then execute on that device.

## Running it

- **Bare metal only.** There is no Dockerfile. Install Node 24 (pinned through Volta) and pnpm, then run `pnpm install`. The `postinstall` hook runs the setup pipeline (Python env, llama.cpp build tools, tool dependencies, profile prompts). Then run `pnpm build && pnpm start`, or `pnpm dev:server` ([package.json](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/package.json#L51-L74)).
- **Profiles.** Each profile has its own config, `.env`, encrypted connections, memory and logs, and one server can serve several profiles at once.
- **Model.** Configure one with `/model <provider> <model>` in the chat, or with `llm.default`. Voice (ASR, TTS, wake word) starts an extra Python TCP server.

## Strengths and caveats

- **Strength: serious agent runtime.** Progressive tool loading, bounded recovery, completion review, artifact-backed large outputs and resumable checkpoints are more than most personal-assistant projects ship.
- **Strength: tool hygiene.** Tools declare schemas, settings and connections in one manifest, run in isolated worker processes, and can be pinned to a remote device.
- **Strength: local-first.** Everything is stored in per-profile SQLite and files, and local llama.cpp/SGLang are supported.
- **Caveat: no action approval gate.** Shell, file and desktop tools execute when the model calls them. Safety is a system prompt, allow/deny lists and connection prerequisites. There is no per-action confirmation.
- **Caveat: autonomous by default.** The pulse and the self-model are enabled in the sample config. Leon may act on its own every 30 minutes unless you turn that off.
- **Caveat: big moving target.** The agent loop alone is over 2,000 lines and the setup pipeline compiles native pieces. Expect setup friction and fast churn.
- **Caveat: two paradigms.** Native skills, NLU and slot filling remain beside the agent loop, so there are two ways to build features.

*Sources: code at 2471443, deepwiki-open wiki (12 pages), verified Q&A.*

## How leon-ai/leon answers the Open-source personal assistants questions

### How is the assistant architected? (answered)

Leon's architecture centers on three runtime modes — **agent** (default), **controlled**, and **smart** — selected in `config.sample.yml` via `routing.mode` (`server/src/../config.ts`). In agent mode, the core loop lives in `server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts` (lines 1–54): a continuous, single-transcript provider tool-calling loop where the LLM calls tools, observations are fed back, and recovery/completion-review check progress. The loop is orchestrated by `react-llm-duty.ts` which imports from `agent-constants.ts`, `tool-execution.ts`, `agent-plan.ts`, and `agent-history-manager.ts`. The **Brain** (`server/src/core/brain/brain.ts`, lines 59–77) manages the answer queue, paraphrasing, TTS, and skill execution via handlers (`dialog-action-skill-handler.ts`, `logic-action-skill-handler.ts`). Tool schemas load progressively from the **ToolkitRegistry** (`server/src/core/tool-manager/toolkit-registry.ts`), and the **ToolExecutor** (`server/src/core/tool-manager/tool-executor.ts`) runs calls concurrently in isolated workers. The frontend/backend split uses client-agnostic **Socket.IO** (`server/src/core/socket-server.ts`) — text (and attached files) comes in as utterances, and the response streams back as tokens and final answers. Legacy [`controlled` mode] routes through **NLU** (natural language understanding) which classifies intent and triggers native skill actions (`Skills -> Actions -> Tools -> Functions`). Packages are organized as a pnpm workspace: `server/` (core + Node.js bridge), `app/` (terminal UI), `web-app/` (React web UI), `bridges/` (Node.js & Python bridges), `skills/` (native + agent skills), `tools/` (toolkits with JSON declarations). A user request flows: client emits an utterance via Socket.IO → `socket-server.ts` dispatches to the NLU or directly to the agent loop → based on routing mode, the LLM chooses/calls tools → results return as observations → final answer is queued through `Brain.talk()` and emitted back to the client.


Citations: [config.sample.yml:12-15](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L12-L15) · [server/src/core/brain/brain.ts:59-77](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/brain/brain.ts#L59-L77) · [server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts:44-55](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts#L44-L55) · [server/src/core/tool-manager/tool-executor.ts:56-72](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/tool-manager/tool-executor.ts#L56-L72) · [server/src/core/tool-manager/toolkit-registry.ts:1-59](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/tool-manager/toolkit-registry.ts#L1-L59)

### How are integrations (email, calendar, chat, docs) implemented? (answered)

Leon integrates external services via the **Toolkit** + **Connection** system rather than MCP. Services are declared in JSON toolkit manifests under `/tools/` — organized by domain (e.g. `communication/gmail`, `productivity_collaboration/notion`, `search_web/hosted`, `calendar_scheduling`). Each tool has a `tool.json` (e.g. `tools/communication/gmail/tool.json`, lines 1–136) describing its functions, parameters, and connection requirements. Connections support two auth methods: **OAuth** and **API key**, both managed through `server/src/core/connections/`. The OAuth flow goes through `oauth-manager.ts` (lines 30–88): it generates PKCE challenge/verifier, redirects via browser to the provider's authorization URL, then exchanges the code for tokens at the token URL. Tokens are encrypted at rest in `connection-store.ts` using AES-256-GCM with a per-profile encryption key (`LEON_CONNECTIONS_ENCRYPTION_KEY`, line 13). The `connection-tool.ts` (lines 55–60) exposes a `setup_connection` tool the agent can call to submit credentials, and `connection-catalog.ts` retrieves credentials scoped to only declared fields. Each profile gets isolated encrypted connection storage. Lee also discovers **Fellows** (other AI coding tools like Claude Code, Codex, etc.) via `server/src/core/llm-manager/fellows/` to reuse their existing API keys — `fellows.json` lists their config file locations and key names. Token refresh is handled transparently via `refreshTokenIfNeeded` using the `supports_refresh` field in the OAuth config (Gmail example has it enabled, line 33). The **Satellite** system (`server/src/satellite.ts`) allows running device-bound tools (e.g. Gmail, file system) from a remote Leon server by proxying tool calls through an encrypted Socket.IO tunnel. This is an on-demand architecture: tools declare their connection needs; the agent loads the toolkit, discovers it needs a connection, presents a setup widget, and only then uses the API.


Citations: [tools/communication/gmail/tool.json:1-136](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/tools/communication/gmail/tool.json#L1-L136) · [server/src/core/connections/oauth-manager.ts:30-88](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/connections/oauth-manager.ts#L30-L88) · [server/src/core/connections/connection-store.ts:1-40](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/connections/connection-store.ts#L1-L40) · [server/src/core/connections/connection-catalog.ts:1-98](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/connections/connection-catalog.ts#L1-L98) · [server/src/core/connections/connection-tool.ts:1-60](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/connections/connection-tool.ts#L1-L60)

### How is memory and user context stored and retrieved? (answered)

Leon's memory is **layered** into three scopes — `persistent`, `daily`, and `discussion` — plus context files as a separate grounding source. Core storage is **SQLite via better-sqlite3** using the schema at `server/src/core/memory-manager/sql/schema.sql`: `memory_items` table stores records (lines 1–26) with fields like `scope`, `kind` (`fact`, `preference`, `event`, `summary`, etc.), `content_md`, `importance` (0–1), `confidence`, `expires_at`, and `is_pinned`. Chunks are stored in `memory_chunks` with FTS5 full-text search and optional embeddings. On top of SQLite, **QMD** (`@tobilu/qmd` package) provides retrieval via `server/src/core/memory-manager/qmd-backend.ts` (lines 37–55): content is mirrored into QMD collections, searched with hybrid (lexical + vector) retrieval, and reranked by namespace weight (`memory_persistent` gets 1.35, `discussion` gets 0.65). The flow in `memory-manager.ts` (lines 142–167): recall starts with QMD retrieval for the top-K chunks (default 12), then may run adaptive second passes if results are weak. Extracted durable facts use an LLM schema (`EXTRACT_PERSISTENT_MEMORY_SCHEMA`, lines 58–75). Memory is **injected into prompts** via `renderRecallPrompt()` which formats as "Memory Recall:" with facts and relevant chunks — this string is added to the agent's system prompt context (`memory-manager.ts`, lines 142–167). Turn observations feed daily and discussion memory automatically via `MemoryManager.observeTurn()`. Summarization happens through `summarizer.ts` for daily markdown summaries. Older short-term content is compacted and pruned based on TTLs (5 days for discussion, 30 days active retention, 180 days cold archive). The **SelfModelManager** (`server/src/core/self-model-manager.ts`, lines 1–60) distills repeated patterns into behavioral principles from conversation digests, injected as a compact snapshot. Context files (under `server/src/core/context-manager/context-files/`) are a separate system: files like `ACTIVITY.md`, `HOST_SYSTEM.md`, `GPU_COMPUTE.md`, and `HOME.md` are refreshed periodically and read via `structured_knowledge` tools to ground environment-aware answers.


Citations: [server/src/core/memory-manager/sql/schema.sql:1-26](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/sql/schema.sql#L1-L26) · [server/src/core/memory-manager/memory-manager.ts:25-75](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/memory-manager.ts#L25-L75) · [server/src/core/memory-manager/memory-manager.ts:142-167](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/memory-manager.ts#L142-L167) · [server/src/core/memory-manager/qmd-backend.ts:37-55](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/memory-manager/qmd-backend.ts#L37-L55) · [server/src/core/self-model-manager.ts:1-60](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/self-model-manager.ts#L1-L60)

### How are actions on the user's behalf gated? (answered)

Actions on the user's behalf are gated through several layers. **First, the connection system**: before any external API tool (Gmail, Notion, etc.) can execute, it must present a declared `connection` with required settings (`tools/communication/gmail/tool.json`, lines 13–14: `required_settings: ["access_token"]`). The agent cannot call a tool until the user completes the OAuth or API-key setup via a browser or the `setup_connection` tool — and the runtime validates declared fields, never exposes raw secrets (`connection-tool.ts`, lines 55–60: "Returns only a safe status, never provider errors"). **Tool availability scoping** is explicit: `config.sample.yml` (lines 121–130) defines `availability.tools.allowed` and `availability.tools.disabled` lists — when `allowed` is non-empty, only those tools are available. Native skills use a parallel `availability.skills` block. **In the agent loop**, the system prompt (`agent-loop.ts`, lines 108–154) enforces a safety policy: "Verify required paths, identifiers, accepted values, and prerequisites before side effects", "Do not invent... tool-produced facts", and "Never use computer or browser automation to hide automation, spoof identity, bypass CAPTCHA". Duplicate call detection prevents repeated bad calls (`agent-loop.ts`, line 29: `findDuplicateToolInputMatch`). The `completion_review` system prompt (`agent-loop.ts`, lines 184–196) serves as an **audit trail**: before a final answer is accepted, the LLM self-reviews its evidence against the original request, enforcing completion standards and flagging blockers. Argument validation uses AJV schemas (`server/src/ajv.ts`) with human-readable errors. The tool executor (`tool-executor.ts`) validates and parses inputs before execution. The **PulseManager** (`server/src/core/pulse-manager.ts`, lines 60–65) tracks proactive actions with a suppression policy that learns from owner declines, so Leon does not repeatedly push unwanted proactive behaviors. **Satellite** architecture provides physical isolation: bound tool calls route to the user's local device, and tools become unavailable when Satellite disconnects, preventing remote execution without the device.


Citations: [tools/communication/gmail/tool.json:13-16](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/tools/communication/gmail/tool.json#L13-L16) · [server/src/core/connections/connection-tool.ts:55-60](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/connections/connection-tool.ts#L55-L60) · [config.sample.yml:121-130](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L121-L130) · [server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts:108-154](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts#L108-L154) · [server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts:184-196](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-duties/react-llm-duty/agent-loop.ts#L184-L196) · [server/src/core/pulse-manager.ts:47-65](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/pulse-manager.ts#L47-L65)

### How are LLM providers selected and configured? (answered)

Leon supports **13 LLM providers**, mapped by `LLM_PROVIDERS_MAP` in `llm-provider.ts` (lines 36–50): OpenAI, Anthropic, DeepSeek, Groq, OpenRouter, ZAI, MiniMax, MoonshotAI, Cerebras, HuggingFace, LlamaCPP (local), SGLang (local), and Celeris. Each has a provider class under `server/src/core/llm-manager/llm-providers/` (e.g. `openai-llm-provider.ts` extends `AISDKRemoteLLMProvider`, which is a thin OpenAI/AI SDK responses-API shim). **Local model support** comes via `llamacpp-llm-provider.ts` (node-llama-cpp binary) and `sglang-llm-provider.ts` (OpenAI-compatible server), both marked in `llm-routing.ts` (lines 26–29) as `LOCAL_PROVIDERS`. Configuration surfaces in `config.sample.yml` (lines 16–92): set `llm.default` to a `<provider>/<model>` string, with optional per-mode overrides under `llm.workflow` and `llm.agent`. API keys are read from environment variables (e.g. `LEON_OPENAI_API_KEY`). The provider config catalog is at `llm-provider-account-configs.ts` (lines 19–40). **Tool-calling & structured outputs** are central to the agent loop: `react-llm-duty.ts` builds OpenAI-compatible tool definitions from the toolkit registry, bundles them into the provider call, and the provider must support tool calling (`LEON.md`, line 37: "Agent models must support tool calling"). Structured output uses JSON Schema on completions via `ajv` validation (`llm-provider.ts` line 5, `inference.ts` lines 1–50). For autonomous sub-tasks, the `runInference()` helper (`inference.ts`, lines 26–50) accepts a `jsonSchema` parameter for structured results. The **AI SDK** wrapper (`@ai-sdk/*` packages in `package.json` lines 80–90: `@ai-sdk/anthropic`, `@ai-sdk/openai`, `@ai-sdk/deepseek`, etc.) provides a unified provider abstraction. **Fellow discovery** (`server/src/core/llm-manager/fellows/`) can auto-detect API keys from other AI tools (Claude Code, Codex, OpenCode, Pi, Hermes, T3 Code, OpenClaw) to reuse existing credentials without re-entering them.


Citations: [server/src/core/llm-manager/llm-provider.ts:36-50](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-provider.ts#L36-L50) · [config.sample.yml:16-92](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L16-L92) · [server/src/core/llm-manager/llm-routing.ts:26-29](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-routing.ts#L26-L29) · [server/src/core/llm-manager/llm-provider-account-configs.ts:19-40](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/llm-provider-account-configs.ts#L19-L40) · [server/src/core/llm-manager/inference.ts:26-50](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/core/llm-manager/inference.ts#L26-L50)

### How is it deployed and self-hosted? (answered)

Leon is **self-hosted** and designed to run on your own hardware. The primary deploy path uses Node.js ≥24 (managed via Volta, per `package.json` line 81 and `README.md` line 82). There is **no Dockerfile or docker-compose file** in the repository — deployment is bare-metal via `pnpm install` followed by `pnpm run build && pnpm start` (or `pnpm run dev:server` for development). **Runtime dependencies**: SQLite through `better-sqlite3` (the only database — no separate DB server required); optional Python runtime for `bridges/python` and `tcp_server` components; optional FFmpeg for media processing. The setup pipeline (`scripts/setup/setup.js`, lines 1–60) orchestrates: Node.js & pnpm validation, Python environment creation, CMake/Ninja/jq for local LLM compilation, node-llama-cpp setup, nltk data, PyTorch, QMD LLM indexing, skill dependencies, tool settings, and profile preferences. **Profile isolation**: each usage context gets its own profile directory (`~/.leon/profiles/<profile-id>/`) with isolated config, encrypted connections, memory, logs, and settings (`ARCHITECTURE.md`, lines 19–22). Multiple profiles can run concurrently. **Satellite** (`server/src/satellite.ts`, lines 1–60) is an optional companion process run via `pnpm dev:satellite` or `pnpm start:satellite` on a user device, connecting to a remote Leon server to proxy device-local tool execution. **Required external accounts** depend on the LLM provider: at minimum you need an API key for your chosen provider (Anthropic, OpenAI, etc.) stored via environment variable in the profile `.env` file (referenced from `config.sample.yml` as `env:` fields). For integrations (Gmail, Notion, TickTick), each needs its own OAuth application credentials. The startup sequence runs a `pre-check` (`server/src/pre-check.ts`) that validates the environment, then the main server starts with Fastify for HTTP and Socket.IO for real-time client communication. **Built-in commands** provide runtime configuration: `/model <provider> <model name>` to set the active model, `/connection` for tool connections.


Citations: [package.json:35-37](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/package.json#L35-L37) · [scripts/setup/setup.js:1-60](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/scripts/setup/setup.js#L1-L60) · [server/src/satellite.ts:1-60](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/server/src/satellite.ts#L1-L60) · [config.sample.yml:60-65](https://github.com/leon-ai/leon/blob/24714435b41cb6de7abf7b4a963f2ca4d0bb05ce/config.sample.yml#L60-L65)
