# NousResearch/hermes-agent

> Python agent core behind a CLI, desktop app and 20-platform chat gateway, with guarded shell tools, curated memory and self-written skills.

- Category: [Open-source personal assistants](https://llms-technical-reviews.com/personal-assistants/)
- Repository: https://github.com/NousResearch/hermes-agent (reviewed at commit `950ea4d51eb5667d0258dbead49cc8be82a6e7bf`, 2026-10-06)
- Stars: 251643 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/hermes-agent/

## Overview

Hermes Agent is Nous Research's general-purpose personal agent. One Python agent core runs behind a terminal CLI and TUI, a messaging gateway (Telegram, Discord, Slack, WhatsApp, Signal, Teams, email and about fifteen more platforms), a web dashboard, an Electron desktop app, an OpenAI-compatible API server, and an ACP adapter for editors. The agent drives a real terminal, files, a browser and subagents. It keeps curated memory files about you and itself, and it can write its own skills: the "grows with you" in its own description.

The codebase is very large: hundreds of modules in `agent/` and `tools/`, and a state layer split across dozens of `hermes_state_*.py` files. It is also written defensively. Comments cite issue numbers for nearly every edge case: prompt-cache invariants, credential refresh, gateway reconnects, WAL corruption. Two design rules from the root `AGENTS.md` explain much of it. Per-conversation prompt caching is "sacred", so memory and toolsets are frozen for a session. And the core is a "narrow waist", so new capability is meant to arrive as skills, plugins or MCP servers.

It suits power users who want one always-on agent they can reach from their phone and their terminal. It assumes broad trust: by default it runs shell commands on the host, guarded by an approval layer rather than a sandbox.

## Architecture

```mermaid
flowchart LR
  CLI["CLI / TUI"] --> AG["AIAgent"]
  GW["Gateway + platform plugins"] --> CACHE["Agent cache (per session)"]
  CACHE --> AG
  API["API server / ACP / dashboard"] --> AG
  AG --> LOOP["run_conversation loop"]
  LOOP --> PROV["Provider resolution + transports"]
  LOOP --> TR["run_tool_round"]
  TR --> APR["Approval gate"]
  APR --> TOOLS["Tools: terminal, files, browser, MCP, delegate"]
  TOOLS --> ENV["Terminal backends (local, Docker, SSH, Modal...)"]
  LOOP --> DB["SQLite state (WAL + FTS5)"]
  LOOP --> MEM["MEMORY.md / USER.md + provider plugins"]
  LOOP --> BR["Background review (memory + skills)"]
  CRON["Cron scheduler"] --> AG
```

| Component | Path | Role |
|---|---|---|
| Agent class | `run_agent.py` | `AIAgent`, assembled from mixins; constructor forwards to `agent/agent_init.py` |
| Turn loop | `agent/conversation_loop.py` | `run_conversation` and the phase-based iteration loop |
| Tool round | `agent/turn_tool_round.py` | Validates, dedupes, persists and dispatches tool calls |
| Providers | `hermes_cli/providers.py`, `agent/*_adapter.py` | Provider overlays, API-mode selection, Anthropic/Bedrock/Gemini/Codex adapters |
| Tools | `tools/`, `toolsets.py` | Self-registering tools grouped into toolsets; MCP client; terminal backends in `tools/environments/` |
| Approvals | `tools/approval*.py` | Dangerous-command detection, smart (guardian LLM) approval, deny rules, YOLO |
| Gateway | `gateway/run.py`, `plugins/platforms/` | Long-running process that maps platform messages to cached agents |
| State | `hermes_state.py` + `hermes_state_*.py` | SQLite sessions, messages, usage, FTS5 search |
| Memory and skills | `tools/memory_tool.py`, `tools/skill_manager_tool.py`, `agent/background_review.py` | Curated memory files, agent-written skills, post-turn review fork |
| Scheduler | `cron/` | Scheduled jobs with delivery back to any platform |
| Apps | `apps/desktop/` (Electron), `ui-tui/`, `web/` | Desktop, terminal and web front ends |

## How a request flows

Take a Telegram message, "clean up the Downloads folder on my server":

1. **Receive.** The Telegram platform plugin, which wraps `python-telegram-bot` and requires `TELEGRAM_BOT_TOKEN`, hands the update to the gateway. The gateway keeps one `AIAgent` per session in a cache capped at 128 agents, evicting those idle for over an hour ([gateway/run.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/gateway/run.py#L51-L52)).
2. **Build the turn.** `run_conversation` wraps `_run_conversation_turn`. That function refreshes credentials from `.env`, then `build_turn_context` restores or builds the system prompt (identity, frozen memory snapshot, skills index, context files) and loads history from SQLite ([conversation_loop.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/conversation_loop.py#L1533-L1600)).
3. **Iterate.** The loop runs while `api_call_count < max_iterations` and the shared `IterationBudget` has room. Each pass runs explicit phases: `begin_iteration`, `prepare_iteration`, `assemble_api_request`, a preflight gate (compression and context checks), then an API retry loop and `normalize_model_response` ([conversation_loop.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/conversation_loop.py#L1636-L1682)). The constructor default is `max_iterations=sys.maxsize` ([run_agent.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/run_agent.py#L263-L268)), and `agent.max_turns` defaults to unset, so turns are unbounded unless you configure a cap.
4. **Run tools.** When the model returns tool calls, `run_tool_round` validates names and arguments, expands batches, dedupes calls, caps delegation, and persists the assistant message before executing anything. The docstring calls this "persist-before-execute": a crash mid-command can be resumed ([turn_tool_round.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/turn_tool_round.py#L46-L110)).
5. **Gate.** The `terminal` call (`rm -rf ~/Downloads/*`) passes through `check_all_command_guards`. Hardline patterns and user `approvals.deny` globs block outright. Dangerous patterns go to the configured mode. The default `smart` mode asks a guardian LLM. Escalations send a prompt to the user, here as Telegram inline buttons, with a 300-second timeout ([config_defaults.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/config_defaults.py#L1666-L1700)).
6. **Execute.** The command runs on the configured terminal backend: local, docker, singularity, modal, daytona, vercel_sandbox or ssh. The result is appended and the loop continues until the model answers in text.
7. **Finalize.** `finalize_turn` persists the turn and runs post-turn hooks. Every tenth user turn (`memory.nudge_interval`), a background review fork replays the conversation and decides whether to save memory or skills ([turn_context.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/turn_context.py#L738-L756)).

## Key components

### The agent object

`AIAgent` is assembled from 14 mixins, including streaming delivery, interrupt control, rate-limit credits, session persistence and compression. Its constructor has 84 keyword parameters, from provider routing to a dozen UI callbacks, and forwards all of them to `init_agent` ([run_agent.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/run_agent.py#L241-L312)). The same class serves the CLI, gateway sessions, subagents (`delegate_task`) and the background-review fork. That is why so much per-turn state is reset explicitly at the top of each turn.

### Providers

`HERMES_OVERLAYS` declares 42 provider overlays: Nous Portal (OAuth device code), OpenRouter, OpenAI via the Codex Responses transport, xAI, Qwen OAuth, LM Studio and many others ([providers.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/providers.py#L28-L40)). Anything else resolves through user `providers:` / `custom_providers` entries or models.dev metadata. `determine_api_mode` picks the wire protocol per provider and model: OpenAI chat, Anthropic messages, Codex Responses or Bedrock Converse. Credential pools, fallback chains and a cheaper model for delegated children are all configurable.

### Memory and self-improvement

Built-in memory is two bounded Markdown files: `MEMORY.md` for agent notes (2,200 characters by default) and `USER.md` for your profile (1,375). The `memory` tool edits them with add, replace and remove operations. Both are frozen into the system prompt at session start, so writes made mid-session do not invalidate the prompt cache ([memory_tool.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/memory_tool.py#L1-L47)). One external `MemoryProvider` plugin can be added, such as Honcho, mem0, Supermemory or Hindsight ([config_defaults.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/config_defaults.py#L1312-L1335)). Skills are "procedural memory": `SKILL.md` folders that the agent creates and edits under `~/.hermes/skills/`. The background review forks the agent on the same provider and prompt cache, under a tool whitelist, and writes straight to the memory and skill stores ([background_review.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/background_review.py#L1-L30)). Set `memory.write_approval` to stage those writes for review.

### Session store

`hermes_state.py` is a SQLite store in WAL mode with FTS5 search. Sessions are tagged by source (`cli`, `telegram`, ...), and compression splits a long session into a `parent_session_id` chain ([hermes_state.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_state.py#L1-L6)). The Docker image builds its own SQLite to avoid the upstream WAL-reset bug.

### Approvals

The approval module owns per-session approvals, a permanent allowlist, YOLO state and a denial circuit breaker. After three consecutive guardian denials, the deny message escalates to a hard stop. `HERMES_YOLO_MODE` is read once at import. The comment explains that reading the environment on every call would let a skill flip it at runtime: a prompt-injection path ([approval.py](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/approval.py#L1-L75)). Cron, single-query and unattended contexts default to `deny`, because nobody is there to answer.

## Extending it

- **Skills first.** Bundled skills live in `skills/` and opt-in ones in `optional-skills/`. The agent can install, create and patch skills at runtime.
- **Plugins.** `plugins/` holds platform adapters (each with a `plugin.yaml` declaring its env vars), memory providers, model providers, image and video generation, browser backends, observability and more.
- **MCP.** `tools/mcp_tool*.py` is a full MCP client with OAuth, sampling and health supervision. `mcp_serve.py` exposes Hermes itself as an MCP server.
- **Core tools.** These are self-registering modules discovered by `tools/registry.py` and grouped in `toolsets.py`. The contribution guide sets a high bar for new core tools, because every tool schema is sent on every call.

## Running it

- **Install.** Use the install script, or `uv`/pip with the `hermes` console script (`hermes_cli.main:main`). `hermes setup` configures a provider, and `hermes gateway run` starts the messaging gateway.
- **Docker.** `docker-compose.yml` runs two containers from one image under s6-overlay: `gateway` and a `dashboard` bound to 127.0.0.1:9119. Both use host networking and mount `~/.hermes` ([docker-compose.yml](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/docker-compose.yml#L30-L77)). The OpenAI-compatible API server stays off until you set `API_SERVER_KEY`.
- **Requirements.** You need Python, at least one model provider (or a local LM Studio, llama.cpp or Ollama endpoint), and bot tokens for each platform you enable. Nix users get a flake.

## Strengths and caveats

- **Strength: reach.** One agent core serves CLI, TUI, desktop, about 20 chat platforms, cron jobs and an API, with sessions and memory shared through one profile directory.
- **Strength: operational hardening.** Persist-before-execute tool rounds, credential refresh, compression with session chains, and a custom SQLite build show real production mileage.
- **Strength: layered approvals.** Hardline blocks, deny globs that hold even under YOLO, a guardian model and per-surface defaults are more careful than most agents.
- **Caveat: size and complexity.** The core is spread over hundreds of small modules and mixins. Reading a single code path means following phase functions across many files.
- **Caveat: local shell by default.** Sandboxing is opt-in through a terminal backend. On the default local backend, the guardian LLM's judgement sits between the model and your machine.
- **Caveat: unbounded turns.** There is no default iteration cap, so long tool loops cost money unless you set `agent.max_turns` or a run budget.
- **Caveat: small memory.** Built-in memory is a few kilobytes of Markdown. Real long-term recall needs session search or an external provider plugin.

*Sources: code at 950ea4d, verified Q&A.*

## How NousResearch/hermes-agent answers the Open-source personal assistants questions

### How is the assistant architected? (answered)

**Agent loop and runtime:** The `AIAgent` class (run_agent.py:241) is the central orchestrator, composed of ~15 mixins for streaming, compression, approval, activity tracking, etc. Every user turn flows through `run_conversation()` (conversation_loop.py:1695) which calls `_run_conversation_turn()` (conversation_loop.py:1533). That function drives: pre-flight compression checks, `build_turn_context()` (turn_context.py:1017) to assemble messages and system prompt, `build_api_request()` → `perform_api_call()` to the LLM, `normalize_model_response()` to parse tool calls or text, `run_tool_round()` (turn_tool_round.py:46) to dispatch tools with parallel/segmented execution, `check_api_response()` for finish-reason handling, and `finalize_turn()` for post-turn hooks (memory sync, background review). The loop retries on errors up to `max_iterations`.

**Frontend/backend split:** The CLI (hermes_cli/) provides the interactive TUI via prompt_toolkit, the messaging gateway (gateway/run.py) runs platform adapters that receive messages and call `AIAgent.run_conversation()` per incoming message. The gateway caches agents per session with an LRU cap of 128 and 1h idle TTL. A web dashboard serves config, sessions, and cron management. The Desktop app wraps the CLI in Tauri (Rust).

**Main packages:** `agent/` — core loop, compression, memory, providers; `tools/` — 40+ tool implementations (terminal, browser, file, MCP, skills, etc.); `gateway/` — messaging platform adapters; `hermes_cli/` — CLI, config, profiles, auth; `hermes_state*.py` — SQLite state store.

**Request-to-action flow:** User message → `run_conversation()` → context assembly (system prompt + memory recall + history) → LLM API call → response parsing → if tool_calls: `_execute_tool_calls()` which runs the tool (terminal command, file edit, API call, etc.) → tool result appended to messages → loop back to LLM call until final text response is produced → post-turn memory sync + background review spawn.

> **Editor's note.** Correction: the desktop app in `apps/desktop/` is built with Electron, not Tauri; Tauri is only used by the separate bootstrap installer in `apps/bootstrap-installer/`.

Citations: [run_agent.py:241-312](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/run_agent.py#L241-L312) · [agent/conversation_loop.py:1533-1567](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/conversation_loop.py#L1533-L1567) · [agent/turn_tool_round.py:46-76](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/turn_tool_round.py#L46-L76) · [gateway/run.py:60-70](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/gateway/run.py#L60-L70)

### How are integrations (email, calendar, chat, docs) implemented? (answered)

**Supported services:** Telegram, Discord, Slack, WhatsApp Cloud, Signal, email (IMAP), WeChat Work (yuanbao), QQ Bot, Microsoft Teams, Google Chat, BlueBubbles (Apple Messages), and a generic webhook adapter. Each lives in `gateway/platforms/` as a subclass of `BasePlatformAdapter` (gateway/platforms/base.py:1).

**API clients vs MCP:** All integrations use direct HTTP API clients (Telegram bot API, Discord gateway, WhatsApp Cloud API, Microsoft Graph API, etc.) — there are no MCP-based chat adapters. The project has a separate `tools/mcp_tool*.py` system for MCP tool servers (not messaging).

**OAuth flow:** Calendars and email (Microsoft Graph) route through `tools/microsoft_graph_auth.py` using OAuth 2.0 device-code or authorization-code flows with token refresh. Microsoft Teams gateway auth (`TEAMS_CLIENT_ID`/`TEAMS_CLIENT_SECRET`) is configured via env vars in docker-compose.yml:44-50. Credentials persist to `~/.hermes/.env` or the SQLite credential store.

**Token storage:** API keys and OAuth tokens go to the profile-scoped `.env` file via `credential_persistence.py`/`secret_scope.py`. Gateway platform tokens (Telegram bot tokens, Slack signing secrets, etc.) are read from env vars or `~/.hermes/config.yaml`.

**Sync vs on-demand fetch:** Messaging is push-based — platforms deliver webhooks or long-poll updates to the gateway, which fans them out to per-session agent instances. No periodic polling for new messages; the gateway maintains persistent connections (Telegram getUpdates, Discord gateway, Slack Socket Mode, etc.). The cron scheduler (`hermes_cli/cron.py`) runs periodic jobs that inject prompts into the agent loop.

> **Editor's note.** Correction: most messaging adapters (Telegram, Discord, Slack, WhatsApp, Teams, Google Chat, email, Matrix and others) are plugins under `plugins/platforms/<name>/` with a `plugin.yaml`; `gateway/platforms/` holds the base adapter, the API server and a handful of built-in adapters such as Signal, BlueBubbles and WhatsApp Cloud.

Citations: [gateway/platforms/base.py:1-80](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/gateway/platforms/base.py#L1-L80) · [gateway/platforms/__init__.py:1-5](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/gateway/platforms/__init__.py#L1-L5) · [docker-compose.yml:44-60](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/docker-compose.yml#L44-L60) · [tools/microsoft_graph_auth.py:1-10](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/microsoft_graph_auth.py#L1-L10)

### How is memory and user context stored and retrieved? (answered)

**Storage:** The primary store is SQLite in WAL mode (`hermes_state.py`). It stores session metadata, message history, model config, usage tracking, and FTS5-indexed messages. Memory also lives on disk as `MEMORY.md` (agent-curated notes) and `USER.md` (user profile) under `~/.hermes/memories/`, managed by the `memory` tool (tools/memory_tool.py:1-6). FTS5 full-text search across sessions is provided by `hermes_state_fts.py`.

**What is remembered:** Three tiers — (1) **Session history**: every message in the current conversation (SQLite, per-session rows), searchable across sessions via `/session_search`. (2) **Curated memory**: the `memory` tool writes structured entries to `MEMORY.md`/`USER.md` (add, replace, remove, batch operations). These are FROZEN into the system prompt snapshot at session start. (3) **Background review**: after each turn, a background thread (`background_review.py`) evaluates if new information should be persisted as memory or skills, using an LLM review prompted with `_MEMORY_REVIEW_PROMPT` (run_agent.py:785-786).

**Injection into prompts:** The `MemoryProvider` abstract base (`agent/memory_provider.py:84`) defines `prefetch()` (recall context for the upcoming turn), `queue_prefetch()` (queues background recall), `system_prompt_block()` (static instruction text), and `sync_turn()` (persists a completed turn). The built-in provider reads memory files for the system prompt; external plugins (Honcho, etc.) can replace this via `memory.provider` config.

**Summarization:** Compression (`context_compressor.py`, `conversation_compression.py`) runs mid-session and at compression boundaries, summarizing older history to keep the prompt under context limits. The `on_pre_compress()` hook on `MemoryProvider` lets providers extract insights before compression (memory_provider.py:176-177).


Citations: [tools/memory_tool.py:1-47](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/memory_tool.py#L1-L47) · [agent/background_review.py:1-30](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/agent/background_review.py#L1-L30) · [hermes_state.py:1-30](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_state.py#L1-L30)

### How are actions on the user's behalf gated? (answered)

**Approval / human-in-the-loop:** The approval system (tools/approval.py) gates dangerous actions through three entry points: `check_all_command_guards()`, `check_execute_code_guard()`, and `request_tool_approval()`. Each routes through `prompt_dangerous_approval()` which presents the user with a CLI prompt, gateway transport dialog (Telegram inline buttons, etc.), or MCP elicitation. Commands are classified by `detect_dangerous_command()`/`detect_hardline_command()` (approval_detection.py) — patterns like destructive shell commands, sudo usage, or FS modifications. A session-scoped permanent allowlist (`_permanent_approved` set) and per-session approved set (`_session_approved`) let users remember decisions. The denial breaker circuit (`_denial_tally`, default threshold=3) escalates after repeated smart-approval rejections.

**Permission scopes:** `.hermes/config.yaml` supports `approvals` section with `denial_breaker_threshold`, per-tool allowlist entries, and user-defined deny rules (`_match_user_deny_rule`). The `smart_verdict` module (`approval_smart.py`) delegates ambiguous commands to a guardian LLM for classification.

**YOLO mode:** `HERMES_YOLO_MODE` env var and `/yolo` slash command bypass all approval checks entirely (approval.py:46). The session DB row records yolo mode state (`run_agent.py:339-347`) so resumed sessions maintain the setting.

**Dry-run / draft modes:** Not implemented as a generic feature. The terminal tool infrastructure has sandbox backends (Docker, SSH, Modal, etc.) but no global dry-run mode. Individual tools handle safety through the approval gate rather than draft simulation.

**Audit trail:** Every approval decision (approve/deny/yolo) is logged via shared metrics harness (`record_guardrail_decision`, `record_guardrail_warnings`). SQLite session rows capture the model and provider used per session. Background review actions record their write origin.


Citations: [tools/approval.py:1-75](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/approval.py#L1-L75) · [tools/approval_detection.py:1-10](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/approval_detection.py#L1-L10) · [tools/approval_smart.py:1-10](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/tools/approval_smart.py#L1-L10) · [run_agent.py:339-347](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/run_agent.py#L339-L347)

### How are LLM providers selected and configured? (answered)

**Supported providers:** 40+ providers configured in `hermes_cli/providers.py:28-92` via `HERMES_OVERLAYS` dict plus dynamic lookup from models.dev. Includes: OpenAI (native + Codex/Responses), Anthropic (Messages API), Nous Portal, OpenRouter, xAI/Grok, Google Vertex, AWS Bedrock, Azure Foundry, GitHub Copilot, Qwen, DeepSeek, Alibaba/DashScope, LM Studio, Ollama (cloud + local), vLLM, llama.cpp, Zhipu/GLM, Minimax, StepFun, together, Fireworks, Novita, Nebius, and custom endpoints via `custom_providers` config.

**Config surface:** Providers are resolved through a 7-level chain (`resolve_provider_full()` at providers.py:533): user `providers:` config → lossy-alias registry → built-in overlays → user providers (canonical/raw) → custom_providers → managed llamacpp → models.dev direct. Each `ProviderDef` carries id, transport (`openai_chat`, `anthropic_messages`, `codex_responses`, `bedrock_converse`), base URL, auth type (`api_key`, `oauth_device_code`, `oauth_external`, `external_process`, `aws_sdk`, `vertex`), and env vars. The `determine_api_mode()` function (providers.py:371) picks the wire protocol from host-mandated, provider overlay, or dynamic model-based routing.

**Tool-calling / structured output:** The agent uses OpenAI-compatible function calling for all `chat_completions` providers and Anthropic-style tool_use blocks for `anthropic_messages` providers. The `codex_responses` transport uses OpenAI's Responses API with deterministic tool call IDs (`chat_completion_helpers.py`, `codex_responses_adapter.py`). Structured output (JSON mode) is supported through provider-specific `response_format` parameters.

**Local-model support:** LM Studio (`lmstudio` provider, auto-detected at `127.0.0.1:1234/v1`) with explicit or JIT load modes (`ensure_lmstudio_runtime_loaded()` at run_agent.py:482). llama.cpp managed local runtime (`hermes_cli/local_runtime/`) auto-stages GGUF models. Ollama is supported as a custom endpoint alias. vLLM/sglang work through the `custom` provider.


Citations: [hermes_cli/providers.py:28-92](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/providers.py#L28-L92) · [hermes_cli/providers.py:194-210](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/providers.py#L194-L210) · [hermes_cli/providers.py:533-565](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/providers.py#L533-L565) · [run_agent.py:482-496](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/run_agent.py#L482-L496)

### How is it deployed and self-hosted? (answered)

**Runtime dependencies:** Python 3.11-3.14, SQLite (WAL mode, FTS5, built from source in Docker to avoid the WAL-reset bug), Node.js (for web dashboard and NPM-based MCP tools), FFmpeg (audio transcription/transcoding), ripgrep (FTS5 tokenizer), and system packages (libffi, libyaml, OpenSSL). The project bundles its own Python via `uv` managed by the `pm/` (Package Manager) subsystem.

**Docker/one-click paths:** The `Dockerfile` builds from Debian 13 slim with a custom SQLite build (patched WAL-reset) and `s6-overlay` for PID-1 supervision. `docker-compose.yml` defines two services: `gateway` (the agent + messaging gateway) and `dashboard` (web UI on 127.0.0.1:9119). Both mount `~/.hermes` as `/opt/data` and use host networking. Run with `docker compose up -d`. One-liner install via `curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash` or the PowerShell equivalent (`iex (irm https://hermes-agent.nousresearch.com/install.ps1)`).

**Required external accounts:** At minimum an LLM provider API key (any of 40+ supported). For messaging: bot tokens from Telegram/Discord/Slack. For full functionality: Nous Portal subscription (covers models, web search, image gen, TTS, cloud browser under one OAuth flow via `hermes setup --portal`). Optional: FAL.ai (image gen), OpenAI TTS, Browser Use cloud, Microsoft Graph OAuth (email/calendar), Signal registration.

**Other runtimes:** Seven terminal backends (`terminal_tool_backends.py`): local shell, Docker, SSH, Singularity, Modal, Daytona, Vercel Sandbox. Subagent worktrees (`tools/subagent_worktree.py`) provide git-isolated workspaces for background agents. Cron scheduling (`hermes_cli/cron.py`) with delivery to any messaging platform. Nix flake available for NixOS users.


Citations: [Dockerfile:1-50](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/Dockerfile#L1-L50) · [README.md:36-60](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/README.md#L36-L60) · [hermes_cli/providers.py:28-92](https://github.com/NousResearch/hermes-agent/blob/950ea4d51eb5667d0258dbead49cc8be82a6e7bf/hermes_cli/providers.py#L28-L92)
