LLMs Technical Reviews

NousResearch/hermes-agent

Python agent core behind a CLI, desktop app and 20-platform chat gateway, with guarded shell tools, curated memory and self-written skills.

GitHub ↗★ 252kPythonMITcommit 950ea4d · 2026-10-06homepage ↗

Overview

Hermes Agent is Nous Research’s general-purpose personal agent. One Python agent core runs behind a terminal CLI and TUI, a messaging gateway (Telegram, Discord, Slack, WhatsApp, Signal, Teams, email and about fifteen more platforms), a web dashboard, an Electron desktop app, an OpenAI-compatible API server, and an ACP adapter for editors. The agent drives a real terminal, files, a browser and subagents. It keeps curated memory files about you and itself, and it can write its own skills: the “grows with you” in its own description.

The codebase is very large: hundreds of modules in agent/ and tools/, and a state layer split across dozens of hermes_state_*.py files. It is also written defensively. Comments cite issue numbers for nearly every edge case: prompt-cache invariants, credential refresh, gateway reconnects, WAL corruption. Two design rules from the root AGENTS.md explain much of it. Per-conversation prompt caching is “sacred”, so memory and toolsets are frozen for a session. And the core is a “narrow waist”, so new capability is meant to arrive as skills, plugins or MCP servers.

It suits power users who want one always-on agent they can reach from their phone and their terminal. It assumes broad trust: by default it runs shell commands on the host, guarded by an approval layer rather than a sandbox.

Architecture

flowchart LR
  CLI["CLI / TUI"] --> AG["AIAgent"]
  GW["Gateway + platform plugins"] --> CACHE["Agent cache (per session)"]
  CACHE --> AG
  API["API server / ACP / dashboard"] --> AG
  AG --> LOOP["run_conversation loop"]
  LOOP --> PROV["Provider resolution + transports"]
  LOOP --> TR["run_tool_round"]
  TR --> APR["Approval gate"]
  APR --> TOOLS["Tools: terminal, files, browser, MCP, delegate"]
  TOOLS --> ENV["Terminal backends (local, Docker, SSH, Modal...)"]
  LOOP --> DB["SQLite state (WAL + FTS5)"]
  LOOP --> MEM["MEMORY.md / USER.md + provider plugins"]
  LOOP --> BR["Background review (memory + skills)"]
  CRON["Cron scheduler"] --> AG
Component Path Role
Agent class run_agent.py AIAgent, assembled from mixins; constructor forwards to agent/agent_init.py
Turn loop agent/conversation_loop.py run_conversation and the phase-based iteration loop
Tool round agent/turn_tool_round.py Validates, dedupes, persists and dispatches tool calls
Providers hermes_cli/providers.py, agent/*_adapter.py Provider overlays, API-mode selection, Anthropic/Bedrock/Gemini/Codex adapters
Tools tools/, toolsets.py Self-registering tools grouped into toolsets; MCP client; terminal backends in tools/environments/
Approvals tools/approval*.py Dangerous-command detection, smart (guardian LLM) approval, deny rules, YOLO
Gateway gateway/run.py, plugins/platforms/ Long-running process that maps platform messages to cached agents
State hermes_state.py + hermes_state_*.py SQLite sessions, messages, usage, FTS5 search
Memory and skills tools/memory_tool.py, tools/skill_manager_tool.py, agent/background_review.py Curated memory files, agent-written skills, post-turn review fork
Scheduler cron/ Scheduled jobs with delivery back to any platform
Apps apps/desktop/ (Electron), ui-tui/, web/ Desktop, terminal and web front ends

How a request flows

Take a Telegram message, “clean up the Downloads folder on my server”:

  1. Receive. The Telegram platform plugin, which wraps python-telegram-bot and requires TELEGRAM_BOT_TOKEN, hands the update to the gateway. The gateway keeps one AIAgent per session in a cache capped at 128 agents, evicting those idle for over an hour (gateway/run.py).
  2. Build the turn. run_conversation wraps _run_conversation_turn. That function refreshes credentials from .env, then build_turn_context restores or builds the system prompt (identity, frozen memory snapshot, skills index, context files) and loads history from SQLite (conversation_loop.py).
  3. Iterate. The loop runs while api_call_count < max_iterations and the shared IterationBudget has room. Each pass runs explicit phases: begin_iteration, prepare_iteration, assemble_api_request, a preflight gate (compression and context checks), then an API retry loop and normalize_model_response (conversation_loop.py). The constructor default is max_iterations=sys.maxsize (run_agent.py), and agent.max_turns defaults to unset, so turns are unbounded unless you configure a cap.
  4. Run tools. When the model returns tool calls, run_tool_round validates names and arguments, expands batches, dedupes calls, caps delegation, and persists the assistant message before executing anything. The docstring calls this “persist-before-execute”: a crash mid-command can be resumed (turn_tool_round.py).
  5. Gate. The terminal call (rm -rf ~/Downloads/*) passes through check_all_command_guards. Hardline patterns and user approvals.deny globs block outright. Dangerous patterns go to the configured mode. The default smart mode asks a guardian LLM. Escalations send a prompt to the user, here as Telegram inline buttons, with a 300-second timeout (config_defaults.py).
  6. Execute. The command runs on the configured terminal backend: local, docker, singularity, modal, daytona, vercel_sandbox or ssh. The result is appended and the loop continues until the model answers in text.
  7. Finalize. finalize_turn persists the turn and runs post-turn hooks. Every tenth user turn (memory.nudge_interval), a background review fork replays the conversation and decides whether to save memory or skills (turn_context.py).

Key components

The agent object

AIAgent is assembled from 14 mixins, including streaming delivery, interrupt control, rate-limit credits, session persistence and compression. Its constructor has 84 keyword parameters, from provider routing to a dozen UI callbacks, and forwards all of them to init_agent (run_agent.py). The same class serves the CLI, gateway sessions, subagents (delegate_task) and the background-review fork. That is why so much per-turn state is reset explicitly at the top of each turn.

Providers

HERMES_OVERLAYS declares 42 provider overlays: Nous Portal (OAuth device code), OpenRouter, OpenAI via the Codex Responses transport, xAI, Qwen OAuth, LM Studio and many others (providers.py). Anything else resolves through user providers: / custom_providers entries or models.dev metadata. determine_api_mode picks the wire protocol per provider and model: OpenAI chat, Anthropic messages, Codex Responses or Bedrock Converse. Credential pools, fallback chains and a cheaper model for delegated children are all configurable.

Memory and self-improvement

Built-in memory is two bounded Markdown files: MEMORY.md for agent notes (2,200 characters by default) and USER.md for your profile (1,375). The memory tool edits them with add, replace and remove operations. Both are frozen into the system prompt at session start, so writes made mid-session do not invalidate the prompt cache (memory_tool.py). One external MemoryProvider plugin can be added, such as Honcho, mem0, Supermemory or Hindsight (config_defaults.py). Skills are “procedural memory”: SKILL.md folders that the agent creates and edits under ~/.hermes/skills/. The background review forks the agent on the same provider and prompt cache, under a tool whitelist, and writes straight to the memory and skill stores (background_review.py). Set memory.write_approval to stage those writes for review.

Session store

hermes_state.py is a SQLite store in WAL mode with FTS5 search. Sessions are tagged by source (cli, telegram, …), and compression splits a long session into a parent_session_id chain (hermes_state.py). The Docker image builds its own SQLite to avoid the upstream WAL-reset bug.

Approvals

The approval module owns per-session approvals, a permanent allowlist, YOLO state and a denial circuit breaker. After three consecutive guardian denials, the deny message escalates to a hard stop. HERMES_YOLO_MODE is read once at import. The comment explains that reading the environment on every call would let a skill flip it at runtime: a prompt-injection path (approval.py). Cron, single-query and unattended contexts default to deny, because nobody is there to answer.

Extending it

  • Skills first. Bundled skills live in skills/ and opt-in ones in optional-skills/. The agent can install, create and patch skills at runtime.
  • Plugins. plugins/ holds platform adapters (each with a plugin.yaml declaring its env vars), memory providers, model providers, image and video generation, browser backends, observability and more.
  • MCP. tools/mcp_tool*.py is a full MCP client with OAuth, sampling and health supervision. mcp_serve.py exposes Hermes itself as an MCP server.
  • Core tools. These are self-registering modules discovered by tools/registry.py and grouped in toolsets.py. The contribution guide sets a high bar for new core tools, because every tool schema is sent on every call.

Running it

  • Install. Use the install script, or uv/pip with the hermes console script (hermes_cli.main:main). hermes setup configures a provider, and hermes gateway run starts the messaging gateway.
  • Docker. docker-compose.yml runs two containers from one image under s6-overlay: gateway and a dashboard bound to 127.0.0.1:9119. Both use host networking and mount ~/.hermes (docker-compose.yml). The OpenAI-compatible API server stays off until you set API_SERVER_KEY.
  • Requirements. You need Python, at least one model provider (or a local LM Studio, llama.cpp or Ollama endpoint), and bot tokens for each platform you enable. Nix users get a flake.

Strengths and caveats

  • Strength: reach. One agent core serves CLI, TUI, desktop, about 20 chat platforms, cron jobs and an API, with sessions and memory shared through one profile directory.
  • Strength: operational hardening. Persist-before-execute tool rounds, credential refresh, compression with session chains, and a custom SQLite build show real production mileage.
  • Strength: layered approvals. Hardline blocks, deny globs that hold even under YOLO, a guardian model and per-surface defaults are more careful than most agents.
  • Caveat: size and complexity. The core is spread over hundreds of small modules and mixins. Reading a single code path means following phase functions across many files.
  • Caveat: local shell by default. Sandboxing is opt-in through a terminal backend. On the default local backend, the guardian LLM’s judgement sits between the model and your machine.
  • Caveat: unbounded turns. There is no default iteration cap, so long tool loops cost money unless you set agent.max_turns or a run budget.
  • Caveat: small memory. Built-in memory is a few kilobytes of Markdown. Real long-term recall needs session search or an external provider plugin.

Sources: code at 950ea4d, verified Q&A.

How it answers the Open-source personal assistants questions

Each answer was drafted by a code-reading agent at commit 950ea4d. Its citations were checked mechanically. Compare with the other open-source personal assistants →

How is the assistant architected?

answered

Agent loop and runtime: The AIAgent class (run_agent.py:241) is the central orchestrator, composed of ~15 mixins for streaming, compression, approval, activity tracking, etc. Every user turn flows through run_conversation() (conversation_loop.py:1695) which calls _run_conversation_turn() (conversation_loop.py:1533). That function drives: pre-flight compression checks, build_turn_context() (turn_context.py:1017) to assemble messages and system prompt, build_api_request() → perform_api_call() to the LLM, normalize_model_response() to parse tool calls or text, run_tool_round() (turn_tool_round.py:46) to dispatch tools with parallel/segmented execution, check_api_response() for finish-reason handling, and finalize_turn() for post-turn hooks (memory sync, background review). The loop retries on errors up to max_iterations.

Frontend/backend split: The CLI (hermes_cli/) provides the interactive TUI via prompt_toolkit, the messaging gateway (gateway/run.py) runs platform adapters that receive messages and call AIAgent.run_conversation() per incoming message. The gateway caches agents per session with an LRU cap of 128 and 1h idle TTL. A web dashboard serves config, sessions, and cron management. The Desktop app wraps the CLI in Tauri (Rust).

Main packages: agent/ — core loop, compression, memory, providers; tools/ — 40+ tool implementations (terminal, browser, file, MCP, skills, etc.); gateway/ — messaging platform adapters; hermes_cli/ — CLI, config, profiles, auth; hermes_state*.py — SQLite state store.

Request-to-action flow: User message → run_conversation() → context assembly (system prompt + memory recall + history) → LLM API call → response parsing → if tool_calls: _execute_tool_calls() which runs the tool (terminal command, file edit, API call, etc.) → tool result appended to messages → loop back to LLM call until final text response is produced → post-turn memory sync + background review spawn.

Editor's note. Correction: the desktop app in apps/desktop/ is built with Electron, not Tauri; Tauri is only used by the separate bootstrap installer in apps/bootstrap-installer/.

How are integrations (email, calendar, chat, docs) implemented?

answered

Supported services: Telegram, Discord, Slack, WhatsApp Cloud, Signal, email (IMAP), WeChat Work (yuanbao), QQ Bot, Microsoft Teams, Google Chat, BlueBubbles (Apple Messages), and a generic webhook adapter. Each lives in gateway/platforms/ as a subclass of BasePlatformAdapter (gateway/platforms/base.py:1).

API clients vs MCP: All integrations use direct HTTP API clients (Telegram bot API, Discord gateway, WhatsApp Cloud API, Microsoft Graph API, etc.) — there are no MCP-based chat adapters. The project has a separate tools/mcp_tool*.py system for MCP tool servers (not messaging).

OAuth flow: Calendars and email (Microsoft Graph) route through tools/microsoft_graph_auth.py using OAuth 2.0 device-code or authorization-code flows with token refresh. Microsoft Teams gateway auth (TEAMS_CLIENT_ID/TEAMS_CLIENT_SECRET) is configured via env vars in docker-compose.yml:44-50. Credentials persist to ~/.hermes/.env or the SQLite credential store.

Token storage: API keys and OAuth tokens go to the profile-scoped .env file via credential_persistence.py/secret_scope.py. Gateway platform tokens (Telegram bot tokens, Slack signing secrets, etc.) are read from env vars or ~/.hermes/config.yaml.

Sync vs on-demand fetch: Messaging is push-based — platforms deliver webhooks or long-poll updates to the gateway, which fans them out to per-session agent instances. No periodic polling for new messages; the gateway maintains persistent connections (Telegram getUpdates, Discord gateway, Slack Socket Mode, etc.). The cron scheduler (hermes_cli/cron.py) runs periodic jobs that inject prompts into the agent loop.

Editor's note. Correction: most messaging adapters (Telegram, Discord, Slack, WhatsApp, Teams, Google Chat, email, Matrix and others) are plugins under plugins/platforms/<name>/ with a plugin.yaml; gateway/platforms/ holds the base adapter, the API server and a handful of built-in adapters such as Signal, BlueBubbles and WhatsApp Cloud.

How is memory and user context stored and retrieved?

answered

Storage: The primary store is SQLite in WAL mode (hermes_state.py). It stores session metadata, message history, model config, usage tracking, and FTS5-indexed messages. Memory also lives on disk as MEMORY.md (agent-curated notes) and USER.md (user profile) under ~/.hermes/memories/, managed by the memory tool (tools/memory_tool.py:1-6). FTS5 full-text search across sessions is provided by hermes_state_fts.py.

What is remembered: Three tiers — (1) Session history: every message in the current conversation (SQLite, per-session rows), searchable across sessions via /session_search. (2) Curated memory: the memory tool writes structured entries to MEMORY.md/USER.md (add, replace, remove, batch operations). These are FROZEN into the system prompt snapshot at session start. (3) Background review: after each turn, a background thread (background_review.py) evaluates if new information should be persisted as memory or skills, using an LLM review prompted with _MEMORY_REVIEW_PROMPT (run_agent.py:785-786).

Injection into prompts: The MemoryProvider abstract base (agent/memory_provider.py:84) defines prefetch() (recall context for the upcoming turn), queue_prefetch() (queues background recall), system_prompt_block() (static instruction text), and sync_turn() (persists a completed turn). The built-in provider reads memory files for the system prompt; external plugins (Honcho, etc.) can replace this via memory.provider config.

Summarization: Compression (context_compressor.py, conversation_compression.py) runs mid-session and at compression boundaries, summarizing older history to keep the prompt under context limits. The on_pre_compress() hook on MemoryProvider lets providers extract insights before compression (memory_provider.py:176-177).

How are actions on the user's behalf gated?

answered

Approval / human-in-the-loop: The approval system (tools/approval.py) gates dangerous actions through three entry points: check_all_command_guards(), check_execute_code_guard(), and request_tool_approval(). Each routes through prompt_dangerous_approval() which presents the user with a CLI prompt, gateway transport dialog (Telegram inline buttons, etc.), or MCP elicitation. Commands are classified by detect_dangerous_command()/detect_hardline_command() (approval_detection.py) — patterns like destructive shell commands, sudo usage, or FS modifications. A session-scoped permanent allowlist (_permanent_approved set) and per-session approved set (_session_approved) let users remember decisions. The denial breaker circuit (_denial_tally, default threshold=3) escalates after repeated smart-approval rejections.

Permission scopes: .hermes/config.yaml supports approvals section with denial_breaker_threshold, per-tool allowlist entries, and user-defined deny rules (_match_user_deny_rule). The smart_verdict module (approval_smart.py) delegates ambiguous commands to a guardian LLM for classification.

YOLO mode: HERMES_YOLO_MODE env var and /yolo slash command bypass all approval checks entirely (approval.py:46). The session DB row records yolo mode state (run_agent.py:339-347) so resumed sessions maintain the setting.

Dry-run / draft modes: Not implemented as a generic feature. The terminal tool infrastructure has sandbox backends (Docker, SSH, Modal, etc.) but no global dry-run mode. Individual tools handle safety through the approval gate rather than draft simulation.

Audit trail: Every approval decision (approve/deny/yolo) is logged via shared metrics harness (record_guardrail_decision, record_guardrail_warnings). SQLite session rows capture the model and provider used per session. Background review actions record their write origin.

How are LLM providers selected and configured?

answered

Supported providers: 40+ providers configured in hermes_cli/providers.py:28-92 via HERMES_OVERLAYS dict plus dynamic lookup from models.dev. Includes: OpenAI (native + Codex/Responses), Anthropic (Messages API), Nous Portal, OpenRouter, xAI/Grok, Google Vertex, AWS Bedrock, Azure Foundry, GitHub Copilot, Qwen, DeepSeek, Alibaba/DashScope, LM Studio, Ollama (cloud + local), vLLM, llama.cpp, Zhipu/GLM, Minimax, StepFun, together, Fireworks, Novita, Nebius, and custom endpoints via custom_providers config.

Config surface: Providers are resolved through a 7-level chain (resolve_provider_full() at providers.py:533): user providers: config → lossy-alias registry → built-in overlays → user providers (canonical/raw) → custom_providers → managed llamacpp → models.dev direct. Each ProviderDef carries id, transport (openai_chat, anthropic_messages, codex_responses, bedrock_converse), base URL, auth type (api_key, oauth_device_code, oauth_external, external_process, aws_sdk, vertex), and env vars. The determine_api_mode() function (providers.py:371) picks the wire protocol from host-mandated, provider overlay, or dynamic model-based routing.

Tool-calling / structured output: The agent uses OpenAI-compatible function calling for all chat_completions providers and Anthropic-style tool_use blocks for anthropic_messages providers. The codex_responses transport uses OpenAI's Responses API with deterministic tool call IDs (chat_completion_helpers.py, codex_responses_adapter.py). Structured output (JSON mode) is supported through provider-specific response_format parameters.

Local-model support: LM Studio (lmstudio provider, auto-detected at 127.0.0.1:1234/v1) with explicit or JIT load modes (ensure_lmstudio_runtime_loaded() at run_agent.py:482). llama.cpp managed local runtime (hermes_cli/local_runtime/) auto-stages GGUF models. Ollama is supported as a custom endpoint alias. vLLM/sglang work through the custom provider.

How is it deployed and self-hosted?

answered

Runtime dependencies: Python 3.11-3.14, SQLite (WAL mode, FTS5, built from source in Docker to avoid the WAL-reset bug), Node.js (for web dashboard and NPM-based MCP tools), FFmpeg (audio transcription/transcoding), ripgrep (FTS5 tokenizer), and system packages (libffi, libyaml, OpenSSL). The project bundles its own Python via uv managed by the pm/ (Package Manager) subsystem.

Docker/one-click paths: The Dockerfile builds from Debian 13 slim with a custom SQLite build (patched WAL-reset) and s6-overlay for PID-1 supervision. docker-compose.yml defines two services: gateway (the agent + messaging gateway) and dashboard (web UI on 127.0.0.1:9119). Both mount ~/.hermes as /opt/data and use host networking. Run with docker compose up -d. One-liner install via curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash or the PowerShell equivalent (iex (irm https://hermes-agent.nousresearch.com/install.ps1)).

Required external accounts: At minimum an LLM provider API key (any of 40+ supported). For messaging: bot tokens from Telegram/Discord/Slack. For full functionality: Nous Portal subscription (covers models, web search, image gen, TTS, cloud browser under one OAuth flow via hermes setup --portal). Optional: FAL.ai (image gen), OpenAI TTS, Browser Use cloud, Microsoft Graph OAuth (email/calendar), Signal registration.

Other runtimes: Seven terminal backends (terminal_tool_backends.py): local shell, Docker, SSH, Singularity, Modal, Daytona, Vercel Sandbox. Subagent worktrees (tools/subagent_worktree.py) provide git-isolated workspaces for background agents. Cron scheduling (hermes_cli/cron.py) with delivery to any messaging platform. Nix flake available for NixOS users.