LLMs Technical Reviews

plastic-labs/honcho

Memory server where background LLM workers distil peer messages into conclusions that agents query through a chat endpoint.

GitHub ↗★ 7.5kPythonAGPL-3.0commit 11b22bf · 2026-10-06homepage ↗

Overview

Honcho is Plastic Labs’ memory server for agents. You run it as a FastAPI service backed by Postgres with pgvector. Your application writes conversation messages to it, and you later ask it questions about the participants. It is not an in-process library. The Python and TypeScript SDKs in sdks/ are HTTP clients for the /v3 REST API.

The data model is the main idea. Everything in a workspace is a peer, and humans and agents are treated the same way. Peers exchange messages in sessions. Memory is stored per (observer, observed) pair, not per user. Each pair owns a Collection of Document rows. The code calls these rows observations, and the API calls them conclusions. By default every peer observes only itself. A peer with observe_others enabled in a session also builds its own model of the other peers in that session. So “what Alice believes about Bob” is stored separately from Honcho’s own view of Bob.

The work happens in the background. A separate deriver process reads a Postgres-backed queue, runs one structured LLM call per batch of messages to extract atomic facts, and deduplicates them on write. A dreamer later runs tool-using agents that write deductive and inductive conclusions on top of those facts. At read time, the dialectic agent answers a natural-language question with a tool loop over conclusions and raw messages. Honcho is aimed at teams that want a long-lived user model shared by several agents, and that can run Postgres and two Python processes for it.

Architecture

flowchart LR
  APP["App / SDK / MCP"] --> API["FastAPI /v3 routers"]
  API --> MSG["crud.create_messages"]
  API --> CHAT["peers.chat route"]
  MSG --> PG[("Postgres + pgvector")]
  MSG --> ENQ["deriver.enqueue"]
  ENQ --> Q[("queue table")]
  Q --> QM["QueueManager (deriver process)"]
  QM --> DER["minimal deriver: one LLM call"]
  QM --> SUM["summarizer"]
  QM --> DRM["dreamer: deduction + induction agents"]
  DER --> RM["RepresentationManager"]
  DRM --> RM
  RM --> DOCS["documents per observer/observed"]
  DOCS --> PG
  DOCS -.-> VS["optional external vector store"]
  CHAT --> DIA["DialecticAgent tool loop"]
  DIA --> DOCS
  DIA --> SRCH["hybrid message search (RRF)"]
  SRCH --> PG
Component Path Role
API app src/main.py, src/routers/ FastAPI app; workspaces, peers, sessions, scopes, messages, conclusions, keys and webhooks under /v3
Data model src/models.py Workspace, Peer, Session, Message, Collection, Document, DocumentSource, QueueItem tables
Enqueue src/deriver/enqueue.py Turns new messages into representation, summary, dream and deletion queue rows
Queue worker src/deriver/queue_manager.py, consumer.py Claims work units from Postgres, batches messages by token budget, dispatches by task type
Deriver src/deriver/deriver.py, prompts.py One structured-output LLM call per batch that extracts explicit facts
Representation store src/crud/representation.py, src/crud/document.py Embeds, deduplicates and saves conclusions; builds the “working representation” for reads
Dreamer src/dreamer/ Scheduled consolidation: deduction and induction specialists, optional surprisal sampling
Dialectic src/dialectic/core.py, chat.py Tool-using agent behind POST /peers/{id}/chat
Agent tools src/utils/agent_tools.py Tool schemas and handlers shared by the dialectic and dreamer agents
Search src/utils/search.py Semantic plus full-text message search fused with Reciprocal Rank Fusion
Vector stores src/vector_store/ pgvector by default; Turbopuffer, LanceDB, Qdrant, ChromaDB adapters
LLM layer src/llm/ Provider-neutral call wrapper with OpenAI, Anthropic and Gemini backends and a tool loop
Clients sdks/python, sdks/typescript, mcp/, honcho-cli/ SDKs, MCP server, CLI that can start a local Docker stack

How a request flows

Two paths matter: writing a message, and asking a question.

Write path

  1. POST /v3/workspaces/{w}/sessions/{s}/messages calls crud.create_messages. It then schedules enqueue(payloads) as a FastAPI background task, plus an immediate message-embedding task when EMBED_MESSAGES is on (messages.py).
  2. enqueue first cancels any pending dreams for the sender, because the peer is active again. handle_session then resolves the configuration hierarchy (message, then session, then workspace) and inserts QueueItem rows (enqueue.py).
  3. generate_queue_records adds a summary task every N messages. It also builds one representation task whose observers list holds the sender (if observe_me is on) and every other active peer with observe_others (enqueue.py).
  4. In the deriver process, QueueManager.get_and_claim_work_units claims a representation work unit only when its unprocessed messages total at least REPRESENTATION_BATCH_WORK_UNIT_TARGET_TOKENS (default 512) or the oldest item is older than REPRESENTATION_BATCH_MAX_AGE_SECONDS (default 1800), unless FLUSH_ENABLED is set (queue_manager.py, config.py).
  5. process_work_unit drains the unit in token-capped batches and calls process_representation_batch (queue_manager.py).
  6. process_representation_tasks_batch wraps each message in a <message ... target="true|false"> tag and builds minimal_deriver_prompt. It makes a single honcho_llm_call with response_model=PromptRepresentation (deriver.py, prompts.py).
  7. The same observations are saved into every observer’s collection through RepresentationManager.save_representation (deriver.py). That method batch-embeds the texts, calls crud.create_documents with deduplication, and then checks whether a dream is due (representation.py).

Read path

  1. POST /v3/workspaces/{w}/peers/{p}/chat validates scope and session options, checks that a peer-scoped key may read the named sessions, and calls agentic_chat (or agentic_chat_stream for SSE). The observer is the path peer, and the observed peer is target or the observer itself (peers.py).
  2. agentic_chat loads the configuration and both peer cards, closes its DB session, and builds a DialecticAgent (chat.py).
  3. The agent prefetches relevant observations: separate searches for explicit and higher-level conclusions, 25 of each, or 10 at minimal. With a session id, it also injects recent session history into the system prompt (core.py).
  4. answer runs honcho_llm_call with the level’s tool set, tool choice and iteration cap, then returns plain text, or JSON when the caller passed a response_format (core.py).

Key components

Deriver (extraction)

The deriver is intentionally small. It makes one LLM call per batch and asks only for explicit facts. The prompt tells the model to extract facts only from target="true" messages and to use other peers’ messages only as context. It also asks for absolute dates and resolved pronouns, and it escapes <message inside content so a message cannot forge its own tag (prompts.py). Workspace or session custom_instructions are inserted into this prompt. Peer cards, deductions and summaries are handled elsewhere.

Deduplication and reinforcement

create_documents always runs exact-match deduplication on normalized content, level and session. A duplicate inside the batch is dropped. A duplicate of an existing row increments that row’s times_derived instead of inserting a new one (document.py). With DERIVER.DEDUPLICATE (on by default), it also looks up near neighbours. _semantic_dup_decision then keeps whichever text carries more unique tokens, so it either replaces the old row or rejects the new one (document.py). times_derived is later used as a retrieval signal.

Storage

A Collection is unique on (observer, observed, workspace_name). Composite foreign keys tie both names to real peers (models.py). A Document carries level (explicit, deductive, inductive, contradiction), times_derived, a pgvector embedding, a soft-delete deleted_at, and a sync_state for external vector stores (models.py). Embeddings use an HNSW cosine index. Reasoning-tree edges live in document_sources, which deliberately has no foreign key on source_id (models.py). The defaults are OpenAI text-embedding-3-small at 1536 dimensions.

Working representation

_get_working_representation_internal fills a budget of up to 100 observations. It takes a semantic slice (default a third of the budget), then a most-reinforced slice ordered by times_derived, and then fills the remainder with the most recent documents. An empty session allowlist returns nothing instead of everything (representation.py).

Dialectic agent

Reasoning levels map to tool sets and iteration caps. minimal gets only search_memory and search_messages. The other levels add get_observation_context, grep_messages, two date-based message tools and get_reasoning_chain (agent_tools.py). Under a session allowlist, get_reasoning_chain is removed, because reasoning chains cross session boundaries (core.py). The default iteration caps are 1, 5, 2, 4 and 10 for minimal, low, medium, high and max. Every level defaults to gpt-5.4-mini over the OpenAI transport (config.py). So by default, “higher” levels differ mostly in tool rounds and tool choice, not in model.

search combines pgvector (or external-store) similarity with Postgres full-text search and merges the two ranked lists with reciprocal_rank_fusion (k=60) (search.py). A peer_perspective filter limits results to sessions the peer was in, and to the time window of its membership (search.py).

Dreamer

check_and_schedule_dream counts only explicit documents, so dream output cannot trigger more dreams. It schedules a dream after 50 new explicit facts, at least 8 hours after the last dream, delayed by an idle timeout (dream_scheduler.py). run_dream runs the deduction specialist, then the induction specialist. Each is a tool loop with a hardcoded cap, 12 rounds for deduction and 10 for induction (specialists.py, L740-L745). Both can search memory and messages and create conclusions. Only deduction can delete conclusions and rewrite the peer card (agent_tools.py). Surprisal sampling over k-d trees and other tree types can seed both with “hints”, but SURPRISAL.ENABLED defaults to false (orchestrator.py, config.py).

Scoping and auth

JWT claims form a hierarchy: admin, workspace w, peer p, session s (security.py). Scopes group sessions into a recall boundary. Under a session allowlist, only explicit conclusions are served. The reason is that dreamer output reads across all sessions but is stamped with a single session, and the code says so (representation.py).

Extending it

  • Configuration hierarchy. Reasoning, summary, dream and peer-card settings resolve from message to session to workspace. custom_instructions steers the deriver and the dreamer without code changes.
  • Models. Each stage (deriver, each dialectic level, summary, deduction, induction, embeddings) has its own MODEL_CONFIG with a transport of openai, anthropic or gemini. These can be set in config.toml or through DERIVER_, DIALECTIC_LEVELS__<level>__ and similar environment variables. An overrides block sets a custom base_url and key, so OpenAI-compatible local endpoints work through the openai transport.
  • Vector store. Set VECTOR_STORE_TYPE to turbopuffer, lancedb, qdrant or chromadb, and implement VectorStore (upsert_many, query, delete_many, delete_namespace) for anything else (vector_store/init.py). Namespaces are hashed per (workspace, observer, observed).
  • Structured answers. Pass response_format (JSON Schema) to chat. It is converted to a Pydantic model, and the answer comes back as JSON.
  • Webhooks and clients. Webhook endpoints receive queue-driven deliveries. The mcp/ server runs as a Cloudflare Worker, over stdio or over Streamable HTTP. session.context() in both SDKs renders history plus summary for to_openai / to_anthropic. Coding-agent plugins (Claude Code, Codex, Cursor, OpenCode, OpenClaw, Hermes) are separate packages built on the same API.

Running it

  • Required services. Postgres with the pgvector extension, plus an API key for the configured provider (OpenAI by default, for both chat models and embeddings). Redis is optional and serves as a cache.
  • Two processes. The API (fastapi on src/main.py) and the deriver (python -m src.deriver). Both share the database, which also serves as the job queue. docker-compose.yml.example wires api, deriver, database and redis, and binds ports to localhost.
  • CLI. honcho start in honcho-cli brings up the same four services from a published image and can run a config wizard (--setup basic|advanced).
  • Auth. AUTH_USE_AUTH defaults to false. Set it and AUTH_JWT_SECRET before exposing the API (config.py).

Strengths and caveats

  • Strength: a real multi-party model. Observer/observed collections and per-session observe_others allow different agents to hold different views of the same person. Few memory systems store this.
  • Strength: cheap, careful ingestion. One structured call per token-budgeted batch, target-only extraction, injection-safe tagging, and exact plus semantic deduplication with a reinforcement counter.
  • Strength: honest scoping. Empty allowlists fail closed, and synthesized conclusions are excluded from session-scoped reads. The code says openly that per-conclusion provenance is not there yet.
  • Strength: operational depth. A Postgres queue with work-unit ownership, retries, a reconciler for vector sync and soft-delete cleanup, plus Prometheus metrics and Sentry.
  • Caveat: memory is not immediate. Facts appear only after a work unit reaches 512 tokens or 30 minutes pass, unless you enable FLUSH_ENABLED. Dreams need 50 facts and an 8-hour gap.
  • Caveat: every read costs an LLM call. Chat is an agent loop, so latency and cost scale with the reasoning level. The default levels share one model, so “high” means more tool rounds, not a stronger model.
  • Caveat: heavy footprint. You run Postgres, pgvector and two Python services, and the codebase is large (for example, a 2,800-line agent_tools.py). This is infrastructure, not a drop-in library.
  • Caveat: insecure by default, and AGPL. Auth is off until you configure a JWT secret. The AGPL-3.0 license matters if you modify Honcho and offer it as a network service.

Sources: code at 11b22bf, verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit 11b22bf. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Memory extraction is driven by the Deriver (src/deriver/deriver.py), which processes batches of incoming messages via a single structured-output LLM call — the 'minimal deriver' approach. Messages are enqueued by src/deriver/enqueue.py on message creation, consumed by src/deriver/queue_manager.py, and dispatched through consumer.py's process_item() to process_representation_tasks_batch(). The prompt (minimal_deriver_prompt in src/deriver/prompts.py) wraps each message in an XML tag with target="true" or target="false" and instructs the model to extract only explicit atomic facts about the target peer from target="true" messages; other messages serve solely as interpretive context. The output is a PromptRepresentation Pydantic schema enforcing structured JSON with an explicit list of ExplicitObservationBase objects. The Deriver then converts these into a Representation and saves them via RepresentationManager.save_representation() (src/crud/representation.py:74). Deduplication at write time is two-phase: exact-content dedup (normalized content + level + session for explicit documents) drops identical text within a batch or against existing documents, incrementing the times_derived counter on the existing row instead of inserting a duplicate (src/crud/document.py:677-705). Semantic dedup (cosine distance < 0.05) optionally replaces or rejects near-duplicate text via token-set comparison (src/crud/document.py:1369-1422). Custom instructions per workspace/peer can be threaded into the extraction prompt via reasoning configuration, capped at DERIVER__MAX_CUSTOM_INSTRUCTIONS_TOKENS (default 2000).

How are memories stored?

answered

Memories (called observations internally, conclusions in the API) are stored as Document rows in Postgres (src/models.py:390-521), each with fields: content (the fact text), level (explicit / deductive / inductive / contradiction), embedding (pgvector column at configured dimensionality), times_derived (integer reinforcement counter), session_name, observer, observed, workspace_name, and deleted_at for soft-delete. Documents belong to a Collection (src/models.py:342-382) uniquely keyed by (observer, observed, workspace_name) — so every observer+observed pair gets its own collection, enabling self-representation (observer == observed) and cross-peer modeling. The Documents table has an HNSW index on the embedding column for cosine-distance ANN search (src/models.py:511-519). An orthogonal MessageEmbedding table (src/models.py:284-335) stores chunk-level embeddings of raw messages with sync_state tracking, decoupled from message creation. Embeddings are generated via embedding_client (src/embedding_client.py) using configurable providers (OpenAI by default for embeddings). The pgvector path embeds + queries directly in Postgres. For external vector stores, the VectorStore ABC (src/vector_store/__init__.py:53-195) defines an interface with upsert_many, query, delete_many, and delete_namespace. Implementations include Turbopuffer (src/vector_store/turbopuffer.py), LanceDB (src/vector_store/lancedb.py), Qdrant (src/vector_store/qdrant.py), and ChromaDB (src/vector_store/chroma.py). Namespaces isolate collections per (type, workspace, observer, observed) with a hashed suffix. Both stores are dual-written to in migration mode (settings.VECTOR_STORE.MIGRATED); a sync_state column (pending/synced/failed) tracks consistency and a reconciliation backstop heals missed writes.

How are memories retrieved and injected into the prompt?

answered

Retrieval is multi-strategy. For conclusions (the memory store), RepresentationManager._get_working_representation_internal() (src/crud/representation.py:317-411) blends up to three query strategies within a configurable token budget: (1) semantic search via _query_documents_semantic() — cosine-distance ANN through the pgvector HNSW index or external vector store, filtered by level and session allowlist; (2) most-derived — documents sorted by times_derived descending, surfacing the most reinforced facts; (3) recent — by created_at descending as a fill-in for unused capacity. The allocation splits ~1/3 semantic, ~1/3 most-derived, ~1/3 recent, with over/underflow reclaim. For message search, src/utils/search.py:317-458 implements a hybrid search combining semantic (pgvector/ANN or external vector store) and full-text search (Postgres GIN-indexed to_tsvector('english', content) with ILIKE fallback) fused via Reciprocal Rank Fusion (RRF, search.py:36-75). The Dialectic agent (src/dialectic/core.py and src/dialectic/chat.py) is the primary retrieval consumer: it runs a tool loop armed with 7 tools from src/utils/agent_tools.py — search_memory (semantic conclusion search), search_messages (hybrid message search), get_observation_context, grep_messages, get_messages_by_date_range, search_messages_temporal, and get_reasoning_chain (tree traversal from a conclusion to its premises and derived conclusions). The agent iteratively selects tools to gather context and synthesizes a response once it has enough evidence. At the minimal reasoning level, only search_memory and search_messages are available. Results are presented to the LLM as formatted observations with timestamps, IDs, and derivation metadata.

How are memories updated, consolidated or forgotten?

answered

Memory lifecycle has four phases. Creation: The Deriver extracts observations on message creation; each observation gets times_derived=1. On repeat writes, instead of inserting duplicates the times_derived counter is incremented (src/crud/document.py:693-695), strengthening the fact's reinforcement signal. Consolidation via Dreaming: The Dreamer (src/dreamer/orchestrator.py) runs scheduled 'dreams' using two specialist agents. The DeductionSpecialist (src/dreamer/specialists.py) reads existing explicit observations and produces deductive conclusions via a tool-using agent loop (tools: get_recent_observations, search_memory, search_messages, create_observations_deductive, delete_observations, update_peer_card). The InductionSpecialist produces inductive generalizations (patterns, preferences, personality) from both explicit and deductive conclusions. Surprisal-based prioritization (src/dreamer/surprisal.py) pre-filters observations: it builds a tree structure from embeddings, computes geometric surprisal scores ranking how anomalous or informative each observation is, and feeds the highest-surprisal observations as hints to the specialist agents — focusing the expensive dream cycle on the most 'interesting' or novel findings. Reasoning trees (src/dreamer/trees/) link each conclusion to its premises via document_sources table and source_ids references, enabling get_reasoning_chain traversal. Deletion and cleanup: Documents are soft-deleted via deleted_at timestamp (src/crud/document.py:901-948). The ReconcilerScheduler (src/reconciler/scheduler.py) periodically runs cleanup_soft_deleted_documents() which hard-deletes rows soft-deleted >5 minutes ago and removes their vectors from the external store. The reconciler also heals missed vector syncs and cleans stale queue items. Summarization: Parallel to memory extraction, the summarizer.py creates two-tier session summaries (short every ~20 messages, long every ~60) stored on the session metadata, providing a compressing lifecyle for long-running conversations.

Editor's note. Correction: surprisal sampling is disabled by default (DREAM.SURPRISAL.ENABLED=false), and src/dreamer/trees/ holds the spatial trees used for surprisal scoring; conclusion-to-premise links are stored in the document_sources table.

How is memory scoped and isolated?

answered

Memory scoping has four layers. Workspace isolation: Nearly every table includes workspace_name as a composite foreign key, making cross-workspace data leakage structurally impossible at the schema level (src/models.py — all Document, Message, Collection, QueueItem, etc. tables). This is enforced by the FK constraints themselves, not by query filters. Observer/Observed (Collection) isolation: Documents live in collections keyed by (workspace_name, observer, observed). This means every peer pair has its own vector namespace — peer A's view of peer B is a separate collection from C's view of B. Self-representation (observer == observed) is just the diagonal of this matrix. All CRUD operations query through this triple-key filter (src/crud/document.py:81-87). Authentication scoping via JWT: src/security.py:31-63 defines JWTParams with a hierarchy — admin (ad), workspace (w), peer (p), session (s). A narrower token never falls back to wider access: {w: ws, p: alice} only acts on alice, never on a sibling peer. The require_auth FastAPI dependency (security.py:143-194) checks these claims per-route. Session-scoped peer access is gated by allow_member_read=True exclusively on read-only routes, enforced by an explicit allowlist in tests (tests/routes/test_auth_route_policy.py). Scopes: A Scope is a named grouping of sessions that acts as a visibility boundary on recall. Chat, representation, session context, and workspace search answered through a scope see only the contents of that scope's member sessions (CLAUDE.md:126-131). Adding a session to a scope copies its explicit conclusions into the scope's collection; removing one reconciles copies back out. Session allowlists (session_allowlist parameter on queries) restrict retrieval of both conclusions and messages to specified sessions — and only explicit-level conclusions (level "explicit") are considered safe to scope under this mechanism (src/utils/representation.py:24), because higher-level consolidated observations may have been synthesized from sources outside the allowlisted sessions. An empty allowlist fails closed.

Editor's note. Correction: workspace isolation comes from workspace_name filters in the CRUD queries; the composite foreign keys enforce referential integrity, not access isolation.

How do agents integrate with it, and what is self-hostable?

answered

Honcho integrates via multiple channels. REST API: The core FastAPI server exposes /v3/* routes covering workspaces, peers, sessions, messages, conclusions, keys, and webhooks (src/routers/). The chat endpoint at POST /v3/.../peers/{peer_id}/chat is the main memory-query surface, returning reasoning-grounded natural-language answers. SDKs: Two first-party SDKs live in this repo — sdks/python/ (published as honcho-ai on PyPI) and sdks/typescript/ (published as @honcho-ai/sdk on npm). Both expose Honcho(workspace_id, api_key) clients with methods for peer/session/message CRUD, chat, search, and context injection with .to_openai()/.to_anthropic() formatters. MCP: The platform hosts an MCP HTTP server at mcp.honcho.dev (or self-hosted), usable by any MCP client via claude mcp add honcho --transport http. Agent framework plugins: First-party memory plugins ship for Claude Code (marketplace plugin), Codex (npm install), Cursor (shell installer), OpenCode, OpenClaw, Hermes (built-in upstream), and DeepSeek Harness (Cordis plugin). All read the same ~/.honcho/config.json. Webhooks: src/webhooks/webhook_delivery.py delivers event-driven webhook payloads as configured (src/webhooks/events.py), enabling external service integration. Self-hosting: The entire stack is open-source (AGPL-3.0) and self-hostable via honcho start (Docker Compose-based CLI) or manual uv run fastapi dev src/main.py + uv run python -m src.deriver. Required services: Postgres with pgvector and an LLM provider API key (Gemini for Deriver by default; Anthropic for Dialectic medium+; OpenAI for embeddings). Optional: Redis (cache), external vector stores (Turbopuffer, LanceDB, Qdrant, ChromaDB). The API and deriver are separate processes sharing a database.

Editor's note. Correction: at this commit every LLM stage (deriver, all dialectic levels, summary, dream specialists) defaults to OpenAI gpt-5.4-mini, and embeddings to OpenAI text-embedding-3-small; Gemini and Anthropic are optional transports, not defaults.