plastic-labs/honcho
Memory server where background LLM workers distil peer messages into conclusions that agents query through a chat endpoint.
Overview
Honcho is Plastic Labs’ memory server for agents. You run it as a FastAPI service backed by Postgres with pgvector. Your application writes conversation messages to it, and you later ask it questions about the participants. It is not an in-process library. The Python and TypeScript SDKs in sdks/ are HTTP clients for the /v3 REST API.
The data model is the main idea. Everything in a workspace is a peer, and humans and agents are treated the same way. Peers exchange messages in sessions. Memory is stored per (observer, observed) pair, not per user. Each pair owns a Collection of Document rows. The code calls these rows observations, and the API calls them conclusions. By default every peer observes only itself. A peer with observe_others enabled in a session also builds its own model of the other peers in that session. So “what Alice believes about Bob” is stored separately from Honcho’s own view of Bob.
The work happens in the background. A separate deriver process reads a Postgres-backed queue, runs one structured LLM call per batch of messages to extract atomic facts, and deduplicates them on write. A dreamer later runs tool-using agents that write deductive and inductive conclusions on top of those facts. At read time, the dialectic agent answers a natural-language question with a tool loop over conclusions and raw messages. Honcho is aimed at teams that want a long-lived user model shared by several agents, and that can run Postgres and two Python processes for it.
Architecture
flowchart LR
APP["App / SDK / MCP"] --> API["FastAPI /v3 routers"]
API --> MSG["crud.create_messages"]
API --> CHAT["peers.chat route"]
MSG --> PG[("Postgres + pgvector")]
MSG --> ENQ["deriver.enqueue"]
ENQ --> Q[("queue table")]
Q --> QM["QueueManager (deriver process)"]
QM --> DER["minimal deriver: one LLM call"]
QM --> SUM["summarizer"]
QM --> DRM["dreamer: deduction + induction agents"]
DER --> RM["RepresentationManager"]
DRM --> RM
RM --> DOCS["documents per observer/observed"]
DOCS --> PG
DOCS -.-> VS["optional external vector store"]
CHAT --> DIA["DialecticAgent tool loop"]
DIA --> DOCS
DIA --> SRCH["hybrid message search (RRF)"]
SRCH --> PG
| Component | Path | Role |
|---|---|---|
| API app | src/main.py, src/routers/ |
FastAPI app; workspaces, peers, sessions, scopes, messages, conclusions, keys and webhooks under /v3 |
| Data model | src/models.py |
Workspace, Peer, Session, Message, Collection, Document, DocumentSource, QueueItem tables |
| Enqueue | src/deriver/enqueue.py |
Turns new messages into representation, summary, dream and deletion queue rows |
| Queue worker | src/deriver/queue_manager.py, consumer.py |
Claims work units from Postgres, batches messages by token budget, dispatches by task type |
| Deriver | src/deriver/deriver.py, prompts.py |
One structured-output LLM call per batch that extracts explicit facts |
| Representation store | src/crud/representation.py, src/crud/document.py |
Embeds, deduplicates and saves conclusions; builds the “working representation” for reads |
| Dreamer | src/dreamer/ |
Scheduled consolidation: deduction and induction specialists, optional surprisal sampling |
| Dialectic | src/dialectic/core.py, chat.py |
Tool-using agent behind POST /peers/{id}/chat |
| Agent tools | src/utils/agent_tools.py |
Tool schemas and handlers shared by the dialectic and dreamer agents |
| Search | src/utils/search.py |
Semantic plus full-text message search fused with Reciprocal Rank Fusion |
| Vector stores | src/vector_store/ |
pgvector by default; Turbopuffer, LanceDB, Qdrant, ChromaDB adapters |
| LLM layer | src/llm/ |
Provider-neutral call wrapper with OpenAI, Anthropic and Gemini backends and a tool loop |
| Clients | sdks/python, sdks/typescript, mcp/, honcho-cli/ |
SDKs, MCP server, CLI that can start a local Docker stack |
How a request flows
Two paths matter: writing a message, and asking a question.
Write path
POST /v3/workspaces/{w}/sessions/{s}/messagescallscrud.create_messages. It then schedulesenqueue(payloads)as a FastAPI background task, plus an immediate message-embedding task whenEMBED_MESSAGESis on (messages.py).enqueuefirst cancels any pending dreams for the sender, because the peer is active again.handle_sessionthen resolves the configuration hierarchy (message, then session, then workspace) and insertsQueueItemrows (enqueue.py).generate_queue_recordsadds a summary task every N messages. It also builds one representation task whoseobserverslist holds the sender (ifobserve_meis on) and every other active peer withobserve_others(enqueue.py).- In the deriver process,
QueueManager.get_and_claim_work_unitsclaims a representation work unit only when its unprocessed messages total at leastREPRESENTATION_BATCH_WORK_UNIT_TARGET_TOKENS(default 512) or the oldest item is older thanREPRESENTATION_BATCH_MAX_AGE_SECONDS(default 1800), unlessFLUSH_ENABLEDis set (queue_manager.py, config.py). process_work_unitdrains the unit in token-capped batches and callsprocess_representation_batch(queue_manager.py).process_representation_tasks_batchwraps each message in a<message ... target="true|false">tag and buildsminimal_deriver_prompt. It makes a singlehoncho_llm_callwithresponse_model=PromptRepresentation(deriver.py, prompts.py).- The same observations are saved into every observer’s collection through
RepresentationManager.save_representation(deriver.py). That method batch-embeds the texts, callscrud.create_documentswith deduplication, and then checks whether a dream is due (representation.py).
Read path
POST /v3/workspaces/{w}/peers/{p}/chatvalidates scope and session options, checks that a peer-scoped key may read the named sessions, and callsagentic_chat(oragentic_chat_streamfor SSE). The observer is the path peer, and the observed peer istargetor the observer itself (peers.py).agentic_chatloads the configuration and both peer cards, closes its DB session, and builds aDialecticAgent(chat.py).- The agent prefetches relevant observations: separate searches for explicit and higher-level conclusions, 25 of each, or 10 at
minimal. With a session id, it also injects recent session history into the system prompt (core.py). answerrunshoncho_llm_callwith the level’s tool set, tool choice and iteration cap, then returns plain text, or JSON when the caller passed aresponse_format(core.py).
Key components
Deriver (extraction)
The deriver is intentionally small. It makes one LLM call per batch and asks only for explicit facts. The prompt tells the model to extract facts only from target="true" messages and to use other peers’ messages only as context. It also asks for absolute dates and resolved pronouns, and it escapes <message inside content so a message cannot forge its own tag (prompts.py). Workspace or session custom_instructions are inserted into this prompt. Peer cards, deductions and summaries are handled elsewhere.
Deduplication and reinforcement
create_documents always runs exact-match deduplication on normalized content, level and session. A duplicate inside the batch is dropped. A duplicate of an existing row increments that row’s times_derived instead of inserting a new one (document.py). With DERIVER.DEDUPLICATE (on by default), it also looks up near neighbours. _semantic_dup_decision then keeps whichever text carries more unique tokens, so it either replaces the old row or rejects the new one (document.py). times_derived is later used as a retrieval signal.
Storage
A Collection is unique on (observer, observed, workspace_name). Composite foreign keys tie both names to real peers (models.py). A Document carries level (explicit, deductive, inductive, contradiction), times_derived, a pgvector embedding, a soft-delete deleted_at, and a sync_state for external vector stores (models.py). Embeddings use an HNSW cosine index. Reasoning-tree edges live in document_sources, which deliberately has no foreign key on source_id (models.py). The defaults are OpenAI text-embedding-3-small at 1536 dimensions.
Working representation
_get_working_representation_internal fills a budget of up to 100 observations. It takes a semantic slice (default a third of the budget), then a most-reinforced slice ordered by times_derived, and then fills the remainder with the most recent documents. An empty session allowlist returns nothing instead of everything (representation.py).
Dialectic agent
Reasoning levels map to tool sets and iteration caps. minimal gets only search_memory and search_messages. The other levels add get_observation_context, grep_messages, two date-based message tools and get_reasoning_chain (agent_tools.py). Under a session allowlist, get_reasoning_chain is removed, because reasoning chains cross session boundaries (core.py). The default iteration caps are 1, 5, 2, 4 and 10 for minimal, low, medium, high and max. Every level defaults to gpt-5.4-mini over the OpenAI transport (config.py). So by default, “higher” levels differ mostly in tool rounds and tool choice, not in model.
Hybrid message search
search combines pgvector (or external-store) similarity with Postgres full-text search and merges the two ranked lists with reciprocal_rank_fusion (k=60) (search.py). A peer_perspective filter limits results to sessions the peer was in, and to the time window of its membership (search.py).
Dreamer
check_and_schedule_dream counts only explicit documents, so dream output cannot trigger more dreams. It schedules a dream after 50 new explicit facts, at least 8 hours after the last dream, delayed by an idle timeout (dream_scheduler.py). run_dream runs the deduction specialist, then the induction specialist. Each is a tool loop with a hardcoded cap, 12 rounds for deduction and 10 for induction (specialists.py, L740-L745). Both can search memory and messages and create conclusions. Only deduction can delete conclusions and rewrite the peer card (agent_tools.py). Surprisal sampling over k-d trees and other tree types can seed both with “hints”, but SURPRISAL.ENABLED defaults to false (orchestrator.py, config.py).
Scoping and auth
JWT claims form a hierarchy: admin, workspace w, peer p, session s (security.py). Scopes group sessions into a recall boundary. Under a session allowlist, only explicit conclusions are served. The reason is that dreamer output reads across all sessions but is stamped with a single session, and the code says so (representation.py).
Extending it
- Configuration hierarchy. Reasoning, summary, dream and peer-card settings resolve from message to session to workspace.
custom_instructionssteers the deriver and the dreamer without code changes. - Models. Each stage (deriver, each dialectic level, summary, deduction, induction, embeddings) has its own
MODEL_CONFIGwith atransportofopenai,anthropicorgemini. These can be set inconfig.tomlor throughDERIVER_,DIALECTIC_LEVELS__<level>__and similar environment variables. Anoverridesblock sets a custombase_urland key, so OpenAI-compatible local endpoints work through theopenaitransport. - Vector store. Set
VECTOR_STORE_TYPEtoturbopuffer,lancedb,qdrantorchromadb, and implementVectorStore(upsert_many,query,delete_many,delete_namespace) for anything else (vector_store/init.py). Namespaces are hashed per (workspace, observer, observed). - Structured answers. Pass
response_format(JSON Schema) to chat. It is converted to a Pydantic model, and the answer comes back as JSON. - Webhooks and clients. Webhook endpoints receive queue-driven deliveries. The
mcp/server runs as a Cloudflare Worker, over stdio or over Streamable HTTP.session.context()in both SDKs renders history plus summary forto_openai/to_anthropic. Coding-agent plugins (Claude Code, Codex, Cursor, OpenCode, OpenClaw, Hermes) are separate packages built on the same API.
Running it
- Required services. Postgres with the pgvector extension, plus an API key for the configured provider (OpenAI by default, for both chat models and embeddings). Redis is optional and serves as a cache.
- Two processes. The API (
fastapionsrc/main.py) and the deriver (python -m src.deriver). Both share the database, which also serves as the job queue.docker-compose.yml.examplewiresapi,deriver,databaseandredis, and binds ports to localhost. - CLI.
honcho startinhoncho-clibrings up the same four services from a published image and can run a config wizard (--setup basic|advanced). - Auth.
AUTH_USE_AUTHdefaults to false. Set it andAUTH_JWT_SECRETbefore exposing the API (config.py).
Strengths and caveats
- Strength: a real multi-party model. Observer/observed collections and per-session
observe_othersallow different agents to hold different views of the same person. Few memory systems store this. - Strength: cheap, careful ingestion. One structured call per token-budgeted batch, target-only extraction, injection-safe tagging, and exact plus semantic deduplication with a reinforcement counter.
- Strength: honest scoping. Empty allowlists fail closed, and synthesized conclusions are excluded from session-scoped reads. The code says openly that per-conclusion provenance is not there yet.
- Strength: operational depth. A Postgres queue with work-unit ownership, retries, a reconciler for vector sync and soft-delete cleanup, plus Prometheus metrics and Sentry.
- Caveat: memory is not immediate. Facts appear only after a work unit reaches 512 tokens or 30 minutes pass, unless you enable
FLUSH_ENABLED. Dreams need 50 facts and an 8-hour gap. - Caveat: every read costs an LLM call. Chat is an agent loop, so latency and cost scale with the reasoning level. The default levels share one model, so “high” means more tool rounds, not a stronger model.
- Caveat: heavy footprint. You run Postgres, pgvector and two Python services, and the codebase is large (for example, a 2,800-line
agent_tools.py). This is infrastructure, not a drop-in library. - Caveat: insecure by default, and AGPL. Auth is off until you configure a JWT secret. The AGPL-3.0 license matters if you modify Honcho and offer it as a network service.
Sources: code at 11b22bf, verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit 11b22bf. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemory extraction is driven by the Deriver (src/deriver/deriver.py), which processes batches of incoming messages via a single structured-output LLM call — the 'minimal deriver' approach. Messages are enqueued by src/deriver/enqueue.py on message creation, consumed by src/deriver/queue_manager.py, and dispatched through consumer.py's process_item() to process_representation_tasks_batch(). The prompt (minimal_deriver_prompt in src/deriver/prompts.py) wraps each message in an XML tag with target="true" or target="false" and instructs the model to extract only explicit atomic facts about the target peer from target="true" messages; other messages serve solely as interpretive context. The output is a PromptRepresentation Pydantic schema enforcing structured JSON with an explicit list of ExplicitObservationBase objects. The Deriver then converts these into a Representation and saves them via RepresentationManager.save_representation() (src/crud/representation.py:74). Deduplication at write time is two-phase: exact-content dedup (normalized content + level + session for explicit documents) drops identical text within a batch or against existing documents, incrementing the times_derived counter on the existing row instead of inserting a duplicate (src/crud/document.py:677-705). Semantic dedup (cosine distance < 0.05) optionally replaces or rejects near-duplicate text via token-set comparison (src/crud/document.py:1369-1422). Custom instructions per workspace/peer can be threaded into the extraction prompt via reasoning configuration, capped at DERIVER__MAX_CUSTOM_INSTRUCTIONS_TOKENS (default 2000).
How are memories stored?
answeredMemories (called observations internally, conclusions in the API) are stored as Document rows in Postgres (src/models.py:390-521), each with fields: content (the fact text), level (explicit / deductive / inductive / contradiction), embedding (pgvector column at configured dimensionality), times_derived (integer reinforcement counter), session_name, observer, observed, workspace_name, and deleted_at for soft-delete. Documents belong to a Collection (src/models.py:342-382) uniquely keyed by (observer, observed, workspace_name) — so every observer+observed pair gets its own collection, enabling self-representation (observer == observed) and cross-peer modeling. The Documents table has an HNSW index on the embedding column for cosine-distance ANN search (src/models.py:511-519). An orthogonal MessageEmbedding table (src/models.py:284-335) stores chunk-level embeddings of raw messages with sync_state tracking, decoupled from message creation. Embeddings are generated via embedding_client (src/embedding_client.py) using configurable providers (OpenAI by default for embeddings). The pgvector path embeds + queries directly in Postgres. For external vector stores, the VectorStore ABC (src/vector_store/__init__.py:53-195) defines an interface with upsert_many, query, delete_many, and delete_namespace. Implementations include Turbopuffer (src/vector_store/turbopuffer.py), LanceDB (src/vector_store/lancedb.py), Qdrant (src/vector_store/qdrant.py), and ChromaDB (src/vector_store/chroma.py). Namespaces isolate collections per (type, workspace, observer, observed) with a hashed suffix. Both stores are dual-written to in migration mode (settings.VECTOR_STORE.MIGRATED); a sync_state column (pending/synced/failed) tracks consistency and a reconciliation backstop heals missed writes.
How are memories retrieved and injected into the prompt?
answeredRetrieval is multi-strategy. For conclusions (the memory store), RepresentationManager._get_working_representation_internal() (src/crud/representation.py:317-411) blends up to three query strategies within a configurable token budget: (1) semantic search via _query_documents_semantic() — cosine-distance ANN through the pgvector HNSW index or external vector store, filtered by level and session allowlist; (2) most-derived — documents sorted by times_derived descending, surfacing the most reinforced facts; (3) recent — by created_at descending as a fill-in for unused capacity. The allocation splits ~1/3 semantic, ~1/3 most-derived, ~1/3 recent, with over/underflow reclaim. For message search, src/utils/search.py:317-458 implements a hybrid search combining semantic (pgvector/ANN or external vector store) and full-text search (Postgres GIN-indexed to_tsvector('english', content) with ILIKE fallback) fused via Reciprocal Rank Fusion (RRF, search.py:36-75). The Dialectic agent (src/dialectic/core.py and src/dialectic/chat.py) is the primary retrieval consumer: it runs a tool loop armed with 7 tools from src/utils/agent_tools.py — search_memory (semantic conclusion search), search_messages (hybrid message search), get_observation_context, grep_messages, get_messages_by_date_range, search_messages_temporal, and get_reasoning_chain (tree traversal from a conclusion to its premises and derived conclusions). The agent iteratively selects tools to gather context and synthesizes a response once it has enough evidence. At the minimal reasoning level, only search_memory and search_messages are available. Results are presented to the LLM as formatted observations with timestamps, IDs, and derivation metadata.
How are memories updated, consolidated or forgotten?
answeredMemory lifecycle has four phases. Creation: The Deriver extracts observations on message creation; each observation gets times_derived=1. On repeat writes, instead of inserting duplicates the times_derived counter is incremented (src/crud/document.py:693-695), strengthening the fact's reinforcement signal. Consolidation via Dreaming: The Dreamer (src/dreamer/orchestrator.py) runs scheduled 'dreams' using two specialist agents. The DeductionSpecialist (src/dreamer/specialists.py) reads existing explicit observations and produces deductive conclusions via a tool-using agent loop (tools: get_recent_observations, search_memory, search_messages, create_observations_deductive, delete_observations, update_peer_card). The InductionSpecialist produces inductive generalizations (patterns, preferences, personality) from both explicit and deductive conclusions. Surprisal-based prioritization (src/dreamer/surprisal.py) pre-filters observations: it builds a tree structure from embeddings, computes geometric surprisal scores ranking how anomalous or informative each observation is, and feeds the highest-surprisal observations as hints to the specialist agents — focusing the expensive dream cycle on the most 'interesting' or novel findings. Reasoning trees (src/dreamer/trees/) link each conclusion to its premises via document_sources table and source_ids references, enabling get_reasoning_chain traversal. Deletion and cleanup: Documents are soft-deleted via deleted_at timestamp (src/crud/document.py:901-948). The ReconcilerScheduler (src/reconciler/scheduler.py) periodically runs cleanup_soft_deleted_documents() which hard-deletes rows soft-deleted >5 minutes ago and removes their vectors from the external store. The reconciler also heals missed vector syncs and cleans stale queue items. Summarization: Parallel to memory extraction, the summarizer.py creates two-tier session summaries (short every ~20 messages, long every ~60) stored on the session metadata, providing a compressing lifecyle for long-running conversations.
How is memory scoped and isolated?
answeredMemory scoping has four layers. Workspace isolation: Nearly every table includes workspace_name as a composite foreign key, making cross-workspace data leakage structurally impossible at the schema level (src/models.py — all Document, Message, Collection, QueueItem, etc. tables). This is enforced by the FK constraints themselves, not by query filters. Observer/Observed (Collection) isolation: Documents live in collections keyed by (workspace_name, observer, observed). This means every peer pair has its own vector namespace — peer A's view of peer B is a separate collection from C's view of B. Self-representation (observer == observed) is just the diagonal of this matrix. All CRUD operations query through this triple-key filter (src/crud/document.py:81-87). Authentication scoping via JWT: src/security.py:31-63 defines JWTParams with a hierarchy — admin (ad), workspace (w), peer (p), session (s). A narrower token never falls back to wider access: {w: ws, p: alice} only acts on alice, never on a sibling peer. The require_auth FastAPI dependency (security.py:143-194) checks these claims per-route. Session-scoped peer access is gated by allow_member_read=True exclusively on read-only routes, enforced by an explicit allowlist in tests (tests/routes/test_auth_route_policy.py). Scopes: A Scope is a named grouping of sessions that acts as a visibility boundary on recall. Chat, representation, session context, and workspace search answered through a scope see only the contents of that scope's member sessions (CLAUDE.md:126-131). Adding a session to a scope copies its explicit conclusions into the scope's collection; removing one reconciles copies back out. Session allowlists (session_allowlist parameter on queries) restrict retrieval of both conclusions and messages to specified sessions — and only explicit-level conclusions (level "explicit") are considered safe to scope under this mechanism (src/utils/representation.py:24), because higher-level consolidated observations may have been synthesized from sources outside the allowlisted sessions. An empty allowlist fails closed.
How do agents integrate with it, and what is self-hostable?
answeredHoncho integrates via multiple channels. REST API: The core FastAPI server exposes /v3/* routes covering workspaces, peers, sessions, messages, conclusions, keys, and webhooks (src/routers/). The chat endpoint at POST /v3/.../peers/{peer_id}/chat is the main memory-query surface, returning reasoning-grounded natural-language answers. SDKs: Two first-party SDKs live in this repo — sdks/python/ (published as honcho-ai on PyPI) and sdks/typescript/ (published as @honcho-ai/sdk on npm). Both expose Honcho(workspace_id, api_key) clients with methods for peer/session/message CRUD, chat, search, and context injection with .to_openai()/.to_anthropic() formatters. MCP: The platform hosts an MCP HTTP server at mcp.honcho.dev (or self-hosted), usable by any MCP client via claude mcp add honcho --transport http. Agent framework plugins: First-party memory plugins ship for Claude Code (marketplace plugin), Codex (npm install), Cursor (shell installer), OpenCode, OpenClaw, Hermes (built-in upstream), and DeepSeek Harness (Cordis plugin). All read the same ~/.honcho/config.json. Webhooks: src/webhooks/webhook_delivery.py delivers event-driven webhook payloads as configured (src/webhooks/events.py), enabling external service integration. Self-hosting: The entire stack is open-source (AGPL-3.0) and self-hostable via honcho start (Docker Compose-based CLI) or manual uv run fastapi dev src/main.py + uv run python -m src.deriver. Required services: Postgres with pgvector and an LLM provider API key (Gemini for Deriver by default; Anthropic for Dialectic medium+; OpenAI for embeddings). Optional: Redis (cache), external vector stores (Turbopuffer, LanceDB, Qdrant, ChromaDB). The API and deriver are separate processes sharing a database.