LLMs Technical Reviews
Home / Agent memory layers / claude-mem

thedotmack/claude-mem

Coding-agent memory plugin: hooks capture tool calls, an observer LLM writes XML observations into SQLite and Chroma for next-session recall.

GitHub ↗★ 97kTypeScriptApache-2.0commit 0a803bd · 2026-10-06homepage ↗

Overview

claude-mem is a memory plugin for coding agents. It started as a Claude Code plugin and now installs into Codex, Cursor, Windsurf, OpenCode, OpenClaw, Kimi, Antigravity and others through npx claude-mem install --ide <name>. It hooks into the agent’s lifecycle. Every tool call the agent makes is captured, a second “observer” LLM session turns those tool calls into short structured observations, and at the start of the next session a timeline of recent observations is injected back into the agent’s context.

The important design choice is that memory is written by an LLM. A separate observer conversation, by default a Claude Agent SDK query() running claude-haiku-4-5, is fed each tool call (name, input, output). It replies in a fixed XML format (<observation> with type, title, facts, narrative, concepts, files). claude-mem then parses that XML with regular expressions and stores the result. The parsing is mechanical. The compression is not, and it costs model tokens on every tool call the agent makes. The observer can also run on Gemini, OpenRouter, Codex or any OpenAI-compatible endpoint.

The default deployment is local. A long-lived Bun worker owns a SQLite database (with FTS5) and a Chroma vector store, which it runs as a chroma-mcp child process through uvx. An opt-in “server” runtime moves the same pipeline onto Postgres, a BullMQ/Redis queue and API-key auth for teams. The codebase is big and very defensive. Much of it handles quota exhaustion, crashed child processes, hook timeouts and data that must not be lost when the worker is down.

Architecture

flowchart LR
  A["Coding agent (Claude Code, Codex, Cursor...)"] --> H["Hook CLI: worker-service hook"]
  H --> SP["Hook spool (files on disk)"]
  SP --> W["Bun worker (Express HTTP)"]
  W --> Q["SessionManager queue"]
  Q --> OBS["Observer LLM (Agent SDK / Gemini / OpenRouter)"]
  OBS --> P["parseAgentXml"]
  P --> DB["SQLite + FTS5"]
  DB --> CH["ChromaSync -> chroma-mcp"]
  DB --> CS["CloudSync (optional hub)"]
  DB --> CTX["ContextBuilder"]
  CTX --> A
  MCP["MCP server: search / timeline / get_observations"] --> W
  A --> MCP
  UI["Web viewer (React, SSE)"] --> W
Component Path Role
Hook wiring plugin/hooks/hooks.json Registers SessionStart, UserPromptSubmit, PostToolUse, Stop, SessionEnd commands that call the bundled worker script
Hook handlers src/cli/handlers/ context, session-init, observation, summarize, session-end; each returns a HookResult and never blocks the agent
Platform adapters src/cli/adapters/ Normalise hook payloads from Claude Code, Codex, Cursor, Windsurf, Kimi and others
Worker src/services/worker-service.ts Long-lived Bun process; Express routes, spool drainer, generator lifecycle
Observer providers src/services/worker/ClaudeProvider.ts, GeminiProvider.ts, OpenRouterProvider.ts, CodexProvider.ts, OpenAICompatProvider.ts Run the observer conversation that writes observations
Prompts and parser src/sdk/prompts.ts, src/sdk/parser.ts XML output contract, built from the active “mode”; regex parser
Response processing src/services/worker/agents/ResponseProcessor.ts Classifies output, stores rows, triggers Chroma sync, SSE broadcast
SQLite store src/services/sqlite/SessionStore.ts, src/storage/sqlite/schema.ts Observations, summaries, prompts, tool uses, FTS5 tables, migrations
Vector sync src/services/sync/ChromaSync.ts, ChromaMcpManager.ts Splits each observation into Chroma documents; supervises chroma-mcp
Context injection src/services/context/ Builds the SessionStart timeline under a token budget
Search src/services/worker/SearchManager.ts, search/strategies/ Chroma semantic search with SQLite filtering and fallback
MCP servers src/servers/mcp-server.ts, src/server/mcp/recall-mcp-server.ts Local stdio tools for the agent; read-only hosted recall endpoint
Server runtime src/server/ Optional Postgres + BullMQ + BetterAuth deployment
Modes plugin/modes/*.json Prompt and observation-type profiles (code, translated variants, law-study, others)

How a request flows

Take one Edit tool call inside a Claude Code session:

  1. Hook fires. PostToolUse runs worker-service.cjs hook claude-code observation. observationHandler drops untracked projects and subagent calls, then either posts to the server runtime or spools the event (observation.ts).
  2. Spool, then nudge. spoolHookEvent writes the event to a durable on-disk spool and fires a 250 ms best-effort POST /api/spool/nudge at the worker. The hook does not wait for processing (spool-hook-event.ts).
  3. Drain. The worker’s HookSpoolDrainer hands each entry to the same ingestObservation function its HTTP route uses, and keeps an entry if ingest fails (hook-spool-drain.ts).
  4. Ingest. ingestObservation checks whether the user prompt was marked private, strips memory tags (and, when enabled, redacts secrets) from tool input and output, writes a raw tool_uses backup row, queues the message and makes sure an observer generator is running (shared.ts).
  5. Observe. ClaudeProvider.startSession waits for a concurrency slot, builds hardened Agent SDK options with isolated credentials, and streams queued tool calls into query({ prompt: messageGenerator }) (ClaudeProvider.ts). The init prompt carries the mode’s instructions and an XML skeleton (prompts.ts).
  6. Parse. processAgentResponse first classifies quota, auth, context-overflow and transport failures so the batch is kept rather than lost. Then it calls parseAgentXml (ResponseProcessor.ts, parser.ts).
  7. Store. storeObservations inserts rows in one transaction with ON CONFLICT(memory_session_id, content_hash) DO NOTHING. The queued messages are confirmed only after the write succeeds (SessionStore.ts, ResponseProcessor.ts).
  8. Index and broadcast. Each new row is synced to Chroma (fire-and-forget), cloud sync is notified, and the viewer gets an SSE event (ResponseProcessor.ts).
  9. Next session. On SessionStart, contextHandler asks the worker for context (it skips resume, so the prompt cache stays valid). buildContextOutput renders a header, a timeline of recent observations, the latest summary and the previous assistant message, returned as additionalContext (context.ts, ContextBuilder.ts).

Key components

Observer and modes

The observer is a real conversation, not a stateless call. The same session keeps receiving tool calls, and it is recycled when its context fills up. The output contract comes from a mode JSON (plugin/modes/code.json by default). The mode defines the observation types, field guidance and placeholder text, so switching mode changes what gets remembered. The prompt puts mode-only text first so provider prompt caches hit across sessions (prompts.ts). Provider choice goes through one dispatch rule in provider-dispatch.ts, with optional fallback when a quota runs out.

Storage and dedup

SQLite is the source of truth. The content hash is a 16-hex SHA-256 of (memory_session_id, title, narrative) (store.ts), so it only catches exact repeats within one observer session. A duplicate “reinforces” the existing row instead of being dropped. Cross-session dedup on normalised titles (Tier-0 merge, Tier-1 IDF-cosine review) exists, but it is off by default: CLAUDE_MEM_DEDUP_ENABLED: 'false' (SettingsDefaultsManager.ts).

Context ranking

By default SessionStart injects the N most recent observations (50 by default) for the project. If you set CLAUDE_MEM_REINFORCE_ALPHA > 0, rankByStrength fetches a wider pool and re-ranks its older part with an ACT-R power-law score over each row’s reinforcement dates. The newest quarter is always kept (ObservationCompiler.ts, rank.ts). Nothing is ever deleted by age. Decay affects ranking only.

SearchManager.search has three paths. With no query text, it filters SQLite directly. With query text and Chroma available, it runs a semantic query with where filters for project (including merged project keys) and platform source, then refills any empty category from SQLite. Without Chroma, it falls back to FTS5 (SearchManager.ts). In Chroma, each observation becomes several documents: the narrative, the text and each fact, with ids like obs_<id>_narrative and metadata copied from the row (ChromaSync.ts).

MCP tools

The local stdio MCP server teaches a three-layer pattern: search returns a compact index, timeline shows neighbours of an anchor, get_observations fetches full rows, and get_tool_uses returns the raw tool I/O as a last resort. It also has corpus, work-state and “smart” code-outline tools. All of them proxy to the worker’s HTTP API (mcp-server.ts). A separate, smaller recall-mcp-server exposes only search, context and recent, scoped to an API key’s team, for remote use (recall-mcp-server.ts).

Cloud sync

CloudSync is not Git-based. It treats rows with synced_at IS NULL as a push queue and POSTs canonical JSON operations to a per-user sync hub (workers/sync-hub). Rows are stamped only on an exact ack (CloudSync.ts). It is optional and tied to the paid tier.

Extending it

  • Modes. Add or copy a JSON in plugin/modes/ to change observation types and prompt wording for a domain. This is the cheapest way to change what gets remembered.
  • Observer provider. Set CLAUDE_MEM_PROVIDER to claude, gemini, openrouter, codex or openai-compatible, and CLAUDE_MEM_MODEL for the model. Adding a new provider means a class next to GeminiProvider.ts that feeds results into processAgentResponse.
  • New hosts. Installers live in src/services/integrations/ (Cursor, Codex, Windsurf, OpenCode, OpenClaw, Kimi, T3Code and others) with matching payload adapters in src/cli/adapters/.
  • Privacy. Tag-stripping in src/utils/tag-stripping.ts is the one choke point for removing private content and optional secret redaction.
  • HTTP API. The worker registers route classes (MemoryRoutes, SearchRoutes, MemoryIngestRoutes, DataRoutes and others) in registerRoutes (worker-service.ts). POST /api/memory/save and memory-file ingest are the write entry points for other tools.

Running it

  • Install. npx claude-mem install (add --ide <name> for other hosts). The plugin needs Node.js and auto-installs Bun and uv. uv supplies Python for chroma-mcp.
  • Local data. One SQLite file plus a Chroma directory under ~/.claude-mem/. The worker port is derived from the user id (37700 + uid % 100), so several accounts on one machine do not collide.
  • Model access. The default observer uses your Claude Code login or API key with claude-haiku-4-5. Expect token use roughly proportional to tool-call volume, with quota guards that pause the observer when the subscription window is exhausted.
  • Chroma off. CLAUDE_MEM_CHROMA_ENABLED=false gives SQLite-only (FTS5) search.
  • Server runtime. CLAUDE_MEM_RUNTIME=server with CLAUDE_MEM_SERVER_DATABASE_URL (Postgres) and, in Docker, CLAUDE_MEM_REDIS_URL for BullMQ (create-server-service.ts). docker-compose.yml is provided.

Strengths and caveats

  • Strength: zero-effort capture. Hooks record everything the agent does, with no tool calls or prompting needed from the agent. Context comes back automatically at the next session start.
  • Strength: durable pipeline. Spool files, exactly-once hand-off markers, batch confirmation only after storage, and classification of quota and auth failures mean an outage delays memory rather than losing it.
  • Strength: token-aware retrieval. The index → timeline → full-row MCP pattern and the budgeted SessionStart block keep recall cheap for the main agent.
  • Caveat: memory costs model calls. Every tool call goes through an observer LLM. On a subscription that competes with your own usage, and the code has many guards because of it.
  • Caveat: recency, not relevance, at session start. The injected block is the latest N observations for the project. Semantic recall happens only when the agent calls the MCP search tool. ACT-R ranking and cross-session dedup exist but are off by default.
  • Caveat: no forgetting. There is no TTL or consolidation pass. Rows accumulate until something deletes them explicitly through the worker API (DELETE /api/observation/:id, per-session delete) or the viewer (DataRoutes.ts).
  • Caveat: heavy moving parts. A Bun worker, a Python chroma-mcp child, hook shell shims and per-host installers make the system complex to debug. A large share of the codebase handles its own failure modes.

Sources: code at 0a803bd, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (31 pages), verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit 0a803bd. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Extraction runs in two modes. Per-event extraction: The observationHandler (src/cli/handlers/observation.ts:32) captures each tool-use via a PostToolUse hook and spools it as an agent_event. The ProviderObservationGenerator (src/server/generation/ProviderObservationGenerator.ts:160) loads batched events and calls buildServerGenerationPrompt (src/server/generation/providers/shared/prompt-builder.ts:46) to assemble them into a single-turn LLM prompt. Events are wrapped in <agent_event> XML with id, event_type, source_adapter, privacy-stripped payload, and timestamp. The prompt asks the LLM to return <observation> blocks with type, title, subtitle, facts, narrative, concepts, files_read, files_modified — the type enum comes from the active 'mode' config (observable types like discovery, progress, blocker, decision). Session-summary extraction (processSessionSummaryResponse at src/server/generation/processGeneratedResponse.ts:199) feeds the full session event list and asks for a <summary> with request/investigated/learned/completed/next_steps fields. Both responses are parsed by parseAgentXml (src/sdk/parser.ts:62). Deduplication is handled at write time by a generation_key UNIQUE constraint (src/storage/postgres/observations.ts:103): the key combines generation_job_id, parsed index, and a content digest hash, so retries collapse to one row. The prompt builder also strips <private> and <system-reminder> tags before sending data to the LLM, and the parser salvages unstructured text when content fields are empty (src/sdk/parser.ts:170-191).

How are memories stored?

answered

Storage is layered with three backends. SQLite is the primary local store using bun:sqlite. The memory_items table (src/storage/sqlite/schema.ts:85-105) stores each memory with columns for id, project_id, kind (observation/summary/prompt/manual), type, title, subtitle, text, narrative, facts (JSON), concepts (JSON), files_read, files_modified, and metadata. An FTS5 virtual table (memory_items_fts, src/storage/sqlite/schema.ts:171-182) indexes title/subtitle/text/narrative/facts/concepts with porter unicode61 tokenization. The MemoryItemsRepository (src/storage/sqlite/memory-items.ts:111) implements CRUD operations and FTS5-based search via the MATCH keyword (src/storage/sqlite/memory-items.ts:260-291). Postgres is the server store (src/storage/postgres/schema.ts:223-238). The observations table adds team_id for multi-tenancy, a generated content_search TSVECTOR column using to_tsvector('english', content). Observations use embeddings stored as JSONB, and the generation_key column has a UNIQUE partial index (src/storage/postgres/schema.ts:298-300) for idempotent write deduplication. An observation_sources table links each observation to its upstream events. Chroma is an optional vector database (src/services/sync/ChromaSync.ts) that stores observations as embeddings for semantic search. The embedding function is configurable via the CLAUDE_MEM_CHROMA_EMBEDDING_FUNCTION setting (src/services/sync/ChromaSync.ts:356-362). Chroma runs as a local MCP subprocess (ChromaMcpManager). The memory schema is defined in Zod at src/core/schemas/memory-item.ts:5-44 with full runtime validation — a MemoryItem carries fields for id, projectId, serverSessionId, kind, type, title, subtitle, text, narrative, facts[], concepts[], filesRead[], filesModified[], and metadata.

How are memories retrieved and injected into the prompt?

answered

Retrieval operates at multiple levels. SessionStart context injection: Every new Claude Code session triggers the contextHandler (src/cli/handlers/context.ts:66), which calls /api/context/inject. This returns a rendered timeline of observations and summaries limited by a token budget (~10k chars). The ContextBuilder (src/services/context/ContextBuilder.ts) calls queryObservationsMulti which fetches observations by recency filtered by the active mode's observation types and concepts (src/services/context/ObservationCompiler.ts:82-96). An optional ACT-R reinforcement ranking (src/services/reinforcement/rank.ts:85-106) re-ranks older observations by blended recency+reinforcement score when CLAUDE_MEM_REINFORCE_ALPHA > 0. Full-text search: On SQLite, the MemoryItemsRepository.search method (src/storage/sqlite/memory-items.ts:260-291) builds FTS5 queries with term tokenization, NFKC unicode normalization and alternative expansion. On Postgres, the PostgresObservationRepository.search method (src/storage/postgres/observations.ts:176-258) uses websearch_to_tsquery('english', ...) with ts_rank ordering and GIN indexes. Hybrid search: The HybridSearchStrategy (src/services/worker/search/strategies/HybridSearchStrategy.ts:15) combines SQLite keyword matches with Chroma vector similarity ranking for file-aware search. MCP tools: The recall MCP server (src/server/mcp/recall-mcp-server.ts:140) exposes three read-only tools — search (ranked FTS results), context (search results plus concatenated context string), and recent (newest-first listing). These go through the same Postgres search as the REST API and are scoped/audited per API key. The rendered context for the model includes a header, day-grouped timeline, and the last prior assistant message when enabled.

How are memories updated, consolidated or forgotten?

answered

Update/merge: Memory items support full field-level replacement via MemoryItemsRepository.update (src/storage/sqlite/memory-items.ts:181-243), which merges provided fields over existing values. The dedup system (src/services/dedup/nearDuplicate.ts:49-76) classifies observation pairs into two tiers: Tier-0 exact — normalized titles match → silent auto-merge; Tier-1 candidate — IDF-weighted TF-IDF cosine similarity above a configurable threshold and no IDF-veto fires → flagged as a candidate for review but never auto-merged (src/services/dedup/tfidfCosine.ts:20-34). The IDF veto (src/services/dedup/idfVeto.ts) prevents merging when a rare discriminating token appears in only one of the two observations. Consolidation: Session summaries condense the entire session arc into request/investigated/learned/completed/next_steps fields (processSessionSummaryResponse at src/server/generation/processGeneratedResponse.ts:199-256). These coexist alongside per-event observations rather than replacing them. Cleanup and deletion: The Postgres schema enforces cascade deletes — deleting a team or project cascades to sessions, events, observations, and generation jobs (src/storage/postgres/schema.ts:91-238). The CleanupV12_4_3 utility (src/services/infrastructure/CleanupV12_4_3.ts:27-236) runs one-time cleanup of stale observer sessions with automatic database backup before deletion. FTS maintenance (src/services/infrastructure/FtsMaintenance.ts:60-84) merges FTS5 b-trees incrementally to reclaim space from deleted rows. No explicit TTL/decay: The system relies on recency-based selection and optional ACT-R reinforcement scoring (src/services/reinforcement/rank.ts:50-63) to naturally demote older observations, rather than expiring them. Observations remain in storage indefinitely but are excluded from the context budget window by newer content. Generation jobs have configurable max_attempts with exponential backoff (src/server/generation/processGeneratedResponse.ts:556-560).

How is memory scoped and isolated?

answered

Memory is scoped by a team → project → session hierarchy with multi-tenancy at the team level. Team/Project isolation: Every Postgres observation carries team_id and project_id, and every query enforces scoped WHERE clauses (src/storage/postgres/observations.ts:199-200). API keys are bound to a team (and optionally a project: src/storage/postgres/schema.ts:122-131). The auth middleware requires memories:read or memories:write scopes (src/server/routes/v1/ServerV1Routes.ts:62-71). The ProviderObservationGenerator validates scope bullMQ payloads against canonical Postgres rows at execution time (src/server/generation/ProviderObservationGenerator.ts:202-227) and refuses to run on mismatch. Session scoping: Observations link to server_session_id, and the server_sessions table records platform_source (claude-code, opencode, cursor, etc.), enabling platform-aware filtering (src/storage/postgres/schema.ts:149-167). Search queries support filtering by platform source (src/storage/postgres/observations.ts:223-240). Multi-tenancy: Teams are the tenant boundary with team_members controlling roles (owner, admin, member, viewer: src/storage/sqlite/schema.ts:44-53). Foreign keys cascade from team through projects, sessions, events, and observations — deleting a team cascades all its data (src/storage/postgres/schema.ts:91-256). Privacy controls: Event payloads pass through stripTags before extraction (src/server/generation/providers/shared/prompt-builder.ts:149), removing <private>, <claude-mem-context>, <system-reminder> tags. Subagent observations are dropped early via shouldSkipAgentObservation (src/cli/handlers/observation.ts:59-67), and excluded projects are skipped via shouldTrackProject. The excludeSubagents filter (src/storage/postgres/observations.ts:208-221) keeps subagent-derived rows out of context injection when configured.

How do agents integrate with it, and what is self-hostable?

answered

Agents integrate through multiple channels. Claude Code hooks are the primary path: sessionInit starts a tracked server session, observationHandler captures each PostToolUse event, contextHandler injects a rendered timeline at SessionStart, and sessionEnd triggers summary generation (src/cli/handlers/). The adapter layer (src/cli/adapters/claude-code.ts:39) normalizes platform-specific input shapes into a common NormalizedHookInput. MCP servers provide programmatic access: a local stdio MCP server (src/servers/mcp-server.ts) supplies search, code reading, and memory tools; a remote-recall MCP server (src/server/mcp/recall-mcp-server.ts:140) exposes search, context, and recent tools over Streamable HTTP MCP — both go through the same Postgres search and audit path. REST API: POST /v1/search and POST /v1/context (src/server/routes/v1/ServerV1PostgresRoutes.ts:934-984) provide full-text search over generated observations. POST /api/memory/save (src/services/worker/http/routes/MemoryRoutes.ts:26) allows manual memory creation. OpenCode plugin: A native OpenCode IDE plugin (src/integrations/opencode-plugin/index.ts) integrates with OpenCode's event system. Platform adapters: Supports Claude Code, Codex, Cursor, OpenCode, Antigravity CLI, Windsurf, and Kimi (src/cli/adapters/). Self-hosting: The entire system is self-hostable and Apache 2.0 licensed. SQLite works out of the box with only Bun. Postgres is optional for multi-tenant deployments. Chroma vector search runs as a local MCP subprocess. The installer (npx claude-mem install) handles setup across all supported platforms. Required services depend on the backend: Bun+SQLite for local-only, Postgres for the server, Chroma for vector search (all optional). There are no hosted-only parts — the full codebase and all features are available for self-hosting.

Editor's note. Correction: not everything is self-hostable. Cloud sync to the sync hub is tied to the paid tier; the local worker, hooks, MCP server and search run without it.