MemMachine/MemMachine
FastAPI memory server: raw episodes in a graph or vector store, plus LLM-extracted profile facts updated in the background.
Overview
MemMachine is a self-hostable memory server for agents. It keeps two kinds of memory side by side, and they are built very differently.
- Episodic memory stores what was said. Every message becomes an episode and is kept verbatim. Long-term search embeds “derivatives” of each episode (the whole message by default, or one per sentence), finds matches, and returns the original episodes with their neighbours, reranked. A short-term layer keeps the most recent messages in memory, plus an LLM-written rolling summary. No LLM decides what to remember. Everything is kept.
- Semantic (“profile”) memory stores what is true about someone. A background job sends batches of new messages to an LLM together with the current profile. The LLM answers with
add/deletecommands on a two-leveltag → feature → valuestore, and the server periodically asks the LLM to consolidate crowded tags.
The server (packages/server) is a FastAPI app with a REST API under /api/v2 and an MCP endpoint. Clients are a Python SDK (packages/client, with LangGraph helpers), a TypeScript client, and integrations for LangChain, LlamaIndex, CrewAI, AWS Strands, Dify, n8n, FastGPT and OpenClaw. Storage backends are configured in one YAML file and cover Neo4j or NebulaGraph (graph + vector), Qdrant, Milvus or SQLite (vector), and PostgreSQL/pgvector or SQLite for episodes and profiles.
Architecture
flowchart LR
CL["Client / MCP / integrations"] --> API["FastAPI /api/v2"]
API --> MM["MemMachine facade"]
MM --> ES["Episode store (SQL)"]
MM --> EPI["EpisodicMemory per org/project"]
EPI --> STM["Short-term: deque + LLM summary"]
EPI --> LTM["LongTermMemory"]
LTM --> DECL["Declarative: graph + vectors"]
LTM --> EVT["Event: vector store + segments"]
MM --> SSM["SemanticSessionManager"]
SSM --> HIST["Pending history per set_id"]
HIST --> ING["Background IngestionService"]
ING --> LLM["LLM: add/delete commands"]
ING --> SEM["Semantic storage (pgvector / Neo4j / vector store)"]
| Component | Path | Role |
|---|---|---|
| REST router and service | server/api_v2/router.py, service.py |
Projects, memories (add, search, list, delete), semantic sets, categories and tags, config |
| MCP | server/api_v2/mcp.py, mcp_stdio.py, mcp_http.py |
add_memory, search_memory, delete_memory tools, with ids from headers or env |
| Facade | main/memmachine.py |
Fans writes and queries out to episodic and semantic memory, resolves defaults |
| Episodic memory | episodic_memory/episodic_memory.py |
Combines short-term and long-term stores per session key |
| Short-term memory | episodic_memory/short_term_memory/ |
Recent episodes in a deque, LLM summary on overflow |
| Long-term memory | episodic_memory/long_term_memory/, declarative_memory/, event_memory/ |
Two backends: graph-based “declarative” (default) and segment-based “event” |
| Semantic memory | semantic_memory/ |
Set ids, categories and tags, ingestion, consolidation, storage drivers |
| Retrieval agent | retrieval_agent/ |
Optional query decomposition and multi-tool search (agent_mode) |
| Common | common/ |
Embedders, language models (OpenAI, LiteLLM, Bedrock), rerankers, vector and graph stores |
All paths are relative to packages/server/src/memmachine_server/.
How a request flows
Take POST /api/v2/memories and then POST /api/v2/memories/search:
- Scope. Each request carries
org_idandproject_id. The service builds a session key{org_id}/{project_id}from them (service.py). - Persist the raw episode.
_add_messages_tomaps messages toEpisodeEntryobjects (content, producer, role, timestamp, metadata).MemMachine.add_episodeswrites them to the SQL episode store first (service.py, memmachine.py). - Episodic write. In parallel, the episodes go to the session’s
EpisodicMemory, which adds them to short-term and long-term memory at the same time (episodic_memory.py). With the declarative backend, each message becomes a derivative ("<source>: <content>", or one per sentence whenmessage_sentence_chunkingis on, which is off by default). The derivative is embedded and stored as a node with aDERIVED_FROM_<session>edge to the episode node (declarative_memory.py). - Semantic enqueue. If semantic memory is enabled,
SemanticSessionManager.add_messageresolves the episode’s semantic set ids from its metadata and records the episode id as pending history for each set. No LLM call happens yet (semantic_session_manager.py). - Background extraction. Every 2 seconds,
_background_ingestion_tasklooks for sets with at least 5 pending messages, or pending messages older than 5 minutes. It runsIngestionServiceon them, retrying with exponential backoff up to 60 s on errors (semantic_memory.py). - Search.
query_searchparses the filter string and runs episodic and semantic search concurrently (memmachine.py).EpisodicMemory.query_memoryqueries short-term and long-term memory together (default limit 20) and removes long-term hits already present in short-term (episodic_memory.py). Semantic search embeds the query per set and returns rankedSemanticFeaturerows. - Return. The response holds short-term episodes and summary, scored long-term episodes, and semantic features, as JSON. Building the prompt is the caller’s job.
EpisodicMemory.formalize_query_with_context, which would wrap results in<Summary>/<Episodes>/<Query>tags, is defined but not called anywhere (episodic_memory.py).
Key components
Declarative long-term memory
Search embeds the query and finds up to min(5 × limit, 200) similar derivative nodes. It follows DERIVED_FROM edges to the source episodes. Then it expands each episode into a context window (one third backward, two thirds forward of expand_context), scores whole contexts with the configured reranker, and unifies overlapping windows (declarative_memory.py). Matching on small derivatives and returning whole conversational context is a sound design for “what did we discuss” questions. A reranker is mandatory for this backend: without one, long-term memory is disabled with only a log warning (memmachine.py).
Event long-term memory
The alternative backend segments events (TextSegmenter, 500-character chunks), derives whole-text or sentence derivatives, and stores vectors in a VectorStore (Qdrant, Milvus, SQLite + sqlite-vec) and segments in a SQL SegmentStore partitioned per session (event_memory.py). LongTermMemory.search_scored hides the difference and applies score thresholds in the right direction for the metric (long_term_memory.py).
Short-term memory and summary
New episodes are appended to a deque. When total content plus summary exceeds message_capacity (64,000 characters by default), the oldest episodes are popped and an LLM consolidator folds the batch into the rolling summary (short_term_memory.py). On query, the summary is returned with as many recent episodes as fit. Filters apply to episodes, but the summary is returned unfiltered (short_term_memory.py).
Semantic ingestion
For each pending message and each category of the set, _process_single_set loads up to 50 existing features. llm_feature_update returns commands, and _apply_commands embeds and inserts ADD values (with the episode as citation) or deletes the matching (feature, tag) for DELETE (semantic_ingestion.py). The default update prompt asks the model to act like an “edge detection” layer and emit many atomic facts, so profiles grow fast. When a tag holds 20 or more features, consolidation asks the LLM which memories to keep and what merged memories to add (semantic_ingestion.py). SemanticService accepts a consolidation_threshold parameter but never passes it to IngestionService, so the effective threshold is always the ingestion default of 20.
Scoping
Episodic memory is partitioned by org/project, not by user. Per-user or per-agent separation inside a project comes from metadata filters. The Python client merges its user_id/agent_id/session_id metadata into a filter on each search, and the server applies whatever filter it receives. Semantic memory is scoped by set ids, derived from configurable metadata tags (for example org + project + user). Nothing on the server enforces per-user access.
Extending it
- Backends.
common/defines abstractEmbedder,LanguageModel,Reranker(BM25, cross-encoder, Cohere, Bedrock, embedder, RRF hybrid, identity),VectorStoreandVectorGraphStore. They are selected by name inconfiguration.yml. - Profile schema. Semantic categories, tags and set types are runtime objects managed through the API. Each category carries its own update and consolidation prompts, so what gets extracted is configured, not hard-coded.
- Retrieval agent.
agent_mode=trueon search routes long-term retrieval throughretrieval_agent/, which can decompose a query and merge sub-query results with a configurable LLM and reranker. - Clients. The Python
Memoryobject and the MCP tools wrap the same REST API. The integrations underintegrations/are thin adapters over the client.
Running it
- Docker compose.
docker-compose.ymlstartspgvector/pgvector:pg18,neo4j:5.23-community,qdrantand the MemMachine image.memmachine-compose.shandsample_configs/provide declarative (Neo4j) and event (Qdrant) configurations. - Required. A SQL database for episodes and semantic data (PostgreSQL with pgvector, or SQLite), an embedder, a reranker for the declarative backend, and an LLM for summaries and profile extraction. The short-term summary model defaults to
gpt-4.1. LiteLLM and Bedrock adapters allow other providers. - MCP. The HTTP app mounts MCP at
/mcp. A stdio mode is also available, and both read org, project and user from headers orMM_*env vars.
Strengths and caveats
- Strength: lossless episodic store. Raw conversations are kept and retrieved with surrounding context, so answers can cite what was actually said instead of a lossy extraction.
- Strength: configurable profiles. Categories, tags and prompts are data, with citations back to source episodes and an explicit consolidation step.
- Strength: backend choice. Graph or segment backends, and several vector and SQL stores, all behind one config.
- Caveat: weak tenant boundaries inside a project. User isolation is a client-side filter. The short-term summary is shared by everyone writing to the project and returned to every query.
- Caveat: profile lag and growth. Facts appear only after the ingestion trigger (5 messages or 5 minutes). Updates see at most 50 existing features, and the “extract many facts” prompt relies on consolidation to stay manageable.
- Caveat: silent degradation. Missing embedder, reranker or store config disables long-term memory with a warning instead of failing startup.
- Caveat: injection is up to you. The server returns structured results only. Prompt assembly helpers exist but are not used.
Sources: code at c1c04bc, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (42 pages), verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit c1c04bc. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemMachine extracts memories in two parallel pipelines. Episodic memory stores entire conversation turns (episodes) as-is, via EpisodicMemory.add_memory_episodes() (packages/server/src/memmachine_server/episodic_memory/episodic_memory.py:208). Each episode carries content (the message text), producer_id, producer_role, metadata, and episode_type. No LLM is involved — the raw text plus its metadata is committed to short-term and long-term stores. Profile/semantic memory goes further: an asynchronous IngestionService (packages/server/src/memmachine_server/semantic_memory/semantic_ingestion.py:68) polls for uningested episodes, then calls llm_feature_update() (packages/server/src/memmachine_server/semantic_memory/semantic_llm.py:69) with a two-part prompt. The prompt contains the existing feature set (formatted as tag→feature→value triples) and the new message content. The LLM emits SemanticCommand objects — ADD or DELETE — which are applied to extract atomic facts like name, preferences, or inferred demographics. The update prompt (packages/server/src/memmachine_server/semantic_memory/util/semantic_prompt_template.py:6) instructs the LLM to be a "wide, early layer" performing "edge detection" — extracting multiple distinct facts in parallel, even making low-confidence inferences. Deduplication happens at consolidation time (not write time): when a tag accumulates >=20 features (configurable consolidation_threshold), _deduplicate_features() (semantic_ingestion.py:447) merges overlapping entries using a second LLM call (llm_consolidate_features()), which can delete redundant memories and produce consolidated replacements. Conflict handling is explicit: the LLM issues delete-then-add pairs to update a value, and the consolidation prompt deliberately deletes unspecified memories (a keep-list design, not a delete-list).
How are memories stored?
answeredMemMachine uses a layered, pluggable storage architecture. Episodes (conversation turns) are persisted in EpisodeStorage — a SQLAlchemy-backed relational store (episode_sqlalchemy_store.py) using PostgreSQL or SQLite. Long-term episodic memory has two backend options: a declarative backend using VectorGraphStore, which wraps Neo4j (neo4j_vector_graph_store.py) or NebulaGraph Enterprise (nebula_graph_vector_graph_store.py), storing episodes as graph nodes with vector embeddings and DERIVED_FROM edges; or an event backend using a VectorStore (Qdrant, Milvus, SQLite+sqlite-vec, or SQLite+usearch/hnswlib) for vector embeddings plus a SegmentStore for time-ordered segment data. Short-term memory is an in-memory deque with capacity-based eviction and persistence hooks to SessionDataManager (short_term_memory.py:123). Profile/semantic memory stores extracted facts — each a SemanticFeature with fields (set_id, category, tag, feature_name, value, embedding, citations) — in SemanticStorage, which has implementations for PostgreSQL+pgvector (sqlalchemy_pgvector_semantic.py) and for vector stores with a separate citation table (vector_store_semantic_storage.py). Embeddings are generated by the Embedder abstraction, with concrete classes for OpenAI (openai_embedder.py), Sentence Transformers (sentence_transformer_embedder.py), and Amazon Bedrock (amazon_bedrock_embedder.py). The DatabasesConf class (database_conf.py:587) registers eight storage backends — Neo4j, PostgreSQL, SQLite, NebulaGraph, Qdrant, Milvus, SQLiteVectorStore, and SQLiteVecVectorStore — all configurable from a single YAML configuration file. Re-ranking is handled by the Reranker abstraction with BM25, cross-encoder, Cohere, and identity implementations.
How are memories retrieved and injected into the prompt?
answeredRetrieval is orchestrated by EpisodicMemory.query_memory() (packages/server/src/memmachine_server/episodic_memory/episodic_memory.py:352). It concurrently searches short-term memory (an in-memory deque of recent episodes, filtered by property and capacity) and long-term memory (scored vector search), then deduplicates by episode UID prioritizing short-term copies. The query string is embedded via the configured Embedder (OpenAI, Sentence Transformers, or Bedrock). Long-term search uses search_scored() (long_term_memory.py:264) which dispatches to either the declarative backend (Neo4j graph vector search) or the event backend (Qdrant/Milvus/SQLite vector ANN search). Results are optionally re-ranked by pluggable Reranker implementations: BM25, cross-encoder, Cohere, or identity passthrough. An expand_context parameter pulls chronologically adjacent episodes around each match. Score thresholds are applied directionally — higher-is-better metrics (cosine + reranker scores) drop below-threshold results; lower-is-better metrics (raw euclidean) drop above-threshold results. Semantic memory search (SemanticService.search(), semantic_memory.py:195) embeds the query per set_id (each set can have its own embedder), runs parallel vector searches across sets, and yields SemanticFeature objects. The final prompt injection is done via formalize_query_with_context() (episodic_memory.py:478), which wraps search results in <Summary> and <Episodes> XML tags and appends the original <Query>. An optional retrieval agent (MemMachineAgent in retrieval_agent/agents/memmachine_retriever.py) can orchestrate multi-tool retrieval — decomposing queries, running sub-queries against different memory tools, and consolidating results.
How are memories updated, consolidated or forgotten?
answeredMemMemory manages memory lifecycle through three distinct mechanisms: update, consolidation, and eviction/forgetting. Updates are LLM-driven: when a new episode arrives, IngestionService._process_single_set() (semantic_ingestion.py:122) loads existing features for each category, then calls llm_feature_update() with both the old profile and new message. The LLM returns semantically-meaningful mutations as add/delete commands. Consolidation runs asynchronously for semantic memory when a tag accumulates >=20 features (the consolidation_threshold). _deduplicate_features() (semantic_ingestion.py:447) calls llm_consolidate_features() with all features grouped by tag, plus an LLM prompt that can merge, split, and delete redundant entries. The prompt instructs: "If enough memories share similar features... delete all of them and create a single new memory containing a list." The consolidation is a keep-list design — memories not in keep_memories are deleted. In short-term memory, consolidation is summarization (ShortTermMemoryConsolidator, short_term_memory.py:467): when the deque's total message length exceeds message_capacity (default 64000 chars), the oldest episodes are evicted and a background _run_summary_loop() incrementally generates a summary via LLM. The summary is prepended on retrieval. If the summary prompt exceeds the context window, episodes are binary-split recursively until they fit. Forgetting/deletion is explicit: episodes can be deleted by UID (delete_episodes()) in both short-term and long-term memory. Session-level drops (drop_session_partition()) on the event backend delete the entire VectorStore collection and SegmentStore partition. Semantic features can be deleted individually, by set+category filter, or by set_id. The background ingestion uses a backoff-retry loop (semantic_memory.py:783) that catches transient failures, and marks messages as ingested only after successful processing. There is no built-in TTL-based decay for episodic memory, and no explicit versioning/history for features — updates replace semantically, and the storage layer stores only the latest value per feature.
How is memory scoped and isolated?
answeredMemory scoping is hierarchical and metadata-driven. At the top level, organizations and projects form the identity boundary: every API call passes org_id and project_id, which together construct the session key {org_id}/{project_id} (packages/server/src/memmachine_server/server/api_v2/service.py:44). Within a project, the client creates a Memory instance with a metadata dictionary that typically includes group_id, agent_id, user_id, and session_id (packages/client/src/memmachine_client/memory.py:116). When adding or searching memories, this metadata is automatically merged into a SQL-like filter string (e.g. metadata.user_id='alice' AND metadata.agent_id='travel_agent'), scoping every operation to that context. Semantic memory uses the set_id concept — a string keyed by metadata tags (like ["org_id", "project_id", "user_id"]). A SemanticSetType defines the tag schema; get_semantic_set_id() derives or creates the set from concrete metadata values (memory.py:1105). Set types can be org-level (shared across projects) or project-scoped. Within a set, categories define what facts to extract, and all features carry a set_id plus category_name for filtering. The MCP server passes org_id, proj_id (project), and user_id via HTTP headers or environment variables (MM_ORG_ID, MM_PROJ_ID, MM_USER_ID) (mcp.py:127), stored in ContextVars for thread-safe access. There is no built-in cryptographic access control — scoping is by key-based partition isolation rather than authentication/authorization — and the README mentions no ACL system. On the event backend, separate VectorStore collections and SegmentStore partitions (long_term_memory.py:140) provide physical data isolation per session. The declarative backend uses Neo4j collection-scoped queries filtered by session_id. For multi-tenancy, the DatabasesConf supports named database instances, but tenant isolation is at the application layer (org/project keys filtering), not at the database layer.
How do agents integrate with it, and what is self-hostable?
answeredAgents integrate via three interfaces. Python SDK (memmachine-client, PyPI package): the Memory class provides add(), search(), list(), delete_episodic(), delete_semantic(), and full semantic memory CRUD (add_feature(), add_semantic_category(), etc.). It communicates with the server over HTTP REST at /api/v2/memories. REST API (FastAPI server, api_v2/router.py): endpoints for adding (POST /api/v2/memories), searching (POST /api/v2/memories/search), listing (POST /api/v2/memories/list), configuring episodic memory, managing semantic categories/tags/sets, and health checks. MCP Server (server/mcp_stdio.py, server/mcp_http.py): wraps memory operations as MCP tools (add-memory, search-memories, delete-memories) for direct integration with Claude Desktop, Cursor, and any MCP client. The MCP server accepts org-id, proj-id, and user-id via HTTP headers or env vars. Framework integrations ship as Python packages: LangChain (integrations/langchain/memory.py), LangGraph (client/src/memmachine_client/langgraph.py), CrewAI (integrations/crewai/tool.py), LlamaIndex (integrations/llamaindex/mem_machine_memory.py), AWS Strands (integrations/aws_strands_agent_sdk/), Dify (integrations/dify/), and n8n and FastGPT. The server is fully self-hostable via Docker compose (docker-compose.yml): it requires PostgreSQL (with pgvector), plus one of Neo4j (declarative profile) or Qdrant (event profile), and optionally a vector store for semantic memory. All services are configured through a single configuration.yml. The project ships a memmachine-compose.sh helper script and sample_configs/ with event and declarative templates. OpenAI API keys are needed for LLM-based extraction and summarization, unless using Ollama or LiteLLM (configured via litellm_language_model.py). The Dockerfile packages the server, and docker-compose.yml wires healthchecks, persistent volumes, and networking.