mem0ai/mem0
Python and TypeScript memory layer: an LLM extracts facts from chats into a vector store, recalled by semantic, BM25 and entity scoring.
Overview
Mem0 is a memory library for LLM applications. You pass it a conversation. It asks an LLM to pull out short facts (“prefers window seats”, “works at Acme as a data engineer”), embeds each fact and stores it in a vector database. Later you pass a query and get back the best-matching facts, which you put into your own prompt. Mem0 does not call your agent and does not build the prompt for you. It is a store with an LLM on the write path.
The repository has two products that share one API. The open-source SDK (Memory and AsyncMemory in Python, and an oss module in the TypeScript package) runs fully on your own infrastructure. MemoryClient is a thin HTTP client for the hosted platform at api.mem0.ai. Some features exist only on the hosted side: the OSS classes raise ValueError for decay, timestamp and reference_date, and project.update() is not supported at all (main.py).
At this SHA the open-source write path is add-only. Mem0 never merges, rewrites or deletes an existing memory when new information arrives. It skips exact duplicates and stores everything else as a new record. Graph memory (Neo4j and similar) has also been removed from the SDK: there is no mem0/graphs/ package and no graph_store key in MemoryConfig. “Entities” now live in a second vector collection.
Architecture
flowchart LR
A["Your app / agent"] --> M["Memory (main.py)"]
S["REST server (server/)"] --> M
M --> L["LLM: extraction prompt"]
M --> E["Embedder"]
M --> V["Vector store: memories"]
M --> N["Vector store: entities"]
M --> Q["SQLite: history + messages"]
M --> R["Reranker (optional)"]
K["MemoryClient"] --> H["api.mem0.ai"]
P["Agent plugins / MCP"] --> H
| Component | Path | Role |
|---|---|---|
| Memory core | mem0/memory/main.py |
Memory and AsyncMemory: add, search, get_all, update, delete, history |
| Prompts | mem0/configs/prompts.py |
ADDITIVE_EXTRACTION_PROMPT and the user-prompt builder |
| Scoring | mem0/utils/scoring.py |
BM25 normalisation and additive score fusion |
| Entity extraction | mem0/utils/entity_extraction.py |
spaCy-based proper nouns, quoted text and noun compounds |
| History DB | mem0/memory/storage.py |
SQLiteManager with history and messages tables |
| Providers | mem0/vector_stores/, mem0/llms/, mem0/embeddings/, mem0/reranker/ |
Pluggable backends behind factories in mem0/utils/factory.py |
| Hosted client | mem0/client/ |
MemoryClient for the platform API |
| REST server | server/ |
FastAPI wrapper with JWT/API-key auth and a Next.js dashboard |
| Integrations | integrations/ |
Coding-agent plugins, Vercel AI SDK provider, n8n, Zapier, Strands |
How a request flows
Write (add)
Memory.add()rejects the platform-onlytimestamp, validates the IDs with_build_filters_and_metadataand turns a string or dict into a message list (main.py). At least one ofuser_id,agent_idorrun_idis required, and identity keys inmetadataare dropped (L314-L404).- With
infer=False, each non-system message is embedded and stored as is. Otherwise_add_to_vector_storeruns its phased pipeline (L881-L1220). - It loads the last 10 messages for this scope from SQLite, embeds the new messages and fetches the 10 nearest existing memories. Their UUIDs are replaced by small integers before they go to the LLM.
- One LLM call with
ADDITIVE_EXTRACTION_PROMPT(plusAGENT_CONTEXT_SUFFIXwhen onlyagent_idis set) andresponse_format=json_objectreturns amemorylist. Onlytextandattributed_toare used from each item. - The texts are batch-embedded. Each text is MD5-hashed and dropped if the hash matches one of the 10 neighbours or an earlier item in the same batch.
- The survivors are inserted in one batch (one by one on failure), with
data,hash,text_lemmatized, timestamps and the scope IDs in the payload. AnADDrow goes to the history table. - spaCy extracts entities from every new memory. Each entity is matched to an existing one by normalised text or by cosine ≥ 0.95, and its
linked_memory_idslist is updated; otherwise a new entity row is inserted (L1100-L1204).
Read (search)
search()requires afiltersdict with at least one scope ID and rejects top-leveluser_idarguments (L1393-L1536)._search_vector_storeembeds the query and over-fetchesmax(top_k*4, 60)semantic hits. If the store implementskeyword_search, it also runs BM25 on the lemmatised query (L1642-L1745)._compute_entity_boostslooks up up to 8 query entities in the entity store and gives each linked memory a boost of up to 0.5, smaller when the entity links many memories (L1747-L1827).score_and_rankdrops candidates under the semantic threshold (0.1), adds semantic + BM25 + boost and divides by the maximum possible sum (scoring.py). Expired memories are skipped first. An optional reranker runs last.
Key components
Extraction prompt
The extraction prompt is long: it lists the kinds of facts to keep, quality rules and many few-shot examples. The user prompt carries sections for recent messages, existing memories, new messages, observation date and custom instructions (prompts.py). Its docstring says the LLM may return linked_memory_ids, but _add_to_vector_store does not store them.
Storage
A memory is one vector plus a flat payload. There is no separate memory schema beyond that payload, so any vector store that can filter on payload fields works. About two dozen stores are supported, and 15 of them implement keyword_search. The default is Qdrant with OpenAI for both the LLM and the embedder. At start-up Memory warns when the chosen store has no keyword search, because hybrid scoring then falls back to semantic only (L487-L552). The entity collection is created lazily under <collection>_entities.
History and messages
SQLiteManager keeps a history table (memory id, old and new text, event, timestamps, actor, role) and a messages table trimmed to the last 10 rows per session scope (storage.py). It is a local file (~/.mem0/history.db by default), even when the vectors live in a remote database.
Lifecycle
update() re-embeds the new text, recomputes hash and lemmas, keeps created_at, and relinks entities when the text changed (L1829-L1881, L2052-L2114). delete() removes the vector and strips the id from entity rows. delete_all() lists and deletes in batches (L1883-L1958). expiration_date only hides a memory from reads (L442-L451); nothing purges it.
REST server
server/main.py builds one shared Memory from DEFAULT_CONFIG (pgvector, OpenAI gpt-5-mini, text-embedding-3-small) and exposes /memories, /search, history and reset routes (main.py). Auth is a Bearer JWT, an X-API-Key or a legacy ADMIN_API_KEY, and /reset and delete-all need an admin. A caller’s identity is not tied to the user_id they pass.
Extending it
- Providers: add a module under
mem0/<category>/, subclass itsbase.py, add a config and register it in the factory. For hybrid search, a vector store implementskeyword_search. - Prompting:
custom_instructionsinMemoryConfig, orprompt=on a singleadd(), is appended to the extraction prompt. - Reranking: set
rerankerin the config (Cohere, HuggingFace, sentence-transformers, LLM, Zero Entropy) and passrerank=True. - Agents: the plugins in
integrations/expose asearch_memoriesMCP tool plus capture hooks. They talk toapi.mem0.aiunlessMEM0_API_URLpoints elsewhere (memory_core.py).
Running it
- Library:
pip install mem0ai, set an OpenAI key (or configure other providers), thenMemory()orMemory.from_config({...}). Entity extraction needs spaCy; without it, it returns no entities and the entity boost does nothing. - Server:
server/docker-compose.yamlruns the FastAPI app,pgvector/pgvector:pg17and the dashboard.JWT_SECRETandPOSTGRES_PASSWORDmust be set, orAUTH_DISABLED=truefor local use (docker-compose.yaml). There is no Neo4j service. - Telemetry:
MEM0_TELEMETRYis on unless set to false, in the SDK and in the server.
Strengths and caveats
- Strength: a small, readable core. One file holds the whole add and search path, and every provider sits behind the same small interface.
- Strength: retrieval is better than plain cosine: BM25 and entity links can lift a memory, and
explain=Truereturns the per-signal scores. - Strength: wide backend choice. Any of about two dozen vector stores can hold everything except history.
- Caveat: no conflict handling on write. “Lives in Berlin” and a later “moved to Lisbon” both stay stored, and only hash-identical text is deduplicated. The
add()docstring still says the LLM decides to add, update or delete. - Caveat: dedup only checks the 10 nearest memories, so a duplicate further away can slip through.
- Caveat: isolation is a payload filter, not access control.
get,update,deleteandhistorytake only a memory id, and the REST server does not bind the authenticated caller to auser_id. - Caveat: history and recent messages sit in a local SQLite file, which complicates running several stateless workers.
- Caveat: BM25 and entity boosts only re-rank semantic candidates. A memory that only matches by keyword is never returned.
Sources: code at c93420c, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (19 pages), verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit c93420c. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemories are extracted from conversational messages by an LLM call guided by the ADDITIVE_EXTRACTION_PROMPT system prompt (mem0/configs/prompts.py:468-944). The pipeline (_add_to_vector_store in mem0/memory/main.py:881-1220) works in phases:
- Context gathering — the last 10 messages for the session scope are fetched from the local SQLite message store to resolve pronoun references.
- Existing memory retrieval — the current messages are embedded and used to search the vector store (top-10), returning nearby memories for deduplication reference.
- LLM extraction — the
ADDITIVE_EXTRACTION_PROMPTinstructs the model to produce ADD-only operations: a JSON array of{"id", "text", "attributed_to", "linked_memory_ids"}objects. The prompt covers 12 information categories (personal preferences, important details, plans, professional info, etc.) and includes 11 few-shot examples with detailed quality criteria. For agent-scoped memories (agent_id present, no user_id),AGENT_CONTEXT_SUFFIX(mem0/configs/prompts.py:947-957) adjusts framing to "agent was informed that...". A separateUSER_MEMORY_EXTRACTION_PROMPTandAGENT_MEMORY_EXTRACTION_PROMPT(mem0/configs/prompts.py:63-174) are available for the deprecated non-additive pipeline. - Hash deduplication — each extracted text is MD5-hashed and compared against existing hashes in the vector store payloads and against other texts in the same batch. Duplicates are silently skipped.
- Entity extraction — spaCy-based
extract_entities_batch(mem0/utils/entity_extraction.py) pulls proper nouns, quoted text, and noun compounds from every extracted memory. These entities are deduplicated globally across the batch and upserted into the entity vector store, with linked_memory_ids tracking which memories reference which entities. Entities are matched by exact normalized text first, then by semantic similarity (threshold ≥0.95), avoiding duplicates.
When infer=False, messages are stored verbatim with no LLM extraction — each non-system message becomes one memory record.
text and attributed_to from each extracted item. The linked_memory_ids the prompt asks for are not stored. Hash deduplication compares only against the 10 nearest existing memories.How are memories stored?
answeredMemories are stored primarily in a pluggable vector store with auxiliary SQLite storage for change history and per-session messages.
Vector store — the core persistence layer. Each memory is stored as a vector embedding plus a payload dict. The payload fields (mem0/memory/main.py:1029-1041) include: data (the extracted memory text), hash (MD5 for dedup), text_lemmatized (for BM25 keyword search), created_at/updated_at (ISO timestamps), user_id/agent_id/run_id (scoping), actor_id (who spoke), role (user/assistant), attributed_to, and expiration_date (optional TTL). Any caller-supplied metadata keys survive as additional payload fields. The vector store interface (mem0/vector_stores/base.py) standardizes insert, search, delete, update, list, and optionally keyword_search for BM25. The VectorStoreFactory (mem0/utils/factory.py:180-218) supports 25+ backends: Qdrant, Pinecone, Chroma, Weaviate, Milvus, MongoDB, Redis, Elasticsearch, pgvector, Supabase, Faiss, S3Vectors, and more.
Embeddings — configured via EmbedderFactory (mem0/utils/factory.py:152-177) with 12 providers (OpenAI, Azure OpenAI, Gemini, HuggingFace, FastEmbed, Together, AWS Bedrock, Ollama, etc.). The embedding model is called for query encoding and for each memory text at write time. embed_batch is used for batch efficiency.
Entity store — a second vector store collection (named {collection}_entities or {collection}-entities) storing entities extracted from memories. Each entity record has: data (entity text), entity_type (PROPER, QUOTED, TOPIC, IDENTIFIER), and linked_memory_ids (list of memory UUIDs that mention this entity). This is lazily initialized on first entity operation.
SQLite — local database (mem0/memory/storage.py) with two tables: history tracking every ADD/UPDATE/DELETE on each memory (with old/new text, timestamps, actor_id, role), and messages storing recent conversation messages per session scope (trimmed to 10 most recent).
How are memories retrieved and injected into the prompt?
answeredRetrieval uses a hybrid scoring pipeline in _search_vector_store (mem0/memory/main.py:1642-1745):
- Query processing — the search query is lemmatized for BM25 via
lemmatize_for_bm25, and entities are extracted from it viaextract_entitiesfor entity-based boosting. - Semantic search — the query is embedded and the vector store's
search()retrieves candidates (over-fetched:max(limit * 4, 60)). - Keyword search — if the vector store supports BM25 (Qdrant, Elasticsearch, pgvector), a parallel keyword search runs with the lemmatized query. Raw BM25 scores are sigmoid-normalized into [0, 1] via
normalize_bm25(mem0/utils/scoring.py:43-54), with query-length-adaptive midpoint and steepness. - Entity boost — if the query contains entities,
_compute_entity_boosts(mem0/memory/main.py:1747-1827) searches the entity store for each entity (viaThreadPoolExecutor), and any memory linked to a matching entity (threshold ≥0.5) receives a boost of up to 0.5, attenuated by how many memories the entity links to. - Additive scoring —
score_and_rank(mem0/utils/scoring.py:60-139) computescombined = (semantic + bm25 + entity_boost) / max_possible, clamped to [0, 1]. The divisor adapts: 1.0 for semantic-only, 2.0 with BM25, 2.5 with BM25+entity, 1.5 with entity only. A semantic threshold gate (default 0.1) excludes low-confidence vectors before combination. Results are sorted descending and trimmed to top_k. - Filtering — before scoring, expired memories (payload
expiration_datebefore today) are excluded unlessshow_expired=True. Metadata filters support advanced operators: comparison (eq,ne,gt,gte,lt,lte,in,nin,contains,icontains) and compound logic (AND,OR,NOT) via_process_metadata_filters(mem0/memory/main.py:1538-1613). - Reranking — if a
RerankerConfigis set andrerank=True, an optional reranker stage (Cohere, HuggingFace, Sentence Transformer, LLM-based, Zero Entropy) reorders results before return.
Results are returned as a dict with a "results" list, each item containing id, memory text, hash, timestamps, scope identifiers, score, and any additional metadata.
How are memories updated, consolidated or forgotten?
answeredUpdate/merge — update() (mem0/memory/main.py:1829-1881) retrieves the existing memory by ID, re-embeds the new text, freshens the payload (recomputing hash and lemmatized text, updating updated_at), and calls vector_store.update(). Identity keys (user_id/agent_id/run_id) are immutable — metadata attempts to change them are silently dropped with a warning via _strip_identity_keys (mem0/memory/main.py:143-162). If the text changed, the old entity links are removed and new entities extracted via _link_entities_for_memory. The change is recorded in the SQLite history table as an UPDATE event with the old text preserved.
Deletion — delete() (mem0/memory/main.py:1883-1902) retrieves the existing memory, removes it from the vector store, and strips its memory_id from linked entity records via _remove_memory_from_entity_store. The history entry records the event as a DELETE with is_deleted=1. delete_all() (mem0/memory/main.py:1904-1958) lists memories in batches of 1000 and deletes each one, with a guard against infinite loops on stores that cap list() results.
Expiration/TTL — the expiration_date parameter on add() and update() sets a YYYY-MM-DD date in the payload. _payload_is_expired() (mem0/memory/main.py:442-451) checks whether that date is before today (UTC). Expired memories are excluded from search() and get_all() unless show_expired=True. There is no automated background cleanup — expired records remain in the store but are hidden.
No decay or consolidation — the OSS SDK explicitly rejects decay features. Calling project.update(decay=True) raises ValueError with the message that decay is a hosted-platform-only feature. Neither consolidation, summarization, nor automated forgetting across memories exists in the self-hosted path.
Versioning/history — history(memory_id) (mem0/memory/main.py:1960-1973) queries the SQLite history table for every ADD, UPDATE, and DELETE event on a given memory, ordered by timestamp. Each record stores the old and new text, the event type, timestamps, actor_id, and role.
Reset — reset() (mem0/memory/main.py:2144-2174) deletes the vector store collection and SQLite database, then recreates both from scratch. The entity store (if initialized) is also reset.
How is memory scoped and isolated?
answeredMemory scoping is enforced through three mandatory entity identifiers: user_id, agent_id, and run_id. At least one is required for every operation (_build_filters_and_metadata in mem0/memory/main.py:314-404). All three may be combined for composite scoping (e.g., a specific user interacting with a specific agent in a specific run).
Write-time scoping — when add() is called, the provided entity IDs are validated (trimmed, whitespace-checked, non-empty enforced via _validate_and_trim_entity_id at line 175-209) and stored directly in the vector store payload. Caller-supplied metadata dicts are stripped of identity keys by _strip_identity_keys (mem0/memory/main.py:143-162) — you cannot set user_id inside metadata to override or change the scope. On the update path, re-sending the existing values is silently accepted, but changed values are dropped with a warning.
Query-time scoping — search(), get_all(), and delete_all() require filters containing at least one entity ID. Top-level user_id/agent_id/run_id kwargs are explicitly rejected (_reject_top_level_entity_params at line 165-172); callers must pass filters={"user_id": "..."} instead. Filters are passed directly to the vector store's search() and list() methods, meaning the vector store itself enforces the scope at query time.
Session scope & messages — a deterministic session scope string is built from sorted entity IDs (_build_session_scope at mem0/memory/main.py:412-419), used as the key for storing recent conversation messages in the SQLite messages table.
Actor filtering — an optional actor_id narrows queries to memories from a specific speaker. This is separate from entity scoping and resolved via filters["actor_id"].
Multi-tenancy — the three-identifier system provides natural multi-tenancy. Different users, agents, or runs have fully isolated memory spaces because every store operation writes the identifiers into the payload and every query filters on them. The server's entities REST endpoint (server/routers/entities.py:43-67) scans all payloads to aggregate entities by type, supporting admin-level cross-entity queries.
No tenant-level access control in the OSS SDK itself — that lives in the hosted platform. The FastAPI server (server/main.py:1-80) adds JWT auth and API key authentication.
get, update, delete and history take only a memory id and do not check scope. The REST server authenticates the caller but does not bind them to the user_id in the request.How do agents integrate with it, and what is self-hostable?
answeredSDKs — Mem0 provides four entry points in the Python SDK (mem0/__init__.py:1-6): Memory (self-hosted sync), AsyncMemory (self-hosted async), MemoryClient (hosted platform sync), and AsyncMemoryClient (hosted platform async). The TypeScript SDK (mem0-ts/) mirrors this with a MemoryClient for the hosted API and OSS memory for self-hosted. The CLI packages (cli/python/ and cli/node/) provide command-line access.
REST API — the server/ directory (server/main.py) is a FastAPI application wrapping the Python SDK. It exposes the full CRUD surface (add, search, get, get_all, update, delete, delete_all, history, reset) behind JWT authentication. The server includes admin auth, API key management, rate limiting, and request logging. It runs on Docker Compose with PostgreSQL+pgvector and Neo4j — the SDK code is mounted for hot reload.
MCP server — multiple coding-agent plugins expose Mem0 as an MCP tool (search_memories), notably integrations/cursor-plugin/, integrations/claude-code-plugin/, integrations/codex-plugin/, integrations/kimi-plugin/, integrations/antigravity-plugin/, and the shared integrations/mem0-agent-plugin/ (integrations/mem0-agent-plugin/core/mcp_server.py:1-60). The shared Python core in integrations/agent-plugin-core/python/mcp_server.py provides the runtime that all native plugins use.
Framework plugins — integrations/vercel-ai-sdk/ wraps Mem0 as a Vercel AI SDK provider (@mem0/vercel-ai-provider). integrations/openclaw/, integrations/pi-agent-plugin/, integrations/deepseek-plugin/ are editor/agent integrations. integrations/n8n-nodes-mem0/ is an n8n community node. integrations/zapier-mem0/ is a Zapier app. integrations/mem0-strands/ is a Python Strands MemoryStore.
Self-hostable parts — the entire Python SDK (Memory/AsyncMemory) is fully self-hostable and open-source (Apache 2.0). Required services depend on the vector store choice: you need a running vector database (Qdrant, pgvector, Elasticsearch, etc.) and an LLM provider API key. The Docker Compose stack bundles FastAPI + pgvector + Neo4j for a turnkey solution. The hosted MemoryClient calls api.mem0.ai and requires an API key. The OSS SDK rejects certain platform-only features (decay, temporal search with reference_date/timestamp) at runtime.
server/docker-compose.yaml runs the FastAPI app, pgvector/pgvector:pg17 and the dashboard. It has no Neo4j service, and graph memory is no longer in the SDK. Rate limits (slowapi) apply only to the auth routes. The coding-agent plugins expose only a search_memories MCP tool and call the hosted api.mem0.ai unless MEM0_API_URL is set, so they do not wrap the self-hosted Memory class.