# Agent memory layers

> Libraries and services that give LLM agents long-term memory — extracting, storing, updating and recalling what matters across sessions.

Agent memory layers let an LLM application keep what it learned in one session and use it in the next. On write, something decides what to keep. On read, a search or an agent decides what goes back into the prompt. The ten projects here store different units: short facts about a user, time-stamped edges in a knowledge graph, verbatim conversation episodes, or observations about a coding agent's tool calls.

Four axes separate them. The write path can be one LLM call, a chain of calls, or none at all. The read path can return ranked records, a finished answer, or a block injected into your prompt automatically. Some projects consolidate or invalidate facts in the background, and others only append. Isolation can be enforced access control or a filter on an id you pass. When choosing, also check which databases you must run, and whether extraction runs in the open code or in a vendor's hosted service.


## Projects

- [thedotmack/claude-mem](https://llms-technical-reviews.com/p/claude-mem/) — Coding-agent memory plugin: hooks capture tool calls, an observer LLM writes XML observations into SQLite and Chroma for next-session recall. (★97034, TypeScript)
- [mem0ai/mem0](https://llms-technical-reviews.com/p/mem0/) — Python and TypeScript memory layer: an LLM extracts facts from chats into a vector store, recalled by semantic, BM25 and entity scoring. (★66671, Python)
- [vectorize-io/hindsight](https://llms-technical-reviews.com/p/hindsight/) — Postgres-based agent memory server: LLM fact extraction, four-arm recall with RRF and cross-encoder, LLM consolidation, agentic reflect. (★46262, Python)
- [getzep/graphiti](https://llms-technical-reviews.com/p/graphiti/) — Temporal knowledge-graph memory for agents; LLM-extracted facts with validity dates, stored in Neo4j/FalkorDB, hybrid-searched. (★31488, Python)
- [topoteretes/cognee](https://llms-technical-reviews.com/p/cognee/) — Python memory engine that LLM-extracts a knowledge graph from data into graph and vector stores, then answers queries over it. (★31476, Python)
- [supermemoryai/supermemory](https://llms-technical-reviews.com/p/supermemory/) — Client side of a hosted memory API: remote MCP server with widget, AI SDK/OpenAI/Mastra middleware and SDKs; the engine is closed-source. (★31121, TypeScript)
- [MemoriLabs/Memori](https://llms-technical-reviews.com/p/memori/) — Python/TS SDK that wraps your LLM client, extracts facts via Memori's hosted API, and stores and recalls them in your own SQL or Mongo DB. (★17077, Python)
- [MemTensor/MemOS](https://llms-technical-reviews.com/p/memos/) — Python memory server (graph DB + vectors, LLM extraction, async scheduler) plus a separate local SQLite plugin for agent harnesses. (★11730, TypeScript)
- [plastic-labs/honcho](https://llms-technical-reviews.com/p/honcho/) — Memory server where background LLM workers distil peer messages into conclusions that agents query through a chat endpoint. (★7488, Python)
- [MemMachine/MemMachine](https://llms-technical-reviews.com/p/memmachine/) — FastAPI memory server: raw episodes in a graph or vector store, plus LLM-extracted profile facts updated in the background. (★3055, Python)

In the research queue (not yet published): letta-ai/letta.

## Comparison questions

- [How are memories extracted from interactions?](https://llms-technical-reviews.com/memory/q/extraction/) — What gets stored (facts, events, preferences); LLM extraction prompts; deduplication and conflict handling at write time.
- [How are memories stored?](https://llms-technical-reviews.com/memory/q/storage/) — Vector, graph, key-value or SQL backends; the memory schema; embeddings used; pluggable stores.
- [How are memories retrieved and injected into the prompt?](https://llms-technical-reviews.com/memory/q/retrieval/) — Search strategy (semantic, keyword, graph, temporal); ranking and filtering; how results reach the LLM context.
- [How are memories updated, consolidated or forgotten?](https://llms-technical-reviews.com/memory/q/lifecycle/) — Update/merge logic; summarisation or consolidation; decay, TTL and deletion; versioning or history.
- [How is memory scoped and isolated?](https://llms-technical-reviews.com/memory/q/scoping/) — User / agent / session / tenant scoping; multi-tenancy; access control; privacy controls.
- [How do agents integrate with it, and what is self-hostable?](https://llms-technical-reviews.com/memory/q/integration/) — SDKs, REST API, MCP server, framework plugins; required services; open vs hosted-only parts.

Full comparison: https://llms-technical-reviews.com/compare/memory/