# Agent memory layers: comparison

> Libraries and services that give LLM agents long-term memory — extracting, storing, updating and recalling what matters across sessions.

Canonical page: https://llms-technical-reviews.com/compare/memory/

## How are memories extracted from interactions?

Graphiti and Hindsight extract most carefully. Graphiti turns each message into entity-to-entity facts with validity dates, and Hindsight turns input into selective, dated facts with causal links. Mem0 and Honcho are the cheapest fact extractors. MemMachine and claude-mem keep or compress raw interactions instead of distilling facts.

**Atomic facts from an LLM.** [Mem0](/p/mem0/) makes one add-only call. The call sees the new messages, the last 10 messages and the 10 nearest memories, and an MD5 hash check against those neighbours is the only deduplication. [Honcho](/p/honcho/) waits until a batch reaches 512 tokens or 30 minutes. Its deriver then makes one structured call that extracts only explicit facts from messages tagged `target="true"`. An exact repeat increments `times_derived` on the existing row. [Hindsight](/p/hindsight/) asks for what/when/where/who/why facts with entities and causal relations. Its default concise prompt drops anything not worth recalling "in 6 months", and a background consolidation job deduplicates later.

**Graphs.** [Graphiti](/p/graphiti/) extracts entities, resolves them cheapest first (exact name, then MinHash with Jaccard >= 0.9, then an LLM) and only then extracts edges. A contradicting edge closes the old one. [Cognee](/p/cognee/) asks for a typed `KnowledgeGraph` and a summary for every chunk.

**Raw first, structure later.** [MemMachine](/p/memmachine/) stores every message verbatim with no LLM. A background job then turns batches of new messages into `add`/`delete` commands on a `tag → feature → value` profile. [MemOS](/p/memos/) stores raw chunks at write time, and a scheduler task later replaces them with extracted facts, tool traces, skills and preferences. Contradiction merging runs only with `MOS_ENABLE_REORGANIZE`, which is off by default. In [claude-mem](/p/claude-mem/), an observer LLM (`claude-haiku-4-5` by default) writes an XML observation for each tool call, so every tool call costs model tokens.

**Not in the repository.** [Memori](/p/memori/) sends every turn to its hosted `sdk/augmentation` endpoint, even in BYODB mode, and streamed responses are never augmented. [Supermemory](/p/supermemory/)'s extraction runs inside a closed-source engine.

Pick: Graphiti when facts change and you need to know when.
Pick: Mem0 or Honcho for cheap per-user facts.
Pick: MemMachine when the verbatim conversation must survive.

Per-project answers: https://llms-technical-reviews.com/memory/q/extraction/index.md

## How are memories stored?

If you want one database, Hindsight puts everything in a single Postgres: vectors, keyword index, entity links and the job queue. If you need a real knowledge graph, choose Graphiti (with a graph database server) or Cognee (with three stores, all local files by default).

**Postgres or plain SQL.** [Hindsight](/p/hindsight/) keeps raw facts and consolidated beliefs in one `memory_units` table, told apart by `fact_type` (`world`, `experience`, `observation`). With no configuration it starts an embedded Postgres. [Honcho](/p/honcho/) stores conclusions as `Document` rows in a collection for each (observer, observed) pair, with an HNSW index. Turbopuffer, LanceDB, Qdrant or Chroma can take the vectors. [Memori](/p/memori/) writes plain relational tables to your PostgreSQL, MySQL, SQLite, Oracle, MongoDB or a similar database. Embeddings are stored as blobs with no vector index, and recall scans up to 1,000 of them in memory. [claude-mem](/p/claude-mem/) uses one local SQLite file with FTS5, plus Chroma run as a `chroma-mcp` child process. Its optional server mode moves to Postgres.

**Vector store plus side tables.** [Mem0](/p/mem0/) stores each memory as one vector with a flat payload, in any of about two dozen vector stores (Qdrant by default). Entities go in a second collection, and history goes in a local SQLite file. [MemMachine](/p/memmachine/) writes raw episodes to SQL. Long-term vectors go to Neo4j or NebulaGraph, or to Qdrant, Milvus or SQLite, and profile features go to pgvector.

**Graph database first.** [Graphiti](/p/graphiti/) stores each fact as an edge with its own embedding and four timestamps, in Neo4j, FalkorDB or Neptune. Its Kuzu extra is marked deprecated. [Cognee](/p/cognee/) writes to a graph, a vector and a relational store at once, but its defaults (embedded Kuzu/Ladybug, LanceDB, SQLite) are local files. [MemOS](/p/memos/) keeps typed memory items as graph nodes. Its Docker setup runs Neo4j Community and Qdrant.

**Not inspectable.** [Supermemory](/p/supermemory/)'s storage sits behind its hosted API or a closed local binary. The client types show only that memories carry `version`, `isLatest` and `isForgotten` flags.

Pick: Hindsight or Honcho if Postgres is what you already run.
Pick: Mem0 to reuse an existing vector database.
Pick: Graphiti or Cognee when the relationships must be queryable.

Per-project answers: https://llms-technical-reviews.com/memory/q/storage/index.md

## How are memories retrieved and injected into the prompt?

Hindsight has the most complete ranking pipeline for returning records. Cognee and Honcho return an LLM-written answer by default. Memori, claude-mem and Supermemory insert memory into the prompt for you.

**Ranked records.** [Hindsight](/p/hindsight/) runs four arms for each fact type: semantic, BM25, graph links and a temporal arm that uses a date range parsed from the query. It fuses them with RRF (k = 60), rescores with a cross-encoder, diversifies with MMR and trims to `max_tokens`. [Graphiti](/p/graphiti/) searches edges and nodes by BM25, cosine and BFS, but episodes by BM25 only. Its filters on `valid_at`/`invalid_at` allow "as of" queries. [Mem0](/p/mem0/) adds BM25 and entity boosts to semantic candidates, so a memory that matches only by keyword is never returned. [MemMachine](/p/memmachine/) matches small derivatives and returns whole episodes with their neighbours. Without a reranker, the declarative backend disables long-term memory and only logs a warning. Its `formalize_query_with_context` helper is never called, so you build the prompt yourself. [MemOS](/p/memos/) defaults to graph plus vector recall, and its BM25, full-text and chain-of-thought paths are off.

**Answers.** [Cognee](/p/cognee/) routes most queries to `HYBRID_COMPLETION`, which searches chunks, summaries, entities and edges and ends in one completion. [Honcho](/p/honcho/)'s chat endpoint runs a tool-using agent over conclusions and messages. Every reasoning level defaults to the same `gpt-5.4-mini`, so higher levels mainly add tool rounds.

**Automatic injection.** [Memori](/p/memori/) patches your LLM client and adds a `<memori_context>` block before each call. It scores by `0.85·cosine + 0.15·BM25` over at most 1,000 stored embeddings. At session start, [claude-mem](/p/claude-mem/) injects the 50 most recent observations, not the most relevant ones. Semantic search runs only when the agent calls the MCP `search` tool. [Supermemory](/p/supermemory/)'s middleware injects a profile of stable and recent facts, but the ranking happens on the server and cannot be inspected.

Pick: Hindsight for the best ranked recall you can run yourself.
Pick: Cognee or Honcho when you want an answer, not rows.
Pick: Memori to add memory to existing LLM calls without changing your code.

Per-project answers: https://llms-technical-reviews.com/memory/q/retrieval/index.md

## How are memories updated, consolidated or forgotten?

Graphiti handles changing facts best: a contradicting fact closes the old edge with an `invalid_at` date instead of deleting it. Hindsight and Honcho go furthest at consolidation, with background LLM jobs that merge facts into higher-level beliefs.

**Background consolidation.** After each retain, [Hindsight](/p/hindsight/) asks an LLM to create, update or delete "observations", and tracks `proof_count` and source ids for each one. Recency affects ranking only, and nothing expires. [Honcho](/p/honcho/) "dreams" after 50 new explicit facts and at least 8 hours since the last dream. A deduction agent can delete conclusions and rewrite the peer card, and an induction agent adds generalisations. Its surprisal sampling is off by default. [Cognee](/p/cognee/) re-extracts only the changed chunks on `update()`, and its `improve()` pass feeds sessions and feedback back into the graph. [MemMachine](/p/memmachine/) asks an LLM which features to keep once a tag holds 20. The `consolidation_threshold` setting is never passed through, so the threshold is always 20. Features keep no history.

**Time-stamped invalidation.** [Graphiti](/p/graphiti/) closes contradicted edges with `invalid_at` and `expired_at`. Its `remove_episode` is partial: it does not restore edges that the removed episode had invalidated.

**Append and count.** [Mem0](/p/mem0/) only adds memories. It logs explicit updates and deletes in SQLite, and decay is a hosted-only feature. [Memori](/p/memori/) increments `num_times` when a fact repeats exactly, has no contradiction handling, and deletes per entity, all or nothing. [claude-mem](/p/claude-mem/) never forgets by age. Its ACT-R ranking and cross-session dedup are both off by default. [MemOS](/p/memos/) merges conflicting memories only when its reorganizer is enabled, which it is not by default. A feedback endpoint can rewrite memories through an LLM.

**Hidden.** [Supermemory](/p/supermemory/)'s types show version chains, `isForgotten` and `forgetAfter`, but the logic runs on the closed server.

Pick: Graphiti when you need to ask what was true at a given time.
Pick: Hindsight or Honcho for memory that consolidates on its own.
Pick: Mem0 or Memori when your application owns corrections.

Per-project answers: https://llms-technical-reviews.com/memory/q/lifecycle/index.md

## How is memory scoped and isolated?

Cognee and Honcho enforce isolation in the memory layer itself. Most other projects filter on an id that the caller passes and trust the application to pass the right one.

**Access control built in.** [Cognee](/p/cognee/) stores read, write, delete and share rights as ACL rows per dataset. When the backends support it, each user and dataset gets its own graph and vector database, and authentication is forced on. [Honcho](/p/honcho/) uses JWT claims for admin, workspace, peer and session, and keeps a separate collection for each (observer, observed) pair. An empty session allowlist returns nothing, and only explicit conclusions are served under one. Auth is off until you set `AUTH_USE_AUTH`. [MemOS](/p/memos/) scopes each request to memory cubes (one per user by default). It has scoped, expiring API keys, but `AUTH_ENABLED` is off by default, and a database per user is opt-in. In [claude-mem](/p/claude-mem/)'s server mode, API keys are bound to a team (and optionally a project) with read and write scopes. Local mode is single-user, scoped by project.

**Partition by id.** [Hindsight](/p/hindsight/) filters every query by `bank_id` and supports tag modes (`any`, `all`, the strict variants and `exact`) inside a bank. A Postgres schema per tenant exists only when `HINDSIGHT_API_TENANT_EXTENSION` is set. [Graphiti](/p/graphiti/) uses `group_id`. On FalkorDB that becomes a separate graph, but on Neo4j it is only a property filter, and neither server has auth. [Mem0](/p/mem0/) requires `user_id`, `agent_id` or `run_id` on writes and searches, but `get`, `update` and `delete` take only a memory id. [Memori](/p/memori/) skips memory entirely without an `entity_id`. Only BYODB augmentation hashes the ids, and cloud mode sends them raw. [Supermemory](/p/supermemory/) scopes by container tag (default `sm_project_default`) and leaves enforcement to its API.

**Weak inside a project.** [MemMachine](/p/memmachine/) partitions by organisation and project. User and agent separation is a metadata filter built by the client that the server does not enforce, and the rolling summary is shared by the whole project.

Pick: Cognee or Honcho when several users share one deployment.
Pick: Hindsight for clean per-agent banks with tag filters.
Pick: Mem0 or Graphiti when your backend already controls which ids reach memory.

Per-project answers: https://llms-technical-reviews.com/memory/q/scoping/index.md

## How do agents integrate with it, and what is self-hostable?

Hindsight is the easiest full server to self-host: one process with an embedded Postgres, local embeddings and a local reranker, so it needs only an LLM key. Memori and Supermemory are the exceptions to self-hosting, because their core memory step runs on vendor services.

**Servers you run.** [Hindsight](/p/hindsight/) ships REST, MCP, generated Python, TypeScript, Rust and Go clients, and hooks for many coding agents. [Honcho](/p/honcho/) runs as an API process and a deriver process over Postgres with pgvector. It has Python and TypeScript SDKs, MCP and coding-agent plugins, uses `gpt-5.4-mini` for every stage by default, and is AGPL-3.0. [MemMachine](/p/memmachine/) serves `/api/v2` plus MCP. Its Compose file starts Postgres, Neo4j and Qdrant, and it has adapters for LangChain, CrewAI, LlamaIndex, Dify and n8n. [MemOS](/p/memos/) runs a FastAPI server over Neo4j and Qdrant. Its MCP server wraps the older in-process `MOS` API, and a separate TypeScript SQLite plugin serves agent harnesses. [Cognee](/p/cognee/) offers an SDK, a server, a CLI and `cognee-mcp`, all running on local files by default.

**Libraries.** [Mem0](/p/mem0/)'s `Memory` class runs locally. Its coding-agent plugins call `api.mem0.ai` unless `MEM0_API_URL` is set. [Graphiti](/p/graphiti/) is `graphiti_core` plus FastAPI and MCP wrappers (stdio, SSE or streamable HTTP). It needs a graph database, and neither wrapper has auth. [Memori](/p/memori/) wraps your OpenAI, Anthropic, Gemini or Bedrock client. In BYODB mode your data stays in your database, but extraction always calls Memori's hosted API.

**Coding-agent plugin.** [claude-mem](/p/claude-mem/) installs hooks into Claude Code, Codex, Cursor and others. Behind them run a local Bun worker, SQLite and `chroma-mcp`.

**Client for a hosted engine.** [Supermemory](/p/supermemory/)'s repository holds the MCP Worker, AI SDK middleware and SDKs. The engine is closed-source, both as the hosted API and as the downloadable local binary.

Pick: Hindsight for a single self-hosted server.
Pick: claude-mem for zero-effort memory in a coding agent.
Pick: Mem0 or Graphiti to embed memory in your own Python service.

Per-project answers: https://llms-technical-reviews.com/memory/q/integration/index.md

## Projects

- [thedotmack/claude-mem](https://llms-technical-reviews.com/p/claude-mem/index.md) — Coding-agent memory plugin: hooks capture tool calls, an observer LLM writes XML observations into SQLite and Chroma for next-session recall.
- [mem0ai/mem0](https://llms-technical-reviews.com/p/mem0/index.md) — Python and TypeScript memory layer: an LLM extracts facts from chats into a vector store, recalled by semantic, BM25 and entity scoring.
- [vectorize-io/hindsight](https://llms-technical-reviews.com/p/hindsight/index.md) — Postgres-based agent memory server: LLM fact extraction, four-arm recall with RRF and cross-encoder, LLM consolidation, agentic reflect.
- [getzep/graphiti](https://llms-technical-reviews.com/p/graphiti/index.md) — Temporal knowledge-graph memory for agents; LLM-extracted facts with validity dates, stored in Neo4j/FalkorDB, hybrid-searched.
- [topoteretes/cognee](https://llms-technical-reviews.com/p/cognee/index.md) — Python memory engine that LLM-extracts a knowledge graph from data into graph and vector stores, then answers queries over it.
- [supermemoryai/supermemory](https://llms-technical-reviews.com/p/supermemory/index.md) — Client side of a hosted memory API: remote MCP server with widget, AI SDK/OpenAI/Mastra middleware and SDKs; the engine is closed-source.
- [MemoriLabs/Memori](https://llms-technical-reviews.com/p/memori/index.md) — Python/TS SDK that wraps your LLM client, extracts facts via Memori's hosted API, and stores and recalls them in your own SQL or Mongo DB.
- [MemTensor/MemOS](https://llms-technical-reviews.com/p/memos/index.md) — Python memory server (graph DB + vectors, LLM extraction, async scheduler) plus a separate local SQLite plugin for agent harnesses.
- [plastic-labs/honcho](https://llms-technical-reviews.com/p/honcho/index.md) — Memory server where background LLM workers distil peer messages into conclusions that agents query through a chat endpoint.
- [MemMachine/MemMachine](https://llms-technical-reviews.com/p/memmachine/index.md) — FastAPI memory server: raw episodes in a graph or vector store, plus LLM-extracted profile facts updated in the background.