How are memories extracted from interactions?
What gets stored (facts, events, preferences); LLM extraction prompts; deduplication and conflict handling at write time.
Verdict
Graphiti and Hindsight extract most carefully. Graphiti turns each message into entity-to-entity facts with validity dates, and Hindsight turns input into selective, dated facts with causal links. Mem0 and Honcho are the cheapest fact extractors. MemMachine and claude-mem keep or compress raw interactions instead of distilling facts.
Atomic facts from an LLM. Mem0 makes one add-only call. The call sees the new messages, the last 10 messages and the 10 nearest memories, and an MD5 hash check against those neighbours is the only deduplication. Honcho waits until a batch reaches 512 tokens or 30 minutes. Its deriver then makes one structured call that extracts only explicit facts from messages tagged target="true". An exact repeat increments times_derived on the existing row. Hindsight asks for what/when/where/who/why facts with entities and causal relations. Its default concise prompt drops anything not worth recalling “in 6 months”, and a background consolidation job deduplicates later.
Graphs. Graphiti extracts entities, resolves them cheapest first (exact name, then MinHash with Jaccard >= 0.9, then an LLM) and only then extracts edges. A contradicting edge closes the old one. Cognee asks for a typed KnowledgeGraph and a summary for every chunk.
Raw first, structure later. MemMachine stores every message verbatim with no LLM. A background job then turns batches of new messages into add/delete commands on a tag → feature → value profile. MemOS stores raw chunks at write time, and a scheduler task later replaces them with extracted facts, tool traces, skills and preferences. Contradiction merging runs only with MOS_ENABLE_REORGANIZE, which is off by default. In claude-mem, an observer LLM (claude-haiku-4-5 by default) writes an XML observation for each tool call, so every tool call costs model tokens.
Not in the repository. Memori sends every turn to its hosted sdk/augmentation endpoint, even in BYODB mode, and streamed responses are never augmented. Supermemory’s extraction runs inside a closed-source engine.
Pick: Graphiti when facts change and you need to know when. Pick: Mem0 or Honcho for cheap per-user facts. Pick: MemMachine when the verbatim conversation must survive.
Per-project answers
thedotmack/claude-mem
answeredExtraction runs in two modes. Per-event extraction: The observationHandler (src/cli/handlers/observation.ts:32) captures each tool-use via a PostToolUse hook and spools it as an agent_event. The ProviderObservationGenerator (src/server/generation/ProviderObservationGenerator.ts:160) loads batched events and calls buildServerGenerationPrompt (src/server/generation/providers/shared/prompt-builder.ts:46) to assemble them into a single-turn LLM prompt. Events are wrapped in <agent_event> XML with id, event_type, source_adapter, privacy-stripped payload, and timestamp. The prompt asks the LLM to return <observation> blocks with type, title, subtitle, facts, narrative, concepts, files_read, files_modified — the type enum comes from the active 'mode' config (observable types like discovery, progress, blocker, decision). Session-summary extraction (processSessionSummaryResponse at src/server/generation/processGeneratedResponse.ts:199) feeds the full session event list and asks for a <summary> with request/investigated/learned/completed/next_steps fields. Both responses are parsed by parseAgentXml (src/sdk/parser.ts:62). Deduplication is handled at write time by a generation_key UNIQUE constraint (src/storage/postgres/observations.ts:103): the key combines generation_job_id, parsed index, and a content digest hash, so retries collapse to one row. The prompt builder also strips <private> and <system-reminder> tags before sending data to the LLM, and the parser salvages unstructured text when content fields are empty (src/sdk/parser.ts:170-191).
mem0ai/mem0
answeredMemories are extracted from conversational messages by an LLM call guided by the ADDITIVE_EXTRACTION_PROMPT system prompt (mem0/configs/prompts.py:468-944). The pipeline (_add_to_vector_store in mem0/memory/main.py:881-1220) works in phases:
- Context gathering — the last 10 messages for the session scope are fetched from the local SQLite message store to resolve pronoun references.
- Existing memory retrieval — the current messages are embedded and used to search the vector store (top-10), returning nearby memories for deduplication reference.
- LLM extraction — the
ADDITIVE_EXTRACTION_PROMPTinstructs the model to produce ADD-only operations: a JSON array of{"id", "text", "attributed_to", "linked_memory_ids"}objects. The prompt covers 12 information categories (personal preferences, important details, plans, professional info, etc.) and includes 11 few-shot examples with detailed quality criteria. For agent-scoped memories (agent_id present, no user_id),AGENT_CONTEXT_SUFFIX(mem0/configs/prompts.py:947-957) adjusts framing to "agent was informed that...". A separateUSER_MEMORY_EXTRACTION_PROMPTandAGENT_MEMORY_EXTRACTION_PROMPT(mem0/configs/prompts.py:63-174) are available for the deprecated non-additive pipeline. - Hash deduplication — each extracted text is MD5-hashed and compared against existing hashes in the vector store payloads and against other texts in the same batch. Duplicates are silently skipped.
- Entity extraction — spaCy-based
extract_entities_batch(mem0/utils/entity_extraction.py) pulls proper nouns, quoted text, and noun compounds from every extracted memory. These entities are deduplicated globally across the batch and upserted into the entity vector store, with linked_memory_ids tracking which memories reference which entities. Entities are matched by exact normalized text first, then by semantic similarity (threshold ≥0.95), avoiding duplicates.
When infer=False, messages are stored verbatim with no LLM extraction — each non-system message becomes one memory record.
text and attributed_to from each extracted item. The linked_memory_ids the prompt asks for are not stored. Hash deduplication compares only against the 10 nearest existing memories.vectorize-io/hindsight
answeredWhen content is retained, the pipeline first chunks oversized input at sentence boundaries (chunk_text, fact_extraction.py:2580-2584), then runs LLM-based fact extraction. The extractor uses one of four prompt templates — concise (default), verbose, verbatim, or custom — assembled by _build_extraction_prompt_and_schema (fact_extraction.py:1561-1712). The base prompt (_BASE_FACT_EXTRACTION_PROMPT, line 1058) asks the LLM to extract structured facts with fields: what, when, where, who, why, fact_type (world/assistant), entities, and optional causal_relations. The default concise mode (_CONCISE_GUIDELINES, line 1121) emphasises selectivity: "Would this be useful to recall in 6 months? If no, skip it." Chunks are extracted in parallel via asyncio.gather (fact_extraction.py:2600-2621). Each chunk returns a FactExtractionResponse (pydantic model) parsed via structured output/JSON mode depending on the LLM provider. Extracted facts are then embedded (via the configurable Embeddings class supporting local SentenceTransformers, ONNX, OpenAI, Cohere, etc.; embeddings.py:1-80), written as memory_units rows in PostgreSQL, and entity resolution runs via entity_resolver.py using trigram similarity for fuzzy dedup within the batch (entity_resolver.py:79-80). A trigram-similarity threshold of 1.0 (identical trigram sets) deduplicates names differing only in punctuation or case. Configurable entity labels (entity_labels in config) constrain extraction to a taxonomy. The LLM is also instructed to convert relative temporal expressions to absolute dates (line 1098-1101) and maintain coreference resolution (line 1075-1080).
getzep/graphiti
answeredMemories are extracted from interactions through an LLM-driven pipeline in graphiti_core/utils/maintenance/node_operations.py and edge_operations.py. Each ingested interaction is wrapped as an EpisodicNode and passed through extract_nodes() (line 70) which prompts the LLM with the episode content, previous episodes for context, and a type catalog (Entity plus any user-defined Pydantic entity types). The prompt returns ExtractedEntity objects — name, entity_type_id, and episode_indices. These become EntityNode objects with labels and an initially empty summary. Simultaneously, extract_edges() (edge_operations.py:121) prompts the LLM again, this time with the extracted nodes and the episode, returning ExtractedEdge objects (source, target, relation_type, fact, temporal markers). These become EntityEdge objects. Both paths support message/text/JSON source types and multi-episode batches with per-node/per-edge episode attribution.
At write-time deduplication is two-stage. First, _collapse_exact_duplicate_extracted_nodes() (node_operations.py:348) merges same-message duplicates that share a normalized name, keeping the more specific label set. Then resolve_extracted_nodes() (line 644) performs a two-tier semantic dedup: (1) cosine-similarity search against existing nodes — an exact match short-circuits via _resolve_with_similarity(); (2) unresolved nodes are sent to a dedupe prompt (prompts/dedupe_nodes.py) that returns duplicate-or-new verdicts, and promoted nodes absorb attributes from their extracted counterparts. Edge dedup follows the same pattern: resolve_extracted_edges() (line 332) searches for existing edges between the same endpoint nodes using hybrid search, then the LLM dedupe prompt (resolve_edge in prompts/dedupe_edges.py) selects matching edges or marks contradictions for invalidation. Conflict resolution for edges is temporal: resolve_edge_contradictions() (line 547) uses valid_at/invalid_at dates — a new edge invalidates an old one only if the old edge's valid range predates the new edge's validity.
topoteretes/cognee
answeredMemories are extracted through the cognify pipeline, entered via remember() -> add() -> cognify() (cognee/api/v1/remember/remember.py:867). The task extract_graph_from_data (cognee/tasks/graph/extract_graph_from_data.py:198) sends each chunk via asyncio.gather to extract_content_graph(), which uses Instructor or litellm-native structured output to LLMs, producing a KnowledgeGraph with typed Node and Edge objects (cognee/shared/data_models.py:49-77). What gets stored includes entities, relationships, document summaries, chunk text with embeddings, QA turns, agent trace steps, feedback scores, and skill-run records via typed MemoryEntry types: QAEntry, TraceEntry, FeedbackEntry, SkillRunEntry (cognee/memory/entries.py:24-129). Deduplication at write time collapses duplicate node IDs within a chunk (_remove_duplicate_extracted_nodes_by_id at extract_graph_from_data.py:42); graph adapter writes are idempotent upserts keyed on node id (GraphDBInterface contract, graph_db_interface.py:44-58). An optional ontology mode can drop entities not grounded in an OWL file. The improve loop and contradiction detection provide post-extraction consolidation.
supermemoryai/supermemory
answeredMemories are extracted through two paths. Explicit saves: the add_memory tool (apps/mcp/src/server/tools/add-memory.ts:29-63) and guided-save tool (apps/mcp/src/server/tools/guided-save.ts:8-60) accept content from the LLM and call SupermemoryClient.createMemory(), which POSTs to the API. Implicit profile facts: the getProfile() method (apps/mcp/src/server/client/index.ts:270-305) calls the API's profile endpoint which returns two arrays — static (long-lived preferences, traits) and dynamic (recent context, activity). These facts are extracted server-side by the IngestContentWorkflow (referenced in CLAUDE.md) which does content-type detection, AI-powered summarization, automatic tagging, and vector embedding generation via Cloudflare AI. No client-side deduplication — the response includes raw fact strings; deduplication and conflict resolution happen in the upstream API/ingestion pipeline, not in this MCP server's code.
MemoriLabs/Memori
answeredMemories are extracted from LLM conversations through an augmentation pipeline that runs asynchronously after each LLM invocation. The flow begins in handle_post_response (memori/llm/pipelines/post_invoke.py:94-144), which formats the conversation turn (query + response) into a payload and passes it to the MemoryManager, which calls handle_augmentation. For cloud mode the payload is posted to cloud/augmentation; for BYODB mode it is enqueued onto an async runtime backed by a thread-pool with a semaphore (max 50 workers, memori/memory/augmentation/_runtime.py:17-50). The AdvancedAugmentation class (memori/memory/augmentation/augmentations/memori/_augmentation.py:39-286) is the registered handler. It builds an AugmentationPayload containing conversation messages, a previous summary (if one exists), and metadata (entity/process attribution, LLM provider, framework, storage dialect) and sends it to the cloud API via api.augmentation_async(). The API response carries: entity facts (natural-language statements like "the user prefers Python"), semantic triples (structured subject-predicate-object facts), process attributes, and a conversation summary. The response is parsed into a Memories object (memori/memory/_struct.py:102-112) containing Entity, Process, and Conversation sub-objects. Entity facts are then embedded locally via the configured embedding model (default all-MiniLM-L6-v2) before being written. Deduplication is handled at write time by the storage driver: both entity facts and knowledge-graph triples use ON CONFLICT ... DO UPDATE SET num_times = num_times + 1 (memori/storage/drivers/postgresql/_driver.py:260-280 for facts, memori/storage/drivers/postgresql/_driver.py:554-574 for triples), so repeating the same fact increments a frequency counter rather than creating a duplicate row. For cloud-mode users the entire extraction pipeline runs server-side; for BYODB there is a Rust native core path that performs extraction locally. The entity ID is hashed (SHA-256) before being sent to the cloud API for privacy (memori/memory/augmentation/_models.py:15-18).
MemTensor/MemOS
answeredExtraction uses LLM-driven prompts. BaseTextMemory.extract() (general.py:45-79) concatenates messages into SIMPLE_STRUCT_MEM_READER_PROMPT (mem_reader_prompts.py:1-101) which instructs the LLM to extract facts, events, preferences, plans from user perspective, resolving time/pronoun references, writing third-person. Returns JSON with memory list of key/value/tags/memory_type. PreferenceTextMemory (preference.py:69-79) delegates to ExtractorFactory. Write-time dedup: NaiveAdder (adder.py:53-165) uses LLM judges to compare new vs existing memories. Tree-text: NodeHandler.detect() (handler.py:30-74) finds embedding-similar candidates (0.8 threshold), LLM classifies as contradictory/redundant/independent; resolve() (handler.py:76-128) fuses or hard-updates, archives originals.
NodeHandler conflict detection (0.8 similarity threshold) runs only when MOS_ENABLE_REORGANIZE=true, which is off by default. It is not part of the default write path.plastic-labs/honcho
answeredMemory extraction is driven by the Deriver (src/deriver/deriver.py), which processes batches of incoming messages via a single structured-output LLM call — the 'minimal deriver' approach. Messages are enqueued by src/deriver/enqueue.py on message creation, consumed by src/deriver/queue_manager.py, and dispatched through consumer.py's process_item() to process_representation_tasks_batch(). The prompt (minimal_deriver_prompt in src/deriver/prompts.py) wraps each message in an XML tag with target="true" or target="false" and instructs the model to extract only explicit atomic facts about the target peer from target="true" messages; other messages serve solely as interpretive context. The output is a PromptRepresentation Pydantic schema enforcing structured JSON with an explicit list of ExplicitObservationBase objects. The Deriver then converts these into a Representation and saves them via RepresentationManager.save_representation() (src/crud/representation.py:74). Deduplication at write time is two-phase: exact-content dedup (normalized content + level + session for explicit documents) drops identical text within a batch or against existing documents, incrementing the times_derived counter on the existing row instead of inserting a duplicate (src/crud/document.py:677-705). Semantic dedup (cosine distance < 0.05) optionally replaces or rejects near-duplicate text via token-set comparison (src/crud/document.py:1369-1422). Custom instructions per workspace/peer can be threaded into the extraction prompt via reasoning configuration, capped at DERIVER__MAX_CUSTOM_INSTRUCTIONS_TOKENS (default 2000).
MemMachine/MemMachine
answeredMemMachine extracts memories in two parallel pipelines. Episodic memory stores entire conversation turns (episodes) as-is, via EpisodicMemory.add_memory_episodes() (packages/server/src/memmachine_server/episodic_memory/episodic_memory.py:208). Each episode carries content (the message text), producer_id, producer_role, metadata, and episode_type. No LLM is involved — the raw text plus its metadata is committed to short-term and long-term stores. Profile/semantic memory goes further: an asynchronous IngestionService (packages/server/src/memmachine_server/semantic_memory/semantic_ingestion.py:68) polls for uningested episodes, then calls llm_feature_update() (packages/server/src/memmachine_server/semantic_memory/semantic_llm.py:69) with a two-part prompt. The prompt contains the existing feature set (formatted as tag→feature→value triples) and the new message content. The LLM emits SemanticCommand objects — ADD or DELETE — which are applied to extract atomic facts like name, preferences, or inferred demographics. The update prompt (packages/server/src/memmachine_server/semantic_memory/util/semantic_prompt_template.py:6) instructs the LLM to be a "wide, early layer" performing "edge detection" — extracting multiple distinct facts in parallel, even making low-confidence inferences. Deduplication happens at consolidation time (not write time): when a tag accumulates >=20 features (configurable consolidation_threshold), _deduplicate_features() (semantic_ingestion.py:447) merges overlapping entries using a second LLM call (llm_consolidate_features()), which can delete redundant memories and produce consolidated replacements. Conflict handling is explicit: the LLM issues delete-then-add pairs to update a value, and the consolidation prompt deliberately deletes unspecified memories (a keep-list design, not a delete-list).