LLMs Technical Reviews

How are memories retrieved and injected into the prompt?

Search strategy (semantic, keyword, graph, temporal); ranking and filtering; how results reach the LLM context.

Verdict

Hindsight has the most complete ranking pipeline for returning records. Cognee and Honcho return an LLM-written answer by default. Memori, claude-mem and Supermemory insert memory into the prompt for you.

Ranked records. Hindsight runs four arms for each fact type: semantic, BM25, graph links and a temporal arm that uses a date range parsed from the query. It fuses them with RRF (k = 60), rescores with a cross-encoder, diversifies with MMR and trims to max_tokens. Graphiti searches edges and nodes by BM25, cosine and BFS, but episodes by BM25 only. Its filters on valid_at/invalid_at allow “as of” queries. Mem0 adds BM25 and entity boosts to semantic candidates, so a memory that matches only by keyword is never returned. MemMachine matches small derivatives and returns whole episodes with their neighbours. Without a reranker, the declarative backend disables long-term memory and only logs a warning. Its formalize_query_with_context helper is never called, so you build the prompt yourself. MemOS defaults to graph plus vector recall, and its BM25, full-text and chain-of-thought paths are off.

Answers. Cognee routes most queries to HYBRID_COMPLETION, which searches chunks, summaries, entities and edges and ends in one completion. Honcho’s chat endpoint runs a tool-using agent over conclusions and messages. Every reasoning level defaults to the same gpt-5.4-mini, so higher levels mainly add tool rounds.

Automatic injection. Memori patches your LLM client and adds a <memori_context> block before each call. It scores by 0.85·cosine + 0.15·BM25 over at most 1,000 stored embeddings. At session start, claude-mem injects the 50 most recent observations, not the most relevant ones. Semantic search runs only when the agent calls the MCP search tool. Supermemory’s middleware injects a profile of stable and recent facts, but the ranking happens on the server and cannot be inspected.

Pick: Hindsight for the best ranked recall you can run yourself. Pick: Cognee or Honcho when you want an answer, not rows. Pick: Memori to add memory to existing LLM calls without changing your code.

Per-project answers

thedotmack/claude-mem

answered

Retrieval operates at multiple levels. SessionStart context injection: Every new Claude Code session triggers the contextHandler (src/cli/handlers/context.ts:66), which calls /api/context/inject. This returns a rendered timeline of observations and summaries limited by a token budget (~10k chars). The ContextBuilder (src/services/context/ContextBuilder.ts) calls queryObservationsMulti which fetches observations by recency filtered by the active mode's observation types and concepts (src/services/context/ObservationCompiler.ts:82-96). An optional ACT-R reinforcement ranking (src/services/reinforcement/rank.ts:85-106) re-ranks older observations by blended recency+reinforcement score when CLAUDE_MEM_REINFORCE_ALPHA > 0. Full-text search: On SQLite, the MemoryItemsRepository.search method (src/storage/sqlite/memory-items.ts:260-291) builds FTS5 queries with term tokenization, NFKC unicode normalization and alternative expansion. On Postgres, the PostgresObservationRepository.search method (src/storage/postgres/observations.ts:176-258) uses websearch_to_tsquery('english', ...) with ts_rank ordering and GIN indexes. Hybrid search: The HybridSearchStrategy (src/services/worker/search/strategies/HybridSearchStrategy.ts:15) combines SQLite keyword matches with Chroma vector similarity ranking for file-aware search. MCP tools: The recall MCP server (src/server/mcp/recall-mcp-server.ts:140) exposes three read-only tools — search (ranked FTS results), context (search results plus concatenated context string), and recent (newest-first listing). These go through the same Postgres search as the REST API and are scoped/audited per API key. The rendered context for the model includes a header, day-grouped timeline, and the last prior assistant message when enabled.

mem0ai/mem0

answered

Retrieval uses a hybrid scoring pipeline in _search_vector_store (mem0/memory/main.py:1642-1745):

  1. Query processing — the search query is lemmatized for BM25 via lemmatize_for_bm25, and entities are extracted from it via extract_entities for entity-based boosting.
  2. Semantic search — the query is embedded and the vector store's search() retrieves candidates (over-fetched: max(limit * 4, 60)).
  3. Keyword search — if the vector store supports BM25 (Qdrant, Elasticsearch, pgvector), a parallel keyword search runs with the lemmatized query. Raw BM25 scores are sigmoid-normalized into [0, 1] via normalize_bm25 (mem0/utils/scoring.py:43-54), with query-length-adaptive midpoint and steepness.
  4. Entity boost — if the query contains entities, _compute_entity_boosts (mem0/memory/main.py:1747-1827) searches the entity store for each entity (via ThreadPoolExecutor), and any memory linked to a matching entity (threshold ≥0.5) receives a boost of up to 0.5, attenuated by how many memories the entity links to.
  5. Additive scoring — score_and_rank (mem0/utils/scoring.py:60-139) computes combined = (semantic + bm25 + entity_boost) / max_possible, clamped to [0, 1]. The divisor adapts: 1.0 for semantic-only, 2.0 with BM25, 2.5 with BM25+entity, 1.5 with entity only. A semantic threshold gate (default 0.1) excludes low-confidence vectors before combination. Results are sorted descending and trimmed to top_k.
  6. Filtering — before scoring, expired memories (payload expiration_date before today) are excluded unless show_expired=True. Metadata filters support advanced operators: comparison (eq, ne, gt, gte, lt, lte, in, nin, contains, icontains) and compound logic (AND, OR, NOT) via _process_metadata_filters (mem0/memory/main.py:1538-1613).
  7. Reranking — if a RerankerConfig is set and rerank=True, an optional reranker stage (Cohere, HuggingFace, Sentence Transformer, LLM-based, Zero Entropy) reorders results before return.

Results are returned as a dict with a "results" list, each item containing id, memory text, hash, timestamps, scope identifiers, score, and any additional metadata.

vectorize-io/hindsight

answered

Recall (recall_async, memory_engine.py:8676-8735) uses 4-way parallel retrieval per fact type. The orchestration runs through retrieve_all_fact_types_parallel (search/retrieval.py:100-245), which first extracts temporal constraints from the query text (temporal_extraction.py), then calls the store's unified recall_unified method. The four arms are: (1) Semantic — ANN vector search using the configured vector index (HNSW/DiskANN) on the embedding column; (2) BM25 keyword — PostgreSQL @@ tsquery on search_vector (or vchord_bm25/pg_search depending on extension), built by retrieve_semantic_bm25_combined_sql (memories/pg/recall.py:28-100), which emits a single UNION query per fact_type with configurable per-fact-type partial HNSW indexes; (3) Graph — link expansion through memory_links using three signals: entity links (shared entities via unit_entities), semantic links (precomputed kNN graph), and causal links (caused_by), implemented in memories/pg/link_expansion.py:27-47; (4) Temporal — date-range filtering on occurred_start/occurred_end. Results from all arms are fused via Reciprocal Rank Fusion (fusion.py:29-109, RRF formula score(d) = Σ 1/(k+rank(d)) with k=60). After fusion, a cross-encoder reranker (reranking.py:1-38) rescales scores with neural reranking (local CrossEncoder or remote TEI/Cohere). The reranker applies recency boost (_RECENCY_ALPHA=0.2, linear or exponential decay), temporal proximity boost, and proof-count boost. Maximum Marginal Diversity (MMD) is then applied for diversification, and results are trimmed to the max_tokens budget. The reflect endpoint (reflect_async, memory_engine.py:15069-15148) is an agentic loop that sequentially calls tools (lookup mental models, recall facts, search observations, expand chunks) over multiple iterations.

getzep/graphiti

answered

Retrieval is a four-scope hybrid search orchestrated by search() in graphiti_core/search/search.py (line 98). It executes edge, node, episode, and community searches in parallel (via semaphore_gather), each with its own SearchConfig defining which search methods and reranker to use. The config recipes in search_config_recipes.py provide presets: COMBINED_HYBRID_SEARCH_CROSS_ENCODER uses BM25 + cosine similarity + BFS with a cross-encoder reranker.

Each scope supports three search methods defined in search_config.py: cosine_similarity (vector search on name_embedding/fact_embedding), bm25 (fulltext index search), and bfs (graph traversal from origin nodes). Results from each method are fused by a reranker: rrf (reciprocal rank fusion, the default), mmr (maximal marginal relevance for diversity), cross_encoder (a separate scoring model like OpenAI or BGE reranker that re-ranks candidate facts/names), node_distance (ranks by shortest path to a center_node_uuid), or episode_mentions (ranks by mention frequency).

SearchFilters (search_filters.py:55) enables filtering by node labels, edge types (fact types), temporal ranges (valid_at, invalid_at, created_at, expired_at), edge UUIDs, and arbitrary property filters. The group_ids parameter scopes all searches to specific partitions.

For injection into the LLM context, the search() / search_() methods on the Graphiti class (graphiti.py:1653-1755) return SearchResults containing edges (as EntityEdge objects with fact text), nodes, episodes, and communities. The Python SDK makes these directly available. The MCP server exposes search_memory_facts (returns facts as text) and search_nodes (returns entity nodes). The REST API at server/graph_service/routers/retrieve.py exposes /search and /get-memory endpoints that compose query from recent messages, call graphiti.search(), format edges as FactResult objects, and return them as JSON.

Editor's note. Correction: search methods differ by scope. Edges and nodes support cosine_similarity, bm25 and bfs; communities support bm25 and cosine_similarity; episodes support bm25 only (search_config.py).

topoteretes/cognee

answered

Retrieval is handled by recall() (cognee/api/v1/recall/recall.py:347) which wraps authorized_search() with 20+ SearchType strategies (cognee/modules/search/types/SearchType.py:4) including HYBRID_COMPLETION (default), GRAPH_COMPLETION, RAG_COMPLETION, CHUNKS, CHUNKS_LEXICAL, TRIPLET_COMPLETION, and TEMPORAL. The free rule-based router (cognee/api/v1/recall/query_router.py:47) picks the strategy when query_type is omitted: quoted phrases to CHUNKS_LEXICAL, coding keywords to CODING_RULES, default to HYBRID. Results reach LLM context through build_completion_prompts (cognee/modules/retrieval/utils/completion.py:21): system prompt = task template only; user prompt = history + question/context + guidance block. Session entries can short-circuit the graph search (_search_session at recall.py:167). Ranking uses brute-force triplet search over graph edges with personalization weights. Empty graphs do not reach the LLM (skip_completion_on_empty_context flag, base_retriever.py:40).

Editor's note. Clarification: brute-force triplet search is the ranking for GRAPH_COMPLETION. The default HYBRID_COMPLETION runs a chunk/summary vector lane and an entity lane (entity and EdgeType_relationship_name vectors, expanded through the graph) in parallel, then makes one completion (hybrid_retriever.py _retrieve_one).

supermemoryai/supermemory

answered

Retrieval happens through three channels. 1. MCP context prompt (apps/mcp/src/server/prompts/context.ts:12-115): registered as the 'context' prompt, it calls getClient().getProfile() to load static and dynamic fact arrays (capped at CONTEXT_FACT_LIMIT = 8 each, line 12), formats them via formatFactSection (apps/mcp/src/server/space-presentation.ts:89-101), and injects them as a user-role message into the LLM's context window. The format adds a '+N more' line when facts exceed the limit (lines 98-99). 2. get_profile tool (apps/mcp/src/server/tools/get-profile.ts:1-63): returns the full fact arrays as structured content with no cap, plus a text summary. 3. search_memory tool (apps/mcp/src/server/tools/search-memory.ts:1-70): performs hybrid semantic+keyword search via SupermemoryClient.search() (client/index.ts:247-267), which accepts query, limit, threshold (similarity floor), and searchMode: 'hybrid'. Results include similarity scores and the matched text. The context prompt also appends recently-active spaces (lines 71-82) for temporal awareness. There is no separate ranking or filtering layer in this MCP server — the API returns already-ranked results.

MemoriLabs/Memori

answered

Retrieval is a two-phase process executed inside inject_recalled_facts() (memori/llm/pipelines/recall_injection.py:107-226), which is called by every Invoke.invoke() pipeline before the LLM request is sent (memori/llm/invoke/invoke.py:28-31).

Query extraction. The last user message is extracted from the kwargs by extract_user_query() (memori/llm/helpers/query_extraction.py:44-94), which supports all supported provider message formats (OpenAI-style messages, Anthropic, Google contents/request, xAI).

Semantic search. The query is embedded using the same model as write-time (all-MiniLM-L6-v2) via Recall._embed_query() (memori/memory/recall.py:213-218). For BYODB, the entity's stored fact embeddings are loaded from memori_entity_fact (up to MEMORI_RECALL_EMBEDDINGS_LIMIT, default 1000) into a FAISS IndexFlatIP index. Cosine similarity is computed via find_similar_embeddings() (memori/search/_faiss.py:93-133), which normalizes both the query and stored embeddings with L2 normalization before using inner-product search.

Lexical reranking. Candidate facts are reranked by a BM25-style scorer (memori/search/_lexical.py:74-124) that tokenizes query and fact text, removes stopwords, computes TF-IDF scores, and normalizes them. The final rank_score is a weighted combination: w_cos * cosine_sim + w_lex * bm25_score. For short queries (≤2 tokens) lexical weight increases from 0.15 to 0.30 (memori/search/_lexical.py:127-149).

Thresholding and injection. Facts below recall_relevance_threshold (default 0.1, memori/_config.py:87-88) are dropped. The surviving facts are formatted as bullet points with optional timestamps and conversation summaries, then wrapped in a <memori_context> tag. The context is injected into the provider-appropriate location — the system field for Anthropic/Bedrock, a system message for OpenAI, instructions for Google — so the LLM sees: "Only use the relevant context if it is relevant to the user's query." For cloud mode, the search is performed by the Memori Cloud API endpoint cloud/recall; results are parsed and filtered identically. Conversation history is also injected separately via inject_conversation_messages() (memori/llm/pipelines/conversation_injection.py:156-226), which replays prior messages from the current conversation into the LLM context.

MemTensor/MemOS

answered

Hybrid search: semantic vector + graph + BM25. GeneralTextMemory.search() (general.py:121-137) embeds query, vector search, sort by score. PreferenceTextMemory.search() (preference.py:81-95) auto-filters status=activated. GraphMemoryRetriever (recall.py:61-100) runs 3 parallel paths: graph dispatch plan (typed edges), Neo4j search_by_embedding() (neo4j.py:840-880, pre-filtered cosine), EnhancedBM25. Reranker scores union. search_text_memories() (search_service.py:68-80) is shared entry: builds SearchContext with session_id/filter/info. CompositeCubeView.search_memories() (composite_cube.py:46-83) parallel-searches all cubes, aggregates into MOSSearchResult (text_mem/act_mem/para_mem/pref_mem/tool_mem/skill_mem). Results injected into LLM context via MOS.search() -> MemChat.

plastic-labs/honcho

answered

Retrieval is multi-strategy. For conclusions (the memory store), RepresentationManager._get_working_representation_internal() (src/crud/representation.py:317-411) blends up to three query strategies within a configurable token budget: (1) semantic search via _query_documents_semantic() — cosine-distance ANN through the pgvector HNSW index or external vector store, filtered by level and session allowlist; (2) most-derived — documents sorted by times_derived descending, surfacing the most reinforced facts; (3) recent — by created_at descending as a fill-in for unused capacity. The allocation splits ~1/3 semantic, ~1/3 most-derived, ~1/3 recent, with over/underflow reclaim. For message search, src/utils/search.py:317-458 implements a hybrid search combining semantic (pgvector/ANN or external vector store) and full-text search (Postgres GIN-indexed to_tsvector('english', content) with ILIKE fallback) fused via Reciprocal Rank Fusion (RRF, search.py:36-75). The Dialectic agent (src/dialectic/core.py and src/dialectic/chat.py) is the primary retrieval consumer: it runs a tool loop armed with 7 tools from src/utils/agent_tools.py — search_memory (semantic conclusion search), search_messages (hybrid message search), get_observation_context, grep_messages, get_messages_by_date_range, search_messages_temporal, and get_reasoning_chain (tree traversal from a conclusion to its premises and derived conclusions). The agent iteratively selects tools to gather context and synthesizes a response once it has enough evidence. At the minimal reasoning level, only search_memory and search_messages are available. Results are presented to the LLM as formatted observations with timestamps, IDs, and derivation metadata.

MemMachine/MemMachine

answered

Retrieval is orchestrated by EpisodicMemory.query_memory() (packages/server/src/memmachine_server/episodic_memory/episodic_memory.py:352). It concurrently searches short-term memory (an in-memory deque of recent episodes, filtered by property and capacity) and long-term memory (scored vector search), then deduplicates by episode UID prioritizing short-term copies. The query string is embedded via the configured Embedder (OpenAI, Sentence Transformers, or Bedrock). Long-term search uses search_scored() (long_term_memory.py:264) which dispatches to either the declarative backend (Neo4j graph vector search) or the event backend (Qdrant/Milvus/SQLite vector ANN search). Results are optionally re-ranked by pluggable Reranker implementations: BM25, cross-encoder, Cohere, or identity passthrough. An expand_context parameter pulls chronologically adjacent episodes around each match. Score thresholds are applied directionally — higher-is-better metrics (cosine + reranker scores) drop below-threshold results; lower-is-better metrics (raw euclidean) drop above-threshold results. Semantic memory search (SemanticService.search(), semantic_memory.py:195) embeds the query per set_id (each set can have its own embedder), runs parallel vector searches across sets, and yields SemanticFeature objects. The final prompt injection is done via formalize_query_with_context() (episodic_memory.py:478), which wraps search results in <Summary> and <Episodes> XML tags and appends the original <Query>. An optional retrieval agent (MemMachineAgent in retrieval_agent/agents/memmachine_retriever.py) can orchestrate multi-tool retrieval — decomposing queries, running sub-queries against different memory tools, and consolidating results.

Editor's note. Correction: formalize_query_with_context() is defined but never called; the search API returns short-term episodes and summary, scored long-term episodes and semantic features as JSON, and the caller builds the prompt.

← How are memories stored? · How are memories updated, consolidated or forgotten? →