LLMs Technical Reviews
Home / Agent memory layers / supermemory

supermemoryai/supermemory

Client side of a hosted memory API: remote MCP server with widget, AI SDK/OpenAI/Mastra middleware and SDKs; the engine is closed-source.

GitHub ↗★ 31kTypeScriptMITcommit ac21804 · 2026-10-05homepage ↗

Overview

Supermemory is a hosted memory API. You send it content and conversations, and you get back extracted memories, a per-user “profile” (stable facts plus recent context) and hybrid search over memories and documents. Readers should know one thing first: the memory engine is not in this repository. Extraction, contradiction handling, forgetting, embeddings, connectors and search all run behind api.supermemory.ai (/v3/documents, /v4/search, /v4/profile, /v4/conversations). The project’s own self-hosting docs say the downloadable local server “is built from a separate, non-public codebase” (overview.mdx).

So the open-source repo is the client side of that product. It contains the remote MCP server (a Cloudflare Worker with an embedded React widget), the @supermemory/tools middleware for the Vercel AI SDK, OpenAI SDK, Mastra and VoltAgent, Python integrations (OpenAI Agents, Pipecat, LiveKit, Cartesia, Agent Framework), a memory-graph visualisation component, a Raycast extension, an agent skill, the docs site, and a small Next.js app that is now mostly a redirect page to the console. Anything this review says about internals is limited to that client code. Claims about how memories are extracted or ranked cannot be checked from here.

Architecture

flowchart LR
  H["MCP host (Claude, ChatGPT, Cursor)"] --> W["MCP Worker: Hono /mcp"]
  W --> AUTH["validateApiKey / OAuth"]
  W --> SRV["createSupermemoryServer"]
  SRV --> T["Tools, prompts, widget resource"]
  SRV --> DO["SpaceState Durable Object"]
  T --> CL["SupermemoryClient (supermemory SDK)"]
  APP["Your app: AI SDK / OpenAI / Mastra"] --> MW["withSupermemory middleware"]
  MW --> API["Hosted API (closed source)"]
  CL --> API
  PY["Python integrations"] --> API
Component Path Role
MCP Worker entry apps/mcp/src/server/index.ts Hono app: OAuth metadata, bearer auth, /mcp handler, direct file-upload proxy
MCP server factory apps/mcp/src/server/server.ts Builds an McpServer per request and registers tools, resources, the widget and the context prompt
Tools apps/mcp/src/server/tools/ search_memory, get_profile, add_memory, save-memory, guided-save, list/get documents and memories, spaces, graph, uploads
API client apps/mcp/src/server/client/index.ts Wraps the supermemory SDK; container-tag defaulting, forget, search, profile
Space state apps/mcp/src/server/space-state.ts Durable Object holding the active space per actor and one-shot upload sessions
RBAC apps/mcp/src/server/auth/rbac.ts Read/write permission per container tag from the session
Widget apps/mcp/src/widget/ React app served as an MCP UI resource (memory browser, graph, save flows)
Framework middleware packages/tools/src/ withSupermemory for AI SDK, OpenAI, Mastra, VoltAgent; Claude memory-tool adapter
Python integrations packages/*-python/ Same idea for Python agent frameworks
Graph UI packages/memory-graph/ Canvas visualisation of documents, memories and version links

How a request flows

Two paths matter: an MCP host calling a tool, and an app using the middleware.

MCP search_memory:

  1. Authenticate. handleMcpRequest reads the bearer token and resolves it as an API key or OAuth token against the API. It returns 503 with Retry-After on a transient auth-backend failure and 401 otherwise (index.ts).
  2. Build the server. createSupermemoryServer makes a fresh McpServer for this actor. It looks up the actor’s SpaceState Durable Object, and container-tag resolution falls back to sm_project_default (server.ts).
  3. Call the API. The tool calls SupermemoryClient.search, which forwards q, limit, threshold and the tag to client.search.memories with searchMode: "hybrid". Ranking happens entirely on the server (client/index.ts).

AI SDK withSupermemory(model, { containerTag, customId }):

  1. Before the model call. transformParamsWithMemory takes the last user message and calls buildMemoriesText. The result is cached per turn and injected into the system prompt (middleware.ts). In the default profile mode no query text is sent. query and full modes send the last user message (memory-client.ts).
  2. Fetch context. supermemoryProfileSearch POSTs to /v4/profile with include: ["static", "dynamic"] and an optional q (memory-client.ts).
  3. After the response. With addMemory: "always" (the default), saveMemoryAfterResponse sends the whole turn to /v4/conversations under customId. The server extracts memories from it (middleware.ts, vercel/index.ts).

Key components

MCP server

There are sixteen tools, registered in one place (tools/index.ts). Writes go through client.add and come back as queued, which shows that ingestion is asynchronous on the server. forgetMemory first tries an exact-content forget. On a 404 it searches with a 0.85 similarity threshold and forgets the closest memory hit, never a chunk (client/index.ts). The context prompt builds a compact block: the active space, up to 8 stable and 8 recent profile facts, and recently active spaces (context.ts).

Spaces and access

A “space” is a container tag. SpaceState stores the actor’s active tag and handles short-lived upload sessions. It stores only a hash of each upload token, deletes the session on first use, and expires it with an alarm (space-state.ts). effectiveContainerTagAccess gives write by default, falls back to per-tag membership for restricted sessions, and downgrades to read for read-scoped or out-of-scope tags (rbac.ts). Enforcement still happens at the API. This code only decides what the widget offers.

Memory model (as seen by clients)

The shared types show what the API returns. Memories carry version, isLatest, parent and root ids, isForgotten, isStatic and isInference, and documents can have forgetAfter. So versioning, temporal forgetting and inferred memories exist server-side. The client neither implements nor configures them.

Extending it

  • Your own app. Use the supermemory SDK directly, or wrap a model with withSupermemory (mode: profile | query | full, addMemory: always | never, custom prompt template).
  • Agents with tools. supermemoryTools and the OpenAI function-calling tools expose search and add as model-callable tools. A Claude memory-tool adapter maps Anthropic’s memory tool onto Supermemory documents.
  • MCP. Add a file in apps/mcp/src/server/tools/ and register it in registerAllTools. The widget resource and structured-content output schemas are in resources/ and tools/output-schemas.ts.

Running it

  • Hosted. Get an API key from the console and point any SDK or the MCP URL at the hosted service. Connectors (Drive, Notion, Gmail, OneDrive) and MCP are hosted features.
  • MCP Worker yourself. bun run build then wrangler deploy in apps/mcp. The only bindings are the two Durable Objects and API_URL (wrangler.jsonc). It still needs a Supermemory API behind it.
  • Local engine. npx supermemory local downloads a closed-source single binary that serves the same API on port 6767 with an embedded graph store, local bge-base-en-v1.5 embeddings and your own LLM key. It is free within a “lite” licence. The docs list connectors and the Supermemory MCP as platform-only.

Strengths and caveats

  • Strength: easy adoption. One middleware wrap or one MCP URL gives an app per-user profile context plus automatic conversation capture. The profile endpoint avoids writing retrieval prompts yourself.
  • Strength: polished MCP surface. OAuth discovery, a UI widget, space selection, RBAC-aware writes and direct file uploads through single-use tokens are more complete than most memory MCPs.
  • Caveat: the core is not open. Extraction, conflict handling, decay and ranking cannot be inspected or changed. The “open-source memory engine” framing does not match this repository.
  • Caveat: vendor dependency. Every path ends at the Supermemory API, either hosted or the closed local binary, so data residency and pricing are the vendor’s terms.
  • Caveat: limited tuning. Client-side knobs are the container tag, the mode, a similarity threshold and the prompt template. Anything deeper needs the API’s own settings.

Sources: code at ac21804, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (38 pages), verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit ac21804. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Memories are extracted through two paths. Explicit saves: the add_memory tool (apps/mcp/src/server/tools/add-memory.ts:29-63) and guided-save tool (apps/mcp/src/server/tools/guided-save.ts:8-60) accept content from the LLM and call SupermemoryClient.createMemory(), which POSTs to the API. Implicit profile facts: the getProfile() method (apps/mcp/src/server/client/index.ts:270-305) calls the API's profile endpoint which returns two arrays — static (long-lived preferences, traits) and dynamic (recent context, activity). These facts are extracted server-side by the IngestContentWorkflow (referenced in CLAUDE.md) which does content-type detection, AI-powered summarization, automatic tagging, and vector embedding generation via Cloudflare AI. No client-side deduplication — the response includes raw fact strings; deduplication and conflict resolution happen in the upstream API/ingestion pipeline, not in this MCP server's code.

Editor's note. Correction: extraction is not implemented anywhere in this repository; the IngestContentWorkflow is only described in CLAUDE.md and runs inside the closed-source hosted API. Client paths that feed it are add/save tools and the middleware's /v4/conversations upload after each model response.

How are memories stored?

answered

Storage is primarily API-driven through the supermemory SDK. The SupermemoryClient (apps/mcp/src/server/client/index.ts:138-160) wraps the SDK and exposes methods that call REST endpoints. The memory schema is defined in apps/mcp/src/shared/types.ts:139-152: each MemoryEntry has id, memory (string), version, isLatest, isForgotten, isStatic, isInference, createdAt, updatedAt, sourceCount, documentIds, plus a history array of prior versions (line 125-137). Container tags (apps/mcp/src/shared/types.ts:39-53) scope memories into named spaces with documentCount and memoryCount. The SDK search (apps/mcp/src/server/client/index.ts:252-256) uses searchMode: 'hybrid', indicating a combination of vector (embeddings) and keyword retrieval. Memory entries have parent/root tracking via parentMemoryId and rootMemoryId (line 85-86), enabling version chains. The upstream document schema (documentsApiResponseSchema, line 113-118) includes type, title, summary, and content fields. No pluggable store interface exists in this repo — the MCP server delegates fully to the hosted API.

How are memories retrieved and injected into the prompt?

answered

Retrieval happens through three channels. 1. MCP context prompt (apps/mcp/src/server/prompts/context.ts:12-115): registered as the 'context' prompt, it calls getClient().getProfile() to load static and dynamic fact arrays (capped at CONTEXT_FACT_LIMIT = 8 each, line 12), formats them via formatFactSection (apps/mcp/src/server/space-presentation.ts:89-101), and injects them as a user-role message into the LLM's context window. The format adds a '+N more' line when facts exceed the limit (lines 98-99). 2. get_profile tool (apps/mcp/src/server/tools/get-profile.ts:1-63): returns the full fact arrays as structured content with no cap, plus a text summary. 3. search_memory tool (apps/mcp/src/server/tools/search-memory.ts:1-70): performs hybrid semantic+keyword search via SupermemoryClient.search() (client/index.ts:247-267), which accepts query, limit, threshold (similarity floor), and searchMode: 'hybrid'. Results include similarity scores and the matched text. The context prompt also appends recently-active spaces (lines 71-82) for temporal awareness. There is no separate ranking or filtering layer in this MCP server — the API returns already-ranked results.

How are memories updated, consolidated or forgotten?

answered

Memory lifecycle is tracked through version chains and state flags. Versioning: each MemoryEntry has version, isLatest, parentMemoryId, and rootMemoryId (apps/mcp/src/shared/types.ts:140-147). The history array (line 151) stores prior versions as MemoryEntryHistory objects (lines 125-137). Forgetting: the isForgotten boolean flag marks memories as forgotten (apps/mcp/src/shared/types.ts:143). The forgetMemory method (apps/mcp/src/server/client/index.ts:181-239) first tries exact-match deletion via the SDK, then falls back to a similarity search (threshold 0.85, line 200) to find a semantically matching memory to forget. TTL/expiration: forgetAfter (nullable string) and forgetReason (nullable string) fields on DocumentMemoryEntry (apps/mcp/src/shared/types.ts:81-82) suggest scheduled expiration, though the MCP server does not implement the expiry logic itself — that happens in the upstream API. Update/merge: the add_memory tool (apps/mcp/src/server/tools/add-memory.ts) always creates new memory; there is no merge or consolidation logic in this MCP server. The updatesMemoryId and nextVersionId fields on the graph type (apps/mcp/src/widget/views/Graph.tsx:99-100) confirm a linked-list version pattern. Decay and deletion are server-side concerns delegated to the hosted API.

How is memory scoped and isolated?

answered

Memory scoping uses a container-tag (space) model. Space isolation: every tool operation is scoped by containerTag — a string identifier for a named space (apps/mcp/src/shared/types.ts:39-53). The SupermemoryClient constructor takes an optional containerTag which defaults to 'sm_project_default' (apps/mcp/src/server/client/index.ts:158-159). When no explicit tag exists, getProfile() returns empty facts (line 271-278). Multi-tenancy: the ActorContext carries userId, organizationId, bearerToken, and oauthClientId (apps/mcp/src/server/index.ts:227-232). The SpaceState DurableObject (apps/mcp/src/server/space-state.ts:26-73) persists per-actor state keyed by ${org}:${user}. Access control: the rbac.ts module provides effectiveContainerTagAccess() returning read/write permissions per tag. Session scope (sessionScopeSchema in shared/types.ts:13-21) discriminates full vs scoped access with optional tag and rate-limit constraints. Privacy: there are no client-side privacy controls (redaction, content filtering) — the API is trusted to enforce permissions server-side.

How do agents integrate with it, and what is self-hostable?

answered

Integration is primarily through the MCP server (apps/mcp/src/server/server.ts:45-111) which exposes 15+ tools registered in tools/index.ts:19-36. The server uses the @modelcontextprotocol/server SDK and registers a widget UI resource (resources/widget.ts:12-47) — a React app served as text/html that renders inside MCP hosts (Cursor, Claude Code, ChatGPT). The widget communicates with the server through MCP tool calls and structured content. REST API: the SupermemoryClient also talks directly to https://api.supermemory.ai (configurable via API_URL env var, server.ts:59) for batch/document operations. SDK: the supermemory npm package provides the client library. Self-hosting: the server deploys on Cloudflare Workers (wrangler.jsonc), requires a Hyperdrive (DB) binding, AI binding for embeddings, and KV storage. The widget is a standalone React app compiled by scripts/build-widget.ts. The MCP server itself is open-source and self-hostable; the upstream API (console.supermemory.ai) is a hosted service with its own billing.

Editor's note. Correction: the MCP Worker's wrangler.jsonc binds only two Durable Objects (MCP_SERVER, SPACE_STATE) and API_URL; there is no Hyperdrive, AI or KV binding. Separately, the memory engine itself is not open source: the local server (npx supermemory local) is a closed binary, and the repo also ships @supermemory/tools middleware (withSupermemory for AI SDK/OpenAI/Mastra) and Python framework integrations.