supermemoryai/supermemory
Client side of a hosted memory API: remote MCP server with widget, AI SDK/OpenAI/Mastra middleware and SDKs; the engine is closed-source.
Overview
Supermemory is a hosted memory API. You send it content and conversations, and you get back extracted memories, a per-user “profile” (stable facts plus recent context) and hybrid search over memories and documents. Readers should know one thing first: the memory engine is not in this repository. Extraction, contradiction handling, forgetting, embeddings, connectors and search all run behind api.supermemory.ai (/v3/documents, /v4/search, /v4/profile, /v4/conversations). The project’s own self-hosting docs say the downloadable local server “is built from a separate, non-public codebase” (overview.mdx).
So the open-source repo is the client side of that product. It contains the remote MCP server (a Cloudflare Worker with an embedded React widget), the @supermemory/tools middleware for the Vercel AI SDK, OpenAI SDK, Mastra and VoltAgent, Python integrations (OpenAI Agents, Pipecat, LiveKit, Cartesia, Agent Framework), a memory-graph visualisation component, a Raycast extension, an agent skill, the docs site, and a small Next.js app that is now mostly a redirect page to the console. Anything this review says about internals is limited to that client code. Claims about how memories are extracted or ranked cannot be checked from here.
Architecture
flowchart LR
H["MCP host (Claude, ChatGPT, Cursor)"] --> W["MCP Worker: Hono /mcp"]
W --> AUTH["validateApiKey / OAuth"]
W --> SRV["createSupermemoryServer"]
SRV --> T["Tools, prompts, widget resource"]
SRV --> DO["SpaceState Durable Object"]
T --> CL["SupermemoryClient (supermemory SDK)"]
APP["Your app: AI SDK / OpenAI / Mastra"] --> MW["withSupermemory middleware"]
MW --> API["Hosted API (closed source)"]
CL --> API
PY["Python integrations"] --> API
| Component | Path | Role |
|---|---|---|
| MCP Worker entry | apps/mcp/src/server/index.ts |
Hono app: OAuth metadata, bearer auth, /mcp handler, direct file-upload proxy |
| MCP server factory | apps/mcp/src/server/server.ts |
Builds an McpServer per request and registers tools, resources, the widget and the context prompt |
| Tools | apps/mcp/src/server/tools/ |
search_memory, get_profile, add_memory, save-memory, guided-save, list/get documents and memories, spaces, graph, uploads |
| API client | apps/mcp/src/server/client/index.ts |
Wraps the supermemory SDK; container-tag defaulting, forget, search, profile |
| Space state | apps/mcp/src/server/space-state.ts |
Durable Object holding the active space per actor and one-shot upload sessions |
| RBAC | apps/mcp/src/server/auth/rbac.ts |
Read/write permission per container tag from the session |
| Widget | apps/mcp/src/widget/ |
React app served as an MCP UI resource (memory browser, graph, save flows) |
| Framework middleware | packages/tools/src/ |
withSupermemory for AI SDK, OpenAI, Mastra, VoltAgent; Claude memory-tool adapter |
| Python integrations | packages/*-python/ |
Same idea for Python agent frameworks |
| Graph UI | packages/memory-graph/ |
Canvas visualisation of documents, memories and version links |
How a request flows
Two paths matter: an MCP host calling a tool, and an app using the middleware.
MCP search_memory:
- Authenticate.
handleMcpRequestreads the bearer token and resolves it as an API key or OAuth token against the API. It returns 503 withRetry-Afteron a transient auth-backend failure and 401 otherwise (index.ts). - Build the server.
createSupermemoryServermakes a freshMcpServerfor this actor. It looks up the actor’sSpaceStateDurable Object, and container-tag resolution falls back tosm_project_default(server.ts). - Call the API. The tool calls
SupermemoryClient.search, which forwardsq,limit,thresholdand the tag toclient.search.memorieswithsearchMode: "hybrid". Ranking happens entirely on the server (client/index.ts).
AI SDK withSupermemory(model, { containerTag, customId }):
- Before the model call.
transformParamsWithMemorytakes the last user message and callsbuildMemoriesText. The result is cached per turn and injected into the system prompt (middleware.ts). In the defaultprofilemode no query text is sent.queryandfullmodes send the last user message (memory-client.ts). - Fetch context.
supermemoryProfileSearchPOSTs to/v4/profilewithinclude: ["static", "dynamic"]and an optionalq(memory-client.ts). - After the response. With
addMemory: "always"(the default),saveMemoryAfterResponsesends the whole turn to/v4/conversationsundercustomId. The server extracts memories from it (middleware.ts, vercel/index.ts).
Key components
MCP server
There are sixteen tools, registered in one place (tools/index.ts). Writes go through client.add and come back as queued, which shows that ingestion is asynchronous on the server. forgetMemory first tries an exact-content forget. On a 404 it searches with a 0.85 similarity threshold and forgets the closest memory hit, never a chunk (client/index.ts). The context prompt builds a compact block: the active space, up to 8 stable and 8 recent profile facts, and recently active spaces (context.ts).
Spaces and access
A “space” is a container tag. SpaceState stores the actor’s active tag and handles short-lived upload sessions. It stores only a hash of each upload token, deletes the session on first use, and expires it with an alarm (space-state.ts). effectiveContainerTagAccess gives write by default, falls back to per-tag membership for restricted sessions, and downgrades to read for read-scoped or out-of-scope tags (rbac.ts). Enforcement still happens at the API. This code only decides what the widget offers.
Memory model (as seen by clients)
The shared types show what the API returns. Memories carry version, isLatest, parent and root ids, isForgotten, isStatic and isInference, and documents can have forgetAfter. So versioning, temporal forgetting and inferred memories exist server-side. The client neither implements nor configures them.
Extending it
- Your own app. Use the
supermemorySDK directly, or wrap a model withwithSupermemory(mode: profile | query | full,addMemory: always | never, custom prompt template). - Agents with tools.
supermemoryToolsand the OpenAI function-calling tools expose search and add as model-callable tools. A Claude memory-tool adapter maps Anthropic’s memory tool onto Supermemory documents. - MCP. Add a file in
apps/mcp/src/server/tools/and register it inregisterAllTools. The widget resource and structured-content output schemas are inresources/andtools/output-schemas.ts.
Running it
- Hosted. Get an API key from the console and point any SDK or the MCP URL at the hosted service. Connectors (Drive, Notion, Gmail, OneDrive) and MCP are hosted features.
- MCP Worker yourself.
bun run buildthenwrangler deployinapps/mcp. The only bindings are the two Durable Objects andAPI_URL(wrangler.jsonc). It still needs a Supermemory API behind it. - Local engine.
npx supermemory localdownloads a closed-source single binary that serves the same API on port 6767 with an embedded graph store, localbge-base-en-v1.5embeddings and your own LLM key. It is free within a “lite” licence. The docs list connectors and the Supermemory MCP as platform-only.
Strengths and caveats
- Strength: easy adoption. One middleware wrap or one MCP URL gives an app per-user profile context plus automatic conversation capture. The profile endpoint avoids writing retrieval prompts yourself.
- Strength: polished MCP surface. OAuth discovery, a UI widget, space selection, RBAC-aware writes and direct file uploads through single-use tokens are more complete than most memory MCPs.
- Caveat: the core is not open. Extraction, conflict handling, decay and ranking cannot be inspected or changed. The “open-source memory engine” framing does not match this repository.
- Caveat: vendor dependency. Every path ends at the Supermemory API, either hosted or the closed local binary, so data residency and pricing are the vendor’s terms.
- Caveat: limited tuning. Client-side knobs are the container tag, the mode, a similarity threshold and the prompt template. Anything deeper needs the API’s own settings.
Sources: code at ac21804, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (38 pages), verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit ac21804. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemories are extracted through two paths. Explicit saves: the add_memory tool (apps/mcp/src/server/tools/add-memory.ts:29-63) and guided-save tool (apps/mcp/src/server/tools/guided-save.ts:8-60) accept content from the LLM and call SupermemoryClient.createMemory(), which POSTs to the API. Implicit profile facts: the getProfile() method (apps/mcp/src/server/client/index.ts:270-305) calls the API's profile endpoint which returns two arrays — static (long-lived preferences, traits) and dynamic (recent context, activity). These facts are extracted server-side by the IngestContentWorkflow (referenced in CLAUDE.md) which does content-type detection, AI-powered summarization, automatic tagging, and vector embedding generation via Cloudflare AI. No client-side deduplication — the response includes raw fact strings; deduplication and conflict resolution happen in the upstream API/ingestion pipeline, not in this MCP server's code.
How are memories stored?
answeredStorage is primarily API-driven through the supermemory SDK. The SupermemoryClient (apps/mcp/src/server/client/index.ts:138-160) wraps the SDK and exposes methods that call REST endpoints. The memory schema is defined in apps/mcp/src/shared/types.ts:139-152: each MemoryEntry has id, memory (string), version, isLatest, isForgotten, isStatic, isInference, createdAt, updatedAt, sourceCount, documentIds, plus a history array of prior versions (line 125-137). Container tags (apps/mcp/src/shared/types.ts:39-53) scope memories into named spaces with documentCount and memoryCount. The SDK search (apps/mcp/src/server/client/index.ts:252-256) uses searchMode: 'hybrid', indicating a combination of vector (embeddings) and keyword retrieval. Memory entries have parent/root tracking via parentMemoryId and rootMemoryId (line 85-86), enabling version chains. The upstream document schema (documentsApiResponseSchema, line 113-118) includes type, title, summary, and content fields. No pluggable store interface exists in this repo — the MCP server delegates fully to the hosted API.
How are memories retrieved and injected into the prompt?
answeredRetrieval happens through three channels. 1. MCP context prompt (apps/mcp/src/server/prompts/context.ts:12-115): registered as the 'context' prompt, it calls getClient().getProfile() to load static and dynamic fact arrays (capped at CONTEXT_FACT_LIMIT = 8 each, line 12), formats them via formatFactSection (apps/mcp/src/server/space-presentation.ts:89-101), and injects them as a user-role message into the LLM's context window. The format adds a '+N more' line when facts exceed the limit (lines 98-99). 2. get_profile tool (apps/mcp/src/server/tools/get-profile.ts:1-63): returns the full fact arrays as structured content with no cap, plus a text summary. 3. search_memory tool (apps/mcp/src/server/tools/search-memory.ts:1-70): performs hybrid semantic+keyword search via SupermemoryClient.search() (client/index.ts:247-267), which accepts query, limit, threshold (similarity floor), and searchMode: 'hybrid'. Results include similarity scores and the matched text. The context prompt also appends recently-active spaces (lines 71-82) for temporal awareness. There is no separate ranking or filtering layer in this MCP server — the API returns already-ranked results.
How are memories updated, consolidated or forgotten?
answeredMemory lifecycle is tracked through version chains and state flags. Versioning: each MemoryEntry has version, isLatest, parentMemoryId, and rootMemoryId (apps/mcp/src/shared/types.ts:140-147). The history array (line 151) stores prior versions as MemoryEntryHistory objects (lines 125-137). Forgetting: the isForgotten boolean flag marks memories as forgotten (apps/mcp/src/shared/types.ts:143). The forgetMemory method (apps/mcp/src/server/client/index.ts:181-239) first tries exact-match deletion via the SDK, then falls back to a similarity search (threshold 0.85, line 200) to find a semantically matching memory to forget. TTL/expiration: forgetAfter (nullable string) and forgetReason (nullable string) fields on DocumentMemoryEntry (apps/mcp/src/shared/types.ts:81-82) suggest scheduled expiration, though the MCP server does not implement the expiry logic itself — that happens in the upstream API. Update/merge: the add_memory tool (apps/mcp/src/server/tools/add-memory.ts) always creates new memory; there is no merge or consolidation logic in this MCP server. The updatesMemoryId and nextVersionId fields on the graph type (apps/mcp/src/widget/views/Graph.tsx:99-100) confirm a linked-list version pattern. Decay and deletion are server-side concerns delegated to the hosted API.
How is memory scoped and isolated?
answeredMemory scoping uses a container-tag (space) model. Space isolation: every tool operation is scoped by containerTag — a string identifier for a named space (apps/mcp/src/shared/types.ts:39-53). The SupermemoryClient constructor takes an optional containerTag which defaults to 'sm_project_default' (apps/mcp/src/server/client/index.ts:158-159). When no explicit tag exists, getProfile() returns empty facts (line 271-278). Multi-tenancy: the ActorContext carries userId, organizationId, bearerToken, and oauthClientId (apps/mcp/src/server/index.ts:227-232). The SpaceState DurableObject (apps/mcp/src/server/space-state.ts:26-73) persists per-actor state keyed by ${org}:${user}. Access control: the rbac.ts module provides effectiveContainerTagAccess() returning read/write permissions per tag. Session scope (sessionScopeSchema in shared/types.ts:13-21) discriminates full vs scoped access with optional tag and rate-limit constraints. Privacy: there are no client-side privacy controls (redaction, content filtering) — the API is trusted to enforce permissions server-side.
How do agents integrate with it, and what is self-hostable?
answeredIntegration is primarily through the MCP server (apps/mcp/src/server/server.ts:45-111) which exposes 15+ tools registered in tools/index.ts:19-36. The server uses the @modelcontextprotocol/server SDK and registers a widget UI resource (resources/widget.ts:12-47) — a React app served as text/html that renders inside MCP hosts (Cursor, Claude Code, ChatGPT). The widget communicates with the server through MCP tool calls and structured content. REST API: the SupermemoryClient also talks directly to https://api.supermemory.ai (configurable via API_URL env var, server.ts:59) for batch/document operations. SDK: the supermemory npm package provides the client library. Self-hosting: the server deploys on Cloudflare Workers (wrangler.jsonc), requires a Hyperdrive (DB) binding, AI binding for embeddings, and KV storage. The widget is a standalone React app compiled by scripts/build-widget.ts. The MCP server itself is open-source and self-hostable; the upstream API (console.supermemory.ai) is a hosted service with its own billing.