How do agents integrate with it, and what is self-hostable?
SDKs, REST API, MCP server, framework plugins; required services; open vs hosted-only parts.
Verdict
Hindsight is the easiest full server to self-host: one process with an embedded Postgres, local embeddings and a local reranker, so it needs only an LLM key. Memori and Supermemory are the exceptions to self-hosting, because their core memory step runs on vendor services.
Servers you run. Hindsight ships REST, MCP, generated Python, TypeScript, Rust and Go clients, and hooks for many coding agents. Honcho runs as an API process and a deriver process over Postgres with pgvector. It has Python and TypeScript SDKs, MCP and coding-agent plugins, uses gpt-5.4-mini for every stage by default, and is AGPL-3.0. MemMachine serves /api/v2 plus MCP. Its Compose file starts Postgres, Neo4j and Qdrant, and it has adapters for LangChain, CrewAI, LlamaIndex, Dify and n8n. MemOS runs a FastAPI server over Neo4j and Qdrant. Its MCP server wraps the older in-process MOS API, and a separate TypeScript SQLite plugin serves agent harnesses. Cognee offers an SDK, a server, a CLI and cognee-mcp, all running on local files by default.
Libraries. Mem0’s Memory class runs locally. Its coding-agent plugins call api.mem0.ai unless MEM0_API_URL is set. Graphiti is graphiti_core plus FastAPI and MCP wrappers (stdio, SSE or streamable HTTP). It needs a graph database, and neither wrapper has auth. Memori wraps your OpenAI, Anthropic, Gemini or Bedrock client. In BYODB mode your data stays in your database, but extraction always calls Memori’s hosted API.
Coding-agent plugin. claude-mem installs hooks into Claude Code, Codex, Cursor and others. Behind them run a local Bun worker, SQLite and chroma-mcp.
Client for a hosted engine. Supermemory’s repository holds the MCP Worker, AI SDK middleware and SDKs. The engine is closed-source, both as the hosted API and as the downloadable local binary.
Pick: Hindsight for a single self-hosted server. Pick: claude-mem for zero-effort memory in a coding agent. Pick: Mem0 or Graphiti to embed memory in your own Python service.
Per-project answers
thedotmack/claude-mem
answeredAgents integrate through multiple channels. Claude Code hooks are the primary path: sessionInit starts a tracked server session, observationHandler captures each PostToolUse event, contextHandler injects a rendered timeline at SessionStart, and sessionEnd triggers summary generation (src/cli/handlers/). The adapter layer (src/cli/adapters/claude-code.ts:39) normalizes platform-specific input shapes into a common NormalizedHookInput. MCP servers provide programmatic access: a local stdio MCP server (src/servers/mcp-server.ts) supplies search, code reading, and memory tools; a remote-recall MCP server (src/server/mcp/recall-mcp-server.ts:140) exposes search, context, and recent tools over Streamable HTTP MCP — both go through the same Postgres search and audit path. REST API: POST /v1/search and POST /v1/context (src/server/routes/v1/ServerV1PostgresRoutes.ts:934-984) provide full-text search over generated observations. POST /api/memory/save (src/services/worker/http/routes/MemoryRoutes.ts:26) allows manual memory creation. OpenCode plugin: A native OpenCode IDE plugin (src/integrations/opencode-plugin/index.ts) integrates with OpenCode's event system. Platform adapters: Supports Claude Code, Codex, Cursor, OpenCode, Antigravity CLI, Windsurf, and Kimi (src/cli/adapters/). Self-hosting: The entire system is self-hostable and Apache 2.0 licensed. SQLite works out of the box with only Bun. Postgres is optional for multi-tenant deployments. Chroma vector search runs as a local MCP subprocess. The installer (npx claude-mem install) handles setup across all supported platforms. Required services depend on the backend: Bun+SQLite for local-only, Postgres for the server, Chroma for vector search (all optional). There are no hosted-only parts — the full codebase and all features are available for self-hosting.
mem0ai/mem0
answeredSDKs — Mem0 provides four entry points in the Python SDK (mem0/__init__.py:1-6): Memory (self-hosted sync), AsyncMemory (self-hosted async), MemoryClient (hosted platform sync), and AsyncMemoryClient (hosted platform async). The TypeScript SDK (mem0-ts/) mirrors this with a MemoryClient for the hosted API and OSS memory for self-hosted. The CLI packages (cli/python/ and cli/node/) provide command-line access.
REST API — the server/ directory (server/main.py) is a FastAPI application wrapping the Python SDK. It exposes the full CRUD surface (add, search, get, get_all, update, delete, delete_all, history, reset) behind JWT authentication. The server includes admin auth, API key management, rate limiting, and request logging. It runs on Docker Compose with PostgreSQL+pgvector and Neo4j — the SDK code is mounted for hot reload.
MCP server — multiple coding-agent plugins expose Mem0 as an MCP tool (search_memories), notably integrations/cursor-plugin/, integrations/claude-code-plugin/, integrations/codex-plugin/, integrations/kimi-plugin/, integrations/antigravity-plugin/, and the shared integrations/mem0-agent-plugin/ (integrations/mem0-agent-plugin/core/mcp_server.py:1-60). The shared Python core in integrations/agent-plugin-core/python/mcp_server.py provides the runtime that all native plugins use.
Framework plugins — integrations/vercel-ai-sdk/ wraps Mem0 as a Vercel AI SDK provider (@mem0/vercel-ai-provider). integrations/openclaw/, integrations/pi-agent-plugin/, integrations/deepseek-plugin/ are editor/agent integrations. integrations/n8n-nodes-mem0/ is an n8n community node. integrations/zapier-mem0/ is a Zapier app. integrations/mem0-strands/ is a Python Strands MemoryStore.
Self-hostable parts — the entire Python SDK (Memory/AsyncMemory) is fully self-hostable and open-source (Apache 2.0). Required services depend on the vector store choice: you need a running vector database (Qdrant, pgvector, Elasticsearch, etc.) and an LLM provider API key. The Docker Compose stack bundles FastAPI + pgvector + Neo4j for a turnkey solution. The hosted MemoryClient calls api.mem0.ai and requires an API key. The OSS SDK rejects certain platform-only features (decay, temporal search with reference_date/timestamp) at runtime.
server/docker-compose.yaml runs the FastAPI app, pgvector/pgvector:pg17 and the dashboard. It has no Neo4j service, and graph memory is no longer in the SDK. Rate limits (slowapi) apply only to the auth routes. The coding-agent plugins expose only a search_memories MCP tool and call the hosted api.mem0.ai unless MEM0_API_URL is set, so they do not wrap the self-hosted Memory class.vectorize-io/hindsight
answeredHindsight is self-hostable as a single process: uv run hindsight-api starts a FastAPI server with embedded PostgreSQL (pg0), local SentenceTransformers embeddings, and local CrossEncoder reranking — no external services required. The REST API (api/http.py) exposes REST endpoints for all operations (retain, recall, reflect, CRUD on memories/documents/entities/mental models/directives/banks). An auto-generated OpenAPI spec feeds 4 SDKs: Python (hindsight_client), TypeScript, Rust, and Go (hindsight-clients/). The Python SDK has a high-level Hindsight class with synchronous convenience methods (hindsight_client/hindsight_client.py:80-142). An MCP server (api/mcp.py:1-100) exposes 30+ tools (retain, recall, reflect, list_banks, etc.) over the Model Context Protocol HTTP transport, registered in mcp_tools.py:602-701. Framework integrations (hindsight-integrations/) span AG2, AutoGen, CrewAI, LangGraph, Pydantic AI, Claude Agent SDK, OpenAI Agents SDK (via ai-sdk), Aider, and agno. The Coding Agents integration (hindsight-integrations/coding-agents/) is a single npm package supporting 22 agent harnesses (Claude Code, Codex CLI, Cursor, Copilot, Grok Build, Dcode, opencode, Kilo, Devin CLI, Cline, Qwen Code, Kimi Code, etc.) with automatic session hooks that recall context before each prompt and retain transcripts afterward. A portable Agent Plugin (agent-plugin/plugin.json) exposes MCP tools for agents with plugin loaders. The CLI (hindsight-cli/) is a Rust binary with commands for memory, documents, entities, banks, mental models, audit logs, and filesystem mount of the knowledge base. A control plane (hindsight-control-plane/) is a Next.js admin UI with dashboards for banks, memories, documents, entities, knowledge base, audit logs, and operations. The embed package (hindsight-embed/) provides hindsight-local-mcp for daemonless usage via uvx. All server configuration is via environment variables with sensible defaults; hierarchical config overrides per-bank for LLM settings and operation tuning.
getzep/graphiti
answeredAgents can integrate with Graphiti through three tiers. Python SDK: the Graphiti class (graphiti.py) is the primary interface — add_episode(), search(), add_triplet(), summarize_saga(), build_communities(). It depends on pydantic, an LLM client, an embedder, a cross-encoder, and a graph driver. REST API (server/graph_service/main.py) is a FastAPI application with /messages (episode ingestion via background worker), /search, /get-memory, /entity-edge, and /clear endpoints. It wraps the Python SDK and is deployable via uvicorn (uvicorn graph_service.main:app). MCP Server (mcp_server/) exposes all core operations as MCP tools: add_memory, add_triplet, search_memory_facts, search_nodes, summarize_saga, build_communities, delete_episode, clear_graph, and more. It supports both Neo4j and FalkorDB backends via a configuration system with LLM/embedder/reranker provider factories. The MCP server can run in stdio or SSE mode and is containerized with Docker Compose.
All three tiers are fully self-hostable — nothing requires a hosted service. Required infrastructure: a graph database (Neo4j 5.26+, FalkorDB 1.1.2+, Kuzu, or Neptune) and API keys for the chosen LLM/embedding provider. The default stack uses OpenAI for all three (LLM, embeddings, reranking), but each is pluggable: LLMClient implementations exist for OpenAI, Anthropic, Gemini, Groq; EmbedderClient for OpenAI, Voyage, Gemini, Azure OpenAI; CrossEncoderClient for OpenAI, BGE, Gemini. The LLMRuntime subsystem (llm_client/llm_runtime.py) supports multi-model prompt routing. OpenTelemetry tracing is optional but integrated. The Kuzu driver is a lightweight embedded option for single-process deployments. There is no hosted/cloud tier visible in this repository.
topoteretes/cognee
answeredCognee offers four integration surfaces. The Python SDK is the primary entry (exported from cognee/init.py with await cognee.remember()/recall()/improve()/forget()). The REST API is a FastAPI server at /api/v1/ (cognee/api/client.py:362) with routers for remember, recall, improve, forget, add, cognify, search, datasets, users, permissions, skills, sessions, and integrations (GitHub, Gmail, Google Drive, Linear, Slack). The MCP server (cognee-mcp/src/server.py:355) provides remember/recall/forget/cognify_status/improve tools, connects in-process or via --api-url (REST) or --serve-url (Cognee Cloud), and supports per-client agent-scoped datasets. The CLI (cognee-cli) exposes all operations. Everything is self-hostable: minimal setup needs no external services (local defaults: SQLite, LanceDB, LadybugDB) and no API keys when using the GLiNER demo extractor with fastembed. Docker Compose deploys the full stack. Cognee Cloud is the hosted multi-tenant SaaS option; the same SDK works locally or via cognee.serve(url=...).
supermemoryai/supermemory
answeredIntegration is primarily through the MCP server (apps/mcp/src/server/server.ts:45-111) which exposes 15+ tools registered in tools/index.ts:19-36. The server uses the @modelcontextprotocol/server SDK and registers a widget UI resource (resources/widget.ts:12-47) — a React app served as text/html that renders inside MCP hosts (Cursor, Claude Code, ChatGPT). The widget communicates with the server through MCP tool calls and structured content. REST API: the SupermemoryClient also talks directly to https://api.supermemory.ai (configurable via API_URL env var, server.ts:59) for batch/document operations. SDK: the supermemory npm package provides the client library. Self-hosting: the server deploys on Cloudflare Workers (wrangler.jsonc), requires a Hyperdrive (DB) binding, AI binding for embeddings, and KV storage. The widget is a standalone React app compiled by scripts/build-widget.ts. The MCP server itself is open-source and self-hostable; the upstream API (console.supermemory.ai) is a hosted service with its own billing.
MemoriLabs/Memori
answeredMemori offers multiple integration surfaces:
Python SDK (pip install memori). The primary SDK at memori/ provides a Memori builder that registers with any supported LLM client (OpenAI, Anthropic, Google, xAI, Bedrock, LiteLLM). The core integration point is an LLM Invoke class (memori/llm/invoke/invoke.py:24-155) that wraps every LLM call in a three-stage pipeline: conversation-history injection, fact recall injection (inject_recalled_facts) and post-response handling (handle_post_response). The LLM client registration happens through adapters — e.g., memori/llm/adapters/openai/, memori/llm/adapters/anthropic/ — which monkey-patch or wrap the client's .create() / .invoke() method to route through Memori's pipeline.
TypeScript SDK (npm install @memorilabs/memori). Listed in the README with analogous LLM registration pattern.
BYODB (Bring Your Own Database). The self-hostable mode allows running against PostgreSQL, MongoDB, SQLite, MySQL, TiDB, OceanBase, CockroachDB, or Oracle. Storage is pluggable: Registry.register_adapter() maps connection types to adapters (SQLAlchemy, MongoDB native, Django, DB-API), and Registry.register_driver() maps the detected dialect to the driver implementation (memori/storage/_registry.py:18-67). An optional Rust native core (memori_python) handles local embedding and augmentation.
MCP Server. Memori provides an MCP endpoint at https://api.memorilabs.ai/mcp/ that works with Claude Code, Cursor, Codex, and Warp. The MCP transport uses HTTP with API-key headers for entity and process attribution. This does not require any SDK integration — just a claude mcp add command.
Framework integrations. A Hermes Agent memory provider (integrations/hermes/) and an OpenClaw plugin (@memorilabs/openclaw-memori) allow drop-in use with those agent frameworks.
Services required. Cloud mode requires only a MEMORI_API_KEY. BYODB mode requires a running database and optionally a TEI server or native Rust core for local embeddings. The entire codebase is Apache 2.0 licensed — there are no closed-source gates in the Python SDK itself, though the cloud augmentation API is a hosted service.
MemTensor/MemOS
answered3 channels. Python SDK: MOS.simple()/MOS(config) (main.py:24-80), register_mem_cube()/add()/search()/chat()/create_user(). PyPI as MemoryOS. REST API: FastAPI at memos.api.start_api (server_api.py:39-73), /product/* (add/search/chat/feedback) + /admin/* key mgmt. server_router.py (server_router.py:68-100) class-based handlers with DI. MCP Server: MOSMCPServer (mcp_serve.py:125-430) wraps MOS as FastMCP tools: chat/create_user/create_cube/register_cube/search_memories/add_memory/feedback/delete_all. Exposed to MCP hosts via fastmcp. DreamPlugin (dream/plugin.py:36-78) shows hook system: DREAM_EXECUTE/SEARCH_MEMORY_RESULTS/ADD_AFTER/MEMORY_ITEMS_AFTER_FINE_EXTRACT. plugin_manager auto-discovers. Self-hostable: Python 3.10+, Qdrant (embedded) or Milvus, optionally Neo4j/PolarDB/PG+pgvector. Ollama/Sentence Transformers for local embedders + LLMs. SQLite for users. Dockerfile. Only cloud-only: Ark embedder, OpenAI key (default path) - all swappable. NEO4J_USE_MULTI_DB=false (mcp_serve.py:27) for single-tenant local. OpenClaw/Hermes plugins in apps/.
memos.api.server_api:app. There is no memos.api.start_api module at this commit.plastic-labs/honcho
answeredHoncho integrates via multiple channels. REST API: The core FastAPI server exposes /v3/* routes covering workspaces, peers, sessions, messages, conclusions, keys, and webhooks (src/routers/). The chat endpoint at POST /v3/.../peers/{peer_id}/chat is the main memory-query surface, returning reasoning-grounded natural-language answers. SDKs: Two first-party SDKs live in this repo — sdks/python/ (published as honcho-ai on PyPI) and sdks/typescript/ (published as @honcho-ai/sdk on npm). Both expose Honcho(workspace_id, api_key) clients with methods for peer/session/message CRUD, chat, search, and context injection with .to_openai()/.to_anthropic() formatters. MCP: The platform hosts an MCP HTTP server at mcp.honcho.dev (or self-hosted), usable by any MCP client via claude mcp add honcho --transport http. Agent framework plugins: First-party memory plugins ship for Claude Code (marketplace plugin), Codex (npm install), Cursor (shell installer), OpenCode, OpenClaw, Hermes (built-in upstream), and DeepSeek Harness (Cordis plugin). All read the same ~/.honcho/config.json. Webhooks: src/webhooks/webhook_delivery.py delivers event-driven webhook payloads as configured (src/webhooks/events.py), enabling external service integration. Self-hosting: The entire stack is open-source (AGPL-3.0) and self-hostable via honcho start (Docker Compose-based CLI) or manual uv run fastapi dev src/main.py + uv run python -m src.deriver. Required services: Postgres with pgvector and an LLM provider API key (Gemini for Deriver by default; Anthropic for Dialectic medium+; OpenAI for embeddings). Optional: Redis (cache), external vector stores (Turbopuffer, LanceDB, Qdrant, ChromaDB). The API and deriver are separate processes sharing a database.
MemMachine/MemMachine
answeredAgents integrate via three interfaces. Python SDK (memmachine-client, PyPI package): the Memory class provides add(), search(), list(), delete_episodic(), delete_semantic(), and full semantic memory CRUD (add_feature(), add_semantic_category(), etc.). It communicates with the server over HTTP REST at /api/v2/memories. REST API (FastAPI server, api_v2/router.py): endpoints for adding (POST /api/v2/memories), searching (POST /api/v2/memories/search), listing (POST /api/v2/memories/list), configuring episodic memory, managing semantic categories/tags/sets, and health checks. MCP Server (server/mcp_stdio.py, server/mcp_http.py): wraps memory operations as MCP tools (add-memory, search-memories, delete-memories) for direct integration with Claude Desktop, Cursor, and any MCP client. The MCP server accepts org-id, proj-id, and user-id via HTTP headers or env vars. Framework integrations ship as Python packages: LangChain (integrations/langchain/memory.py), LangGraph (client/src/memmachine_client/langgraph.py), CrewAI (integrations/crewai/tool.py), LlamaIndex (integrations/llamaindex/mem_machine_memory.py), AWS Strands (integrations/aws_strands_agent_sdk/), Dify (integrations/dify/), and n8n and FastGPT. The server is fully self-hostable via Docker compose (docker-compose.yml): it requires PostgreSQL (with pgvector), plus one of Neo4j (declarative profile) or Qdrant (event profile), and optionally a vector store for semantic memory. All services are configured through a single configuration.yml. The project ships a memmachine-compose.sh helper script and sample_configs/ with event and declarative templates. OpenAI API keys are needed for LLM-based extraction and summarization, unless using Ollama or LiteLLM (configured via litellm_language_model.py). The Dockerfile packages the server, and docker-compose.yml wires healthchecks, persistent volumes, and networking.