LLMs Technical Reviews

MemTensor/MemOS

Python memory server (graph DB + vectors, LLM extraction, async scheduler) plus a separate local SQLite plugin for agent harnesses.

GitHub ↗★ 12kTypeScriptApache-2.0commit a7367d0 · 2026-09-22homepage ↗

Overview

MemOS (“Memory Operating System”) from MemTensor is two products in one repository, and they share no code.

The first is the Python service in src/memos (PyPI MemoryOS, import memos). It is a FastAPI server that turns chat messages into memory items with LLM prompts, stores them as nodes in a graph database (Neo4j, Neo4j Community + Qdrant, PolarDB, Postgres or NebulaGraph) with embeddings, and searches them through several parallel recall paths. The memories are typed: long-term facts, user memories, preferences, tool traces, skills and raw-file chunks. A scheduler does the expensive LLM extraction in the background. The library also contains the older MOS/GeneralMemCube API and the KV-cache and LoRA “activation” and “parametric” memory types from the MemOS paper. The REST server does not wire up those two types.

The second is apps/memos-local-plugin, a TypeScript package (the largest code base in the repo, and why GitHub reports the project as TypeScript). It is a local-first memory plugin for OpenClaw, Hermes Agent and DeepSeek Harness, built on one SQLite file. It records step-level traces, gives them value with a reward signal, induces policies and “world models”, and crystallizes skills. The README’s newer claims (“self-evolving”, skill reuse) describe this plugin, not the Python server.

This overview follows the Python service in detail and covers the plugin in its own section.

Architecture

flowchart LR
  C["Client / agent"] --> API["FastAPI server_api"]
  API --> ADD["AddHandler"]
  API --> SRCH["SearchHandler"]
  ADD --> VIEW["SingleCubeView / CompositeCubeView"]
  SRCH --> VIEW
  VIEW --> READER["MultiModalStructMemReader"]
  READER --> LLM["Extraction LLM"]
  VIEW --> TREE["SimpleTreeTextMemory"]
  TREE --> MGR["MemoryManager"]
  MGR --> GDB["Graph DB + vectors"]
  VIEW --> SCHED["MemScheduler (local or Redis queue)"]
  SCHED --> READER
  SCHED --> TREE
  TREE --> SEARCHER["Searcher: parallel paths"]
  SEARCHER --> GDB
  SEARCHER --> NET["Internet retriever"]
Component Path Role
REST app src/memos/api/server_api.py, api/routers/server_router.py /product/add, /product/search, chat, feedback, delete, scheduler status
Component wiring src/memos/api/handlers/component_init.py Builds the graph DB, LLMs, embedder, reader, text memory, searcher, feedback and scheduler from env config
Cube views src/memos/multi_mem_cube/ Per-cube add and search. A cube id defaults to the user id
Mem reader src/memos/mem_reader/multi_modal_struct.py Parses messages and files into chunks (fast) or LLM-extracted memories (fine)
Tree text memory src/memos/memories/textual/tree.py, tree_text_memory/ Graph-backed memory: manager, reorganizer, searcher, recall, reranker
Graph stores src/memos/graph_dbs/ Neo4j, Neo4j Community, PolarDB, Postgres, NebulaGraph backends
Scheduler src/memos/mem_scheduler/ Task queue and handlers (mem_read, add, feedback, organize)
Plugins src/memos/plugins/ Hook registry. Some hooks, such as memory versioning, have no implementation in the repo
Local plugin apps/memos-local-plugin/ Separate TS memory engine with SQLite, three-tier retrieval and skills

How a request flows

A self-hosted POST /product/add followed by POST /product/search:

  1. Startup. server_router calls handlers.init_server() at import time. That builds every component once: graph DB, LLMs, embedder, MultiModalStructMemReader, SimpleTreeTextMemory, Searcher, SimpleMemFeedback and the scheduler (component_init.py). The ASGI app is memos.api.server_api:app (server_api.py).
  2. Route. /product/add goes to AddHandler.handle_add_memories, which builds a SingleCubeView per writable cube (default: the user id) (add_handler.py).
  3. Fast write. async_mode defaults to "async" (product_models.py). In async mode the reader runs in fast mode: messages are parsed and windowed into raw chunks with no LLM call, then written to the graph straight away (single_cube.py, multi_modal_struct.py).
  4. Schedule. _schedule_memory_tasks submits a mem_read message with the new ids to the scheduler (single_cube.py). The HTTP call returns.
  5. Fine extraction. The mem_read handler loads the raw chunks and calls fine_transfer_simple_mem. It adds the resulting memories, then hard-deletes the raw chunks and their working-memory bindings (mem_read_handler.py, L395-L425). Fine mode runs four extractors in parallel: facts, tool trajectories, skills and preferences (multi_modal_struct.py).
  6. Store. MemoryManager._add_memories_batch writes nodes to the graph store in batches of 5. It queues them for the reorganizer only when reorganize is enabled (manager.py).
  7. Search. /product/search defaults to mode=fast and top_k=10. SingleCubeView dispatches to fast, fine or mixture search (single_cube.py). Searcher._retrieve_paths then runs working memory, long-term + user memory, internet and, when requested, keyword, tool, skill and preference paths in a thread pool (searcher.py).
  8. Return. Results are reranked, deduplicated, trimmed to top_k and grouped by type (text_mem, pref_mem, tool_mem, skill_mem). The search endpoint only returns memories. Prompt injection happens in the chat endpoints or in the caller.

Key components

Two-stage ingestion

The fast/fine split is the central design choice. A write is visible to search almost immediately as raw chunks, and quality improves later when the scheduler swaps the chunks for extracted memories. The price is a window in which search returns raw text. If the scheduler is behind or down, those chunks stay. Synchronous callers can pass async_mode="sync" to get fine extraction inline.

Hybrid recall

For long-term and user memory, GraphMemoryRetriever.retrieve runs graph recall (by parsed keys and tags) and vector recall in parallel, and optionally BM25 and full-text recall. It merges them by node id (recall.py). Graph recall keeps nodes whose key matches a parsed key or whose tags overlap the query’s by at least two. Working memory is read in full rather than searched. Query expansion via chain-of-thought sub-queries is also optional (searcher.py). In the server’s default config, BM25, CoT, full-text and fast-graph are all off (env flags BM25_CALL, VEC_COT_CALL, FULLTEXT_CALL, FAST_GRAPH), so out of the box this is graph + vector recall plus a reranker (config.py).

Organization and lifecycle

MemoryManager keeps per-type capacities. The library defaults are 20 working, 1500 long-term and 480 user memories, and the server raises long-term and user to 1e6 (manager.py). A sync add trims working memory to its limit. Conflict handling, where NodeHandler detects contradictory or redundant pairs and merges them with an LLM, lives in GraphStructureReorganizer. That reorganizer only runs its threads when MOS_ENABLE_REORGANIZE=true (reorganizer.py, L185-L210). The default server therefore does not resolve contradictions. The memory-version pipeline in the reader is a plugin hook with no implementation in this repo.

Feedback

/product/feedback, or /product/add with is_feedback, sends a natural-language correction to SimpleMemFeedback. It uses the searcher and an LLM to edit or replace existing memories, synchronously or as a scheduler task.

The local plugin (TypeScript)

apps/memos-local-plugin/core is a separate engine. Storage is a better-sqlite3 database with float32 vector BLOBs and brute-force cosine search in JS (vector.ts). At turn start it retrieves from three tiers in parallel: skills, then traces and episodes, then world models (retrieve.ts). A task-level reward is back-propagated along a trace as a discounted value with exponential time decay. This value sets which traces surface, and which crystallize into skills (backprop.ts). Adapters run it in-process (OpenClaw, DeepSeek Harness) or over a JSON-RPC bridge (Hermes), with a local HTTP viewer.

Extending it

  • Providers. Each category (LLM, embedder, vector DB, graph DB, chunker, reranker, memory, scheduler) follows the same pattern: a base.py class, a pydantic config in configs/, and an entry in the factory’s backend_to_class.
  • Plugins. memos.plugins discovers plugins at startup and lets them register components and hooks around add, search and the dream pipeline. Handlers are decorated with @hookable.
  • Cubes. writable_cube_ids and readable_cube_ids fan one request out to several cubes, which is how shared and per-user memory are combined.
  • Library use. MOS / GeneralMemCube still offer add, search and chat in-process, and the FastMCP server in api/mcp_serve.py wraps that MOS API.

Running it

  • Docker. docker/docker-compose.yml starts the API, Neo4j 5.26 and Qdrant. Copy docker/.env.example, set LLM and embedder keys, then docker compose up. The default graph backend is neo4j-community, with vectors in Qdrant.
  • Local. make install (Poetry, all extras), then make serve (uvicorn memos.api.server_api:app). Redis is optional and enables the Redis-stream scheduler queue. A SQLite file backs scheduler bookkeeping.
  • Required. An OpenAI-compatible LLM (or Ollama, vLLM, Qwen, DeepSeek), an embedder, and one supported graph store. Internet retrieval needs a Tavily, Bocha or Xinyu key and runs only when a search request sets internet_search.
  • Local plugin. install.sh for OpenClaw/Hermes, or dsh plugin for DeepSeek Harness. It needs Node.js and an LLM/embedding provider, and no server.

Strengths and caveats

  • Strength: rich memory model. Typed memories (facts, user, preference, tool, skill, raw file) with sources, tags, status and history are far more structured than a flat fact list.
  • Strength: write latency. The fast-then-fine pipeline keeps /add cheap and moves LLM work to a queue that can scale out on Redis.
  • Strength: pluggable everywhere. Five graph backends, two vector DBs and ten LLM backends sit behind uniform factories.
  • Caveat: heavy and sprawling. The default server needs a graph DB, a vector DB, several LLM roles and many env flags. Code paths (MOS, cube views, scheduler, plugins) overlap, and many features are off by default.
  • Caveat: contradiction handling is opt-in. Without the reorganizer, new facts are added next to old ones. Nothing merges or retires them.
  • Caveat: two unrelated engines. Benchmarks and “self-evolving” claims cannot be attributed to one code path without checking which product was measured. The TS plugin’s memory does not interoperate with the Python server’s graph.

Sources: code at a7367d0, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (27 pages), verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit a7367d0. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Extraction uses LLM-driven prompts. BaseTextMemory.extract() (general.py:45-79) concatenates messages into SIMPLE_STRUCT_MEM_READER_PROMPT (mem_reader_prompts.py:1-101) which instructs the LLM to extract facts, events, preferences, plans from user perspective, resolving time/pronoun references, writing third-person. Returns JSON with memory list of key/value/tags/memory_type. PreferenceTextMemory (preference.py:69-79) delegates to ExtractorFactory. Write-time dedup: NaiveAdder (adder.py:53-165) uses LLM judges to compare new vs existing memories. Tree-text: NodeHandler.detect() (handler.py:30-74) finds embedding-similar candidates (0.8 threshold), LLM classifies as contradictory/redundant/independent; resolve() (handler.py:76-128) fuses or hard-updates, archives originals.

Editor's note. Correction: NodeHandler conflict detection (0.8 similarity threshold) runs only when MOS_ENABLE_REORGANIZE=true, which is off by default. It is not part of the default write path.

How are memories stored?

answered

Pluggable backends across 3 tiers. Vector DB: Qdrant/Milvus with collection_name, vector_dimension, cosine/euclidean/dot. VecDBItem payload holds TextualMemoryItem + embedding (general.py:81-104). Graph DB: MemoryManager (organize/manager.py:54-88) uses Neo4j/PolarDB/PG+pgvector (graph_dbs/factory.py:14-19) as labeled nodes with typed edges. EmbedderFactory (embedders/factory.py:16-36): Ollama, Sentence Transformers, Ark, OpenAI; CachingEmbedder adds LRU+TTL. TextualMemoryItem schema (item.py:94-212): id, memory text, TreeNodeTextualMemoryMetadata with memory_type (9 categories), embedding, sources, tags, status, user_id, session_id, version, history, confidence. ActivationMemory = KV-cache; ParametricMemory = LoRA. Env-driven config with local-first defaults.

How are memories retrieved and injected into the prompt?

answered

Hybrid search: semantic vector + graph + BM25. GeneralTextMemory.search() (general.py:121-137) embeds query, vector search, sort by score. PreferenceTextMemory.search() (preference.py:81-95) auto-filters status=activated. GraphMemoryRetriever (recall.py:61-100) runs 3 parallel paths: graph dispatch plan (typed edges), Neo4j search_by_embedding() (neo4j.py:840-880, pre-filtered cosine), EnhancedBM25. Reranker scores union. search_text_memories() (search_service.py:68-80) is shared entry: builds SearchContext with session_id/filter/info. CompositeCubeView.search_memories() (composite_cube.py:46-83) parallel-searches all cubes, aggregates into MOSSearchResult (text_mem/act_mem/para_mem/pref_mem/tool_mem/skill_mem). Results injected into LLM context via MOS.search() -> MemChat.

How are memories updated, consolidated or forgotten?

answered

Versioned archival + Dream consolidation + scheduler. NodeHandler.resolve() (handler.py:76-188) merges conflicting memories with MERGED_TO edge, sets originals to archived. Each TextualMemoryMetadata has version + history of ArchivedTextualMemory (item.py:109-128). GraphStructureReorganizer (reorganizer.py:81-107) runs background queue for add/remove/merge/update + periodic clustering under LLM-summarized parents. Scheduler has MEM_ARCHIVE_TASK_LABEL (task_schemas.py:38). Dream plugin (off by default): DreamSignalStore (signal_store.py:8-62) accumulates signals, triggers at 100, runs Contextualizer->MotiveFormation->DirectRecall->ConsolidationReasoning->DiarySummary->Persistence. DreamMemoryLifecycle (dream/types.py:67-80): last_hit_at, hit_count, usefulness_score, invalidated_by_feedback. maintenance.py (dream/maintenance.py:1-56) documents 3 planned strategies (not implemented): stale-by-TTL, low-usefulness archive, invalidation-by-feedback. Feedback API (mem_feedback/feedback.py:315) archives replaced memories.

How is memory scoped and isolated?

answered

Multi-layered: user, cube, session, database. User-level: MOSCore (core.py:38-68) validates via UserManager (user_manager.py:56-100, SQLite/MySQL/Redis) with roles ROOT/ADMIN/USER/GUEST + user_cube_association. Cube-level: GeneralMemCube (mem_cube/general.py:24-48) has separate stores per cube_id. SingleCubeView (single_cube.py:48-100) scopes via UserContext + search_filter. Admin API (admin_router.py) manages per-cube API keys with scopes + expiry. Database-level: PolarDB/Neo4j use_multi_db (api/config.py:868-903) gives each user separate DB; shared mode uses user_name tag. TextualMemoryItem fields: user_id, session_id, visibility private/public/session (item.py:101-148). Search auto-filters user_name + status=activated (preference.py:92-94). CompositeCubeView (composite_cube.py:18-50) fan-out writes + parallel searches across cubes.

How do agents integrate with it, and what is self-hostable?

answered

3 channels. Python SDK: MOS.simple()/MOS(config) (main.py:24-80), register_mem_cube()/add()/search()/chat()/create_user(). PyPI as MemoryOS. REST API: FastAPI at memos.api.start_api (server_api.py:39-73), /product/* (add/search/chat/feedback) + /admin/* key mgmt. server_router.py (server_router.py:68-100) class-based handlers with DI. MCP Server: MOSMCPServer (mcp_serve.py:125-430) wraps MOS as FastMCP tools: chat/create_user/create_cube/register_cube/search_memories/add_memory/feedback/delete_all. Exposed to MCP hosts via fastmcp. DreamPlugin (dream/plugin.py:36-78) shows hook system: DREAM_EXECUTE/SEARCH_MEMORY_RESULTS/ADD_AFTER/MEMORY_ITEMS_AFTER_FINE_EXTRACT. plugin_manager auto-discovers. Self-hostable: Python 3.10+, Qdrant (embedded) or Milvus, optionally Neo4j/PolarDB/PG+pgvector. Ollama/Sentence Transformers for local embedders + LLMs. SQLite for users. Dockerfile. Only cloud-only: Ark embedder, OpenAI key (default path) - all swappable. NEO4J_USE_MULTI_DB=false (mcp_serve.py:27) for single-tenant local. OpenClaw/Hermes plugins in apps/.

Editor's note. Correction: the REST app is memos.api.server_api:app. There is no memos.api.start_api module at this commit.