LLMs Technical Reviews

MemoriLabs/Memori

Python/TS SDK that wraps your LLM client, extracts facts via Memori's hosted API, and stores and recalls them in your own SQL or Mongo DB.

GitHub ↗★ 17kPythonApache-2.0commit 574b1ea · 2026-09-17homepage ↗

Overview

Memori is a memory layer that you attach to an existing LLM client. You call Memori(...).llm.register(client) and set attribution(entity_id, process_id). From then on, every chat.completions.create (or the Anthropic, Gemini, xAI, Bedrock, LangChain or Agno equivalent) goes through Memori. Before the call, Memori injects facts it remembers about the entity. After the call, it stores the turn and starts background “Advanced Augmentation”, which turns the conversation into facts, semantic triples, process attributes and a running summary.

There are two modes. Cloud mode (no connection passed, MEMORI_API_KEY set) sends conversation messages, augmentation and recall to Memori’s hosted API. BYODB (“bring your own database”, a connection factory passed to the constructor) stores everything in your PostgreSQL, CockroachDB, MySQL, TiDB, OceanBase, Oracle, SQLite or MongoDB, and runs embedding and ranking locally. One point decides how you should read the rest: in both modes, the extraction step is a call to Memori’s hosted sdk/augmentation endpoint. The repository contains no extraction prompt and no local extraction LLM. BYODB keeps your data in your database, but it still sends each turn to Memori for processing.

The repo holds the Python SDK (memori/), a TypeScript SDK (memori-ts/), a Rust engine with PyO3 and napi bindings (core/), and integrations for Claude Code, Hermes and OpenClaw (integrations/).

Architecture

flowchart LR
  APP["Your code"] --> CLIENT["Wrapped LLM client"]
  CLIENT --> INV["Invoke pipeline"]
  INV --> RECALL["inject_recalled_facts"]
  INV --> HIST["inject_conversation_messages"]
  INV --> LLM["LLM provider"]
  INV --> POST["handle_post_response"]
  POST --> WRITER["Writer: messages to DB"]
  POST --> AUG["handle_augmentation"]
  AUG --> API["Memori hosted API (sdk/augmentation)"]
  AUG --> RUST["Rust engine (memori_python)"]
  RUST --> API
  API --> WRITES["Facts, triples, summary writes"]
  WRITES --> DB["Your DB (SQL or Mongo)"]
  RECALL --> SEARCH["Dense + BM25 ranking"]
  SEARCH --> DB
Component Path Role
SDK entry memori/__init__.py Memori class: mode selection, attribution, recall, delete_entity_memories
LLM registry and clients memori/llm/_registry.py, memori/llm/clients/ Matches the client type and replaces its create/parse methods with wrappers
Invoke pipeline memori/llm/invoke/invoke.py Sync, async and streaming wrappers: inject, call, post-process
Recall injection memori/llm/pipelines/recall_injection.py Retrieves facts and adds a <memori_context> block to the system prompt
Post-invoke memori/llm/pipelines/post_invoke.py, memori/memory/_writer.py Persists the turn and triggers augmentation
Augmentation memori/memory/augmentation/ Background runtime, AdvancedAugmentation, batched DB writer
Search memori/search/ FAISS cosine search plus BM25 rerank
Storage memori/storage/ Adapter and driver registries, per-dialect drivers and migrations
Rust engine core/src/ fastembed embeddings, retrieval, augmentation worker; storage stays in Python via callbacks
Integrations integrations/, memori-ts/ Claude Code skill, Hermes provider, OpenClaw plugin, TS SDK

How a request flows

Take a BYODB setup with an OpenAI client and mem.attribution("user_123", "support_agent"):

  1. Setup. Passing conn sets byodb=True. The constructor starts the storage manager, the augmentation manager and, unless disabled, the Rust core (init.py). Schema creation is not automatic. You call mem.config.storage.build() once (_manager.py).
  2. Wrap. llm.register(client) finds the OpenAi handler, which saves the original chat.completions.create, beta.chat.completions.parse and responses.create and installs wrappers (direct.py).
  3. Recall. On each call, Invoke.invoke first runs inject_recalled_facts and then inject_conversation_messages (invoke.py). Recall takes the last user message, retrieves facts (Rust core if present, else Python Recall), drops those under recall_relevance_threshold, and appends a <memori_context> block to system, instructions, the Gemini system instruction, or a system message (recall_injection.py).
  4. History. In BYODB, the same pipeline resolves or creates the entity, process, session and conversation rows and replays earlier messages of the current conversation into the request (conversation_injection.py).
  5. Persist the turn. After the provider responds, handle_post_response builds a payload and calls MemoryManager.execute, which strips system messages and writes the turn through Writer (BYODB) or posts it to cloud/conversation/messages (cloud) (post_invoke.py, _manager.py). This write is synchronous.
  6. Augment. handle_augmentation returns early without an entity or process id. Otherwise it posts to cloud/augmentation (cloud), submits to the Rust engine on a thread pool (BYODB with Rust core), or enqueues a Python job (_handler.py).
  7. Extract. On the Python path, AdvancedAugmentation.process reads the stored summary, sends only the last user/assistant pair when a summary exists, hashes entity and process ids with SHA-256, and awaits api.augmentation_async (_augmentation.py). The Rust path does the same call (pipeline.rs).
  8. Write memories. Returned facts are embedded locally. Facts, triples, process attributes and the new summary become write tasks for a batched DB writer (_augmentation.py).

Key components

Interception, not tools

Memori does not give the model a “remember” tool. It patches the client object, so memory happens on every call without code changes beyond registration. The Registry matches client objects by type and LangChain/Agno models by named arguments (_registry.py). Streaming responses are wrapped in iterator classes that write the turn when the stream ends. Those wrappers call only MemoryManager.execute, not handle_augmentation (iterator.py). So at this commit, a stream consumed through Iterator/AsyncIterator (for example stream=True on the OpenAI client) is stored as history but produces no new facts.

Augmentation runtime

The Python path runs on a dedicated asyncio loop in a daemon thread, guarded by a 50-slot semaphore (_runtime.py). Results go to a DB writer that batches up to 100 writes. A QuotaExceededError from the API disables augmentation for that instance, and later enqueues re-raise it (_manager.py). augmentation.wait() lets short scripts block until writes finish.

Storage

Registry.adapter matches the connection object (SQLAlchemy, DB-API, Django, MongoDB, also context-manager pools). Registry.driver picks the driver by dialect (_registry.py). The schema is relational: entity, process, session, conversation, message, memori_entity_fact (text, embedding blob, num_times, uniq hash), and a knowledge graph split into subject, predicate and object tables. Deduplication is exact-match only. A repeated fact hits ON CONFLICT (entity_id, uniq) and increments num_times (postgresql/_driver.py). A conversation is reused while its last message is within session_timeout_minutes (30) (postgresql/_driver.py).

Recall ranking

BYODB recall loads up to MEMORI_RECALL_EMBEDDINGS_LIMIT (1000) stored embeddings for the entity, builds an in-memory FAISS inner-product index over normalised vectors, and takes candidates (_faiss.py, _core.py). A BM25 score over the candidate pool is blended in: 0.85·cosine + 0.15·bm25, or 0.70/0.30 for queries of two tokens or fewer, both overridable by env vars (_lexical.py). Defaults are 5 facts and a 0.1 threshold (_config.py). There is no vector index in the database. Each recall is a brute-force scan of one entity’s most recent rows, which is fine for per-user memories but caps how much a single entity can usefully hold.

Rust engine

RustCoreAdapter.maybe_create enables the engine only in BYODB mode (_adapter.py). The engine owns fastembed embeddings, a worker pool for augmentation jobs and the retrieval pipeline. Storage remains in Python through three callbacks (fetch embeddings, fetch facts by id, write batch). Its run_retrieval mirrors the Python ranking (pipeline.rs).

Extending it

  • New database. Add an adapter with @Registry.register_adapter(matcher) and a driver with @Registry.register_driver(dialect), plus a migration module. CockroachDB reuses the PostgreSQL driver.
  • New LLM client. Register a BaseClient subclass with @Registry.register_client(matcher) and a payload adapter with register_adapter.
  • Manual use. mem.recall(query, limit) returns ranked facts without an LLM call. delete_entity_memories() removes an entity’s facts and triples (BYODB only). Cloud users also get agent_recall, agent_compaction and capture_agent_turn for agent harnesses.
  • Embeddings. MEMORI_EMBEDDINGS_MODEL changes the model (default all-MiniLM-L6-v2). The Python path can also call a TEI server.
  • Agent harnesses. integrations/claude-code is a Bash-invoked skill backed by Memori Cloud. integrations/hermes and integrations/openclaw are memory providers for those agents.

Running it

  • Cloud. pip install memori (or npm install @memorilabs/memori), set MEMORI_API_KEY, register a client, set attribution. Nothing else to host.
  • BYODB. Pass a connection factory such as a SQLAlchemy sessionmaker, call mem.config.storage.build() once, then use the client as normal. The wheel bundles the Rust extension. MEMORI_DISABLE_RUST_CORE=1 falls back to Python and FAISS. Network access to Memori’s API is still needed for extraction, and quota errors from it stop augmentation.
  • Development. docker-compose.yml starts PostgreSQL, MySQL, MongoDB, OceanBase and Oracle containers for the integration test suite.

Strengths and caveats

  • Strength: zero-friction capture. Wrapping the client means existing code gains memory with two lines, across sync, async and streaming calls and many providers.
  • Strength: your database. BYODB supports more database engines than most memory layers, with plain relational tables you can query and audit.
  • Strength: off the hot path. Extraction runs in the background. Only recall and the message write add latency to the call.
  • Caveat: extraction is a hosted service. The prompts and models that decide what becomes a memory are not in the repo. BYODB does not mean offline or self-contained.
  • Caveat: naive lifecycle. Updates are exact-text upserts with a counter. There is no contradiction handling, no decay and no merge of near-duplicates in the client. Deletion is per entity, all or nothing.
  • Caveat: brute-force recall. Ranking reloads up to 1000 embeddings per query and searches them in memory.
  • Caveat: streaming gap. Iterator-wrapped streams skip augmentation, so streaming chat apps may quietly learn nothing.
  • Caveat: implicit injection. Recalled facts and replayed history change your prompts silently. Token cost and prompt behaviour need monitoring.

Sources: code at 574b1ea, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (56 pages), verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit 574b1ea. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Memories are extracted from LLM conversations through an augmentation pipeline that runs asynchronously after each LLM invocation. The flow begins in handle_post_response (memori/llm/pipelines/post_invoke.py:94-144), which formats the conversation turn (query + response) into a payload and passes it to the MemoryManager, which calls handle_augmentation. For cloud mode the payload is posted to cloud/augmentation; for BYODB mode it is enqueued onto an async runtime backed by a thread-pool with a semaphore (max 50 workers, memori/memory/augmentation/_runtime.py:17-50). The AdvancedAugmentation class (memori/memory/augmentation/augmentations/memori/_augmentation.py:39-286) is the registered handler. It builds an AugmentationPayload containing conversation messages, a previous summary (if one exists), and metadata (entity/process attribution, LLM provider, framework, storage dialect) and sends it to the cloud API via api.augmentation_async(). The API response carries: entity facts (natural-language statements like "the user prefers Python"), semantic triples (structured subject-predicate-object facts), process attributes, and a conversation summary. The response is parsed into a Memories object (memori/memory/_struct.py:102-112) containing Entity, Process, and Conversation sub-objects. Entity facts are then embedded locally via the configured embedding model (default all-MiniLM-L6-v2) before being written. Deduplication is handled at write time by the storage driver: both entity facts and knowledge-graph triples use ON CONFLICT ... DO UPDATE SET num_times = num_times + 1 (memori/storage/drivers/postgresql/_driver.py:260-280 for facts, memori/storage/drivers/postgresql/_driver.py:554-574 for triples), so repeating the same fact increments a frequency counter rather than creating a duplicate row. For cloud-mode users the entire extraction pipeline runs server-side; for BYODB there is a Rust native core path that performs extraction locally. The entity ID is hashed (SHA-256) before being sent to the cloud API for privacy (memori/memory/augmentation/_models.py:15-18).

Editor's note. Correction: extraction is never local. In cloud mode the SDK posts to cloud/augmentation; in BYODB both the Python AdvancedAugmentation and the Rust core call Memori's hosted sdk/augmentation endpoint, and only the returned facts are embedded and written locally.

How are memories stored?

answered

Memori stores memories in relational databases (PostgreSQL, MySQL, SQLite, TiDB, OceanBase, Oracle, CockroachDB) via SQLAlchemy/dbapi adapters, or in MongoDB via a document adapter (memori/storage/_registry.py:18-67). The schema is defined through migration files — for PostgreSQL the full DDL is in memori/storage/migrations/_postgresql.py. The core tables are:

  • memori_entity — one row per entity (user/person/thing), keyed by external_id (your application's user ID).
  • memori_process — one row per agent process, similarly keyed.
  • memori_session — ties an entity + process to a session UUID.
  • memori_conversation — a session-scoped conversation with an optional summary TEXT populated by augmentation.
  • memori_conversation_message — individual message rows with role, type and content.
  • memori_entity_fact — the primary memory table: stores content (the fact as text), content_embedding (BYTEA for PostgreSQL/MySQL, BLOB for others), a numeric num_times counter, and a uniq hash for deduplication. A UNIQUE constraint on (entity_id, uniq) prevents duplicates; ON CONFLICT increments the counter.
  • memori_knowledge_graph — stores semantic triples via FK references to memori_subject, memori_predicate, and memori_object tables, plus its own num_times frequency counter.
  • memori_process_attribute — key-value attributes for the process, also with dedup via uniq.
  • memori_entity_fact_mention — links facts back to the conversation where they were mentioned.

Embeddings default to all-MiniLM-L6-v2 (memori/_config.py:62-63), computed either via a native Rust core (memori_python.NativeEmbedder, memori/native/_embeddings.py:24-32) or via a TEI (Text Embeddings Inference) server (memori/embeddings/_tei_embed.py). The store is pluggable: Registry.register_driver() maps dialect strings to driver classes (e.g., postgresql and cockroachdb share the same Postgres driver, memori/storage/drivers/postgresql/_driver.py:792-793), and Registry.register_adapter() maps connection objects to the right adapter (SQLAlchemy, MongoDB, dbapi, Django). Builders auto-detect the dialect and run migrations on startup (memori/storage/_builder.py, memori/storage/_manager.py:28-35).

How are memories retrieved and injected into the prompt?

answered

Retrieval is a two-phase process executed inside inject_recalled_facts() (memori/llm/pipelines/recall_injection.py:107-226), which is called by every Invoke.invoke() pipeline before the LLM request is sent (memori/llm/invoke/invoke.py:28-31).

Query extraction. The last user message is extracted from the kwargs by extract_user_query() (memori/llm/helpers/query_extraction.py:44-94), which supports all supported provider message formats (OpenAI-style messages, Anthropic, Google contents/request, xAI).

Semantic search. The query is embedded using the same model as write-time (all-MiniLM-L6-v2) via Recall._embed_query() (memori/memory/recall.py:213-218). For BYODB, the entity's stored fact embeddings are loaded from memori_entity_fact (up to MEMORI_RECALL_EMBEDDINGS_LIMIT, default 1000) into a FAISS IndexFlatIP index. Cosine similarity is computed via find_similar_embeddings() (memori/search/_faiss.py:93-133), which normalizes both the query and stored embeddings with L2 normalization before using inner-product search.

Lexical reranking. Candidate facts are reranked by a BM25-style scorer (memori/search/_lexical.py:74-124) that tokenizes query and fact text, removes stopwords, computes TF-IDF scores, and normalizes them. The final rank_score is a weighted combination: w_cos * cosine_sim + w_lex * bm25_score. For short queries (≤2 tokens) lexical weight increases from 0.15 to 0.30 (memori/search/_lexical.py:127-149).

Thresholding and injection. Facts below recall_relevance_threshold (default 0.1, memori/_config.py:87-88) are dropped. The surviving facts are formatted as bullet points with optional timestamps and conversation summaries, then wrapped in a <memori_context> tag. The context is injected into the provider-appropriate location — the system field for Anthropic/Bedrock, a system message for OpenAI, instructions for Google — so the LLM sees: "Only use the relevant context if it is relevant to the user's query." For cloud mode, the search is performed by the Memori Cloud API endpoint cloud/recall; results are parsed and filtered identically. Conversation history is also injected separately via inject_conversation_messages() (memori/llm/pipelines/conversation_injection.py:156-226), which replays prior messages from the current conversation into the LLM context.

How are memories updated, consolidated or forgotten?

answered

Updates, consolidation, and deletion follow distinct patterns:

Incremental updates via upsert. Both entity facts and knowledge-graph triples use PostgreSQL's ON CONFLICT mechanism. For entity facts, the unique constraint is (entity_id, uniq), where uniq is a deterministic hash of the fact content (memori/storage/drivers/postgresql/_driver.py:257,278-280). When the same fact is re-extracted from a later conversation, the row is not duplicated — instead num_times is incremented and date_last_time is refreshed. The same applies to knowledge-graph triples (constraint on entity_id, subject_id, predicate_id, object_id, memori/storage/drivers/postgresql/_driver.py:572-574) and process attributes (constraint on process_id, uniq).

Conversation summaries. The AdvancedAugmentation response can include a conversation.summary string, which is written back to memori_conversation.summary via the conversation.update write path (memori/memory/augmentation/augmentations/memori/_augmentation.py:280-285). This summary is then used as context for future augmentations: when selecting messages to send to the API, the augmentation handler reads the existing summary and sends it alongside the latest user/assistant turn rather than the full history (memori/memory/augmentation/augmentations/memori/_augmentation.py:53-84,135-137).

Deletion. There is an explicit Recall.delete_entity_memories() method (memori/memory/recall.py:193-211) that deletes all facts and knowledge-graph entries for a given entity. In BYODB mode, the driver's delete_by_entity() cascades through the schema: deleting from memori_entity_fact (which cascades to memori_entity_fact_mention) and from memori_knowledge_graph (which cascades to orphaned subjects, predicates, and objects).

No TTL or decay. This version of the codebase does not implement automatic time-to-live, decay, or consolidation merging of facts. There is no background job scanning for stale memories. The recall_relevance_threshold (default 0.1) provides a static retrieval-time filter, but facts persist indefinitely once written. Individual scaffolding for versioning (the uuid and date_updated columns on every table) exists in the schema but is not leveraged for history tracking at the application layer.

How is memory scoped and isolated?

answered

Memories are scoped through a three-level hierarchy of entity → session → conversation, with optional process-level isolation.

Entity. The top-level scope is an entity_id — a string provided by the caller (typically a user ID). It is hashed with SHA-256 before being sent to cloud endpoints for privacy (memori/memory/augmentation/_models.py:15-18). In BYODB mode it maps to memori_entity.external_id. Every fact, session, and conversation is ultimately owned by an entity. If entity_id is not set, memory extraction (AdvancedAugmentation.process() at memori/memory/augmentation/augmentations/memori/_augmentation.py:126-127) and recall (inject_recalled_facts at memori/llm/pipelines/recall_injection.py:112-114) are both skipped entirely.

Process. An optional process_id distinguishes which agent or process within an application is interacting on behalf of the entity. This maps to memori_process and enables process-attribute memories (preferences about a specific agent) separate from entity-level facts. When provided, it is included in cloud API payloads as attribution and stored in the session relationship.

Session. A session_id (UUID string) ties together entity + process into a session row in memori_session. Conversations are scoped to a session. A session has a configurable timeout (session_timeout_minutes, default 30, memori/_config.py:84). The conversation.create driver method (memori/storage/drivers/postgresql/_driver.py:37-93) checks the last activity time against this timeout: if the most recent message is within the window, the existing conversation is reused; otherwise a new conversation is created, effectively starting a fresh conversation within the same session.

Multi-tenancy. In BYODB mode, each application provides its own database, so tenant isolation is at the database level. In cloud mode, entity IDs are the isolation boundary — the API key determines the project/account, and entity IDs scope within it. There is no role-based access control or per-entity ACL enforcement in the open-source SDK; these are presumed to be handled by the hosting application.

Privacy. The entity_id is SHA-256 hashed before transmission to the cloud API, preventing direct plaintext ID leaks to the server.

Editor's note. Correction: only the BYODB augmentation payload hashes entity and process ids with SHA-256. In cloud mode, recall, augmentation metadata and agent endpoints send the raw entity_id.

How do agents integrate with it, and what is self-hostable?

answered

Memori offers multiple integration surfaces:

Python SDK (pip install memori). The primary SDK at memori/ provides a Memori builder that registers with any supported LLM client (OpenAI, Anthropic, Google, xAI, Bedrock, LiteLLM). The core integration point is an LLM Invoke class (memori/llm/invoke/invoke.py:24-155) that wraps every LLM call in a three-stage pipeline: conversation-history injection, fact recall injection (inject_recalled_facts) and post-response handling (handle_post_response). The LLM client registration happens through adapters — e.g., memori/llm/adapters/openai/, memori/llm/adapters/anthropic/ — which monkey-patch or wrap the client's .create() / .invoke() method to route through Memori's pipeline.

TypeScript SDK (npm install @memorilabs/memori). Listed in the README with analogous LLM registration pattern.

BYODB (Bring Your Own Database). The self-hostable mode allows running against PostgreSQL, MongoDB, SQLite, MySQL, TiDB, OceanBase, CockroachDB, or Oracle. Storage is pluggable: Registry.register_adapter() maps connection types to adapters (SQLAlchemy, MongoDB native, Django, DB-API), and Registry.register_driver() maps the detected dialect to the driver implementation (memori/storage/_registry.py:18-67). An optional Rust native core (memori_python) handles local embedding and augmentation.

MCP Server. Memori provides an MCP endpoint at https://api.memorilabs.ai/mcp/ that works with Claude Code, Cursor, Codex, and Warp. The MCP transport uses HTTP with API-key headers for entity and process attribution. This does not require any SDK integration — just a claude mcp add command.

Framework integrations. A Hermes Agent memory provider (integrations/hermes/) and an OpenClaw plugin (@memorilabs/openclaw-memori) allow drop-in use with those agent frameworks.

Services required. Cloud mode requires only a MEMORI_API_KEY. BYODB mode requires a running database and optionally a TEI server or native Rust core for local embeddings. The entire codebase is Apache 2.0 licensed — there are no closed-source gates in the Python SDK itself, though the cloud augmentation API is a hosted service.

Editor's note. Correction: the Rust core handles local embedding and retrieval, not augmentation; extraction always goes through Memori's hosted API, so BYODB is not self-contained. Also, iterator-wrapped streaming responses are stored but never augmented at this commit.