vectorize-io/hindsight
Postgres-based agent memory server: LLM fact extraction, four-arm recall with RRF and cross-encoder, LLM consolidation, agentic reflect.
Overview
Hindsight is a memory server for agents. You talk to it through three verbs. Retain stores content: an LLM extracts facts, entities, dates and causal links and writes them to Postgres. Recall retrieves facts with four retrieval arms, fuses them and reranks them. Reflect answers a question with an agentic loop over the stored memory. Each memory lives in a bank (one per user, agent or project). After each retain, a background consolidation job uses an LLM to fold new raw facts into deduplicated “observations”.
The engine is a Python/FastAPI service (hindsight-api-slim) built entirely on Postgres. pgvector (or DiskANN, vchord, ScaNN) handles vectors, a tsvector or BM25 extension handles keywords, and plain tables hold the entity graph and the job queue. With no configuration it starts an embedded Postgres (pg0), local bge-small-en-v1.5 embeddings and a local ms-marco-MiniLM cross-encoder, so the only external dependency is an LLM API key. Around the engine sit a Next.js control plane, a Rust CLI, generated Python/TypeScript/Rust/Go clients, an MCP endpoint, and a very large hindsight-integrations/ folder (over 50 frameworks and coding agents).
The vocabulary has shifted over time, and the code is the reference. Raw facts are memory_units with fact_type world or experience. Consolidated beliefs are memory_units with fact_type='observation'. “Mental models” and “knowledge pages” are named, reflect-generated summaries that can be refreshed. A migration literally renames the old mental_model fact type to observation. Reflect itself is read-only and does not do the consolidating.
Architecture
flowchart LR
C["Clients: SDKs, MCP, CLI, integrations"] --> API["FastAPI: api/http.py, /mcp"]
API --> ENG["MemoryEngine"]
ENG --> RET["Retain orchestrator"]
RET --> FX["LLM fact extraction"]
RET --> EMB["Embeddings"]
RET --> ER["Entity resolver"]
RET --> PG["Postgres: memory_units, links, entities"]
ENG --> RC["Recall: 4 arms per fact type"]
RC --> PG
RC --> FU["RRF fusion"] --> RR["Cross-encoder + boosts"]
ENG --> RF["Reflect agent (tools)"]
RF --> RC
ENG --> Q["async_operations queue"]
Q --> W["Worker poller"]
W --> CO["Consolidation (LLM)"]
CO --> PG
| Component | Path | Role |
|---|---|---|
| HTTP API | hindsight-api-slim/hindsight_api/api/http.py |
/v1/default/banks/{bank_id}/... routes: memories (retain), recall, reflect, mental models, consolidation, documents, entities |
| MCP | hindsight_api/api/mcp.py, mcp_tools.py |
Same operations as MCP tools on the API server |
| Engine | hindsight_api/engine/memory_engine.py |
retain_batch_async, recall_async, reflect_async, task submission; very large single module |
| Retain pipeline | engine/retain/ |
orchestrator.py, fact_extraction.py, entity_processing.py, link_creation.py, fact_storage.py |
| Memories store | engine/memories/base.py, postgres.py, pg/ |
Pluggable interface over every memory table; Postgres is the default |
| Search | engine/search/ |
retrieval.py, fusion.py, reranking.py, tags.py, temporal_extraction.py |
| Consolidation | engine/consolidation/ |
LLM job that creates, updates or deletes observations |
| Reflect | engine/reflect/ |
Tool-calling agent: search_mental_models, search_observations, recall, expand, done |
| Worker | hindsight_api/worker/ |
Polls the DB for tasks with FOR UPDATE SKIP LOCKED |
| Extensions | hindsight_api/extensions/ |
Tenant/auth, memory-defense (redaction), operation validators, alternate memory stores |
| Control plane | hindsight-control-plane/ |
Next.js admin UI |
| Integrations | hindsight-integrations/ |
Claude Code, Codex, Cursor, CrewAI, LangGraph, Pydantic AI, OpenAI Agents and many more |
How a request flows
Take POST /v1/default/banks/alice/memories with one conversation, then a recall:
- Accept.
api_retaineither queues the batch withsubmit_async_retain(async=true) or runsretain_batch_asyncinline (http.py). - Extract.
_extract_and_embedcallsfact_extraction.extract_facts_from_contents. It chunks the text and asks the LLM for structured facts (what/when/where/who/why, fact type, entities, causal relations) in the bank’s extraction mode, which defaults toconcise. Then it embeds each fact with its date added to the text (orchestrator.py). - Resolve entities. Entities are matched to existing canonical entities in a separate phase before the write transaction.
- Write.
_insert_facts_and_linksruns in one transaction. It insertsmemory_units, links units to entities, and creates temporal, semantic (ANN neighbours above a similarity threshold) and causal links. Entity-to-entity edges for the UI graph are deferred and are not needed for retrieval (orchestrator.py). - Schedule consolidation. After the write, the engine submits a consolidation task when observations and auto-consolidation are enabled, plus graph maintenance (memory_engine.py).
- Consolidate. A worker picks up the task and runs
run_consolidation_job. It reads unconsolidated facts, fetches similar existing observations, and asks the LLM to create, update or delete observations, preferring update over a near-duplicate (consolidator.py, prompts.py). - Recall.
recall_asyncembeds the query and callsretrieve_all_fact_types_parallel. That function extracts a date range from the query text, then makes onerecall_unifiedstore call that returns semantic, BM25, graph and temporal hits per fact type (memory_engine.py, retrieval.py). - Fuse, rerank, trim. The arms are merged with RRF (k = 60) (fusion.py). A cross-encoder rescores the merged list, with small multiplicative boosts for recency, temporal proximity and proof count (reranking.py). An MMR step diversifies the results, which are then cut to
max_tokens.
Key components
Retain
Extraction is the expensive and opinionated step. The concise prompt tells the model to skip anything not worth recalling later. It also resolves relative dates to absolute ones and keeps coreferences consistent. Documents with a document_id are upserted, and a delta path re-extracts only changed chunks of a re-sent document. A memory-defense extension can screen or redact content before anything is stored.
Storage
Everything is Postgres. engine/memories/base.py defines the interface for the memory tables (memory_units, memory_links, unit_entities, documents, chunks, entities, entity_cooccurrences, invalidated_memory_units). An alternative store can own a bank’s memories through HINDSIGHT_API_MEMORIES_EXTENSION, while banks and operations stay in Postgres (base.py). Keyword search defaults to native tsvector, and pg_search or vchord BM25 can replace it.
Recall
All four arms run per fact type (world, experience, observation). Graph retrieval walks entity co-membership, semantic kNN links and causal links. The temporal arm uses the date range extracted from the query, or one the caller supplies. Recency decay only changes ranking: linear over 365 days by default, optionally exponential (90-day half-life) or none. Nothing expires. prefer_observations drops raw facts that a returned observation already covers.
Reflect
reflect_async is an agentic loop. The model starts with no retrieved context and calls read-only tools (mental models first, then observations, raw recall, chunk expansion) until it calls done. Tags passed by the caller are enforced on every internal tool call, and the docstring is explicit that reflect persists nothing (memory_engine.py). Mental models are reflect runs over a stored source_query, saved with version history and refreshed on demand or after consolidation.
Scoping and tenancy
Every query is filtered by bank_id. Inside a bank, tags (TEXT[] with GIN indexes) and tags_match modes (any, all, the _strict variants and exact) narrow retain, recall and reflect. Multi-tenancy is an extension: TenantExtension.authenticate returns a Postgres schema per tenant. The built-in API-key implementation has to be switched on with HINDSIGHT_API_TENANT_EXTENSION (tenant.py).
Worker
Long operations (async retain, consolidation, mental-model refresh, graph maintenance) are rows in an operations table. worker/poller.py claims them with FOR UPDATE SKIP LOCKED (poller.py), so no Redis or Celery is needed. The API can run the worker in-process or as separate processes.
Extending it
- Providers.
HINDSIGHT_API_LLM_PROVIDERcovers OpenAI (the default,gpt-4o-minifallback model), Anthropic, Gemini, Groq, Bedrock, Vertex, Ollama, LM Studio, LiteLLM and more. Embeddings and reranker each have local, ONNX, TEI, OpenAI, Cohere and similar backends (config.py, L1283-L1285). - Per-bank config. Extraction mode, custom extraction prompts, entity label taxonomies, the observations mission, consolidation strategies and directives (rules reflect must follow) are resolved hierarchically per bank.
- Extensions. Tenant/auth, memory defense, operation validators, HTTP/MCP extensions and alternative memory stores load from env-configured
module:Classstrings. - Integrations.
hindsight-integrations/claude-codewiresSessionStart,UserPromptSubmit(auto-recall) andStop(auto-retain) hooks.npx @vectorize-io/hindsight-coding-agents install alldoes the same for many other coding agents.
Running it
- One process.
pip install hindsight-apiandhindsight-apistart the server on port 8888 with embedded Postgres (DEFAULT_DATABASE_URL = "pg0"), local embeddings and a local reranker (config.py). Set an LLM key. The Docker image bundles the API and the control plane. - Production. Point
HINDSIGHT_API_DATABASE_URLat a Postgres with pgvector (Alembic migrations run on start), run separate worker processes, and use the Helm chart inhelm/. - Local agent use.
hindsight-embedruns a self-managing background daemon that shuts down after five idle minutes, for CLI and local MCP use without a server.
Strengths and caveats
- Strength: serious retrieval. Four arms, RRF, a cross-encoder and explicit temporal parsing are more than most memory layers do, and they all run inside one Postgres.
- Strength: memory that improves. LLM consolidation merges repeated facts into observations with proof counts and source ids, so recall can return beliefs instead of fragments.
- Strength: operationally simple. No separate vector DB, queue or graph DB. Postgres holds data, indexes and jobs.
- Caveat: LLM-heavy writes. Every retain makes extraction calls, and every retain then triggers consolidation calls. Cost and latency scale with ingest volume.
- Caveat: size and churn.
memory_engine.pyis over 23,000 lines, and the terminology (mental models, observations, knowledge pages) has been renamed more than once. Expect to read code to know what a term means at a given version. - Caveat: no forgetting by age. Recency is only a ranking boost. Removal comes from consolidation deleting observations, explicit deletes, or invalidation.
Sources: code at 8830bbb, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (29 pages), verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit 8830bbb. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredWhen content is retained, the pipeline first chunks oversized input at sentence boundaries (chunk_text, fact_extraction.py:2580-2584), then runs LLM-based fact extraction. The extractor uses one of four prompt templates — concise (default), verbose, verbatim, or custom — assembled by _build_extraction_prompt_and_schema (fact_extraction.py:1561-1712). The base prompt (_BASE_FACT_EXTRACTION_PROMPT, line 1058) asks the LLM to extract structured facts with fields: what, when, where, who, why, fact_type (world/assistant), entities, and optional causal_relations. The default concise mode (_CONCISE_GUIDELINES, line 1121) emphasises selectivity: "Would this be useful to recall in 6 months? If no, skip it." Chunks are extracted in parallel via asyncio.gather (fact_extraction.py:2600-2621). Each chunk returns a FactExtractionResponse (pydantic model) parsed via structured output/JSON mode depending on the LLM provider. Extracted facts are then embedded (via the configurable Embeddings class supporting local SentenceTransformers, ONNX, OpenAI, Cohere, etc.; embeddings.py:1-80), written as memory_units rows in PostgreSQL, and entity resolution runs via entity_resolver.py using trigram similarity for fuzzy dedup within the batch (entity_resolver.py:79-80). A trigram-similarity threshold of 1.0 (identical trigram sets) deduplicates names differing only in punctuation or case. Configurable entity labels (entity_labels in config) constrain extraction to a taxonomy. The LLM is also instructed to convert relative temporal expressions to absolute dates (line 1098-1101) and maintain coreference resolution (line 1075-1080).
How are memories stored?
answeredMemories are stored in PostgreSQL across several tables created by the initial migration (5a366d414dce_initial_schema.py). The central table is memory_units (line 266), which stores each fact as a row with columns: id (UUID PK), bank_id, text, embedding (Vector(384) by default), context, event_date, occurred_start, occurred_end, mentioned_at, fact_type (world/experience/observation), confidence_score, access_count, metadata (JSONB), tags (TEXT[]), plus timestamps. The embedding dimension is auto-detected from the model. Vector indexes support pgvector HNSW, pgvectorscale DiskANN, vchord, or alloydb_scann (initial_schema.py:363-368). For keyword search, a search_vector column is maintained — native PostgreSQL tsvector, vchord_bm25 vector, or TEXT for extension-backed BM25 (initial_schema.py:310-332). Entity storage lives in entities (canonical_name, bank_id, metadata JSONB, first_seen, last_seen, mention_count) with a unique index on (bank_id, LOWER(canonical_name)). unit_entities joins memory_units to entities. memory_links stores precomputed semantic (kNN-graph), causal, and entity co-occurrence links. chunks stores document chunks with their own embeddings. The documents table records source documents. The PostgresMemories class (memories/postgres.py) implements the pluggable MemoriesExtension interface (memories/base.py); alternative storage backends can be loaded via HINDSIGHT_API_MEMORIES_EXTENSION env var. Embeddings are generated by a pluggable Embeddings class (embeddings.py) supporting local (SentenceTransformers, ONNX) and remote providers (OpenAI, Cohere, Google, TEI, LiteLLM).
How are memories retrieved and injected into the prompt?
answeredRecall (recall_async, memory_engine.py:8676-8735) uses 4-way parallel retrieval per fact type. The orchestration runs through retrieve_all_fact_types_parallel (search/retrieval.py:100-245), which first extracts temporal constraints from the query text (temporal_extraction.py), then calls the store's unified recall_unified method. The four arms are: (1) Semantic — ANN vector search using the configured vector index (HNSW/DiskANN) on the embedding column; (2) BM25 keyword — PostgreSQL @@ tsquery on search_vector (or vchord_bm25/pg_search depending on extension), built by retrieve_semantic_bm25_combined_sql (memories/pg/recall.py:28-100), which emits a single UNION query per fact_type with configurable per-fact-type partial HNSW indexes; (3) Graph — link expansion through memory_links using three signals: entity links (shared entities via unit_entities), semantic links (precomputed kNN graph), and causal links (caused_by), implemented in memories/pg/link_expansion.py:27-47; (4) Temporal — date-range filtering on occurred_start/occurred_end. Results from all arms are fused via Reciprocal Rank Fusion (fusion.py:29-109, RRF formula score(d) = Σ 1/(k+rank(d)) with k=60). After fusion, a cross-encoder reranker (reranking.py:1-38) rescales scores with neural reranking (local CrossEncoder or remote TEI/Cohere). The reranker applies recency boost (_RECENCY_ALPHA=0.2, linear or exponential decay), temporal proximity boost, and proof-count boost. Maximum Marginal Diversity (MMD) is then applied for diversification, and results are trimmed to the max_tokens budget. The reflect endpoint (reflect_async, memory_engine.py:15069-15148) is an agentic loop that sequentially calls tools (lookup mental models, recall facts, search observations, expand chunks) over multiple iterations.
How are memories updated, consolidated or forgotten?
answeredLifecycle management has three main paths. Consolidation (consolidation/consolidator.py:1425-1544) runs as a background job after retain, processing unconsolidated memories (memory_units where consolidated_at IS NULL). It groups new facts, fetches existing observations by embedding similarity, and calls the LLM with a consolidation prompt (prompts.py:1-80) that decides whether to CREATE a new observation, UPDATE an existing one (merging new evidence), or DELETE a stale one. Observations are stored in memory_units with fact_type='observation', tracking proof_count and source_memory_ids. The consolidator also runs LLM-based semantic dedup (_dedup_adjudicate, line 275) with a configurable consolidation_dedup_threshold that probes nearest neighbours and asks the LLM if the new observation is truly distinct (using _DEDUP_PROMPT, line 178-197). Dedup at write time via the entity resolver uses trigram similarity clustering (entity_resolver.py:79-80) to merge near-identical entity names within a batch. Mental model (knowledge page) refresh turns a synthetic document into a curated summary: it runs reflect on a configured source_query scope and writes the result as markdown content in the knowledge_pages table. Pages can refresh on a cron schedule or after each consolidation round. Explicit delete (delete_memory_unit, memory_engine.py:10936-10985) cascades to memory_links, unit_entities, and invalidates dependent observations. Direct invalidation (invalidate_memory endpoint) marks a memory unit as invalidated in invalidated_memory_units so consolidation can re-process its dependents. There is no TTL-based decay — the recency boost in reranking (reranking.py:35-53) is a query-time scoring adjustment (linear decay over 365 days, exponential with 90-day half-life, or disabled), not a deletion mechanism. Version history is tracked on mental models (mental_model_versions table) and on the consolidation history JSONB column per observation.
How is memory scoped and isolated?
answeredMemory isolation is primarily by bank — each bank is an isolated memory store ("brain" for one user/agent). All operations require a bank_id parameter. Under the hood, tables like memory_units, entities, and documents carry a bank_id column, and all queries filter on it (e.g., memories/pg/recall.py parameterises every SELECT with WHERE bank_id = $1). Multi-tenancy is implemented via the TenantExtension (extensions/tenant.py:54-100). Each tenant gets a PostgreSQL schema; all queries use fully-qualified table names (schema_name.memory_units) isolated at the database level. The default is ApiKeyTenantExtension (extensions/builtin/tenant.py) which authenticates via an API key and returns the tenant's schema. Tag-based scoping provides a second isolation layer within a bank (search/tags.py:1-23). Tags are TEXT[] arrays on memory_units, filtered via GIN-indexed PostgreSQL operators (&& for overlap, @> for containment). Five matching modes control strictness: any/any_strict (OR), all/all_strict (AND), and exact (set equality for observation scopes). Compound tag_groups support boolean trees with AND/OR/NOT groups at any nesting depth. Each bank operation can carry tags and tags_match, and the reflect agentic loop enforces the caller's tag scope on every internal tool call (memory_engine.py:15184-15193). Admission control (api/admission.py:1-100) bounds concurrency per operation class (recall, retain, reflect) with lane-specific max_in_flight and max_wait_seconds limits, returning 503 when a request cannot be served within its deadline. Memory Defense (retain/orchestrator.py imports) is a configurable policy that can block or redact sensitive content before storage, acting as a privacy/safety gate at write time.
How do agents integrate with it, and what is self-hostable?
answeredHindsight is self-hostable as a single process: uv run hindsight-api starts a FastAPI server with embedded PostgreSQL (pg0), local SentenceTransformers embeddings, and local CrossEncoder reranking — no external services required. The REST API (api/http.py) exposes REST endpoints for all operations (retain, recall, reflect, CRUD on memories/documents/entities/mental models/directives/banks). An auto-generated OpenAPI spec feeds 4 SDKs: Python (hindsight_client), TypeScript, Rust, and Go (hindsight-clients/). The Python SDK has a high-level Hindsight class with synchronous convenience methods (hindsight_client/hindsight_client.py:80-142). An MCP server (api/mcp.py:1-100) exposes 30+ tools (retain, recall, reflect, list_banks, etc.) over the Model Context Protocol HTTP transport, registered in mcp_tools.py:602-701. Framework integrations (hindsight-integrations/) span AG2, AutoGen, CrewAI, LangGraph, Pydantic AI, Claude Agent SDK, OpenAI Agents SDK (via ai-sdk), Aider, and agno. The Coding Agents integration (hindsight-integrations/coding-agents/) is a single npm package supporting 22 agent harnesses (Claude Code, Codex CLI, Cursor, Copilot, Grok Build, Dcode, opencode, Kilo, Devin CLI, Cline, Qwen Code, Kimi Code, etc.) with automatic session hooks that recall context before each prompt and retain transcripts afterward. A portable Agent Plugin (agent-plugin/plugin.json) exposes MCP tools for agents with plugin loaders. The CLI (hindsight-cli/) is a Rust binary with commands for memory, documents, entities, banks, mental models, audit logs, and filesystem mount of the knowledge base. A control plane (hindsight-control-plane/) is a Next.js admin UI with dashboards for banks, memories, documents, entities, knowledge base, audit logs, and operations. The embed package (hindsight-embed/) provides hindsight-local-mcp for daemonless usage via uvx. All server configuration is via environment variables with sensible defaults; hierarchical config overrides per-bank for LLM settings and operation tuning.