topoteretes/cognee
Python memory engine that LLM-extracts a knowledge graph from data into graph and vector stores, then answers queries over it.
Overview
Cognee turns documents, chat turns and agent traces into a knowledge graph with embeddings, then answers questions from that graph. You write with remember(), read with recall(), run a maintenance pass with improve() and delete with forget(). Under these four calls are older, lower-level steps that still exist: add() (ingest raw data), cognify() (extract a graph), search() and memify().
Its unit of memory is different from a fact-store library. Cognee does not keep a list of short facts per user. It chunks every input, asks an LLM for typed nodes and edges per chunk, writes them to a graph database, and embeds chunks, summaries, entity names and edge labels in a vector database. Recall then pulls chunks, entities and relationship “facts” together and usually ends in one LLM completion, so a recall returns an answer, not just rows. There is also a fast session layer: with session_id, remember() writes to a cache, and improve() later moves that content into the graph.
The project is large (SDK, FastAPI server, MCP server, CLI, Next.js UI, connectors, a code-graph pipeline, evals). Everything in the repository runs locally. The hosted Cognee Cloud is just another server that the same client can point at.
Architecture
flowchart LR
A["Agent / app"] --> SDK["remember / recall / improve / forget"]
API["FastAPI /api/v1"] --> SDK
MCP["cognee-mcp"] --> SDK
MCP -.->|"--api-url"| API
SDK --> P["Pipelines (add, cognify)"]
P --> LLM["LLM structured output"]
P --> G["Graph DB (Ladybug default)"]
P --> V["Vector DB (LanceDB default)"]
SDK --> R["Relational DB (users, datasets, ACLs)"]
SDK --> C["Session cache"]
C -->|"improve()"| G
| Component | Path | Role |
|---|---|---|
| Public API | cognee/api/v1/{remember,recall,improve,forget,update}/ |
The four memory verbs plus in-place update |
| Ingestion | cognee/tasks/ingestion/ |
ingest_data saves inputs to file storage and creates Data rows |
| Cognify pipeline | cognee/api/v1/cognify/cognify.py |
Task list: classify, chunk, extract graph + summarise, store |
| Graph extraction | cognee/tasks/graph/, cognee/infrastructure/llm/extraction/ |
Per-chunk LLM call that returns a KnowledgeGraph |
| Data model | cognee/infrastructure/engine/models/DataPoint.py |
Base class for every stored node |
| Storage adapters | cognee/infrastructure/databases/{graph,vector,relational,cache}/ |
Pluggable backends |
| Retrieval | cognee/modules/retrieval/, cognee/modules/search/ |
One retriever class per SearchType |
| Improve | cognee/modules/improve/ |
Nine ordered maintenance stages |
| Users and permissions | cognee/modules/users/ |
Tenants, users, roles, per-dataset ACLs |
| MCP server | cognee-mcp/src/server.py |
MCP tools over the SDK or a remote API |
How a request flows
Write (remember without session_id)
remember()runsadd()and thencognify()on the target dataset (defaultmain_dataset). Withself_improvement=True, it also startsimprove()afterwards (remember.py).ingest_dataturns each input into a text file in Cognee storage and records aDatarow with its location, MIME type and content hash.get_default_tasksbuilds the cognify pipeline:classify_documents,extract_chunks_from_documents,extract_graph_and_summarize,add_data_points. Optional tasks add provenance,detect_contradictionsandresolve_temporal_contradictions(cognify.py).extract_graph_and_summarizeruns graph extraction and chunk summarisation in parallel (extract_graph_and_summarize.py).extract_graph_from_datacallsextract_content_graphonce per chunk withasyncio.gather(extract_graph_from_data.py). That call renders the graph prompt and asksLLMGateway.acreate_structured_outputfor aKnowledgeGraphof nodes (id, name, type, description) and edges (source, target, relationship name, description) (extract_content_graph.py, data_models.py).integrate_chunk_graphsremoves duplicate node ids, optionally grounds entities in an OWL ontology, and skips edges that already exist in the graph (L87-L195).add_data_pointswalks theDataPointtree into nodes and edges, deduplicates them, writes them to the graph and indexes each declared field in the vector store (add_data_points.py).
Read (recall)
recall()resolves a scope. With only asession_id, it searches the session cache first and stops on a hit. Otherwise it searches the graph (recall.py).- If no
query_typeis given,route_queryapplies two regex rules: a fully quoted query goes toCHUNKS_LEXICAL, coding-rules wording goes toCODING_RULES, and everything else goes toHYBRID_COMPLETION(query_router.py). HybridRetriever._retrieve_oneembeds the query once, then runs a chunk lane (chunks and summaries) and an entity lane (entity vectors plusEdgeType_relationship_nameedge vectors expanded through the graph) at the same time (hybrid_retriever.py).- The context is formatted into a prompt and sent to the LLM. If every lane is empty, no LLM call is made.
Key components
DataPoint
Every stored node subclasses DataPoint. metadata.index_fields says which fields to embed (one vector collection per <Type>_<field>). identity_fields give a deterministic UUID5 id, so the same entity merges across ingestions. The base also carries version, valid_to for superseded facts, feedback_weight, importance_weight and provenance fields (DataPoint.py). A custom graph model is simply a set of DataPoint subclasses passed as graph_model.
Search types
SearchType has 20 members, from plain CHUNKS and SUMMARIES to GRAPH_COMPLETION, chain-of-thought and decomposition variants, TEMPORAL, CODE, AGENTIC_COMPLETION and raw CYPHER (SearchType.py). GRAPH_COMPLETION ranks triplets with brute_force_triplet_search, weighted by feedback and user preferences (graph_completion_retriever.py).
Improve
improve() runs nine stages in order: feedback_weights, persist_session_qa, persist_agent_traces, extract_agent_context, distill_sessions, update_user_preferences, build_truth_subspace, triplet_enrichment, global_context_index (registry.py). Each stage first checks whether it can run, without LLM calls. The first seven need session_ids, and a lock stops two runs from touching the same sessions or dataset.
Update and forget
update() keeps the document’s data_id. By default it diffs the new text against the old, re-extracts only the changed chunks, and falls back to a full rebuild when that is not possible (update.py). forget() deletes one item, a dataset or everything the user owns. memory_only=True keeps the raw files (forget.py).
Access control
Multi-tenant mode is the default when the configured graph and vector backends support it. Each user and dataset then gets its own graph and vector databases, and authentication is forced on (context_global_variables.py, get_authenticated_user.py). Read, write, delete and share rights are ACL rows per principal and dataset.
Extending it
- Graph schema: pass your own
graph_model(pydantic orDataPointsubclasses) or acustom_prompttoremember/cognify. - Pipelines: a pipeline is a list of
Taskobjects. You can build one from your own async functions andadd_data_points. - Backends: graph (Ladybug/Kuzu, Neo4j, Neptune, Neptune Analytics, Turso, Postgres), vector (LanceDB, pgvector, Neptune Analytics, Turso), relational (SQLite, Postgres) and session cache (SQLite by default; Redis, Postgres, filesystem and others) are chosen by env vars.
- Retrievers: register a community retriever for a
SearchType. An OWL ontology can ground or filter entities. - MCP:
cognee-mcpexposesremember,recall,forget,improveandcognify_status. With agent scoping on (the default), each MCP client writes to its own<client>_memorydataset (server.py).
Running it
- Library:
pip install cognee, setLLM_API_KEY, thenawait cognee.remember(...)andawait cognee.recall(...). Defaults are OpenAI for the LLM and embeddings, and embedded SQLite, LanceDB and Ladybug on disk, so no database server is needed. - Server: the root
docker-compose.ymlruns the API. MCP, UI, Neo4j, Postgres/pgvector and Redis are optional Compose profiles. - Telemetry: usage events are sent unless
TELEMETRY_DISABLEDis set.
Strengths and caveats
- Strength: relationships are stored, so multi-hop and “how are these connected” questions can follow real edges instead of nearest-neighbour text.
- Strength: permissions are a real part of the model: tenants, users, per-dataset ACLs and, by default, separate databases per user and dataset.
- Strength: edits are not destructive. In-place
update()with chunk diffs,valid_tosupersession and opt-in contradiction detection go further than an append-only store. - Caveat: ingestion is expensive. Every chunk costs an extraction call and a summary call, and the default LLM is a hosted OpenAI model.
- Caveat: recall usually returns an LLM-written answer. For raw context, pin a
query_typesuch asCHUNKSor useonly_context. - Caveat:
CYPHERandNATURAL_LANGUAGEsearch run queries on the graph and only need read permission. They are allowed unlessALLOW_CYPHER_QUERY=false(get_search_type_retriever_instance.py). The query router never picks them by itself. - Caveat: the surface is very large and changes fast. Old and new verbs (
add/cognify/search/memifyandremember/recall/improve) exist side by side.
Sources: code at b32d8af, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit b32d8af. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemories are extracted through the cognify pipeline, entered via remember() -> add() -> cognify() (cognee/api/v1/remember/remember.py:867). The task extract_graph_from_data (cognee/tasks/graph/extract_graph_from_data.py:198) sends each chunk via asyncio.gather to extract_content_graph(), which uses Instructor or litellm-native structured output to LLMs, producing a KnowledgeGraph with typed Node and Edge objects (cognee/shared/data_models.py:49-77). What gets stored includes entities, relationships, document summaries, chunk text with embeddings, QA turns, agent trace steps, feedback scores, and skill-run records via typed MemoryEntry types: QAEntry, TraceEntry, FeedbackEntry, SkillRunEntry (cognee/memory/entries.py:24-129). Deduplication at write time collapses duplicate node IDs within a chunk (_remove_duplicate_extracted_nodes_by_id at extract_graph_from_data.py:42); graph adapter writes are idempotent upserts keyed on node id (GraphDBInterface contract, graph_db_interface.py:44-58). An optional ontology mode can drop entities not grounded in an OWL file. The improve loop and contradiction detection provide post-extraction consolidation.
How are memories stored?
answeredMemories are stored in three parallel pluggable backends. Graph DB through GraphDBInterface (graph_db_interface.py:36): nodes are DataPoints with deterministic UUID5 ids from identity_fields; edges are (source_id, target_id, relationship_name) triples, all upsert-keyed on id. Default is Ladybug (embedded Kuzu) with Neo4j, Neptune, Turso, and Postgres alternatives. Vector DB through VectorDBInterface (vector_db_interface.py:11): one collection per DataPoint type and embedded field pair (e.g. DocumentChunk_text), with idempotent upserts and embedding via the configured engine. Default is LanceDB with PGVector, Neptune Analytics, and Turso alternatives. Relational DB (SQLite default, PostgreSQL) holds users, datasets, permissions, ACLs, and pipeline run status. The DataPoint base (DataPoint.py:28) adds version, created_at, updated_at, bi-temporal valid_to, feedback_weight, importance_weight, ontology_uri, and provenance stamps. The session cache (CACHE_BACKEND=sqlite/postgres/redis) stores QA turns, agent traces, and distilled context as a fast layer.
How are memories retrieved and injected into the prompt?
answeredRetrieval is handled by recall() (cognee/api/v1/recall/recall.py:347) which wraps authorized_search() with 20+ SearchType strategies (cognee/modules/search/types/SearchType.py:4) including HYBRID_COMPLETION (default), GRAPH_COMPLETION, RAG_COMPLETION, CHUNKS, CHUNKS_LEXICAL, TRIPLET_COMPLETION, and TEMPORAL. The free rule-based router (cognee/api/v1/recall/query_router.py:47) picks the strategy when query_type is omitted: quoted phrases to CHUNKS_LEXICAL, coding keywords to CODING_RULES, default to HYBRID. Results reach LLM context through build_completion_prompts (cognee/modules/retrieval/utils/completion.py:21): system prompt = task template only; user prompt = history + question/context + guidance block. Session entries can short-circuit the graph search (_search_session at recall.py:167). Ranking uses brute-force triplet search over graph edges with personalization weights. Empty graphs do not reach the LLM (skip_completion_on_empty_context flag, base_retriever.py:40).
GRAPH_COMPLETION. The default HYBRID_COMPLETION runs a chunk/summary vector lane and an entity lane (entity and EdgeType_relationship_name vectors, expanded through the graph) in parallel, then makes one completion (hybrid_retriever.py _retrieve_one).How are memories updated, consolidated or forgotten?
answeredUpdates use chunk-level incremental diffing (cognee/api/v1/update/update.py:28): new content is diffed against stored text via a paragraph-anchored multi-region diff, only changed chunks are re-chunked and re-extracted, unchanged chunks keep their ids, entities and embeddings. Falls back to full delete+re-add+cognify when preconditions fail. The document keeps its data_id across every path. Consolidation runs through improve() (cognee/api/v1/improve/improve.py:101) with nine ordered stages: feedback_weights, persist_session_qa, persist_agent_traces, extract_agent_context, distill_sessions, update_user_preferences, build_truth_subspace, triplet_enrichment, global_context_index (cognee/modules/improve/registry.py:27-37). Each stage gates first with zero LLM calls; improvements serialize via the improve lock. Deletion via forget() (cognee/api/v1/forget/forget.py:22) supports data_id, dataset, everything=True, and memory_only=True, deleting across graph, vector, and session backends with pipeline status reset for re-cognify. DataPoint.valid_to provides bi-temporal fact supersession.
How is memory scoped and isolated?
answeredMemory is scoped via a multi-tenant hierarchy of Tenant -> User -> Dataset -> Data with Read/Write/Delete/Share permissions enforced through ACLs (cognee/modules/users/models/ACL.py:8). The User model (cognee/modules/users/models/User.py:15) has roles, tenants, and parent_user_id for agent accounts. With ENABLE_BACKEND_ACCESS_CONTROL=true (default), each user+dataset pair gets isolated graph and vector databases through a dataset-database registry. Session records tie sessions to (user_id, dataset_id). recall() supports scope=['session', 'trace', 'session_context', 'graph', 'tools', 'code'] with session_first short-circuit to avoid graph search. The MCP server adds per-client dataset scoping (cursor_vscode_memory, etc.) from clientInfo.name. The teacher-student hierarchy via parent_user_id lets agent accounts inherit permissions. Querying unpermitted datasets raises PermissionDeniedError; name-based lookups only resolve within the caller's own datasets.
ENABLE_BACKEND_ACCESS_CONTROL) is on by default only when the configured graph and vector backends support per-dataset databases (multi_user_support_possible). In that mode authentication is forced on.How do agents integrate with it, and what is self-hostable?
answeredCognee offers four integration surfaces. The Python SDK is the primary entry (exported from cognee/init.py with await cognee.remember()/recall()/improve()/forget()). The REST API is a FastAPI server at /api/v1/ (cognee/api/client.py:362) with routers for remember, recall, improve, forget, add, cognify, search, datasets, users, permissions, skills, sessions, and integrations (GitHub, Gmail, Google Drive, Linear, Slack). The MCP server (cognee-mcp/src/server.py:355) provides remember/recall/forget/cognify_status/improve tools, connects in-process or via --api-url (REST) or --serve-url (Cognee Cloud), and supports per-client agent-scoped datasets. The CLI (cognee-cli) exposes all operations. Everything is self-hostable: minimal setup needs no external services (local defaults: SQLite, LanceDB, LadybugDB) and no API keys when using the GLiNER demo extractor with fastembed. Docker Compose deploys the full stack. Cognee Cloud is the hosted multi-tenant SaaS option; the same SDK works locally or via cognee.serve(url=...).