LLMs Technical Reviews

topoteretes/cognee

Python memory engine that LLM-extracts a knowledge graph from data into graph and vector stores, then answers queries over it.

GitHub ↗★ 31kPythonApache-2.0commit b32d8af · 2026-10-01homepage ↗

Overview

Cognee turns documents, chat turns and agent traces into a knowledge graph with embeddings, then answers questions from that graph. You write with remember(), read with recall(), run a maintenance pass with improve() and delete with forget(). Under these four calls are older, lower-level steps that still exist: add() (ingest raw data), cognify() (extract a graph), search() and memify().

Its unit of memory is different from a fact-store library. Cognee does not keep a list of short facts per user. It chunks every input, asks an LLM for typed nodes and edges per chunk, writes them to a graph database, and embeds chunks, summaries, entity names and edge labels in a vector database. Recall then pulls chunks, entities and relationship “facts” together and usually ends in one LLM completion, so a recall returns an answer, not just rows. There is also a fast session layer: with session_id, remember() writes to a cache, and improve() later moves that content into the graph.

The project is large (SDK, FastAPI server, MCP server, CLI, Next.js UI, connectors, a code-graph pipeline, evals). Everything in the repository runs locally. The hosted Cognee Cloud is just another server that the same client can point at.

Architecture

flowchart LR
  A["Agent / app"] --> SDK["remember / recall / improve / forget"]
  API["FastAPI /api/v1"] --> SDK
  MCP["cognee-mcp"] --> SDK
  MCP -.->|"--api-url"| API
  SDK --> P["Pipelines (add, cognify)"]
  P --> LLM["LLM structured output"]
  P --> G["Graph DB (Ladybug default)"]
  P --> V["Vector DB (LanceDB default)"]
  SDK --> R["Relational DB (users, datasets, ACLs)"]
  SDK --> C["Session cache"]
  C -->|"improve()"| G
Component Path Role
Public API cognee/api/v1/{remember,recall,improve,forget,update}/ The four memory verbs plus in-place update
Ingestion cognee/tasks/ingestion/ ingest_data saves inputs to file storage and creates Data rows
Cognify pipeline cognee/api/v1/cognify/cognify.py Task list: classify, chunk, extract graph + summarise, store
Graph extraction cognee/tasks/graph/, cognee/infrastructure/llm/extraction/ Per-chunk LLM call that returns a KnowledgeGraph
Data model cognee/infrastructure/engine/models/DataPoint.py Base class for every stored node
Storage adapters cognee/infrastructure/databases/{graph,vector,relational,cache}/ Pluggable backends
Retrieval cognee/modules/retrieval/, cognee/modules/search/ One retriever class per SearchType
Improve cognee/modules/improve/ Nine ordered maintenance stages
Users and permissions cognee/modules/users/ Tenants, users, roles, per-dataset ACLs
MCP server cognee-mcp/src/server.py MCP tools over the SDK or a remote API

How a request flows

Write (remember without session_id)

  1. remember() runs add() and then cognify() on the target dataset (default main_dataset). With self_improvement=True, it also starts improve() afterwards (remember.py).
  2. ingest_data turns each input into a text file in Cognee storage and records a Data row with its location, MIME type and content hash.
  3. get_default_tasks builds the cognify pipeline: classify_documents, extract_chunks_from_documents, extract_graph_and_summarize, add_data_points. Optional tasks add provenance, detect_contradictions and resolve_temporal_contradictions (cognify.py).
  4. extract_graph_and_summarize runs graph extraction and chunk summarisation in parallel (extract_graph_and_summarize.py). extract_graph_from_data calls extract_content_graph once per chunk with asyncio.gather (extract_graph_from_data.py). That call renders the graph prompt and asks LLMGateway.acreate_structured_output for a KnowledgeGraph of nodes (id, name, type, description) and edges (source, target, relationship name, description) (extract_content_graph.py, data_models.py).
  5. integrate_chunk_graphs removes duplicate node ids, optionally grounds entities in an OWL ontology, and skips edges that already exist in the graph (L87-L195).
  6. add_data_points walks the DataPoint tree into nodes and edges, deduplicates them, writes them to the graph and indexes each declared field in the vector store (add_data_points.py).

Read (recall)

  1. recall() resolves a scope. With only a session_id, it searches the session cache first and stops on a hit. Otherwise it searches the graph (recall.py).
  2. If no query_type is given, route_query applies two regex rules: a fully quoted query goes to CHUNKS_LEXICAL, coding-rules wording goes to CODING_RULES, and everything else goes to HYBRID_COMPLETION (query_router.py).
  3. HybridRetriever._retrieve_one embeds the query once, then runs a chunk lane (chunks and summaries) and an entity lane (entity vectors plus EdgeType_relationship_name edge vectors expanded through the graph) at the same time (hybrid_retriever.py).
  4. The context is formatted into a prompt and sent to the LLM. If every lane is empty, no LLM call is made.

Key components

DataPoint

Every stored node subclasses DataPoint. metadata.index_fields says which fields to embed (one vector collection per <Type>_<field>). identity_fields give a deterministic UUID5 id, so the same entity merges across ingestions. The base also carries version, valid_to for superseded facts, feedback_weight, importance_weight and provenance fields (DataPoint.py). A custom graph model is simply a set of DataPoint subclasses passed as graph_model.

Search types

SearchType has 20 members, from plain CHUNKS and SUMMARIES to GRAPH_COMPLETION, chain-of-thought and decomposition variants, TEMPORAL, CODE, AGENTIC_COMPLETION and raw CYPHER (SearchType.py). GRAPH_COMPLETION ranks triplets with brute_force_triplet_search, weighted by feedback and user preferences (graph_completion_retriever.py).

Improve

improve() runs nine stages in order: feedback_weights, persist_session_qa, persist_agent_traces, extract_agent_context, distill_sessions, update_user_preferences, build_truth_subspace, triplet_enrichment, global_context_index (registry.py). Each stage first checks whether it can run, without LLM calls. The first seven need session_ids, and a lock stops two runs from touching the same sessions or dataset.

Update and forget

update() keeps the document’s data_id. By default it diffs the new text against the old, re-extracts only the changed chunks, and falls back to a full rebuild when that is not possible (update.py). forget() deletes one item, a dataset or everything the user owns. memory_only=True keeps the raw files (forget.py).

Access control

Multi-tenant mode is the default when the configured graph and vector backends support it. Each user and dataset then gets its own graph and vector databases, and authentication is forced on (context_global_variables.py, get_authenticated_user.py). Read, write, delete and share rights are ACL rows per principal and dataset.

Extending it

  • Graph schema: pass your own graph_model (pydantic or DataPoint subclasses) or a custom_prompt to remember/cognify.
  • Pipelines: a pipeline is a list of Task objects. You can build one from your own async functions and add_data_points.
  • Backends: graph (Ladybug/Kuzu, Neo4j, Neptune, Neptune Analytics, Turso, Postgres), vector (LanceDB, pgvector, Neptune Analytics, Turso), relational (SQLite, Postgres) and session cache (SQLite by default; Redis, Postgres, filesystem and others) are chosen by env vars.
  • Retrievers: register a community retriever for a SearchType. An OWL ontology can ground or filter entities.
  • MCP: cognee-mcp exposes remember, recall, forget, improve and cognify_status. With agent scoping on (the default), each MCP client writes to its own <client>_memory dataset (server.py).

Running it

  • Library: pip install cognee, set LLM_API_KEY, then await cognee.remember(...) and await cognee.recall(...). Defaults are OpenAI for the LLM and embeddings, and embedded SQLite, LanceDB and Ladybug on disk, so no database server is needed.
  • Server: the root docker-compose.yml runs the API. MCP, UI, Neo4j, Postgres/pgvector and Redis are optional Compose profiles.
  • Telemetry: usage events are sent unless TELEMETRY_DISABLED is set.

Strengths and caveats

  • Strength: relationships are stored, so multi-hop and “how are these connected” questions can follow real edges instead of nearest-neighbour text.
  • Strength: permissions are a real part of the model: tenants, users, per-dataset ACLs and, by default, separate databases per user and dataset.
  • Strength: edits are not destructive. In-place update() with chunk diffs, valid_to supersession and opt-in contradiction detection go further than an append-only store.
  • Caveat: ingestion is expensive. Every chunk costs an extraction call and a summary call, and the default LLM is a hosted OpenAI model.
  • Caveat: recall usually returns an LLM-written answer. For raw context, pin a query_type such as CHUNKS or use only_context.
  • Caveat: CYPHER and NATURAL_LANGUAGE search run queries on the graph and only need read permission. They are allowed unless ALLOW_CYPHER_QUERY=false (get_search_type_retriever_instance.py). The query router never picks them by itself.
  • Caveat: the surface is very large and changes fast. Old and new verbs (add/cognify/search/memify and remember/recall/improve) exist side by side.

Sources: code at b32d8af, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the Agent memory layers questions

Each answer was drafted by a code-reading agent at commit b32d8af. Its citations were checked mechanically. Compare with the other agent memory layers →

How are memories extracted from interactions?

answered

Memories are extracted through the cognify pipeline, entered via remember() -> add() -> cognify() (cognee/api/v1/remember/remember.py:867). The task extract_graph_from_data (cognee/tasks/graph/extract_graph_from_data.py:198) sends each chunk via asyncio.gather to extract_content_graph(), which uses Instructor or litellm-native structured output to LLMs, producing a KnowledgeGraph with typed Node and Edge objects (cognee/shared/data_models.py:49-77). What gets stored includes entities, relationships, document summaries, chunk text with embeddings, QA turns, agent trace steps, feedback scores, and skill-run records via typed MemoryEntry types: QAEntry, TraceEntry, FeedbackEntry, SkillRunEntry (cognee/memory/entries.py:24-129). Deduplication at write time collapses duplicate node IDs within a chunk (_remove_duplicate_extracted_nodes_by_id at extract_graph_from_data.py:42); graph adapter writes are idempotent upserts keyed on node id (GraphDBInterface contract, graph_db_interface.py:44-58). An optional ontology mode can drop entities not grounded in an OWL file. The improve loop and contradiction detection provide post-extraction consolidation.

How are memories stored?

answered

Memories are stored in three parallel pluggable backends. Graph DB through GraphDBInterface (graph_db_interface.py:36): nodes are DataPoints with deterministic UUID5 ids from identity_fields; edges are (source_id, target_id, relationship_name) triples, all upsert-keyed on id. Default is Ladybug (embedded Kuzu) with Neo4j, Neptune, Turso, and Postgres alternatives. Vector DB through VectorDBInterface (vector_db_interface.py:11): one collection per DataPoint type and embedded field pair (e.g. DocumentChunk_text), with idempotent upserts and embedding via the configured engine. Default is LanceDB with PGVector, Neptune Analytics, and Turso alternatives. Relational DB (SQLite default, PostgreSQL) holds users, datasets, permissions, ACLs, and pipeline run status. The DataPoint base (DataPoint.py:28) adds version, created_at, updated_at, bi-temporal valid_to, feedback_weight, importance_weight, ontology_uri, and provenance stamps. The session cache (CACHE_BACKEND=sqlite/postgres/redis) stores QA turns, agent traces, and distilled context as a fast layer.

How are memories retrieved and injected into the prompt?

answered

Retrieval is handled by recall() (cognee/api/v1/recall/recall.py:347) which wraps authorized_search() with 20+ SearchType strategies (cognee/modules/search/types/SearchType.py:4) including HYBRID_COMPLETION (default), GRAPH_COMPLETION, RAG_COMPLETION, CHUNKS, CHUNKS_LEXICAL, TRIPLET_COMPLETION, and TEMPORAL. The free rule-based router (cognee/api/v1/recall/query_router.py:47) picks the strategy when query_type is omitted: quoted phrases to CHUNKS_LEXICAL, coding keywords to CODING_RULES, default to HYBRID. Results reach LLM context through build_completion_prompts (cognee/modules/retrieval/utils/completion.py:21): system prompt = task template only; user prompt = history + question/context + guidance block. Session entries can short-circuit the graph search (_search_session at recall.py:167). Ranking uses brute-force triplet search over graph edges with personalization weights. Empty graphs do not reach the LLM (skip_completion_on_empty_context flag, base_retriever.py:40).

Editor's note. Clarification: brute-force triplet search is the ranking for GRAPH_COMPLETION. The default HYBRID_COMPLETION runs a chunk/summary vector lane and an entity lane (entity and EdgeType_relationship_name vectors, expanded through the graph) in parallel, then makes one completion (hybrid_retriever.py _retrieve_one).

How are memories updated, consolidated or forgotten?

answered

Updates use chunk-level incremental diffing (cognee/api/v1/update/update.py:28): new content is diffed against stored text via a paragraph-anchored multi-region diff, only changed chunks are re-chunked and re-extracted, unchanged chunks keep their ids, entities and embeddings. Falls back to full delete+re-add+cognify when preconditions fail. The document keeps its data_id across every path. Consolidation runs through improve() (cognee/api/v1/improve/improve.py:101) with nine ordered stages: feedback_weights, persist_session_qa, persist_agent_traces, extract_agent_context, distill_sessions, update_user_preferences, build_truth_subspace, triplet_enrichment, global_context_index (cognee/modules/improve/registry.py:27-37). Each stage gates first with zero LLM calls; improvements serialize via the improve lock. Deletion via forget() (cognee/api/v1/forget/forget.py:22) supports data_id, dataset, everything=True, and memory_only=True, deleting across graph, vector, and session backends with pipeline status reset for re-cognify. DataPoint.valid_to provides bi-temporal fact supersession.

How is memory scoped and isolated?

answered

Memory is scoped via a multi-tenant hierarchy of Tenant -> User -> Dataset -> Data with Read/Write/Delete/Share permissions enforced through ACLs (cognee/modules/users/models/ACL.py:8). The User model (cognee/modules/users/models/User.py:15) has roles, tenants, and parent_user_id for agent accounts. With ENABLE_BACKEND_ACCESS_CONTROL=true (default), each user+dataset pair gets isolated graph and vector databases through a dataset-database registry. Session records tie sessions to (user_id, dataset_id). recall() supports scope=['session', 'trace', 'session_context', 'graph', 'tools', 'code'] with session_first short-circuit to avoid graph search. The MCP server adds per-client dataset scoping (cursor_vscode_memory, etc.) from clientInfo.name. The teacher-student hierarchy via parent_user_id lets agent accounts inherit permissions. Querying unpermitted datasets raises PermissionDeniedError; name-based lookups only resolve within the caller's own datasets.

Editor's note. Addition: multi-tenant mode (ENABLE_BACKEND_ACCESS_CONTROL) is on by default only when the configured graph and vector backends support per-dataset databases (multi_user_support_possible). In that mode authentication is forced on.

How do agents integrate with it, and what is self-hostable?

answered

Cognee offers four integration surfaces. The Python SDK is the primary entry (exported from cognee/init.py with await cognee.remember()/recall()/improve()/forget()). The REST API is a FastAPI server at /api/v1/ (cognee/api/client.py:362) with routers for remember, recall, improve, forget, add, cognify, search, datasets, users, permissions, skills, sessions, and integrations (GitHub, Gmail, Google Drive, Linear, Slack). The MCP server (cognee-mcp/src/server.py:355) provides remember/recall/forget/cognify_status/improve tools, connects in-process or via --api-url (REST) or --serve-url (Cognee Cloud), and supports per-client agent-scoped datasets. The CLI (cognee-cli) exposes all operations. Everything is self-hostable: minimal setup needs no external services (local defaults: SQLite, LanceDB, LadybugDB) and no API keys when using the GLiNER demo extractor with fastembed. Docker Compose deploys the full stack. Cognee Cloud is the hosted multi-tenant SaaS option; the same SDK works locally or via cognee.serve(url=...).