getzep/graphiti
Temporal knowledge-graph memory for agents; LLM-extracted facts with validity dates, stored in Neo4j/FalkorDB, hybrid-searched.
Overview
Graphiti is Zep’s open-source Python library for building a temporal knowledge graph as agent memory. You feed it “episodes” (a chat message, a block of text or a JSON document, each with a reference time). For each episode it uses an LLM to pull out entities and relationships, merges them into an existing property graph, and keeps track of when each fact was true. Retrieval is a hybrid search over that graph: BM25, vector similarity and graph traversal, fused and reranked.
The unit of memory is a fact edge, not a document chunk. An EntityEdge connects two EntityNodes and carries a natural-language fact, an embedding of that fact, the episodes it came from, and bi-temporal fields (valid_at, invalid_at, expired_at). When a new episode contradicts an old fact, the old edge is not deleted. It is closed off with an invalid_at date. This makes Graphiti a good fit for agents whose world changes (preferences, roles, account state) and where “what was true in March?” is a real question.
The repository holds three things: the core library graphiti_core, a small FastAPI service in server/, and an MCP server in mcp_server/. All of it is self-hostable. You need a graph database (Neo4j by default; FalkorDB, Kuzu and Amazon Neptune are alternatives) and an LLM plus an embedding provider. The defaults are OpenAI for all three model roles.
Architecture
flowchart LR
APP["Agent / app"] --> SDK["Graphiti class"]
MCP["MCP server"] --> SDK
REST["FastAPI server"] --> SDK
SDK --> EXT["extract_nodes / extract_edges"]
EXT --> LLM["LLMClient"]
SDK --> RES["resolve nodes / edges"]
RES --> SRCH["search()"]
RES --> LLM
SDK --> SRCH
SRCH --> EMB["EmbedderClient"]
SRCH --> CE["CrossEncoderClient"]
SRCH --> DRV["GraphDriver"]
RES --> DRV
DRV --> DB["Neo4j / FalkorDB / Kuzu / Neptune"]
| Component | Path | Role |
|---|---|---|
| Facade | graphiti_core/graphiti.py |
Graphiti class: add_episode, add_episode_bulk, search, search_, add_triplet, build_communities, summarize_saga, remove_episode |
| Data model | graphiti_core/nodes.py, graphiti_core/edges.py |
EpisodicNode, EntityNode, CommunityNode, SagaNode; EntityEdge, EpisodicEdge (MENTIONS), CommunityEdge, HasEpisodeEdge, NextEpisodeEdge |
| Extraction and dedup | graphiti_core/utils/maintenance/ |
node_operations.py, edge_operations.py, dedup_helpers.py, community_operations.py |
| Prompts | graphiti_core/prompts/ |
Extraction, dedupe and summary prompts, replaceable through a prompt library |
| Search | graphiti_core/search/ |
search(), method and reranker enums, ready-made recipes, SearchFilters |
| Drivers | graphiti_core/driver/ |
GraphDriver ABC plus Neo4j, FalkorDB, Kuzu and Neptune drivers |
| Model clients | llm_client/, embedder/, cross_encoder/ |
OpenAI, Azure, Anthropic, Gemini, Groq, GLiNER2; OpenAI/Voyage/Gemini/Azure embedders; OpenAI, Gemini and BGE rerankers |
| REST service | server/graph_service/ |
FastAPI app with /messages, /search, /get-memory, /episodes/{group_id} |
| MCP server | mcp_server/src/graphiti_mcp_server.py |
Tools such as add_memory, search_memory_facts, search_nodes, delete_episode; stdio, SSE or HTTP transport |
How a request flows
Take one call to await graphiti.add_episode(name, body, source_description, reference_time, group_id="user-42"):
- Scope.
_resolve_request_scopevalidates thegroup_id. If it differs from the driver’s default database, it clones the driver for that group and copies the clients bundle, so concurrent calls for different groups cannot overwrite each other’s target (graphiti.py). Only FalkorDB actually maps a group to its own graph. The baseGraphDriver.clonereturnsself, so on Neo4j thegroup_idis just a property filter (driver.py, falkordb_driver.py). - Context. It loads the last 10 (
RELEVANT_SCHEMA_LIMIT) episodes of the group (or the explicitprevious_episode_uuids) and builds a newEpisodicNodewithvalid_at=reference_time(graphiti.py). - Extract entities.
extract_nodessends the episode, the previous episodes and a catalogue of entity types to the LLM, drops empty names, buildsEntityNodes and collapses exact duplicates within the same episode (node_operations.py). - Resolve entities.
resolve_extracted_nodesembeds each name and runs a cosine search for up to 15 candidates with score >= 0.6. It then tries a deterministic match (exact normalized name, then MinHash/LSH fuzzy match with Jaccard >= 0.9 for high-entropy names). Only the leftovers go to one batched LLM dedupe call (node_operations.py, dedup_helpers.py). - Extract and resolve facts.
_extract_and_resolve_edgescallsextract_edges, rewrites edge endpoints to the resolved node UUIDs, then callsresolve_extracted_edges(graphiti.py). For every new edge, that function runs two hybrid searches: one limited to edges between the same two nodes (duplicate candidates) and one across the whole group (contradiction candidates). It then callsresolve_extracted_edgeper edge (edge_operations.py). - Invalidate.
resolve_edge_contradictionscloses an older, overlapping edge by setting itsinvalid_atto the new edge’svalid_atand stampingexpired_at(edge_operations.py). - Summarize.
extract_attributes_from_nodesfills typed attributes per entity, then batch-updates entity summaries from the new facts only, then re-embeds the names (node_operations.py). - Persist.
_process_episode_databuilds MENTIONS edges, writes episodes, nodes and edges in bulk, and links sagas withHAS_EPISODEandNEXT_EPISODEedges (graphiti.py). Withupdate_communities=Trueit also updates community nodes.
A single episode therefore costs several LLM calls (entity extraction, edge extraction, dedupe, per-edge resolution and timestamps, attributes, summaries) plus a handful of searches. The docstring says it plainly: run add_episode in a background queue, one episode at a time per group.
Key components
Bi-temporal fact edges
EntityEdge holds fact, fact_embedding, episodes, valid_at, invalid_at, expired_at, reference_time and free-form attributes (edges.py). valid_at/invalid_at describe the world. expired_at records when Graphiti learned the fact had stopped being true. Search filters can use all four dates, so queries “as of” a point in time are possible without a separate history table.
Entity resolution
Deduplication is layered from cheap to expensive: vector candidates, then exact and fuzzy name matching, then the LLM. The thresholds are module constants (NODE_DEDUP_CANDIDATE_LIMIT = 15, NODE_DEDUP_COSINE_MIN_SCORE = 0.6, node_operations.py). Ambiguous exact-name hits (two existing nodes with the same name) go to the LLM instead of being guessed.
Hybrid search
search() embeds the query only when a config needs vectors, then runs edge, node, episode and community searches in parallel (search.py). Edges and nodes support BM25, cosine similarity and BFS. Communities support BM25 and cosine. Episodes support BM25 only. Rerankers are RRF, MMR, cross-encoder, node distance and episode mentions (search_config.py). Graphiti.search is the simple path: edges only, with RRF, or node-distance reranking when you pass center_node_uuid. search_ takes a full SearchConfig and defaults to COMBINED_HYBRID_SEARCH_CROSS_ENCODER (graphiti.py).
Communities and sagas
build_communities clusters entities with a plain label-propagation loop and asks the LLM to summarize each cluster (community_operations.py). Sagas are ordered episode threads. summarize_saga keeps a running summary with watermarks, so later runs only read new or backfilled episodes. SagaNode stores both watermarks (nodes.py).
Service wrappers
The MCP server is the more complete wrapper. It queues add_memory calls per group_id, caches group-scoped driver clones, and swaps in a group-scoped Graphiti copy for methods that still read self.driver (graphiti_mcp_server.py). The FastAPI server is thinner. /messages pushes jobs onto one in-process asyncio.Queue, and /search and /get-memory return edges as fact results (ingest.py, retrieve.py).
Extending it
- Custom ontology. Pass Pydantic models as
entity_typesandedge_types, plus anedge_type_mapof(source_label, target_label) -> [edge types]. Model fields become typedattributesthat the LLM fills.excluded_entity_typesdrops classes you don’t want stored. - Prompt control.
custom_extraction_instructionsadds text to the extraction prompts for each call.prompt_libraryreplaces prompt functions for the whole instance. - Model routing. Implement
LLMClient,EmbedderClientorCrossEncoderClient, or pass anLLMRuntimeto send different prompts to different models (graphiti.py). - Storage. Implement
GraphDriverand its operation interfaces underdriver/operations/. - Direct writes.
add_tripletinserts a known fact without extraction.graphiti.nodesandgraphiti.edgesgive CRUD access by namespace. - Search recipes. Build your own
SearchConfigor start from one insearch_config_recipes.py.
Running it
- Library.
pip install graphiti-core, with extras such asfalkordb,anthropic,google-genai,groq,voyageai,neptune,gliner2andtracing. Callbuild_indices_and_constraints()once, thenadd_episodeandsearch. The default clients readOPENAI_API_KEY. The default OpenAI models aregpt-5.5andgpt-4.1-nanofor small prompts, andEMBEDDING_DIMdefaults to 1024. - Databases. The root
docker-compose.ymlruns the FastAPI service with Neo4j 5.26. FalkorDB also has an embeddedfalkordbliteextra for Python 3.12+. Thekuzuextra is marked deprecated inpyproject.tomlbecause upstream Kuzu is unmaintained. - MCP.
mcp_server/ships YAML configs and Docker Compose files for Neo4j and FalkorDB, and runs over stdio, SSE or streamable HTTP. - Auth. Neither server adds authentication. Put them behind your own gateway.
Strengths and caveats
- Strength: time is a first-class field. Contradictions close old facts instead of overwriting them, and every fact keeps its source episodes. Few memory layers in this category model “true then, false now” this explicitly.
- Strength: dedup is cheap first, LLM last. Exact and MinHash matching settle many entities before any dedupe prompt runs.
- Strength: pluggable at every layer. Four graph backends, several LLM, embedder and reranker clients, typed ontologies and replaceable prompts.
- Caveat: write cost and latency. Ingest is a chain of LLM calls and searches for each episode, and the authors ask for sequential, awaited ingestion per group. That suits background memory building, not a synchronous request path.
- Caveat: shared recipe mutation.
Graphiti.searchsetslimiton the module-levelEDGE_HYBRID_SEARCH_RRFobject. The same object is used for edge dedup inside ingestion, so a call with a smallnum_resultsalso shrinks later dedup candidate lists in the same process. - Caveat:
remove_episodeis partial. It deletes only edges whose first episode is the removed one, and nodes mentioned only there, using the defaultself.driver(graphiti.py). It does not restore edges that the episode had invalidated, and on FalkorDB it needs the group-scoped client that the MCP server builds for you. - Caveat: fragile REST worker. The FastAPI
AsyncWorkercatches onlyCancelledError. One failingadd_episodejob ends the worker task, and queued messages stop being processed until restart. - Caveat: no built-in decay or ACL. Forgetting is explicit (invalidation or deletion), and isolation is the
group_idyou pass in.
Sources: code at 689de29, verified Q&A.
How it answers the Agent memory layers questions
Each answer was drafted by a code-reading agent at commit 689de29. Its citations were checked mechanically. Compare with the other agent memory layers →
How are memories extracted from interactions?
answeredMemories are extracted from interactions through an LLM-driven pipeline in graphiti_core/utils/maintenance/node_operations.py and edge_operations.py. Each ingested interaction is wrapped as an EpisodicNode and passed through extract_nodes() (line 70) which prompts the LLM with the episode content, previous episodes for context, and a type catalog (Entity plus any user-defined Pydantic entity types). The prompt returns ExtractedEntity objects — name, entity_type_id, and episode_indices. These become EntityNode objects with labels and an initially empty summary. Simultaneously, extract_edges() (edge_operations.py:121) prompts the LLM again, this time with the extracted nodes and the episode, returning ExtractedEdge objects (source, target, relation_type, fact, temporal markers). These become EntityEdge objects. Both paths support message/text/JSON source types and multi-episode batches with per-node/per-edge episode attribution.
At write-time deduplication is two-stage. First, _collapse_exact_duplicate_extracted_nodes() (node_operations.py:348) merges same-message duplicates that share a normalized name, keeping the more specific label set. Then resolve_extracted_nodes() (line 644) performs a two-tier semantic dedup: (1) cosine-similarity search against existing nodes — an exact match short-circuits via _resolve_with_similarity(); (2) unresolved nodes are sent to a dedupe prompt (prompts/dedupe_nodes.py) that returns duplicate-or-new verdicts, and promoted nodes absorb attributes from their extracted counterparts. Edge dedup follows the same pattern: resolve_extracted_edges() (line 332) searches for existing edges between the same endpoint nodes using hybrid search, then the LLM dedupe prompt (resolve_edge in prompts/dedupe_edges.py) selects matching edges or marks contradictions for invalidation. Conflict resolution for edges is temporal: resolve_edge_contradictions() (line 547) uses valid_at/invalid_at dates — a new edge invalidates an old one only if the old edge's valid range predates the new edge's validity.
How are memories stored?
answeredThe system uses a property-graph model stored in one of four backends: Neo4j (primary), FalkorDB, Kuzu, or Neptune. Node and edge base classes live in nodes.py and edges.py. EntityNode (nodes.py:499) stores: uuid, name (with name_embedding for semantic search), summary (accumulated entity description), group_id (partition key), labels (e.g. 'Person', 'Entity'), and attributes (arbitrary key-values from custom Pydantic models). EntityEdge (edges.py:263) stores: source_node_uuid, target_node_uuid, name (relation type), fact (textual fact with fact_embedding), episodes (provenance list), temporal fields valid_at/invalid_at/expired_at, and typed attributes.
Episodic input is stored as EpisodicNode (nodes.py:318) with content, source, valid_at, entity_edges, and optional episode_metadata for custom filtering. SagaNode groups ordered episodes into longer conversation threads. A MENTIONS edge (EpisodicEdge) links episodes to entities; a NEXT_EPISODE edge chains sequential episodes within a saga; a HAS_EPISODE edge attaches episodes to a saga.
Embeddings are generated by pluggable EmbedderClient implementations: OpenAI, Voyage, Gemini, Azure OpenAI (embedder/ directory). EntityNode.name_embedding and EntityEdge.fact_embedding are stored directly as properties on the nodes/edges in the graph database. The default dimension is read from EMBEDDING_DIM in embedder/client.py. Community nodes aggregate entity clusters with their own name_embedding for search.
The GraphDriver ABC (driver/driver.py) defines the storage abstraction with execute_query(), session(), transaction support, and specialized operation interfaces (entity_node_ops, search_ops, etc.). Concrete drivers (Neo4jDriver, FalkorDriver, KuzuDriver, NeptuneDriver) map Cypher/SQL-like queries to their respective backends. Fulltext indexes are created on entity names, edge facts, episode content, and community names. Customization is possible in CLAUDE.md: set index names via ENTITY_INDEX_NAME, EDGE_INDEX_NAME, etc.
How are memories retrieved and injected into the prompt?
answeredRetrieval is a four-scope hybrid search orchestrated by search() in graphiti_core/search/search.py (line 98). It executes edge, node, episode, and community searches in parallel (via semaphore_gather), each with its own SearchConfig defining which search methods and reranker to use. The config recipes in search_config_recipes.py provide presets: COMBINED_HYBRID_SEARCH_CROSS_ENCODER uses BM25 + cosine similarity + BFS with a cross-encoder reranker.
Each scope supports three search methods defined in search_config.py: cosine_similarity (vector search on name_embedding/fact_embedding), bm25 (fulltext index search), and bfs (graph traversal from origin nodes). Results from each method are fused by a reranker: rrf (reciprocal rank fusion, the default), mmr (maximal marginal relevance for diversity), cross_encoder (a separate scoring model like OpenAI or BGE reranker that re-ranks candidate facts/names), node_distance (ranks by shortest path to a center_node_uuid), or episode_mentions (ranks by mention frequency).
SearchFilters (search_filters.py:55) enables filtering by node labels, edge types (fact types), temporal ranges (valid_at, invalid_at, created_at, expired_at), edge UUIDs, and arbitrary property filters. The group_ids parameter scopes all searches to specific partitions.
For injection into the LLM context, the search() / search_() methods on the Graphiti class (graphiti.py:1653-1755) return SearchResults containing edges (as EntityEdge objects with fact text), nodes, episodes, and communities. The Python SDK makes these directly available. The MCP server exposes search_memory_facts (returns facts as text) and search_nodes (returns entity nodes). The REST API at server/graph_service/routers/retrieve.py exposes /search and /get-memory endpoints that compose query from recent messages, call graphiti.search(), format edges as FactResult objects, and return them as JSON.
How are memories updated, consolidated or forgotten?
answeredUpdate/merge happens at write time via the dedup pipeline: when resolve_extracted_nodes() finds a semantic match in the graph, the existing node's attributes are merged with the new extraction using an overlay merge (apply_capped_attributes() in attribute_utils.py with merge_mode='overlay'), preserving prior values that the new extraction omitted. Edge resolution (resolve_extracted_edge() in edge_operations.py:638) works similarly but uses merge_mode='replace' for typed edge attributes. Edges accumulate provenance (episodes list) across extractions.
Summarization refines entity descriptions. extract_attributes_from_nodes() (node_operations.py:744) first calls per-entity attribute extraction via an LLM prompt, then batched summary generation via _extract_entity_summaries_batch(). Nodes with short summaries have edge facts appended directly (no LLM call); longer ones are partitioned into flights of 30 and sent to the LLM's extract_summaries_batch prompt. The per-entity summary cap is MAX_SUMMARY_CHARS (configurable in text_utils.py). Saga summarization (Graphiti.summarize_saga(), graphiti.py:499) incrementally updates a saga's running summary using a wall-clock watermark (last_summarized_at) and episode-time watermark (last_summarized_episode_valid_at), so backfilled episodes are caught on subsequent runs without re-summarizing everything.
Invalidation / forgetting uses bi-temporal markers. Edges carry valid_at (when a fact became true) and invalid_at (when it stopped being true). When resolve_extracted_edge() identifies contradictions, the older edge's invalid_at is set to the new edge's valid_at, and expired_at records when the system recognized the change. The resolve_edge_contradictions() function (edge_operations.py:547) handles the temporal algebra — an old edge is invalidated only if its validity window precedes the new edge's window. Nodes can be removed via remove_episode() (graphiti.py:1892), which cascade-deletes edges and nodes that were created solely by the removed episode. Whole groups are deletable via Node.delete_by_group_id(). There is no TTL-based automatic decay — forgetting is explicit through invalidation or deletion.
How is memory scoped and isolated?
answeredAll graph data is scoped by group_id, a string partition key present on every node and edge (Node.group_id, Edge.group_id). It serves as multi-tenancy, separating user/agent/session or tenant data. All queries — extraction, search, maintenance — accept group_ids: list[str] parameters. Search filters (SearchFilters) combine with group_ids for access control at query time. The @handle_multiple_group_ids decorator normalizes single/multi group_id calls.
Group-level isolation is enforced at the database level as well. When a group_id differs from the default database name, _resolve_request_scope() (graphiti.py:1066) clones the driver with .clone(database=group_id), creating a physically separate graph in FalkorDB or a logically partitioned one in Neo4j. This prevents concurrent requests from different groups from interfering (fix for issue #1676). The MCP server caches these group-scoped driver clones in _group_drivers (graphiti_mcp_server.py:188).
Deletion is group-aware: Node.delete_by_group_id() removes all nodes in a partition. build_communities() scopes its community detection to specified group_ids. Communities, sagas, and episodes all carry group_id.
Entity type scoping: users can define custom Pydantic entity type models (e.g., Person, Organization) and pass them as entity_types to add_episode(). The system uses these to classify extracted entities, plus the excluded_entity_types parameter to filter out unwanted types. Similarly, edge_types scopes which relation types are valid between which entity type pairs.
There is no separate user/role-based access control layer — that's delegated to the application layer. The REST server implements no auth middleware; the MCP server relies on the host's transport security.
How do agents integrate with it, and what is self-hostable?
answeredAgents can integrate with Graphiti through three tiers. Python SDK: the Graphiti class (graphiti.py) is the primary interface — add_episode(), search(), add_triplet(), summarize_saga(), build_communities(). It depends on pydantic, an LLM client, an embedder, a cross-encoder, and a graph driver. REST API (server/graph_service/main.py) is a FastAPI application with /messages (episode ingestion via background worker), /search, /get-memory, /entity-edge, and /clear endpoints. It wraps the Python SDK and is deployable via uvicorn (uvicorn graph_service.main:app). MCP Server (mcp_server/) exposes all core operations as MCP tools: add_memory, add_triplet, search_memory_facts, search_nodes, summarize_saga, build_communities, delete_episode, clear_graph, and more. It supports both Neo4j and FalkorDB backends via a configuration system with LLM/embedder/reranker provider factories. The MCP server can run in stdio or SSE mode and is containerized with Docker Compose.
All three tiers are fully self-hostable — nothing requires a hosted service. Required infrastructure: a graph database (Neo4j 5.26+, FalkorDB 1.1.2+, Kuzu, or Neptune) and API keys for the chosen LLM/embedding provider. The default stack uses OpenAI for all three (LLM, embeddings, reranking), but each is pluggable: LLMClient implementations exist for OpenAI, Anthropic, Gemini, Groq; EmbedderClient for OpenAI, Voyage, Gemini, Azure OpenAI; CrossEncoderClient for OpenAI, BGE, Gemini. The LLMRuntime subsystem (llm_client/llm_runtime.py) supports multi-model prompt routing. OpenTelemetry tracing is optional but integrated. The Kuzu driver is a lightweight embedded option for single-process deployments. There is no hosted/cloud tier visible in this repository.