Graphify-Labs/graphify
Coding-assistant skill and CLI that turns a repo into a tree-sitter code graph plus LLM-extracted doc nodes, queried over MCP by traversal.
Overview
graphify builds a knowledge graph of a software project and gives coding agents tools to query it, so they read less raw code. It is mainly distributed as a /graphify skill for AI coding assistants (Claude Code, Codex, Cursor, Gemini CLI, Copilot and about fifteen more), backed by a large Python package (graphifyy on PyPI) with a graphify CLI and an MCP server.
It is not GraphRAG in the Microsoft sense. The core graph comes from deterministic tree-sitter parsing: files, classes, functions, imports, calls and inheritance, plus cross-file resolution passes that bind calls to their definitions. The dispatch table maps more than 100 file suffixes to extractors. An LLM is used only for documents, PDFs and images, and to give communities short names. There are no embeddings and no vector store. Retrieval is keyword scoring that picks seed nodes, then a BFS or DFS over the graph, rendered as text within a token budget.
The project is big (about 78,000 lines of Python in the package) and shows signs of heavy issue-driven hardening: many comments cite issue numbers, there are shrink guards that refuse to overwrite a graph with a smaller one, and every edge carries an EXTRACTED, INFERRED or AMBIGUOUS confidence label. The output is a graphify-out/ folder with graph.json, an interactive graph.html and a GRAPH_REPORT.md.
Architecture
flowchart LR
SK["/graphify skill (host agent)"] --> DT["detect(): classify files"]
CLI["graphify extract / update"] --> DT
DT --> AST["extract(): tree-sitter AST + resolvers"]
DT --> SEM["semantic pass: subagents or llm.py backend"]
AST --> MRG["merge AST + semantic JSON"]
SEM --> MRG
MRG --> BLD["build_from_json -> NetworkX"]
BLD --> CL["cluster(): Leiden + splits"]
CL --> AN["analyze: god nodes, surprises"]
AN --> OUT["graph.json / graph.html / GRAPH_REPORT.md"]
OUT --> MCP["serve.py MCP tools"]
OUT --> QCLI["graphify query / path / explain"]
HK["hooks: git + PreToolUse nudge"] --> CLI
| Component | Path | Role |
|---|---|---|
| Skill | graphify/skill.md (+ per-host variants) |
Step-by-step script the host agent runs: detect, AST, subagent extraction, build, label, export |
| CLI | graphify/__main__.py, graphify/cli.py |
extract, update, query, path, explain, watch, export, install, hooks |
| Detection | graphify/detect.py |
Walks the tree and buckets files into code, document, paper, image, video |
| AST extraction | graphify/extract.py, graphify/extractors/ |
Per-language tree-sitter extractors and cross-file call/import resolution |
| Semantic extraction | graphify/llm.py |
Provider backends, token-budget chunk packing, adaptive retry, evidence binding |
| Build and dedup | graphify/build.py, graphify/dedup.py |
Extraction JSON to NetworkX; exact-id and fuzzy label dedup |
| Clustering | graphify/cluster.py |
Leiden (native Rust) with Louvain fallback, oversized and low-cohesion splits |
| Analysis and report | graphify/analyze.py, graphify/report.py |
God nodes, surprising connections, suggested questions, GRAPH_REPORT.md |
| Exports | graphify/export.py, graphify/exporters/ |
JSON, HTML, Obsidian, SVG, GraphML, Cypher, Neo4j and FalkorDB push |
| Query server | graphify/serve.py |
MCP stdio/HTTP server with traversal, path and community tools |
| Incremental | graphify/watch.py, graphify/cache.py, graphify/hooks.py |
File-hash caches, code-only rebuilds, watch mode, git hooks |
How a request flows
A first /graphify . run inside Claude Code:
- Detect. The skill writes the interpreter path and scan root, then calls
detect()to classify files. - AST pass. For code files it calls
extract(code_files, cache_root=...).extractlooks up an extractor per suffix in_DISPATCH, parses in a process pool, then runs cross-file resolution. Theresolution_context_*arguments let an incremental run bind calls to callees that were not re-parsed (extract.py, L7410-L7458). - Semantic pass. Only docs, papers and images get one. If
GEMINI_API_KEYis set the skill callsextract_corpus_parallel(..., backend="gemini"). Otherwise the host agent dispatches its own subagents with the shipped extraction spec, and each writes a chunk JSON file (skill.md). A code-only repo skips this step and needs no API key. - Merge and build. AST nodes win on id collisions.
build_from_jsonturns the merged dict into anetworkx.Graph, or aDiGraphwith--directed(build.py). - Cluster and analyse.
cluster(G)returns{community_id: [node_ids]}. Thengod_nodes,surprising_connectionsandsuggest_questionsfeedreport.generate. - Label. The host agent reads each community’s members and writes a two-to-five-word name, which is saved to
.graphify_labels.jsonand stamped intograph.json(skill.md). Headless runs usellm.label_communitiesin batches of 100 (llm.py). - Write.
to_jsonrefuses to overwrite an existing graph with fewer nodes unless forced (export.py).graph.htmland the report are written next to it. - Query. An agent then calls the
query_graphMCP tool orgraphify query "...", as described below.
Key components
Query engine
_query_graph_text is the whole retrieval path (serve.py). It splits the question into terms (with jieba for Chinese), scores nodes in one pass using IDF-weighted token matches behind a trigram prefilter, and picks seeds with a gap heuristic that keeps at least one seed per term. Relation verbs such as “calls” are kept out of that per-term guarantee. It then optionally filters edges by relation (explicit or inferred from the question), runs BFS (default) or DFS to depth 3, and renders nodes and edges as text with seeds first, within 2,000 tokens. The MCP handler caps depth at 6 and logs each query (serve.py). The other tools are get_node, get_neighbors, get_community, god_nodes, graph_stats and shortest_path, plus three pull-request tools (list_prs, get_pr_impact, triage_prs). graphify returns context, not answers. The calling agent does the reasoning.
Clustering
cluster() drops isolates (each becomes its own community), can exclude hub nodes above a degree percentile and reattach them by neighbour vote, and partitions the rest with Leiden. It splits any community above 25% of the graph (minimum 10 nodes) with a second Leiden pass. It re-splits communities of 50 or more nodes whose cohesion (intra-community edges over possible pairs) is below 0.05. IDs are ordered by size with a sorted-member tiebreak, so they stay stable between runs (cluster.py, L383-L398). Communities are flat. There are no hierarchical levels and no LLM community reports, only names.
Semantic extraction
extract_corpus_parallel packs files into 60,000-token chunks grouped by directory and runs four chunks in parallel (Ollama runs one at a time) (llm.py, L2160-L2175). A truncated output, a context-overflow error or a timeout splits the chunk in half and retries, up to depth 3 (llm.py). _bind_node_evidence flags LLM-proposed code nodes as unverified when none of their identifiers appear in the file that was read. They are flagged, not dropped (llm.py). Headless backends include Claude, OpenAI, Gemini, Kimi, DeepSeek, Ollama, Azure, Bedrock and a claude-cli shell-out.
Deduplication
deduplicate_entities first collapses exact id collisions with a deterministic survivor ranking. It then merges near-identical labels through MinHash/LSH blocking and Jaro-Winkler verification, with an optional LLM tiebreak. It refuses to run across nodes from more than one repo (dedup.py).
Agent hooks
Besides git hooks that rebuild on commit, graphify installs a PreToolUse guard. When a fresh graph exists, it nudges the agent toward graphify query instead of grep or raw reads. Strict mode denies the first raw read of indexed code in a session. The guard fails open on any error (cli.py).
Extending it
- New language. Add an
extract_<lang>module undergraphify/extractors/, register its suffixes in_DISPATCHandcollect_files, then add it toCODE_EXTENSIONSand_WATCHED_EXTENSIONSand add a fixture test.ARCHITECTURE.mddocuments this, and a test imports every symbol it names. - Library use. Every stage is a plain function on dicts and NetworkX graphs (
detect,extract,build_from_json,cluster,to_json), so you can build a custom pipeline in Python. - Exports. Push to Neo4j or FalkorDB, or write Cypher, GraphML, SVG, an Obsidian vault or a per-community markdown wiki.
- Backends.
graphify provider add <name> --base-url ... --default-model ... --env-key ...registers an extra endpoint for the headless semantic pass. Built-in names cannot be overridden.
Running it
uv tool install graphifyy(or pipx), thengraphify installto register the skill with your assistant, then/graphify .inside the assistant.- Headless and CI:
graphify extract <path> [--backend gemini|claude|openai|deepseek|ollama|kimi] [--code-only]runs the full pipeline with whichever key is set (cli.py). graphify updatere-extracts changed code only, with no LLM. Doc changes need the skill orextractagain (cli.py).graphify watchand the git hooks automate this.python -m graphify.serve graphify-out/graph.jsonstarts the MCP server over stdio, or over Streamable HTTP with--transport httpand an optional--api-key. Everything lives undergraphify-out/in the repo.
Strengths and caveats
- Strength: deterministic, free code graph. For code, no model is involved. Results are reproducible, cached per file hash, and nothing leaves the machine.
- Strength: provenance. Confidence labels on every edge,
unverifiedflags on LLM-invented code symbols, and shrink guards on writes make the graph auditable. - Strength: broad coverage. More than 100 suffixes across mainstream and niche languages (COBOL, Fortran, Pascal, Verilog, Terraform, SQL), and around twenty assistant integrations.
- Caveat: lexical retrieval only. Seeds come from token and trigram matching on labels and attributes. A question phrased in terms that do not appear in identifiers or doc labels can miss, and there is no embedding fallback.
- Caveat: flat communities, no summaries. Unlike GraphRAG-style systems there is no global, community-summary query mode. Community names come from one short LLM pass.
- Caveat: the skill path depends on the host agent. Semantic extraction quality and cost depend on the assistant’s subagents following a long scripted prompt. The headless
extractpath is more predictable. - Caveat: complexity.
extract.pyalone is over 9,000 lines and the extractor engine is over 8,000. It is powerful, but hard to fork or audit compared with smaller tools in this category.
Sources: code at 5c7b847, deepwiki-open wiki (11 pages), verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit 5c7b847. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredChunking. Non-code files (docs, PDFs, images) are processed by the LLM-based semantic pipeline in llm.py. Files are first split at _FILE_CHAR_CAP (20,000 characters) into FileSlice units via expand_oversized_files (llm.py:2638). These units are then packed into chunks by _pack_chunks_by_tokens() (llm.py:2160–2203), using a greedy algorithm that groups by parent directory (so related files share one LLM call) and closes a chunk when a configurable token budget (default 60,000 tokens) would be exceeded, with a hard cap of 20 images per chunk.
Entity and relation extraction. Code files are handled entirely deterministically via tree-sitter AST extractors in graphify/extractors/ — one per language (Python, JS/TS, Rust, Go, Java, C++, etc.). These produce nodes representing classes, functions, imports, and variables, and edges for calls, imports, inherits, references, etc. The LLM semantic pass uses a detailed system prompt (_EXTRACTION_SYSTEM in llm.py:481–511) that defines a JSON schema with nodes, edges, and hyperedges. Every edge is tagged with a confidence tier — EXTRACTED, INFERRED, or AMBIGUOUS — and each node carries a file_type of code, document, paper, image, rationale, or concept. Injection sentinels in source text are neutralised (_neutralise_injection_sentinels, llm.py:575–582). Fabrication is mitigated via _bind_node_evidence (llm.py:703–767): code-typed nodes whose symbol name has no substring match in the file bytes the model actually read are flagged with verification = "unverified".
Entity resolution / deduplication. The dedup.py module implements a pipeline: deduplicate_entities() (dedup.py:557–657) runs exact-ID dedup first (one survivor per node ID, deterministically ranked by _collision_rank), then fuzzy label-based dedup via a cascade: exact normalization → entropy gate (_entropy ≥ 2.5, dedup.py:29–38) → MinHash/LSH blocking at threshold 0.7 (dedup.py:48–53) → Jaro-Winkler verification at 92.0 (dedup.py:247) → same-community boost (+5.0) → union-find merge. An optional LLM pass (dedup_llm_backend) resolves ambiguous pairs in the 75–92 Jaro-Winkler zone.
Schema / ontology. There is no formal ontology. The type vocabulary is limited to file_type on nodes and relation on edges. The ID scheme follows a {path}_{entity} pattern where the path is the repo-relative file path with extension dropped, all segments joined by underscores (llm.py:500).
Where and how is the graph stored?
answeredIn-memory NetworkX. The graph lives in memory as a networkx.Graph (or nx.DiGraph when directed=True) throughout building, clustering, analysis, and query serving. build_from_json() (build.py:877–1038) constructs it: nodes are added with all attributes except id unpacked as kwargs, and edges are added with their metadata.
Serialised to JSON on disk. The canonical serialisation is graph.json, written by to_json() in export.py (line 272). It uses NetworkX's node_link_data format: a JSON object with top-level nodes (array of dicts) and links (array of edge dicts). Community IDs are stamped onto each node at write time. A safety check (export.py:278–331) refuses to overwrite an existing graph with a shrinking node count unless force=True. A backup is snapshotted before overwrite when the graph has semantic or curated content (backup_if_protected, export.py:42–104).
Optional graph database export. The exporters/graphdb.py module provides push_to_neo4j() (graphdb.py:22–98) which uses MERGE statements to upsert nodes and edges into a running Neo4j instance. Node labels are derived from file_type (capitalised). A similar push_to_falkordb exists. These are pure exports — there is no database-backed storage layer.
No vector store or embeddings. The README (line 38) states explicitly: "No embeddings, no vector store: a real graph you traverse." Node attributes include norm_label (diacritic-stripped lowercased label) for textual matching, but no embedding vectors are stored.
Global graph. A ~/.graphify/global-graph.json accumulates graphs from multiple repos via global_add() (global_graph.py:79). It is a flat union with repo-tagged nodes, not a federated index.
Are communities, summaries or hierarchies built over the graph?
answeredCommunity detection — Leiden with Louvain fallback. The cluster() function in cluster.py (lines 232–362) runs Leiden community detection via the native Rust graspologic_native.leiden() when available, falling back to the Python networkx.algorithms.community.louvain_communities implementation. Directed graphs are converted to undirected first. The resolution parameter (default 1.0) controls community granularity: >1.0 yields more, smaller communities.
Hub exclusion. An exclude_hubs_percentile option (cluster.py:248) removes nodes whose degree exceeds a given percentile of the degree distribution before partitioning, then reattaches them by majority-vote neighbour community. This prevents super-hubs (utility modules, entry points) from pulling unrelated subsystems into the same community.
Oversized community splitting. Communities larger than 25% of graph nodes (min 10) are split by running a second Leiden pass on the subgraph (cluster.py:336–342). A subsequent cohesion check (cluster.py:344–353) re-splits communities whose cohesion_score (internal edge density relative to outgoing edges) falls below 0.3, catching remaining doc-hub bridges.
Community IDs are stable. Communities are indexed by descending size, with a tiebreak on sorted node tuples, so identical groupings get identical IDs run-to-run (cluster.py:355–362).
No hierarchical summaries. Communities are flat — there are no hierarchical summaries, no community-summary generation via LLM, and no community-level text aggregation. The community ID is simply stamped onto each node when writing graph.json (export.py:343–347). The get_community MCP tool (serve.py:2008–2016) returns all nodes belonging to a community ID, with no accompanying summary text.
When computed. Communities are computed once per build via cluster() after the graph is assembled. They are not incrementally updated — a full re-cluster runs on each build. The MCP server reconstructs the community dict from stored node attributes on load (_communities_from_graph, serve.py:94–101).
cohesion_score is intra-community edges over possible node pairs, not a ratio against outgoing edges. Communities do get short LLM- or agent-written names (label_communities / skill Step 5), though no summaries.How does query-time retrieval use the graph?
answeredMCP server tools. The query surface is a set of MCP tools defined in serve.py (lines 1959–2114): query_graph (the primary search+traverse tool), get_node, get_neighbors, get_community, shortest_path, god_nodes, and graph_stats. There is no distinct local/global/hybrid mode — the single query_graph tool supports BFS and DFS traversal modes.
Scoring and seed selection. A query in _query_graph_text() (serve.py:1361–1420) first splits the question into search terms via _query_terms() (serve.py:293), which handles Chinese segmentation via jieba when available. These terms are scored against every node in _score_query() (serve.py:554–680), which computes a combined TF-IDF–style score using per-token frequency (via _compute_idf) with a full-query exact-match bonus. Seeds are selected by _pick_seeds() via a gap-based heuristic over the ranked list, guaranteeing at least one seed per distinct query term. Relational-intent verbs ("calls", "uses", etc.) are demoted from the per-term guarantee to prevent an incidental match on a verb string from seating a decoy root (serve.py:1386–1391).
Trigram prefilter for speed. Before scoring each node, a trigram index (_get_trigram_index, serve.py:451–472) prunes the candidate set to nodes whose normalised text trigrams overlap with query trigrams, reducing passes over the full graph. On sparse queries the fallback is a full scan.
Traversal. Starting from selected seeds, the graph is traversed via BFS (breadth search — the default) or DFS (depth-first trace). Depth is capped (default 3, max 6). Edge relations can be filtered via explicit context_filter or auto-inferred from query intent words (_infer_context_filters, serve.py:940). The traversal runs on a filtered subgraph view.
Context assembly. The visited subgraph is rendered to plain text by _subgraph_to_text() — nodes as labelled entries with their source file, edges as directional relation statements — and capped at a token_budget (default 2000 tokens) so the output fits in an LLM context window. Seeds are rendered first so they survive truncation.
No vector search. There is no vector/hybrid search layer. All retrieval is text-based (trigram prefilter → token/scored exact & prefix matching → graph traversal). The output is returned as text directly to the calling agent (the MCP client), not fed into any secondary LLM query pipeline within graphify.
How are updates and incremental indexing handled?
answeredPer-file AST cache. The AST extraction cache (cache.py) is content-addressed: each file's extraction result is keyed by a hash of its bytes and namespaced by package version (cache/ast/v{version}-s{schema}/). On re-extraction, load_cached() returns the cached result for unchanged files; only new or modified files are re-parsed. Stale entries from older versions are cleaned up eagerly (_cleanup_stale_ast_entries, cache.py:45–69).
Per-file semantic cache. LLM extraction results are cached individually under cache/semantic/ (or cache/semantic-{mode}/ for deep mode), keyed by file-content hash. The cache also fingerprints the extraction prompt text (_PROMPT_FP_LEN = 12, cache.py:82), so entries invalidate when the prompt changes without re-billing unchanged files on every patch release. The MCP server checkpoints each chunk's results to the semantic cache as it completes, so an interrupted run loses only in-flight chunks (_checkpoint_chunk, llm.py:2677–2700). Corrupt entries are detected and reported without silently serving stale data.
File-system watch mode. graphify watch (watch.py) monitors the filesystem for changes and writes changed paths to a .pending_changes file. The pending drain mechanism (_drain_pending, watch.py:49–78) reads and deduplicates these paths. A co-operative lock (_rebuild_lock) prevents concurrent rebuilds. Post-commit hooks can append to the pending file without owning the lock.
Incremental merge. The build_merge() function (build.py:1953) takes an existing graph plus new extraction results and produces a merged graph without a full rebuild. Node deduplication and edge rewiring handle ID collisions deterministically. The resolution_context_nodes and resolution_context_edges parameters (extract.py:7417–7457) let incremental re-extraction pass unchanged file metadata as read-only context, so cross-file resolvers can still bind calls to unchanged callees without re-parsing them.
Deletion. The --update CLI flag deletes graph nodes for files that no longer exist on disk before merging new extractions. There is no explicit per-entity deletion API.
Global graph. global_graph.py tracks per-repo commits and allows adding/updating a repo's graph in a cross-repo global graph via global_add().
How are LLM cost and latency controlled during indexing and query?
answeredToken-budget chunk packing. The semantic extraction pipeline (extract_corpus_parallel, llm.py:2567) packs files into chunks using a configurable token_budget (default 60,000 tokens). Each file's token cost is estimated (characters/4 for code, 1,600 fixed per image), and chunks are closed when adding the next file would exceed the budget (llm.py:2160–2203). This prevents single over-large LLM requests that waste money on truncated retries.
Adaptive retry with bisection. When an LLM response hits finish_reason="length" (truncation), _extract_with_adaptive_retry() (llm.py:2319) bisects the chunk and re-extracts each half recursively, up to a configurable max_retry_depth (default 3, controlled by GRAPHIFY_MAX_RETRY_DEPTH). This bounds worst-case per-chunk cost at 2^depth calls. Setting depth to 0 disables all retries — one call per chunk, full stop.
Hollow-response backoff, not bisection. A hollow response (HTTP 200 with empty/unparseable content) is retried with backoff (2s, 8s) rather than bisected, because bisecting a backend issue cannot converge and would cost far more (llm.py:2406–2421). After 3 attempts the chunk fails loudly.
Model choice per stage. The BACKENDS dict (llm.py:104–224) defines model defaults per provider (claude, openai, gemini, kimi, ollama, deepseek, azure, bedrock, claude-cli) with configurable pricing per megatoken. Users can shift cost by selecting cheaper models (e.g., Ollama/local models cost $0). Model overrides are settable per provider via env vars (ANTHROPIC_MODEL, GRAPHIFY_OPENAI_MODEL, etc.).
Caching avoids re-extraction. Both the AST cache (versioned content-addressed) and semantic cache (content-hash + prompt-fingerprinted) skip re-extraction of unchanged files, eliminating LLM calls entirely for files that haven't changed since the last run.
Concurrency control. The thread pool defaults to 4 concurrent chunks (max_concurrency, llm.py:2594). Ollama defaults to serial (1) because concurrent requests cause VRAM pressure and hollow responses (llm.py:2671–2672). The GRAPHIFY_MAX_RETRIES env var controls SDK-level retries on rate limits, defaulting to 6 (llm.py:430).
Timeout guard. GRAPHIFY_API_TIMEOUT (default 600s) caps total request wall-clock time per chunk (llm.py:410–420). A timed-out chunk triggers adaptive bisection rather than silent failure.