LLMs Technical Reviews

Are communities, summaries or hierarchies built over the graph?

Community detection (e.g. Leiden); hierarchical summaries; when they are computed; if not done, say so.

Verdict

GraphRAG and its compact clone nano-graphrag are the reference here: they build hierarchical Leiden communities and an LLM report for each one at index time. Semantica and LLM Graph Builder can also build summarized hierarchies, but you have to trigger them, and in LLM Graph Builder the summary step is broken at the reviewed commit.

Hierarchy plus reports at index time. GraphRAG runs hierarchical Leiden (max_cluster_size 10) on the largest connected component and writes reports deepest level first. Its fallback for oversized communities never receives sub-community reports, so those communities are trimmed instead. nano-graphrag uses the same graspologic call, also on the largest component only. It drops every report and regenerates all of them on each insert.

Hierarchies you trigger. LLM Graph Builder runs GDS Leiden with up to 3 levels and summarizes parent communities from their children. This happens only when the /post_processing job runs with enable_communities. That route passes the embedding provider name in the position of the summary-model argument, so summaries fail unless an LLM config exists under that name. Semantica’s CommunityHierarchyBuilder runs real Louvain or Leiden, and CommunitySummarizer writes reports bottom-up, falling back to an extractive report when no LLM is set. Its simpler CommunityDetector labels greedy modularity as both “Louvain” and “Leiden”.

Flat clusters without summaries. graphify runs Leiden, falling back to Louvain. It splits any community larger than 25% of the graph, and re-splits communities of 50 or more nodes whose cohesion is below 0.05. The LLM or the host agent writes only short names for them.

None. LightRAG has no clustering step, and its “global” mode searches relationship vectors. HippoRAG uses Personalized PageRank instead. TrustGraph does not apply here: it relies on reranked two-hop traversal. AutoFlow has no detection step. Admins can create “synopsis” entities by hand. Vector Graph RAG expands a subgraph at query time only.

Pick: GraphRAG for corpus-wide thematic questions, if you can afford the reports. Pick: nano-graphrag to study or fork the same design in about 1,100 lines. Pick: LightRAG or HippoRAG when questions are about specific entities and the corpus changes often.

Per-project answers

Graphify-Labs/graphify

answered

Community detection — Leiden with Louvain fallback. The cluster() function in cluster.py (lines 232–362) runs Leiden community detection via the native Rust graspologic_native.leiden() when available, falling back to the Python networkx.algorithms.community.louvain_communities implementation. Directed graphs are converted to undirected first. The resolution parameter (default 1.0) controls community granularity: >1.0 yields more, smaller communities.

Hub exclusion. An exclude_hubs_percentile option (cluster.py:248) removes nodes whose degree exceeds a given percentile of the degree distribution before partitioning, then reattaches them by majority-vote neighbour community. This prevents super-hubs (utility modules, entry points) from pulling unrelated subsystems into the same community.

Oversized community splitting. Communities larger than 25% of graph nodes (min 10) are split by running a second Leiden pass on the subgraph (cluster.py:336–342). A subsequent cohesion check (cluster.py:344–353) re-splits communities whose cohesion_score (internal edge density relative to outgoing edges) falls below 0.3, catching remaining doc-hub bridges.

Community IDs are stable. Communities are indexed by descending size, with a tiebreak on sorted node tuples, so identical groupings get identical IDs run-to-run (cluster.py:355–362).

No hierarchical summaries. Communities are flat — there are no hierarchical summaries, no community-summary generation via LLM, and no community-level text aggregation. The community ID is simply stamped onto each node when writing graph.json (export.py:343–347). The get_community MCP tool (serve.py:2008–2016) returns all nodes belonging to a community ID, with no accompanying summary text.

When computed. Communities are computed once per build via cluster() after the graph is assembled. They are not incrementally updated — a full re-cluster runs on each build. The MCP server reconstructs the community dict from stored node attributes on load (_communities_from_graph, serve.py:94–101).

Editor's note. Correction: the cohesion re-split threshold is 0.05 (not 0.3) and only applies to communities of 50+ nodes; cohesion_score is intra-community edges over possible node pairs, not a ratio against outgoing edges. Communities do get short LLM- or agent-written names (label_communities / skill Step 5), though no summaries.

HKUDS/LightRAG

answered

LightRAG does not perform community detection, community summarization, or hierarchical graph summarization. A grep for "leiden", "louvain", "community detection", "community summar", and "hierarchical summary" across the entire lightrag/ Python source tree returns zero matches in the core logic. The only hit for "community" in non-test Python code is inside lightrag/chunker/paragraph_semantic.py, which is about paragraph-level document chunking (a semantic text-splitting technique) and has nothing to do with graph community detection.

Instead of building communities, LightRAG takes a retrieval-time approach to knowledge synthesis. The longest path in the knowledge graph was present in earlier GraphRAG papers (like Microsoft's GraphRAG) which pioneered the Leiden-based community summary approach, but LightRAG opted for a simpler design where the query process itself bridges local and global context:

  • Local mode (lightrag/operate.py:5366-5376): starts from the most vector-similar entities and retrieves their neighbor edges via _get_node_data (entity vector DB → graph node degree → related edges).
  • Global mode (lightrag/operate.py:5377-5386): starts from the most vector-similar relationships and collects their endpoint entities via _get_edge_data (relation vector DB → related entities).
  • Hybrid mode (lightrag/operate.py:5388-5408): runs both and round-robin merges results.
  • Mix mode (lightrag/operate.py:5411-5430): adds direct vector search over document chunks on top.

The context assembly stages (_build_query_context → _perform_kg_search → _apply_token_truncation → _merge_all_chunks → _build_context_str at lightrag/operate.py:6075-6197) feed all retrieved entities, relations, and document chunks into the LLM prompt for final answer synthesis. There is no pre-computed community-level summary at any point.

Editor's note. Correction: the sentence about a 'longest path' in earlier GraphRAG papers is not supported by the code or the cited lines. The rest stands: there is no community detection or hierarchical summary.

microsoft/graphrag

answered

Yes — GraphRAG builds a hierarchical community structure using the Leiden algorithm.

Detection. The create_communities workflow calls cluster_graph, which normalizes edges to undirected, optionally restricts to the largest connected component (use_lcc), then calls hierarchical_leiden via graspologic-native at hierarchical_leiden.py:11. Parameters: max_cluster_size (default 10), resolution 1.0, randomness 0.001, configurable seed.

Hierarchy. _compute_leiden_communities (at cluster_graph.py:51) builds level-to-community and cluster-to-parent mappings. Each community stores level (0=root), community ID, parent, children, plus aggregated entity_ids, relationship_ids, and text_unit_ids (see create_communities.py:86-192).

Report generation. The create_community_reports workflow builds a local context (nodes, edges, claims) per community and calls an LLM to generate structured reports. Each report has summary, full_content (narrative), findings (structured JSON), rank, and rating_explanation. Reports can be weighted by text-unit count of their entities (at community_context.py:189).

When computed. Communities and reports are built during indexing, not at query time.

Editor's note. Addition: all level contexts are built before any report is generated (summarize_communities.py L61-L73), so build_level_context always gets an empty report table. Oversized communities are trimmed instead of being filled with sub-community reports.

semantica-agi/semantica

answered

Semantica has a full community detection and hierarchical summarization pipeline. The CommunityDetector class (semantica/kg/community_detector.py) supports Louvain (via NetworkX greedy_modularity_communities with resolution parameter, and a basic fallback), Leiden (via optional igraph+leidenalg or cdlib, falling back to a native Python implementation with local-moving and refinement phases at lines 894-1034), overlapping k-clique communities, and label propagation (supports chunked processing for large graphs). The CommunityHierarchyBuilder (semantica/kg/community_hierarchy.py:553-1419) builds multi-level partitions by iteratively coarsening the graph via Louvain or Leiden at increasing resolution, then constructs a CommunityHierarchy — a tree of HierarchicalCommunity nodes keyed by SHA-256 content hashes with metrics (internal/external edges, density, conductance) and deterministic parent-child resolution. The CommunitySummarizer (semantica/kg/community_summarizer.py:547-2297) generates structured CommunityReport objects bottom-up via summarize_hierarchy(): for each community it extracts a subgraph, computes centrality scores (degree/betweenness/closeness/eigenvector/PageRank), packs context within a token budget (prioritizing child reports, edges, anchor entities, and source text), and calls an LLM with a Pydantic-typed schema (CommunityReportLLMSchema) to produce a title, summary, findings list, and impact rating. It has a 5-tier LLM unwrap strategy (_call_llm) and falls back to an extractive baseline when no LLM is configured. Reports are cached via SHA-256 content hashing with atomic disk persistence. The hierarchy is computed on demand when the user invokes community detection; it is not automatically triggered during ingestion or graph building.

Editor's note. Correction: CommunityDetector's Louvain is NetworkX greedy modularity (Clauset-Newman-Moore) with the resolution argument effectively fixed at 1.0, and its Leiden method just relabels that result. Real Louvain (nx.louvain_communities) and Leiden (leidenalg, cdlib, then a native fallback) are only in CommunityHierarchyBuilder in community_hierarchy.py.

neo4j-labs/llm-graph-builder

answered

Community detection. Communities are built using the Neo4j Graph Data Science (GDS) library's Leiden algorithm (communities.py:232-247). write_communities() calls gds.leiden.write() on a named graph projection, with includeIntermediateCommunities=True, up to MAX_COMMUNITY_LEVELS=3 hierarchical levels, and minCommunitySize=1. The projection is built over all non-Chunk/non-Document/non-Community nodes and their relationships, with an undirected adjacency weight based on edge count (communities.py:20-34). The Leiden partition IDs are written as a communities array property on each __Entity__ node.

Hierarchical community creation. create_community_properties() (communities.py:468-499) executes a multi-step Cypher pipeline:

  1. A uniqueness constraint on __Community__.id.
  2. CREATE_COMMUNITY_LEVELS creates __Community__ nodes at level 0 (from lowest Leiden partition), then level 1 and 2 parent groups, with PARENT_COMMUNITY edges linking child→parent.
  3. Ranks (distinct document count) and weights (chunk count) are computed per community via CREATE_COMMUNITY_RANKS/CREATE_COMMUNITY_WEIGHTS and their parent variants.

Summaries. create_community_summaries() (communities.py:313-371) generates LLM summaries for each level-0 community and then for parent communities (by summarizing child summaries). It uses a LangChain ChatPromptTemplate | llm | StrOutputParser pipeline, processing communities concurrently via ThreadPoolExecutor (max 10 workers). The LLM receives the community's subgraph (node IDs, types, descriptions and relationship triples) and returns a title and natural-language summary. Summaries are stored as summary and title properties on __Community__ nodes via STORE_COMMUNITY_SUMMARIES.

Embeddings and indexes for communities are generated in create_community_embeddings() (communities.py:374-403) using the configured embedding model in batches of 100. Full-text (community_keyword) and vector (community_vector) indexes are also created for hybrid search.

When they are computed. Communities are a post-extraction step invoked via the create_communities() entry point (communities.py:519-533), which first clears existing communities, creates a new GDS graph projection, runs Leiden, then builds all properties and summaries. This is triggered manually — not automatically on each document ingestion.

OSU-NLP-Group/HippoRAG

answered

Not implemented. The repository does not implement community detection (Leiden, Louvain, or any other algorithm), hierarchical summaries over graph regions, or any form of graph partitioning. A grep for commun, leiden, louvain, hierarch, and summar across the entire src/ directory returned zero matches. The graph is used strictly as a flat, node-labeled structure where edges represent fact co-occurrence, passage–entity membership, and embedding-similarity (synonymy).

What happens instead. When querying, the system computes fact–query similarity scores, selects top facts, propagates their scores to the entities they contain, runs personalized PageRank (igraph.Graph.personalized_pagerank) over the entire entity+passage graph, and ranks passage nodes by their PPR score (HippoRAG.py:2177–2218). There is no hierarchical summarization step and no community-level retrieval. The graph is always treated as one monolithic component.

This is a deliberate architectural difference from approaches like GraphRAG (which uses Leiden clustering + community summaries). HippoRAG's PPR-based retrieval effectively lets relevance propagate along paths through the graph without needing a pre-computed community hierarchy.

gusye1234/nano-graphrag

answered

Community detection. The clustering method on NetworkXStorage (gdb_networkx.py:165-168,230-252) runs the Leiden algorithm via graspologic.partition.hierarchical_leiden. Before running, the graph is passed through stable_largest_connected_component which discards disconnected subgraphs and stabilizes node ordering. The Leiden parameters include max_cluster_size (default 10) and graph_cluster_seed (0xDEADBEEF). The hierarchical partition produces multiple levels; each node is annotated with {"level": level_key, "cluster": cluster_id} entries stored as a JSON-serialized clusters attribute. The Neo4j backend (gdb_neo4j.py:387-440) instead runs Leiden via GDS (Graph Data Science) library procedures, projecting the graph and writing community IDs back into node properties.

Community schema. After clustering, community_schema() (gdb_networkx.py:170-224) iterates all nodes, groups them by (level, cluster) key, collects member nodes, edges (from node adjacency), source-chunk IDs, and computes occurrence as the ratio of chunk IDs to the max across all communities. It then builds the hierarchy by linking each community to its sub-communities (strict subset of nodes at the next level down).

Community reports. generate_community_report (_op.py:625-697) processes communities level-by-level from highest (most fine-grained) to lowest (root). For each community, _pack_single_community_describe (_op.py:464-600) assembles a CSV of entities, relationships, and optional sub-community reports (used for large communities with >100 nodes or edges), truncating each section by token budget. The LLM (best_model_func) is prompted with community_report (prompt.py:63-192) asking for a JSON-structured report with title, summary, an impact severity rating (0-10), and 5-10 detailed findings. Results are stored as both report_string (markdown) and report_json in the community_reports KV store. Reports are recomputed on every document insertion — the code explicitly comments "TODO: don't support incremental update for communities now, so we have to drop all" (graphrag.py:318).

zilliztech/vector-graph-rag

answered

Communities, hierarchical summaries, and community detection (e.g. Leiden) are NOT implemented.

The project intentionally takes a different approach from Microsoft's GraphRAG. Rather than detecting communities and computing hierarchical summaries over them, Vector Graph RAG stores the graph's adjacency as ID lists in metadata and performs subgraph expansion at query time via SubGraph.expand() (graph/knowledge_graph.py:261-361).

There is no offline graph partitioning, no community detection algorithm (Leiden or otherwise), no summary-of-summaries pyramid, and no global search mode. The system has a single retrieval path: entity extraction from the question, vector search for similar entities and relations, subgraph expansion from seed nodes, optional LLM reranking, and passage retrieval.

The SubGraph class (graph/knowledge_graph.py:152-591) implements lazy expansion: it holds entity and relation IDs and fetches neighbor records from Milvus on demand during expand(degree=...). Each expansion step goes from current relations to their connected entities, then from those entities to their connected relations — this is graph traversal at query time, not community detection at index time.

A naive RAG baseline (VectorGraphRAG.query_naive()) is available for comparison, which skips the graph entirely and does direct passage vector search.

pingcap/autoflow

insufficient evidence

AutoFlow does not implement automated community detection (e.g., Leiden algorithm), hierarchical summarization, or any Graph ML-based partitioning of the graph. A search of all Python files in both core/ and backend/ for terms like "community", "leiden", "hierarch", "cluster", and "partition" — none appear in the graph-related code.

What exists instead is a synopsis entity mechanism, exposed as an admin API (POST /admin/knowledge_bases/{kb_id}/graph/entities/synopsis, backend/app/api/admin_routes/knowledge_base/graph/routes.py:56-78). The TiDBGraphEditor.create_synopsis_entity() method (backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_editor.py:163-222) allows a user (or an external process) to manually create an entity of type synopsis, giving it a name, description, topic, and a list of related entity IDs. It then creates "is a part of" relationships linking the synopsis entity to each listed entity. During graph retrieval, synopsis entities are fetched via fetch_similar_entities(..., entity_type=EntityType.synopsis) and added to the result set (tidb_graph_store.py:551-554). This provides a form of user-driven grouping, but it is not an automatic community-detection pass.

There is also no evidence of hierarchical summaries being computed over entity groups. The graph is flat — entities of type original are the only automatically-created kind.

trustgraph-ai/trustgraph

not applicable

This repository does not implement community detection, hierarchical community summaries, or any analogous graph grouping mechanism. There is no Leiden algorithm, Louvain, graph partitioning, or summarization of entity clusters.

Instead, TrustGraph relies on local graph traversal with cross-encoder reranking (retrieval/graph_rag/graph_rag.py:323-496) to navigate the knowledge graph at query time. From seed entities found via vector search, it performs up to max_path_length (2) iterative hops, retrieving edges adjacent to the frontier, scoring them with a cross-encoder reranker, and selecting the top edge_limit (25) edges.

The ontology provides a class hierarchy (extract/kg/ontology/ontology_loader.py:15-26) with subclass_of for domain/range validation, but this is an inheritance hierarchy, not community detection.

← Where and how is the graph stored? · How does query-time retrieval use the graph? →