LLMs Technical Reviews

How is the knowledge graph extracted from documents?

Chunking; entity and relation extraction prompts or models; entity resolution / deduplication; schema or ontology.

Verdict

Most of these tools have an LLM read each chunk and then merge entities only when their normalized names match exactly. The two exceptions are AutoFlow, which compares description embeddings and asks an LLM before it merges, and graphify, which merges near-identical names with fuzzy matching.

GraphRAG-style delimited prompts. GraphRAG merges entities by (title, type), so one name with two types becomes two nodes. It has four default entity types, and its fast method swaps the LLM for noun-phrase extraction. nano-graphrag uses the same prompt family and upper-cases names. Missing relationship endpoints become UNKNOWN nodes. LightRAG merges by name, picks the type by majority vote, runs gleaning at most once, and calls the LLM to summarize descriptions only after 8 fragments pile up.

Triples and OpenIE. HippoRAG runs NER on every chunk first, then triple extraction based on that output. By default it does not chunk: one document becomes one passage. It links aliases with synonymy edges (similarity of 0.8 or more) and does not merge them. Vector Graph RAG makes one JSON-mode call per 1,000-character chunk. Its name normalizer drops every character outside [A-Za-z0-9 ]. TrustGraph runs several extractors on each chunk, including ontology-guided OntoRAG, which checks domain and range. Its IRIs are built from the name, so the same name always maps to the same node.

Schema-driven and framework pipelines. LLM Graph Builder wraps LangChain’s LLMGraphTransformer and accepts optional allowed node and relationship lists. Fuzzy merging happens only when a user approves it in the UI. Unless the user’s email ends in @neo4j.com, chunks beyond MAX_TOKEN_CHUNK_SIZE are silently dropped. AutoFlow makes two DSPy calls per chunk. Semantica defaults to spaCy NER plus pattern triplets. Relation extraction is off unless extract_relations=True, and its conflict step only counts conflicts and changes nothing.

Code first. graphify parses code with tree-sitter and no LLM. It uses an LLM only for documents, PDFs and images, and merges names with MinHash plus Jaro-Winkler (threshold 92).

Pick: GraphRAG or LightRAG for general prose with tunable prompts. Pick: TrustGraph or LLM Graph Builder when a schema or ontology must constrain extraction. Pick: graphify for codebases, and AutoFlow when duplicate entities are the main worry.

Per-project answers

Graphify-Labs/graphify

answered

Chunking. Non-code files (docs, PDFs, images) are processed by the LLM-based semantic pipeline in llm.py. Files are first split at _FILE_CHAR_CAP (20,000 characters) into FileSlice units via expand_oversized_files (llm.py:2638). These units are then packed into chunks by _pack_chunks_by_tokens() (llm.py:2160–2203), using a greedy algorithm that groups by parent directory (so related files share one LLM call) and closes a chunk when a configurable token budget (default 60,000 tokens) would be exceeded, with a hard cap of 20 images per chunk.

Entity and relation extraction. Code files are handled entirely deterministically via tree-sitter AST extractors in graphify/extractors/ — one per language (Python, JS/TS, Rust, Go, Java, C++, etc.). These produce nodes representing classes, functions, imports, and variables, and edges for calls, imports, inherits, references, etc. The LLM semantic pass uses a detailed system prompt (_EXTRACTION_SYSTEM in llm.py:481–511) that defines a JSON schema with nodes, edges, and hyperedges. Every edge is tagged with a confidence tier — EXTRACTED, INFERRED, or AMBIGUOUS — and each node carries a file_type of code, document, paper, image, rationale, or concept. Injection sentinels in source text are neutralised (_neutralise_injection_sentinels, llm.py:575–582). Fabrication is mitigated via _bind_node_evidence (llm.py:703–767): code-typed nodes whose symbol name has no substring match in the file bytes the model actually read are flagged with verification = "unverified".

Entity resolution / deduplication. The dedup.py module implements a pipeline: deduplicate_entities() (dedup.py:557–657) runs exact-ID dedup first (one survivor per node ID, deterministically ranked by _collision_rank), then fuzzy label-based dedup via a cascade: exact normalization → entropy gate (_entropy ≥ 2.5, dedup.py:29–38) → MinHash/LSH blocking at threshold 0.7 (dedup.py:48–53) → Jaro-Winkler verification at 92.0 (dedup.py:247) → same-community boost (+5.0) → union-find merge. An optional LLM pass (dedup_llm_backend) resolves ambiguous pairs in the 75–92 Jaro-Winkler zone.

Schema / ontology. There is no formal ontology. The type vocabulary is limited to file_type on nodes and relation on edges. The ID scheme follows a {path}_{entity} pattern where the path is the repo-relative file path with extension dropped, all segments joined by underscores (llm.py:500).

HKUDS/LightRAG

answered

Chunking. Documents are split into overlapping token-sized windows by the default fixed-token chunker (chunking_by_token_size / chunking_by_fixed_token in lightrag/chunker/token_size.py), which segments on a character delimiter if provided then windows at a configurable chunk_token_size (default 1200) + chunk_overlap_token_size (default 100). Alternative strategies — chunking_by_recursive_character (R), chunking_by_semantic_vector (V), and chunking_by_paragraph_semantic (P) — are selectable per document via process_options in the file pipeline (lightrag/chunker/__init__.py).

Entity and relation extraction. Each chunk is sent to the LLM (role "extract", configured by role_llm_funcs at lightrag/llm_roles.py:52-57) through extract_entities in lightrag/operate.py:3942. Two prompt formats exist: the default text-delimiter format (lightrag/prompt.py:56-207) using tuple_delimiter (<|#|>) and completion_delimiter (<|COMPLETE|>) separators, or JSON structured output (lightrag/prompt.py:175-295) when entity_extraction_use_json is set. Both extract entity_name, entity_type (from a guided type taxonomy with 11 classes like Person, Organization, Concept, etc.), and entity_description, plus binary (source, target, keywords, description) relations. A gleaning step (entity_extract_max_gleaning > 0) re-queries the LLM with the original prompt + prior response to find missed or corrected entities, merging by preferring longer descriptions (lightrag/operate.py:4290-4363).

Entity resolution / deduplication. Entities extracted from different chunks for the same document are merged in merge_nodes_and_edges (lightrag/operate.py:3514). Edges are deduplicated by sorting their endpoint tuple and using a seen set (lightrag/operate.py:3652-3654). On upsert to the graph, the graph store's upsert_node and upsert_edge in NetworkX (lightrag/kg/networkx_impl.py:801-845) are idempotent — calling graph.add_node(id, **data) on an existing node overwrites its attributes. There is no cross-document entity resolution or explicit coreference resolution beyond the LLM's per-prompt consistency instructions ("Ensure consistent naming across the entire extraction process"). Entity names are normalized with normalize_entity_name from lightrag/utils.py.

Schema / ontology. Not user-extensible in the code itself; the 11 entity types (Person, Creature, Organization, Location, Event, Concept, Method, Content, Data, Artifact, NaturalObject) are defined in PROMPTS["default_entity_types_guidance"] at lightrag/prompt.py:20-34 and can be overridden via addon_params['entity_types_guidance']. Relations are unlabeled binary tuples with keyword tags and free-text descriptions.

Editor's note. Correction: entities do merge across documents. _merge_nodes_then_upsert reads the existing node by its normalized name and merges into it, with the entity type picked by majority vote. What is missing is fuzzy or alias resolution. Gleaning runs at most once, whatever entity_extract_max_gleaning is set to.

microsoft/graphrag

answered

Chunking. Documents are split by TokenChunker with configurable token size and overlap, or by SentenceChunker. Chunks are written to the text_units table via create_base_text_units, with a SHA-512 hash as ID and a token count. Optional document metadata can be prepended to each chunk.

Extraction prompt. The prompt at extract_graph.py:6-126 asks the LLM to identify entities of specified types (ORGANIZATION, PERSON, GEO, etc.) with name, type, and description, then identify pairwise relationships with description and numeric strength. Output is a delimited format with ## separators and <|COMPLETE|> termination.

Multi-round gleaning. GraphExtractor (at graph_extractor.py:38) extracts entities per chunk. A gleaning loop (lines 99-120) re-prompts the model up to max_gleanings times with CONTINUE_PROMPT to catch missed items, using LOOP_PROMPT as a stopping gate.

Merging. Results from all chunks are grouped by exact-match on (title, type) for entities and (source, target) for relationships (see _merge_entities and _merge_relationships at extract_graph.py:104-129). The finalize_graph workflow computes undirected node degrees from the deduplicated edge set.

Summarization. Raw descriptions are re-summarized by a second LLM via the prompt at summarize_descriptions.py:6-20, limited to max_summary_length words.

Schema. Final columns defined in schemas.py:70-159: entities = (id, title, type, description, text_unit_ids, frequency, degree); relationships = (id, source, target, description, weight, combined_degree, text_unit_ids).

semantica-agi/semantica

answered

Graph construction flows through the GraphBuilder.build() method in semantica/kg/graph_builder.py. It processes sources that can be raw text, pre-extracted Entity/Relation objects, or dicts with entities/relationships keys. For raw text, the builder calls _extract_from_text which runs three extractors in sequence: NERExtractor (default method "ml", using spaCy; alternatives: "pattern", "regex", "huggingface", "llm"), RelationExtractor (default "pattern", alternatives up to "llm"), and TripletExtractor (default "pattern"). The NERExtractor (semantica/semantic_extract/ner_extractor.py) supports fallback chains — if the primary method returns nothing, it falls through pattern matching to a capitalized-word heuristic as a last resort. The GraphChunker module (semantica/split/kg_chunkers.py) provides EntityAwareChunker, RelationAwareChunker, GraphBasedChunker, OntologyAwareChunker, and HierarchicalChunker — which can preserve entity boundaries or triplet integrity during chunking. LLM enhancement is additive: LLMExtraction.enhance_entities() and enhance_relations() call an LLM provider with a structured prompt, merge results back by matching entity text or exact triples, and always preserve originals on failure (semantica/semantic_extract/llm_extraction.py:192-285). Entity resolution is done by EntityResolver (if merge_entities=True) using fuzzy/exact/ML strategies, and EntityMerger in semantica/deduplication/entity_merger.py detects duplicate groups via DuplicateDetector, then applies configurable strategies (keep_first, keep_most_complete, merge_all, etc.) with provenance tracking. After resolution, _remap_relationship_endpoints() rewrites relation source/target IDs to the surviving canonical entity ID (graph_builder.py:370-454). Conflict detection via ConflictDetector (optional, resolve_conflicts=True) runs after the graph is built and can auto-resolve conflicts. The output is a dict with entities, relationships, and metadata. Temporal edges are supported via add_temporal_edge(). No fixed schema/ontology is enforced at build time; entity types and relation predicates are free-form strings from the extractors, although the ontology/ module provides schema mapping post-hoc.

neo4j-labs/llm-graph-builder

answered

Chunking. Source documents are split by LangChain's TokenTextSplitter (create_chunks.py:42) with configurable token_chunk_size and chunk_overlap. Non-Neo4j users are capped at MAX_TOKEN_CHUNK_SIZE / token_chunk_size chunks (create_chunks.py:78-80). Chunks receive SHA-1 content-derived IDs stored as Chunk nodes with text, position, length, content_offset properties (make_relationships.py:60-146).

Entity and relation extraction. Chunks are combined in groups of chunks_to_combine and sent to the LLM via get_graph_from_llm() (llm.py:249). This calls LangChain's LLMGraphTransformer (llm.py:222-230) which instructs the LLM to return entities (as nodes) and relationships (as edges) from the chunk text. When the LLM supports structured output (tool calling), node_properties and relationship_properties are set to ["description"] and ignore_tool_usage=False (llm.py:211-220); for models that do not (e.g. Groq), property extraction is disabled and the raw text tool mode is used instead. Built-in additional_instructions (constants.py:885-888) tell the LLM to treat dates/numbers as properties, not separate nodes. For Diffbot, a dedicated DiffbotGraphTransformer is used (llm.py:126-130).

Entity resolution / deduplication. Entities are merged via apoc.merge.node using their id property (make_relationships.py:29), which prevents duplicate creation within a batch. An explicit duplicate-detection Cypher query (graphDB_dataAccess.py:473-518) uses embedding cosine similarity, substring containment, and Levenshtein distance (apoc.text.distance) to find near-duplicates, then merges them via apoc.refactor.mergeNodes (graphDB_dataAccess.py:530-538). A graph_schema_consolidation function (post_processing.py:149-186) uses an LLM to semantically merge similar node labels and relationship types.

Schema / ontology. An optional allowedNodes (comma-separated node labels) and allowedRelationships (triples of source, relation, target) constrain what the LLM extracts (llm.py:257-276). The schema can be pre-extracted from a sample text via populate_graph_schema_from_text (main.py:931), which uses structured-output LLM calls to produce triplet lists (schema_extraction.py:61-87).

OSU-NLP-Group/HippoRAG

answered

Chunking. Documents pass through a BaseTextPreprocessor before indexing. The default TextPreprocessor (src/hipporag/preprocessing.py:18–27) wraps each string as one Chunk without splitting. Configurable chunking is available via BaseConfig fields like preprocess_chunk_max_token_size, preprocess_chunk_overlap_token_size, and preprocess_chunk_func (src/hipporag/utils/config_utils.py:105–117) but the base class is abstract — users supply a custom preprocessor.

Entity and relation extraction. Two-stage OpenIE: first NER, then triple extraction. The NER prompt (src/hipporag/prompts/templates/ner.py:1–22) system-asks "extract named entities… respond with a JSON list of entities" with a one-shot example outputting {"named_entities": [...]}. The triple prompt (src/hipporag/prompts/templates/triple_extraction.py:4–50) system-instructs "construct an RDF graph… respond with a JSON list of triples", conditioned on the NER output. Both use the same LLM (default gpt-4o-mini). batch_openie (src/hipporag/information_extraction/openie_openai.py:186–286) runs NER and triple extraction concurrently via ThreadPoolExecutor with openie_max_workers (default 8) threads.

Entity resolution / deduplication. Entities are normalized by computing an MD5 hash of their lowercased, punctuation-stripped text (compute_mdhash_id(content, prefix="entity-")). This is a string-hash deduplication: two textually identical entity strings map to the same node. There is no fuzzy matching or LLM-based coreference; synonyms are discovered separately via embedding similarity (see synonymy edges).

Graph construction. From the extracted triples, extract_entity_nodes derives entity strings and flatten_facts collects all triples (HippoRAG.py:563–565). add_fact_edges (1179–1228) creates edges between entity nodes, weighted by frequency and tracked by source chunk via _fact_edge_source_counts. add_passage_edges (1230–1276) links each chunk node to the entity nodes appearing in it. add_synonymy_edges (1278–1352) runs KNN (synonymy_edge_topk=2047, threshold 0.8) over entity embeddings to link similar entities. The graph schema uses typed edges storing weight, edge_kind (fact/passage/synonym combinations), fact_source_counts, synonym_score, and passage_source (HippoRAG.py:1588–1596).

Editor's note. Correction: batch_openie does not run NER and triple extraction concurrently; it runs NER for all chunks in one thread pool, then triple extraction (conditioned on the NER output) in a second pool, and any failed chunk aborts the batch.

gusye1234/nano-graphrag

answered

Chunking. Documents are split via chunking_by_token_size (_op.py:31-58), the default strategy: each doc is tokenized (tiktoken or HuggingFace), then a sliding window of chunk_token_size (default 1200) tokens with chunk_overlap_token_size (default 100) overlap produces chunks. An alternative chunking_by_seperators method is also implemented, splitting on sentence/paragraph boundaries. Each chunk gets an md5-hash ID prefixed "chunk-".

Entity and relation extraction. The primary extraction function extract_entities (_op.py:282-414) sends each chunk through the LLM designated as best_model_func (default gpt-4o). The prompt entity_extraction (prompt.py:195-294) asks the model to output ("entity"<|>NAME<|>TYPE<|>DESC) and ("relationship"<|>SRC<|>TGT<|>DESC<|>STRENGTH) tuples separated by ## delimiters. Default entity types are organization, person, geo, event. Gleaning runs entity_extract_max_gleaning (default 1) additional passes: the model receives its prior output and a continuation prompt, then decides YES/NO whether more entities remain via entiti_if_loop_extraction (prompt.py:320-322). An alternative DSPy-based pipeline exists in entity_extraction/module.py using dspy.ChainOfThought with Pydantic schemas, and can optionally run critique-then-refine cycles.

Entity resolution / deduplication. _merge_nodes_then_upsert (_op.py:182-227) is called per unique entity name across all chunks. It loads any existing node from the graph, merges descriptions (sorted, deduplicated, joined by <SEP>), picks the majority entity type, aggregates source IDs, and calls _handle_entity_relation_summary which conditionally uses the cheap_model_func to condense descriptions that exceed entity_summary_to_max_tokens (500). Edges are merged similarly in _merge_edges_then_upsert (_op.py:230-279), summing weights, taking the minimum order, and merging descriptions. If an edge mentions nodes not yet in the graph, they are auto-created as type "UNKNOWN". The undirected graph is enforced by sorting edge tuples before storage.

Schema. There is no formal ontology beyond the four default entity types; the prompt passes entity types as a comma-separated string. Node attributes: entity_type, description, source_id. Edge attributes: weight, description, source_id, order.

pingcap/autoflow

answered

Graph construction is a two-phase DSPy-based LLM extraction pipeline operating per chunk.

Chunking: Documents are split via LlamaIndex's SentenceSplitter (plain text, default 1024 tokens with 20-token overlap) or a custom MarkdownNodeParser (markdown), configured per knowledge base (backend/app/rag/build_index.py:83-133). Each chunk becomes a TextNode stored in a per-KB chunks_{namespace} TiDB table (backend/app/models/chunk.py:53-61, fields: text, meta, embedding, document_id, index_status).

Extraction (core library — core/autoflow/knowledge_graph/): The SimpleKGExtractor (core/autoflow/knowledge_graph/extractors/simple.py:11-23) runs two DSPy programs in sequence. First, KnowledgeGraphExtractor (core/autoflow/knowledge_graph/programs/extract_graph.py:118-147) calls an LLM with the ExtractKnowledgeGraph signature (extract_graph.py:85-115) — a structured DSPy Signature whose prompt asks to identify "all meaningful entities and relationships" from database documentation, consolidating similar entities and capturing directionality. The LLM returns a PredictKnowledgeGraph (lists of PredictEntity and PredictRelationship). Second, EntityCovariateExtractor (core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85) re-reads the text with the entity list and enriches each entity's .meta with a covariate JSON tree (topic + attributes).

Extraction (backend — backend/app/rag/indices/knowledge_graph/extractor.py): The SimpleGraphExtractor does the same but additionally creates fallback entities for any source/target mentioned in a relationship that wasn't in the entity list, marking them "need-revised" (extractor.py:173-208).

Entity resolution / deduplication: In TiDBGraphStore.get_or_create_entity() (backend/.../tidb_graph_store.py:367-466), before inserting, the store computes the cosine distance between the new entity's description embedding and existing entities with the same name. If distance is below a threshold (description_similarity_threshold, default 0.9 → max cosine distance 0.1), it considers them the same. If the descriptions/metadata differ, it calls MergeEntities — a DSPy program (tidb_graph_store.py:53-81) that asks the LLM to decide whether two same-named entities are genuinely identical and, if so, merge their descriptions and metadata. Otherwise a new row is created. Relationships also use description-vector similarity to resolve source/target entities (tidb_graph_store.py:231-249).

Schema/Ontology: There is no fixed ontology. The graph is an open, untyped property graph. Entities have name, description, meta (covariates dict), and an optional entity_type (original or synopsis). Relationships have source_entity, target_entity, description, weight, and meta (including chunk_id/document_id).

trustgraph-ai/trustgraph

answered

Chunking. Documents are split by a recursive character splitter (chunking/recursive/chunker.py:76-77) that tries separators `

, , , then empty string (character fallback), with configurable chunk-size (2000) and chunk-overlap (100). ChunkingService (base/chunking_service.py:28-38`) registers these as parameters.

Multiple parallel extractors. After chunking, several processors consume Chunk messages concurrently:

  1. OntoRAG ontology-based extraction (extract/kg/ontology/extract.py): TextProcessor (text_processor.py:158-196) splits the chunk into sentences via NLTK and extracts noun/verb phrases using POS tagging. These segments are embedded and matched against pre-embedded ontology elements via OntologySelector (ontology_selector.py:107-166) using FAISS cosine similarity (threshold 0.3, top-k 10). Below 5 elements the selector is bypassed. Selected ontology subsets are passed to an LLM prompt (extract-with-ontologies) returning entities, relationships, and attributes in JSONL. TripleConverter (triple_converter.py:54-82) converts output to RDF triples, validating domain/range constraints and expanding URIs.

  2. Definitions extraction (extract/kg/definitions/extract.py:158-185): Calls prompt.extract_definitions(text=chunk) for entity-definition pairs.

  3. Relationships extraction (extract/kg/relationships/extract.py:110-117): Calls prompt.extract_relationships(text=chunk) for SPO triples.

  4. Agent-based extraction (extract/kg/agent/extract.py:178-215): Uses an Agent API with configurable templates.

Entity resolution. The OntoRAG pipeline uses EntityRegistry (entity_normalizer.py:113-165) mapping (entity_name, entity_type) to consistent normalized URIs per session. No cross-document deduplication exists.

Schema / ontology. Defined per-workspace via config push (extract/kg/ontology/extract.py:252-331). OntologyClass and OntologyProperty dataclasses (ontology_loader.py:15-85) capture OWL-like schema with labels, subclass-of, domain, range, inverse-of, and cardinality.

Editor's note. Correction: entity IRIs are deterministic (lower-cased, hyphenated name in the definitions/relationships extractors; ontology id + class + name in OntoRAG), so identical names from different documents merge into one node across the collection. What is missing is fuzzy or embedding-based resolution of aliases and variants, not cross-document deduplication.

zilliztech/vector-graph-rag

answered

Documents are processed through a three-stage pipeline: extraction, deduplication, and embedding.

Chunking is handled by TextChunker (loaders/chunker.py:19-98) which splits documents using semantic separators (paragraphs, sentences) at a default chunk size of 1000 characters with 200-character overlap.

Triplet extraction uses TripletExtractor (llm/extractor.py:88-250), which calls an OpenAI GPT model (default gpt-4o-mini) with a system prompt and one-shot example asking for JSON output of [subject, predicate, object] arrays. The prompt uses response_format: {type: json_object} (llm/extractor.py:155). Each document chunk gets one LLM call; triplets are stored in Document.metadata["triplets"]. Users can supply pre-extracted triplets to skip the LLM call.

Deduplication is done by GraphBuilder (graph/builder.py:25-215). Entity names are normalized via processing_phrases() (llm/extractor.py:22-33), which replaces non-alphanumeric characters with spaces and lowercases. A entity_name_to_id dict maps normalized names to UUIDs so repeated entities across chunks resolve to the same node. Relations are deduplicated by their full normalized text (relation_text_to_id).

Building the graph happens in GraphBuilder._process_documents() (graph/builder.py:136-157): it iterates triplets, calls _add_relation() which calls _add_entity() for subject and object, building six adjacency maps (entity-to-relation, entity-to-passage, relation-to-entity, relation-to-passage, passage-to-entity, passage-to-relation).

No formal schema or ontology is enforced beyond the subject-predicate-object structure. There is no community detection, no hierarchical summary, and no graph database (Neo4j etc).

Where and how is the graph stored? →