# How is the knowledge graph extracted from documents?

> Graph RAG — a good answer covers: Chunking; entity and relation extraction prompts or models; entity resolution / deduplication; schema or ontology.

Canonical page: https://llms-technical-reviews.com/graph-rag/q/graph-construction/

## Verdict

Most of these tools have an LLM read each chunk and then merge entities only when their normalized names match exactly. The two exceptions are [AutoFlow](/p/autoflow/), which compares description embeddings and asks an LLM before it merges, and [graphify](/p/graphify/), which merges near-identical names with fuzzy matching.

**GraphRAG-style delimited prompts.** [GraphRAG](/p/graphrag/) merges entities by `(title, type)`, so one name with two types becomes two nodes. It has four default entity types, and its `fast` method swaps the LLM for noun-phrase extraction. [nano-graphrag](/p/nano-graphrag/) uses the same prompt family and upper-cases names. Missing relationship endpoints become `UNKNOWN` nodes. [LightRAG](/p/lightrag/) merges by name, picks the type by majority vote, runs gleaning at most once, and calls the LLM to summarize descriptions only after 8 fragments pile up.

**Triples and OpenIE.** [HippoRAG](/p/hipporag/) runs NER on every chunk first, then triple extraction based on that output. By default it does not chunk: one document becomes one passage. It links aliases with synonymy edges (similarity of 0.8 or more) and does not merge them. [Vector Graph RAG](/p/vector-graph-rag/) makes one JSON-mode call per 1,000-character chunk. Its name normalizer drops every character outside `[A-Za-z0-9 ]`. [TrustGraph](/p/trustgraph/) runs several extractors on each chunk, including ontology-guided OntoRAG, which checks domain and range. Its IRIs are built from the name, so the same name always maps to the same node.

**Schema-driven and framework pipelines.** [LLM Graph Builder](/p/llm-graph-builder/) wraps LangChain's `LLMGraphTransformer` and accepts optional allowed node and relationship lists. Fuzzy merging happens only when a user approves it in the UI. Unless the user's email ends in `@neo4j.com`, chunks beyond `MAX_TOKEN_CHUNK_SIZE` are silently dropped. AutoFlow makes two DSPy calls per chunk. [Semantica](/p/semantica/) defaults to spaCy NER plus pattern triplets. Relation extraction is off unless `extract_relations=True`, and its conflict step only counts conflicts and changes nothing.

**Code first.** graphify parses code with tree-sitter and no LLM. It uses an LLM only for documents, PDFs and images, and merges names with MinHash plus Jaro-Winkler (threshold 92).

Pick: GraphRAG or LightRAG for general prose with tunable prompts.
Pick: TrustGraph or LLM Graph Builder when a schema or ontology must constrain extraction.
Pick: graphify for codebases, and AutoFlow when duplicate entities are the main worry.

## Per-project answers

### Graphify-Labs/graphify (answered)

**Chunking.** Non-code files (docs, PDFs, images) are processed by the LLM-based semantic pipeline in `llm.py`. Files are first split at `_FILE_CHAR_CAP` (20,000 characters) into `FileSlice` units via `expand_oversized_files` (llm.py:2638). These units are then packed into chunks by `_pack_chunks_by_tokens()` (llm.py:2160–2203), using a greedy algorithm that groups by parent directory (so related files share one LLM call) and closes a chunk when a configurable token budget (default 60,000 tokens) would be exceeded, with a hard cap of 20 images per chunk.

**Entity and relation extraction.** Code files are handled entirely deterministically via tree-sitter AST extractors in `graphify/extractors/` — one per language (Python, JS/TS, Rust, Go, Java, C++, etc.). These produce nodes representing classes, functions, imports, and variables, and edges for `calls`, `imports`, `inherits`, `references`, etc. The LLM semantic pass uses a detailed system prompt (`_EXTRACTION_SYSTEM` in llm.py:481–511) that defines a JSON schema with nodes, edges, and hyperedges. Every edge is tagged with a confidence tier — `EXTRACTED`, `INFERRED`, or `AMBIGUOUS` — and each node carries a `file_type` of `code`, `document`, `paper`, `image`, `rationale`, or `concept`. Injection sentinels in source text are neutralised (`_neutralise_injection_sentinels`, llm.py:575–582). Fabrication is mitigated via `_bind_node_evidence` (llm.py:703–767): code-typed nodes whose symbol name has no substring match in the file bytes the model actually read are flagged with `verification = "unverified"`.

**Entity resolution / deduplication.** The `dedup.py` module implements a pipeline: `deduplicate_entities()` (dedup.py:557–657) runs exact-ID dedup first (one survivor per node ID, deterministically ranked by `_collision_rank`), then fuzzy label-based dedup via a cascade: exact normalization → entropy gate (`_entropy` ≥ 2.5, dedup.py:29–38) → MinHash/LSH blocking at threshold 0.7 (dedup.py:48–53) → Jaro-Winkler verification at 92.0 (dedup.py:247) → same-community boost (+5.0) → union-find merge. An optional LLM pass (`dedup_llm_backend`) resolves ambiguous pairs in the 75–92 Jaro-Winkler zone.

**Schema / ontology.** There is no formal ontology. The type vocabulary is limited to `file_type` on nodes and `relation` on edges. The ID scheme follows a `{path}_{entity}` pattern where the path is the repo-relative file path with extension dropped, all segments joined by underscores (llm.py:500).


Citations: [graphify/llm.py:2160-2203](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L2160-L2203) · [graphify/llm.py:481-511](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L481-L511) · [graphify/dedup.py:557-657](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/dedup.py#L557-L657) · [graphify/llm.py:703-767](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/llm.py#L703-L767) · [graphify/dedup.py:245-250](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/dedup.py#L245-L250) · [graphify/extract.py:1-10](https://github.com/Graphify-Labs/graphify/blob/5c7b84792f453582676548185aaec3824d51dfe2/graphify/extract.py#L1-L10)

### HKUDS/LightRAG (answered)

**Chunking.** Documents are split into overlapping token-sized windows by the default fixed-token chunker (`chunking_by_token_size` / `chunking_by_fixed_token` in `lightrag/chunker/token_size.py`), which segments on a character delimiter if provided then windows at a configurable `chunk_token_size` (default 1200) + `chunk_overlap_token_size` (default 100). Alternative strategies — `chunking_by_recursive_character` (R), `chunking_by_semantic_vector` (V), and `chunking_by_paragraph_semantic` (P) — are selectable per document via `process_options` in the file pipeline (`lightrag/chunker/__init__.py`).

**Entity and relation extraction.** Each chunk is sent to the LLM (role `"extract"`, configured by `role_llm_funcs` at `lightrag/llm_roles.py:52-57`) through `extract_entities` in `lightrag/operate.py:3942`. Two prompt formats exist: the default text-delimiter format (`lightrag/prompt.py:56-207`) using `tuple_delimiter` (`<|#|>`) and `completion_delimiter` (`<|COMPLETE|>`) separators, or JSON structured output (`lightrag/prompt.py:175-295`) when `entity_extraction_use_json` is set. Both extract `entity_name`, `entity_type` (from a guided type taxonomy with 11 classes like Person, Organization, Concept, etc.), and `entity_description`, plus binary `(source, target, keywords, description)` relations. A **gleaning** step (`entity_extract_max_gleaning > 0`) re-queries the LLM with the original prompt + prior response to find missed or corrected entities, merging by preferring longer descriptions (`lightrag/operate.py:4290-4363`).

**Entity resolution / deduplication.** Entities extracted from different chunks for the same document are merged in `merge_nodes_and_edges` (`lightrag/operate.py:3514`). Edges are deduplicated by sorting their endpoint tuple and using a `seen` set (`lightrag/operate.py:3652-3654`). On upsert to the graph, the graph store's `upsert_node` and `upsert_edge` in NetworkX (`lightrag/kg/networkx_impl.py:801-845`) are idempotent — calling `graph.add_node(id, **data)` on an existing node overwrites its attributes. There is no cross-document entity resolution or explicit coreference resolution beyond the LLM's per-prompt consistency instructions ("Ensure consistent naming across the entire extraction process"). Entity names are normalized with `normalize_entity_name` from `lightrag/utils.py`.

**Schema / ontology.** Not user-extensible in the code itself; the 11 entity types (Person, Creature, Organization, Location, Event, Concept, Method, Content, Data, Artifact, NaturalObject) are defined in `PROMPTS["default_entity_types_guidance"]` at `lightrag/prompt.py:20-34` and can be overridden via `addon_params['entity_types_guidance']`. Relations are unlabeled binary tuples with keyword tags and free-text descriptions.

> **Editor's note.** Correction: entities do merge across documents. `_merge_nodes_then_upsert` reads the existing node by its normalized name and merges into it, with the entity type picked by majority vote. What is missing is fuzzy or alias resolution. Gleaning runs at most once, whatever `entity_extract_max_gleaning` is set to.

Citations: [lightrag/operate.py:3942-3965](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3942-L3965) · [lightrag/prompt.py:56-117](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L56-L117) · [lightrag/prompt.py:20-34](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/prompt.py#L20-L34) · [lightrag/chunker/token_size.py:133-180](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/chunker/token_size.py#L133-L180) · [lightrag/operate.py:3514-3545](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L3514-L3545) · [lightrag/operate.py:4290-4365](https://github.com/HKUDS/LightRAG/blob/453dce83d6d0354a06e46c8d4029a0895c4e054b/lightrag/operate.py#L4290-L4365)

### microsoft/graphrag (answered)

**Chunking.** Documents are split by `TokenChunker` with configurable token size and overlap, or by `SentenceChunker`. Chunks are written to the `text_units` table via `create_base_text_units`, with a SHA-512 hash as ID and a token count. Optional document metadata can be prepended to each chunk.

**Extraction prompt.** The prompt at `extract_graph.py:6-126` asks the LLM to identify entities of specified types (ORGANIZATION, PERSON, GEO, etc.) with name, type, and description, then identify pairwise relationships with description and numeric strength. Output is a delimited format with `##` separators and `<|COMPLETE|>` termination.

**Multi-round gleaning.** `GraphExtractor` (at `graph_extractor.py:38`) extracts entities per chunk. A gleaning loop (lines 99-120) re-prompts the model up to `max_gleanings` times with `CONTINUE_PROMPT` to catch missed items, using `LOOP_PROMPT` as a stopping gate.

**Merging.** Results from all chunks are grouped by exact-match on `(title, type)` for entities and `(source, target)` for relationships (see `_merge_entities` and `_merge_relationships` at `extract_graph.py:104-129`). The `finalize_graph` workflow computes undirected node degrees from the deduplicated edge set.

**Summarization.** Raw descriptions are re-summarized by a second LLM via the prompt at `summarize_descriptions.py:6-20`, limited to `max_summary_length` words.

**Schema.** Final columns defined in `schemas.py:70-159`: entities = (id, title, type, description, text_unit_ids, frequency, degree); relationships = (id, source, target, description, weight, combined_degree, text_unit_ids).


Citations: [packages/graphrag-chunking/graphrag_chunking/token_chunker.py:14-69](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag-chunking/graphrag_chunking/token_chunker.py#L14-L69) · [packages/graphrag/graphrag/index/workflows/create_base_text_units.py:56-124](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/workflows/create_base_text_units.py#L56-L124) · [packages/graphrag/graphrag/prompts/index/extract_graph.py:6-126](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/prompts/index/extract_graph.py#L6-L126) · [packages/graphrag/graphrag/index/operations/extract_graph/graph_extractor.py:38-188](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/operations/extract_graph/graph_extractor.py#L38-L188) · [packages/graphrag/graphrag/index/operations/extract_graph/extract_graph.py:104-129](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/index/operations/extract_graph/extract_graph.py#L104-L129) · [packages/graphrag/graphrag/data_model/schemas.py:70-159](https://github.com/microsoft/graphrag/blob/769542fbf1d8e5b4c6a8677fefc34621c87894c5/packages/graphrag/graphrag/data_model/schemas.py#L70-L159)

### semantica-agi/semantica (answered)

Graph construction flows through the `GraphBuilder.build()` method in `semantica/kg/graph_builder.py`. It processes sources that can be raw text, pre-extracted Entity/Relation objects, or dicts with `entities`/`relationships` keys. For raw text, the builder calls `_extract_from_text` which runs three extractors in sequence: `NERExtractor` (default method `"ml"`, using spaCy; alternatives: `"pattern"`, `"regex"`, `"huggingface"`, `"llm"`), `RelationExtractor` (default `"pattern"`, alternatives up to `"llm"`), and `TripletExtractor` (default `"pattern"`). The NERExtractor (`semantica/semantic_extract/ner_extractor.py`) supports fallback chains — if the primary method returns nothing, it falls through pattern matching to a capitalized-word heuristic as a last resort. The `GraphChunker` module (`semantica/split/kg_chunkers.py`) provides `EntityAwareChunker`, `RelationAwareChunker`, `GraphBasedChunker`, `OntologyAwareChunker`, and `HierarchicalChunker` — which can preserve entity boundaries or triplet integrity during chunking. LLM enhancement is additive: `LLMExtraction.enhance_entities()` and `enhance_relations()` call an LLM provider with a structured prompt, merge results back by matching entity text or exact triples, and always preserve originals on failure (`semantica/semantic_extract/llm_extraction.py:192-285`). Entity resolution is done by `EntityResolver` (if `merge_entities=True`) using fuzzy/exact/ML strategies, and `EntityMerger` in `semantica/deduplication/entity_merger.py` detects duplicate groups via `DuplicateDetector`, then applies configurable strategies (`keep_first`, `keep_most_complete`, `merge_all`, etc.) with provenance tracking. After resolution, `_remap_relationship_endpoints()` rewrites relation source/target IDs to the surviving canonical entity ID (graph_builder.py:370-454). Conflict detection via `ConflictDetector` (optional, `resolve_conflicts=True`) runs after the graph is built and can auto-resolve conflicts. The output is a dict with `entities`, `relationships`, and `metadata`. Temporal edges are supported via `add_temporal_edge()`. No fixed schema/ontology is enforced at build time; entity types and relation predicates are free-form strings from the extractors, although the `ontology/` module provides schema mapping post-hoc.


Citations: [semantica/kg/graph_builder.py:1-60](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/graph_builder.py#L1-L60) · [semantica/kg/graph_builder.py:505-575](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/kg/graph_builder.py#L505-L575) · [semantica/semantic_extract/ner_extractor.py:1-80](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/semantic_extract/ner_extractor.py#L1-L80) · [semantica/semantic_extract/llm_extraction.py:189-285](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/semantic_extract/llm_extraction.py#L189-L285) · [semantica/split/kg_chunkers.py:1-50](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/split/kg_chunkers.py#L1-L50) · [semantica/deduplication/entity_merger.py:135-200](https://github.com/semantica-agi/semantica/blob/9a71df67bb50800bd74ebef25f4472726c571458/semantica/deduplication/entity_merger.py#L135-L200)

### neo4j-labs/llm-graph-builder (answered)

**Chunking.** Source documents are split by LangChain's `TokenTextSplitter` (`create_chunks.py:42`) with configurable `token_chunk_size` and `chunk_overlap`. Non-Neo4j users are capped at `MAX_TOKEN_CHUNK_SIZE / token_chunk_size` chunks (`create_chunks.py:78-80`). Chunks receive SHA-1 content-derived IDs stored as `Chunk` nodes with `text`, `position`, `length`, `content_offset` properties (`make_relationships.py:60-146`).

**Entity and relation extraction.** Chunks are combined in groups of `chunks_to_combine` and sent to the LLM via `get_graph_from_llm()` (`llm.py:249`). This calls LangChain's `LLMGraphTransformer` (`llm.py:222-230`) which instructs the LLM to return entities (as nodes) and relationships (as edges) from the chunk text. When the LLM supports structured output (tool calling), `node_properties` and `relationship_properties` are set to `["description"]` and `ignore_tool_usage=False` (`llm.py:211-220`); for models that do not (e.g. Groq), property extraction is disabled and the raw text tool mode is used instead. Built-in `additional_instructions` (`constants.py:885-888`) tell the LLM to treat dates/numbers as properties, not separate nodes. For Diffbot, a dedicated `DiffbotGraphTransformer` is used (`llm.py:126-130`).

**Entity resolution / deduplication.** Entities are merged via `apoc.merge.node` using their `id` property (`make_relationships.py:29`), which prevents duplicate creation within a batch. An explicit duplicate-detection Cypher query (`graphDB_dataAccess.py:473-518`) uses embedding cosine similarity, substring containment, and Levenshtein distance (`apoc.text.distance`) to find near-duplicates, then merges them via `apoc.refactor.mergeNodes` (`graphDB_dataAccess.py:530-538`). A `graph_schema_consolidation` function (`post_processing.py:149-186`) uses an LLM to semantically merge similar node labels and relationship types.

**Schema / ontology.** An optional `allowedNodes` (comma-separated node labels) and `allowedRelationships` (triples of source, relation, target) constrain what the LLM extracts (`llm.py:257-276`). The schema can be pre-extracted from a sample text via `populate_graph_schema_from_text` (`main.py:931`), which uses structured-output LLM calls to produce triplet lists (`schema_extraction.py:61-87`).


Citations: [backend/src/create_chunks.py:29-82](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/create_chunks.py#L29-L82) · [backend/src/llm.py:195-292](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L195-L292) · [backend/src/make_relationships.py:12-32](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L12-L32) · [backend/src/shared/constants.py:884-889](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L884-L889) · [backend/src/graphDB_dataAccess.py:470-538](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L470-L538) · [backend/src/shared/schema_extraction.py:61-87](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/schema_extraction.py#L61-L87)

### OSU-NLP-Group/HippoRAG (answered)

**Chunking.** Documents pass through a `BaseTextPreprocessor` before indexing. The default `TextPreprocessor` (src/hipporag/preprocessing.py:18–27) wraps each string as one `Chunk` without splitting. Configurable chunking is available via `BaseConfig` fields like `preprocess_chunk_max_token_size`, `preprocess_chunk_overlap_token_size`, and `preprocess_chunk_func` (src/hipporag/utils/config_utils.py:105–117) but the base class is abstract — users supply a custom preprocessor.

**Entity and relation extraction.** Two-stage OpenIE: first NER, then triple extraction. The NER prompt (src/hipporag/prompts/templates/ner.py:1–22) system-asks _"extract named entities… respond with a JSON list of entities"_ with a one-shot example outputting `{"named_entities": [...]}`. The triple prompt (src/hipporag/prompts/templates/triple_extraction.py:4–50) system-instructs _"construct an RDF graph… respond with a JSON list of triples"_, conditioned on the NER output. Both use the same LLM (default gpt-4o-mini). `batch_openie` (src/hipporag/information_extraction/openie_openai.py:186–286) runs NER and triple extraction concurrently via `ThreadPoolExecutor` with `openie_max_workers` (default 8) threads.

**Entity resolution / deduplication.** Entities are normalized by computing an MD5 hash of their lowercased, punctuation-stripped text (`compute_mdhash_id(content, prefix="entity-")`). This is a string-hash deduplication: two textually identical entity strings map to the same node. There is no fuzzy matching or LLM-based coreference; synonyms are discovered separately via embedding similarity (see synonymy edges).

**Graph construction.** From the extracted triples, `extract_entity_nodes` derives entity strings and `flatten_facts` collects all triples (HippoRAG.py:563–565). `add_fact_edges` (1179–1228) creates edges between entity nodes, weighted by frequency and tracked by source chunk via `_fact_edge_source_counts`. `add_passage_edges` (1230–1276) links each chunk node to the entity nodes appearing in it. `add_synonymy_edges` (1278–1352) runs KNN (`synonymy_edge_topk=2047`, threshold 0.8) over entity embeddings to link similar entities. The graph schema uses typed edges storing `weight`, `edge_kind` (fact/passage/synonym combinations), `fact_source_counts`, `synonym_score`, and `passage_source` (HippoRAG.py:1588–1596).

> **Editor's note.** Correction: `batch_openie` does not run NER and triple extraction concurrently; it runs NER for all chunks in one thread pool, then triple extraction (conditioned on the NER output) in a second pool, and any failed chunk aborts the batch.

Citations: [src/hipporag/preprocessing.py:18-27](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/preprocessing.py#L18-L27) · [src/hipporag/prompts/templates/ner.py:1-22](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/prompts/templates/ner.py#L1-L22) · [src/hipporag/prompts/templates/triple_extraction.py:4-50](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/prompts/templates/triple_extraction.py#L4-L50) · [src/hipporag/information_extraction/openie_openai.py:186-286](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/information_extraction/openie_openai.py#L186-L286) · [src/hipporag/HippoRAG.py:1179-1228](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1179-L1228) · [src/hipporag/HippoRAG.py:1278-1352](https://github.com/OSU-NLP-Group/HippoRAG/blob/2bfd831417202b49cda9da7973e141a456e16872/src/hipporag/HippoRAG.py#L1278-L1352)

### gusye1234/nano-graphrag (answered)

**Chunking.** Documents are split via `chunking_by_token_size` (`_op.py:31-58`), the default strategy: each doc is tokenized (tiktoken or HuggingFace), then a sliding window of `chunk_token_size` (default 1200) tokens with `chunk_overlap_token_size` (default 100) overlap produces chunks. An alternative `chunking_by_seperators` method is also implemented, splitting on sentence/paragraph boundaries. Each chunk gets an md5-hash ID prefixed `"chunk-"`.

**Entity and relation extraction.** The primary extraction function `extract_entities` (`_op.py:282-414`) sends each chunk through the LLM designated as `best_model_func` (default gpt-4o). The prompt `entity_extraction` (`prompt.py:195-294`) asks the model to output `("entity"<|>NAME<|>TYPE<|>DESC)` and `("relationship"<|>SRC<|>TGT<|>DESC<|>STRENGTH)` tuples separated by `##` delimiters. Default entity types are organization, person, geo, event. Gleaning runs `entity_extract_max_gleaning` (default 1) additional passes: the model receives its prior output and a continuation prompt, then decides YES/NO whether more entities remain via `entiti_if_loop_extraction` (`prompt.py:320-322`). An alternative DSPy-based pipeline exists in `entity_extraction/module.py` using dspy.ChainOfThought with Pydantic schemas, and can optionally run critique-then-refine cycles.

**Entity resolution / deduplication.** `_merge_nodes_then_upsert` (`_op.py:182-227`) is called per unique entity name across all chunks. It loads any existing node from the graph, merges descriptions (sorted, deduplicated, joined by `<SEP>`), picks the majority entity type, aggregates source IDs, and calls `_handle_entity_relation_summary` which conditionally uses the `cheap_model_func` to condense descriptions that exceed `entity_summary_to_max_tokens` (500). Edges are merged similarly in `_merge_edges_then_upsert` (`_op.py:230-279`), summing weights, taking the minimum order, and merging descriptions. If an edge mentions nodes not yet in the graph, they are auto-created as type `"UNKNOWN"`. The undirected graph is enforced by sorting edge tuples before storage.

**Schema.** There is no formal ontology beyond the four default entity types; the prompt passes entity types as a comma-separated string. Node attributes: `entity_type`, `description`, `source_id`. Edge attributes: `weight`, `description`, `source_id`, `order`.


Citations: [nano_graphrag/_op.py:31-58](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_op.py#L31-L58) · [nano_graphrag/_op.py:182-227](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_op.py#L182-L227) · [nano_graphrag/_op.py:282-414](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/_op.py#L282-L414) · [nano_graphrag/prompt.py:195-294](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/prompt.py#L195-L294) · [nano_graphrag/prompt.py:314-327](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/prompt.py#L314-L327) · [nano_graphrag/entity_extraction/module.py:236-330](https://github.com/gusye1234/nano-graphrag/blob/acb35c065614eb5a2f5f1be9a56b235f5a2e0a7a/nano_graphrag/entity_extraction/module.py#L236-L330)

### pingcap/autoflow (answered)

Graph construction is a two-phase DSPy-based LLM extraction pipeline operating per chunk.

**Chunking:** Documents are split via LlamaIndex's `SentenceSplitter` (plain text, default 1024 tokens with 20-token overlap) or a custom `MarkdownNodeParser` (markdown), configured per knowledge base (`backend/app/rag/build_index.py:83-133`). Each chunk becomes a `TextNode` stored in a per-KB `chunks_{namespace}` TiDB table (`backend/app/models/chunk.py:53-61`, fields: `text`, `meta`, `embedding`, `document_id`, `index_status`).

**Extraction (core library — `core/autoflow/knowledge_graph/`):** The `SimpleKGExtractor` (`core/autoflow/knowledge_graph/extractors/simple.py:11-23`) runs two DSPy programs in sequence. First, `KnowledgeGraphExtractor` (`core/autoflow/knowledge_graph/programs/extract_graph.py:118-147`) calls an LLM with the `ExtractKnowledgeGraph` signature (`extract_graph.py:85-115`) — a structured DSPy Signature whose prompt asks to identify "all meaningful entities and relationships" from database documentation, consolidating similar entities and capturing directionality. The LLM returns a `PredictKnowledgeGraph` (lists of `PredictEntity` and `PredictRelationship`). Second, `EntityCovariateExtractor` (`core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85`) re-reads the text with the entity list and enriches each entity's `.meta` with a covariate JSON tree (topic + attributes).

**Extraction (backend — `backend/app/rag/indices/knowledge_graph/extractor.py`):** The `SimpleGraphExtractor` does the same but additionally creates fallback entities for any source/target mentioned in a relationship that wasn't in the entity list, marking them `"need-revised"` (`extractor.py:173-208`).

**Entity resolution / deduplication:** In `TiDBGraphStore.get_or_create_entity()` (`backend/.../tidb_graph_store.py:367-466`), before inserting, the store computes the cosine distance between the new entity's description embedding and existing entities with the same name. If distance is below a threshold (`description_similarity_threshold`, default 0.9 → max cosine distance 0.1), it considers them the same. If the descriptions/metadata differ, it calls `MergeEntities` — a DSPy program (`tidb_graph_store.py:53-81`) that asks the LLM to decide whether two same-named entities are genuinely identical and, if so, merge their descriptions and metadata. Otherwise a new row is created. Relationships also use description-vector similarity to resolve source/target entities (`tidb_graph_store.py:231-249`).

**Schema/Ontology:** There is no fixed ontology. The graph is an open, untyped property graph. Entities have `name`, `description`, `meta` (covariates dict), and an optional `entity_type` (`original` or `synopsis`). Relationships have `source_entity`, `target_entity`, `description`, `weight`, and `meta` (including `chunk_id`/`document_id`).


Citations: [core/autoflow/knowledge_graph/programs/extract_covariates.py:56-85](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/knowledge_graph/programs/extract_covariates.py#L56-L85) · [backend/app/rag/indices/knowledge_graph/extractor.py:83-225](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/extractor.py#L83-L225) · [backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py:367-477](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/indices/knowledge_graph/graph_store/tidb_graph_store.py#L367-L477) · [backend/app/rag/build_index.py:82-158](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/backend/app/rag/build_index.py#L82-L158) · [core/autoflow/configs/chunkers/text.py:1-12](https://github.com/pingcap/autoflow/blob/c4cb19d8fa205bdd4cb38d0ac250d273fcc3e5f2/core/autoflow/configs/chunkers/text.py#L1-L12)

### trustgraph-ai/trustgraph (answered)

**Chunking.** Documents are split by a recursive character splitter (`chunking/recursive/chunker.py:76-77`) that tries separators `

`, `
`, ` `, then empty string (character fallback), with configurable chunk-size (2000) and chunk-overlap (100). `ChunkingService` (`base/chunking_service.py:28-38`) registers these as parameters.

**Multiple parallel extractors.** After chunking, several processors consume `Chunk` messages concurrently:

1. **OntoRAG ontology-based extraction** (`extract/kg/ontology/extract.py`): `TextProcessor` (`text_processor.py:158-196`) splits the chunk into sentences via NLTK and extracts noun/verb phrases using POS tagging. These segments are embedded and matched against pre-embedded ontology elements via `OntologySelector` (`ontology_selector.py:107-166`) using FAISS cosine similarity (threshold 0.3, top-k 10). Below 5 elements the selector is bypassed. Selected ontology subsets are passed to an LLM prompt (`extract-with-ontologies`) returning entities, relationships, and attributes in JSONL. `TripleConverter` (`triple_converter.py:54-82`) converts output to RDF triples, validating domain/range constraints and expanding URIs.

2. **Definitions extraction** (`extract/kg/definitions/extract.py:158-185`): Calls `prompt.extract_definitions(text=chunk)` for entity-definition pairs.

3. **Relationships extraction** (`extract/kg/relationships/extract.py:110-117`): Calls `prompt.extract_relationships(text=chunk)` for SPO triples.

4. **Agent-based extraction** (`extract/kg/agent/extract.py:178-215`): Uses an Agent API with configurable templates.

**Entity resolution.** The OntoRAG pipeline uses `EntityRegistry` (`entity_normalizer.py:113-165`) mapping `(entity_name, entity_type)` to consistent normalized URIs per session. No cross-document deduplication exists.

**Schema / ontology.** Defined per-workspace via config push (`extract/kg/ontology/extract.py:252-331`). `OntologyClass` and `OntologyProperty` dataclasses (`ontology_loader.py:15-85`) capture OWL-like schema with labels, subclass-of, domain, range, inverse-of, and cardinality.

> **Editor's note.** Correction: entity IRIs are deterministic (lower-cased, hyphenated name in the definitions/relationships extractors; ontology id + class + name in OntoRAG), so identical names from different documents merge into one node across the collection. What is missing is fuzzy or embedding-based resolution of aliases and variants, not cross-document deduplication.

Citations: [trustgraph-flow/trustgraph/chunking/recursive/chunker.py:76-120](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/chunking/recursive/chunker.py#L76-L120) · [trustgraph-flow/trustgraph/extract/kg/ontology/extract.py:349-435](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/extract.py#L349-L435) · [trustgraph-flow/trustgraph/extract/kg/ontology/entity_normalizer.py:72-153](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/entity_normalizer.py#L72-L153) · [trustgraph-flow/trustgraph/extract/kg/ontology/triple_converter.py:54-82](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/triple_converter.py#L54-L82) · [trustgraph-flow/trustgraph/extract/kg/ontology/text_processor.py:150-196](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/text_processor.py#L150-L196) · [trustgraph-flow/trustgraph/extract/kg/ontology/ontology_selector.py:76-107](https://github.com/trustgraph-ai/trustgraph/blob/ea19308aaeb8458b9d988794eff7d4c0016bd43f/trustgraph-flow/trustgraph/extract/kg/ontology/ontology_selector.py#L76-L107)

### zilliztech/vector-graph-rag (answered)

Documents are processed through a three-stage pipeline: extraction, deduplication, and embedding.

**Chunking** is handled by `TextChunker` (`loaders/chunker.py:19-98`) which splits documents using semantic separators (paragraphs, sentences) at a default chunk size of 1000 characters with 200-character overlap.

**Triplet extraction** uses `TripletExtractor` (`llm/extractor.py:88-250`), which calls an OpenAI GPT model (default `gpt-4o-mini`) with a system prompt and one-shot example asking for JSON output of `[subject, predicate, object]` arrays. The prompt uses `response_format: {type: json_object}` (`llm/extractor.py:155`). Each document chunk gets one LLM call; triplets are stored in `Document.metadata["triplets"]`. Users can supply pre-extracted triplets to skip the LLM call.

**Deduplication** is done by `GraphBuilder` (`graph/builder.py:25-215`). Entity names are normalized via `processing_phrases()` (`llm/extractor.py:22-33`), which replaces non-alphanumeric characters with spaces and lowercases. A `entity_name_to_id` dict maps normalized names to UUIDs so repeated entities across chunks resolve to the same node. Relations are deduplicated by their full normalized text (`relation_text_to_id`).

**Building the graph** happens in `GraphBuilder._process_documents()` (`graph/builder.py:136-157`): it iterates triplets, calls `_add_relation()` which calls `_add_entity()` for subject and object, building six adjacency maps (entity-to-relation, entity-to-passage, relation-to-entity, relation-to-passage, passage-to-entity, passage-to-relation).

**No formal schema or ontology** is enforced beyond the subject-predicate-object structure. There is no community detection, no hierarchical summary, and no graph database (Neo4j etc).


Citations: [src/vector_graph_rag/llm/extractor.py:22-33](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/llm/extractor.py#L22-L33) · [src/vector_graph_rag/llm/extractor.py:88-250](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/llm/extractor.py#L88-L250) · [src/vector_graph_rag/loaders/chunker.py:19-98](https://github.com/zilliztech/vector-graph-rag/blob/07acd794e2fc9ae14f34a0389db8e8a8e1079d2b/src/vector_graph_rag/loaders/chunker.py#L19-L98)
