LLMs Technical Reviews
Home / Graph RAG / llm-graph-builder

neo4j-labs/llm-graph-builder

FastAPI and React app that turns files, URLs and transcripts into a Neo4j knowledge graph via LLMGraphTransformer, plus GraphRAG chat.

GitHub ↗★ 5.3kJupyter NotebookApache-2.0commit 5ff7af3 · 2026-07-30homepage ↗

Overview

LLM Graph Builder is Neo4j Labs’ end-to-end application for turning unstructured content into a knowledge graph in Neo4j and then chatting with it. It is a full product, not a library. A React frontend lets you connect a Neo4j database, upload PDFs and documents, or point at S3, GCS, web pages, Wikipedia or YouTube. You pick an LLM and optional schema, watch extraction progress, inspect the graph, and ask questions in a chat panel with several retrieval modes. A FastAPI backend (backend/score.py) does the work.

Extraction is delegated to LangChain’s LLMGraphTransformer (or Diffbot’s transformer). The app’s own contribution is the plumbing around it: a “lexical graph” of Document and Chunk nodes with embeddings and NEXT_CHUNK/SIMILAR links, HAS_ENTITY edges from chunks to extracted entities, resumable per-document processing, post-processing jobs (KNN similarity, full-text indexes, entity embeddings, LLM schema consolidation, GDS Leiden communities with LLM summaries), and a set of Cypher retrieval queries for chat.

Everything lives in Neo4j: graph, vectors, full-text indexes, chat history and, optionally, per-user token usage. The project fits teams that already run Neo4j and want a working GraphRAG demo or a starting point. It is less suited as an embeddable component, because the logic is spread across a 1,200-line FastAPI module, a large constants file of Cypher strings, and environment variables.

Architecture

flowchart LR
  UI["React frontend"] --> API["FastAPI score.py"]
  API --> SRC["document_sources: file, S3, GCS, web, YouTube, Wikipedia"]
  SRC --> CH["CreateChunksofDocument (TokenTextSplitter)"]
  CH --> EMB["create_chunk_embeddings"]
  CH --> LLM["get_graph_from_llm: LLMGraphTransformer"]
  LLM --> SAVE["save graph docs + HAS_ENTITY"]
  EMB --> NEO["Neo4j: Document, Chunk, __Entity__"]
  SAVE --> NEO
  API --> PP["/post_processing jobs"]
  PP --> COM["GDS Leiden + community summaries"]
  COM --> NEO
  API --> QA["QA_RAG: chat modes"]
  QA --> NEO
Component Path Role
API server backend/score.py FastAPI routes: /upload, /url/scan, /extract, /post_processing, /chat_bot, duplicates, schema, metrics
Orchestration backend/src/main.py Source loading, chunk-list building, batched processing_source / processing_chunks loop, retries
Sources backend/src/document_sources/ Loaders for local files, S3, GCS, web pages, YouTube transcripts, Wikipedia
Chunking backend/src/create_chunks.py TokenTextSplitter, page numbers and YouTube timestamps
Extraction backend/src/llm.py get_llm from LLM_MODEL_CONFIG_* env vars; LLMGraphTransformer / Diffbot wrapper
Graph writes backend/src/make_relationships.py, graphDB_dataAccess.py Chunk nodes and embeddings, HAS_ENTITY, KNN SIMILAR, duplicates, deletes, indexes
Communities backend/src/communities.py GDS projection, Leiden, community levels, LLM summaries, community embeddings
Post-processing backend/src/post_processing.py Full-text and vector indexes, entity embeddings, LLM label consolidation
Chat backend/src/QA_integration.py, src/shared/constants.py Retrieval modes, Cypher retrieval queries, prompts, chat history
Frontend frontend/ React 18 + Neo4j NDL/NVL components for upload, graph view and chat

How a request flows

Uploading a PDF and clicking “Generate Graph”:

  1. Upload and register. /upload stores the file in parts and creates a Document node. /extract then picks a loader by source_type (score.py).
  2. Chunk. processing_source creates the chunk vector index and calls get_chunkId_chunkDoc_list. That function uses CreateChunksofDocument.split_file_into_chunks with LangChain’s TokenTextSplitter, keeping page numbers or YouTube timestamps (create_chunks.py).
  3. Claim and batch. The document is claimed with a status guard so it is not processed twice. Chunks are then processed in batches of UPDATE_GRAPH_CHUNKS_PROCESSED (default 20). Before each batch, the loop checks a cancel flag (main.py).
  4. Embed and extract. For each batch, processing_chunks embeds every chunk (one embed_query call per chunk) and merges it as (:Chunk)-[:PART_OF]->(:Document). It then calls get_graph_from_llm (main.py, make_relationships.py).
  5. LLM transform. get_graph_from_llm concatenates every chunks_to_combine chunks into one LangChain Document. It parses allowedNodes and allowedRelationship (source, type, target triples), then runs LLMGraphTransformer.aconvert_to_graph_documents, which makes one LLM call per combined document. Models with structured output also extract a description property on nodes and relationships. Groq and non-tool models fall back to prompt-only mode (llm.py, L195-L292).
  6. Write. save_graphDocuments_in_neo4j writes graph documents through Neo4jGraph.add_graph_documents(..., baseEntityLabel=True), so every entity also gets the __Entity__ label, and retries on deadlocks. merge_relationship_between_chunk_and_entites then links each source chunk to its entities with apoc.merge.node on {id} plus MERGE (c)-[:HAS_ENTITY]->(n) (make_relationships.py). Entity resolution is exact match on label plus id string.
  7. Post-process. The frontend then calls /post_processing with a task list: materialize_text_chunk_similarities (KNN SIMILAR edges), hybrid and full-text indexes, entity embeddings (only with ENTITY_EMBEDDING=true), graph_schema_consolidation, and enable_communities (score.py).

Key components

Communities

create_communities drops existing communities and projects every non-Chunk, non-Document, non-Community node into GDS, weighting edges by count. It then runs gds.leiden.write with intermediate communities, maxLevels=3 and minCommunitySize=1 (communities.py, L232-L247). Cypher creates __Community__ nodes per level with IN_COMMUNITY and PARENT_COMMUNITY edges, plus rank and weight. create_community_summaries asks the LLM for a title and summary of each base community, then summarises parents from their children, in a thread pool (communities.py). Summaries are embedded and indexed for the global_vector chat mode. Communities are rebuilt only when the job runs, not on each upload.

Chat retrieval modes

QA_RAG loads the chat history from Neo4j. For graph mode it uses GraphCypherQAChain (text-to-Cypher). Every other mode goes through a Neo4jVector retriever configured from CHAT_MODE_CONFIG_MAP (QA_integration.py, constants.py):

  • vector and fulltext: chunk vector search, optionally hybrid with the keyword full-text index.
  • graph_vector and graph_vector_fulltext (default): chunk search, then a neighbourhood expansion from each chunk’s top entities. Entities with no embedding, or with query similarity between 0.3 and 0.9, get one hop (20 paths). Entities above 0.9 get two hops (40 paths). Entities below 0.3 contribute only themselves (constants.py).
  • entity_vector: entity vector search, then each entity’s chunks, communities and relationships.
  • global_vector: hybrid search over community summaries.

Retrieved documents pass an EmbeddingsFilter, are cut to a per-model token limit, and fill a single answer prompt. Modes without a document filter refuse to run while documents are selected in the UI.

Model configuration

get_llm(model) upper-cases the model key and reads LLM_MODEL_CONFIG_<KEY>, a comma-separated string of model name, key and sometimes base URL. It then dispatches on substrings (GEMINI, OPENAI, AZURE, ANTHROPIC, FIREWORKS, GROQ, BEDROCK, OLLAMA, DIFFBOT) to the matching LangChain chat class. Any other key is treated as an OpenAI-compatible endpoint given as model,base_url,api_key (llm.py, L119-L140). Embeddings default to sentence-transformers all-MiniLM-L6-v2. OpenAI, Gemini/Vertex and Bedrock Titan are alternatives.

Duplicate entities

get_duplicate_nodes_list finds same-label entity pairs by substring containment, edit distance under 3, or embedding cosine above 0.97. It only considers nodes that have an embedding. The UI shows the candidates, and the user approves a merge through apoc.refactor.mergeNodes (graphDB_dataAccess.py). This is a manual cleanup step, not part of ingestion.

Extending it

  • Models. Add an LLM_MODEL_CONFIG_<name> env var and the name to the frontend’s model list. A name without a known provider substring is treated as an OpenAI-compatible server, which covers vLLM, LiteLLM and similar setups without code changes.
  • Schema. Pass allowedNodes and allowedRelationship per extraction, generate them from sample text (/populate_graph_schema), or add additional_instructions, which are sanitised and appended to the transformer prompt.
  • Sources. Each loader in document_sources/ returns LangChain Document pages. A new source needs a loader, a main.py entry point and a branch in /extract.
  • Retrieval. Chat modes are data. Add a Cypher retrieval query and a CHAT_MODE_CONFIG_MAP entry to define a new one.

Running it

  • Prerequisite: a Neo4j 5.23+ database with APOC (Aura or self-hosted). GDS is required only for communities. docker-compose.yml starts only the backend and frontend, not Neo4j.
  • docker compose up with a .env that sets the LLM_MODEL_CONFIG_* keys, the embedding provider and the frontend’s model list. You can also run the backend with uvicorn and the frontend with Vite.
  • Optional: TRACK_USER_USAGE with a separate token-tracker database enforces daily and monthly token limits per user. AUTHENTICATION_REQUIRED enforces login.

Strengths and caveats

  • Strength: complete product. Ingestion from many sources, progress tracking, cancel and retry, graph visualisation, schema tools and chat ship together. It is the fastest way in this category to see GraphRAG working on your own documents in Neo4j.
  • Strength: graph and vectors in one store. Chunks, entities and communities carry embeddings and indexes in Neo4j, so hybrid Cypher retrieval needs no second system.
  • Strength: resumable processing. Per-document status, chunk-position resume and entity-safe deletes make re-runs manageable.
  • Caveat: hidden chunk cap. Unless the user’s email ends in @neo4j.com, a document is cut to MAX_TOKEN_CHUNK_SIZE / token_chunk_size chunks (10,000 tokens’ worth by default). The extra chunks are dropped with only a log line (create_chunks.py). Self-hosters must raise the env var.
  • Caveat: weak entity resolution. Entities merge only on identical label plus id. Fuzzy deduplication is a manual UI step that depends on optional entity embeddings.
  • Caveat: no extraction cache. Every run re-sends chunk text to the LLM, and chunk embeddings are computed one call per chunk.
  • Caveat: argument mix-up in the community job. At this commit, /post_processing calls create_communities(uri, user, password, database, email, embedding_provider, embedding_model) positionally. The signature expects model in the sixth position, so the embedding provider name is passed as the summary LLM key (score.py, communities.py). Unless an LLM_MODEL_CONFIG_ entry exists under that name, summary generation fails, and the error is only logged.
  • Caveat: configuration sprawl. Behaviour depends on many environment variables, and prompts and retrieval live in a 900-line constants file. Customising it means editing the app, not configuring a library.

Sources: code at 5ff7af3, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (12 pages), verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit 5ff7af3. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Chunking. Source documents are split by LangChain's TokenTextSplitter (create_chunks.py:42) with configurable token_chunk_size and chunk_overlap. Non-Neo4j users are capped at MAX_TOKEN_CHUNK_SIZE / token_chunk_size chunks (create_chunks.py:78-80). Chunks receive SHA-1 content-derived IDs stored as Chunk nodes with text, position, length, content_offset properties (make_relationships.py:60-146).

Entity and relation extraction. Chunks are combined in groups of chunks_to_combine and sent to the LLM via get_graph_from_llm() (llm.py:249). This calls LangChain's LLMGraphTransformer (llm.py:222-230) which instructs the LLM to return entities (as nodes) and relationships (as edges) from the chunk text. When the LLM supports structured output (tool calling), node_properties and relationship_properties are set to ["description"] and ignore_tool_usage=False (llm.py:211-220); for models that do not (e.g. Groq), property extraction is disabled and the raw text tool mode is used instead. Built-in additional_instructions (constants.py:885-888) tell the LLM to treat dates/numbers as properties, not separate nodes. For Diffbot, a dedicated DiffbotGraphTransformer is used (llm.py:126-130).

Entity resolution / deduplication. Entities are merged via apoc.merge.node using their id property (make_relationships.py:29), which prevents duplicate creation within a batch. An explicit duplicate-detection Cypher query (graphDB_dataAccess.py:473-518) uses embedding cosine similarity, substring containment, and Levenshtein distance (apoc.text.distance) to find near-duplicates, then merges them via apoc.refactor.mergeNodes (graphDB_dataAccess.py:530-538). A graph_schema_consolidation function (post_processing.py:149-186) uses an LLM to semantically merge similar node labels and relationship types.

Schema / ontology. An optional allowedNodes (comma-separated node labels) and allowedRelationships (triples of source, relation, target) constrain what the LLM extracts (llm.py:257-276). The schema can be pre-extracted from a sample text via populate_graph_schema_from_text (main.py:931), which uses structured-output LLM calls to produce triplet lists (schema_extraction.py:61-87).

Where and how is the graph stored?

answered

Graph database (Neo4j). All graph data is stored in a Neo4j 5.23+ instance with APOC. The langchain_neo4j.Neo4jGraph wrapper is used for writes (common_fn.py:281-295) and a raw neo4j.GraphDatabase.driver for reads (graph_query.py:26-37).

Node/edge schema. The core labels are:

  • :Document {fileName, fileSize, fileType, fileSource, status, model, url, createdAt, nodeCount, relationshipCount, ...} — one per uploaded file.
  • :Chunk {id, text, position, length, fileName, content_offset, page_number?, start_time?, end_time?, embedding} — text fragments with vector embeddings.
  • :__Entity__ (base label via baseEntityLabel=True, common_fn.py:285) plus user-defined sub-labels (e.g. :Person:__Entity__, :Organization:__Entity__).
  • :__Community__ {id, level, summary, title, embedding, community_rank, weight} — hierarchical community summaries (communities.py).

Relationships include PART_OF (Chunk→Document), FIRST_CHUNK / NEXT_CHUNK (Document→Chunk / Chunk→Chunk), HAS_ENTITY (Chunk→Entity), SIMILAR (Chunk↔Chunk, KNN-based), IN_COMMUNITY (Entity→Community), and PARENT_COMMUNITY (Community→Community for hierarchy).

How embeddings sit next to the graph. Three Neo4j vector indexes exist (graphDB_dataAccess.py:551-583; post_processing.py:41-45): vector on Chunk.embedding, entity_vector on __Entity__.embedding, and community_vector on __Community__.embedding. All use cosine similarity. Embeddings are generated by a pluggable model (OpenAI, Gemini, Bedrock Titan, or sentence-transformers/all-MiniLM-L6-v2) and stored as array properties directly on the node, not in a separate vector store. A keyword full-text index on Chunk.text and a community_keyword full-text index on __Community__.summary support hybrid search. The KNN update (graphDB_dataAccess.py:184-195) places SIMILAR relationships between chunks whose embedding cosine similarity meets the KNN_MIN_SCORE threshold.

Are communities, summaries or hierarchies built over the graph?

answered

Community detection. Communities are built using the Neo4j Graph Data Science (GDS) library's Leiden algorithm (communities.py:232-247). write_communities() calls gds.leiden.write() on a named graph projection, with includeIntermediateCommunities=True, up to MAX_COMMUNITY_LEVELS=3 hierarchical levels, and minCommunitySize=1. The projection is built over all non-Chunk/non-Document/non-Community nodes and their relationships, with an undirected adjacency weight based on edge count (communities.py:20-34). The Leiden partition IDs are written as a communities array property on each __Entity__ node.

Hierarchical community creation. create_community_properties() (communities.py:468-499) executes a multi-step Cypher pipeline:

  1. A uniqueness constraint on __Community__.id.
  2. CREATE_COMMUNITY_LEVELS creates __Community__ nodes at level 0 (from lowest Leiden partition), then level 1 and 2 parent groups, with PARENT_COMMUNITY edges linking child→parent.
  3. Ranks (distinct document count) and weights (chunk count) are computed per community via CREATE_COMMUNITY_RANKS/CREATE_COMMUNITY_WEIGHTS and their parent variants.

Summaries. create_community_summaries() (communities.py:313-371) generates LLM summaries for each level-0 community and then for parent communities (by summarizing child summaries). It uses a LangChain ChatPromptTemplate | llm | StrOutputParser pipeline, processing communities concurrently via ThreadPoolExecutor (max 10 workers). The LLM receives the community's subgraph (node IDs, types, descriptions and relationship triples) and returns a title and natural-language summary. Summaries are stored as summary and title properties on __Community__ nodes via STORE_COMMUNITY_SUMMARIES.

Embeddings and indexes for communities are generated in create_community_embeddings() (communities.py:374-403) using the configured embedding model in batches of 100. Full-text (community_keyword) and vector (community_vector) indexes are also created for hybrid search.

When they are computed. Communities are a post-extraction step invoked via the create_communities() entry point (communities.py:519-533), which first clears existing communities, creates a new GDS graph projection, runs Leiden, then builds all properties and summaries. This is triggered manually — not automatically on each document ingestion.

How does query-time retrieval use the graph?

answered

Query-time retrieval is handled by QA_RAG() (QA_integration.py:710-748), which supports multiple modes configured in CHAT_MODE_CONFIG_MAP (constants.py:718-780).

Local vector mode (vector, fulltext). Uses a Neo4jVector retriever on the vector index (Chunk.embedding). VECTOR_SEARCH_QUERY (constants.py:302-326) retrieves top-k matching chunks, their documents, and associated chunk details. For fulltext, a keyword full-text index is added for hybrid vector+keyword search.

Graph-augmented modes (graph_vector, graph_vector_fulltext). VECTOR_GRAPH_SEARCH_QUERY (constants.py:329-507) first retrieves top-k chunks, then for each chunk's entities, performs a 2-hop neighborhood traversal. Entity embedding similarity to the query vector determines traversal depth: entities with mid-range scores (0.3-0.9) get 1-hop neighbors (limit 20), entities with high scores (>0.9) get 2-hop neighbors (limit 40). This produces a combined context of chunk text plus entity/relationship subgraph, assembled into a structured text block with Text Content:, Entities: and Relationships: sections (constants.py:482-498).

Entity vector mode (entity_vector). LOCAL_COMMUNITY_SEARCH_QUERY (constants.py:515-558) retrieves top-k __Entity__ nodes by vector similarity, then fetches their associated chunks (up to 3), communities (up to 3), internal relationships, and outside entity connections (up to 10). This gives a local neighborhood view around matching entities.

Global community mode (global_vector). GLOBAL_VECTOR_SEARCH_QUERY (constants.py:681-695) searches the community_vector index (Community.summary), combining full-text (via community_keyword index) and vector search. It returns community summaries as the context, providing a document-set overview rather than chunk-level detail.

Graph (Cypher) mode (graph). Uses GraphCypherQAChain (QA_integration.py:562-583) which generates a Cypher query from the user's natural language question, executes it against Neo4j, and then answers with the results via a QA LLM.

Context assembly. In all RAG modes, retrieved documents pass through format_documents() (QA_integration.py:181-228) which sorts by similarity score, truncates to a model-specific token cutoff (CHAT_TOKEN_CUT_OFF, constants.py:249-253), formats each as "Document start / source / content / Document end", and collects source metadata, entity IDs, and community IDs. The final context is passed to a ChatPromptTemplate with the CHAT_SYSTEM_TEMPLATE (constants.py:256-297). Default mode is graph_vector_fulltext (CHAT_DEFAULT_MODE, constants.py:716).

How are updates and incremental indexing handled?

answered

Retry conditions for existing documents. The system supports three retry modes per document (constants.py:823-825): start_from_beginning (reprocesses all chunks), delete_entities_and_start_from_beginning (deletes entity nodes unique to this document then reprocesses), and start_from_last_processed_position (resumes from the last chunk that lacks an embedding or has no HAS_ENTITY relationship). These are handled in get_chunkId_chunkDoc_list() (main.py:685-744), which either re-splits the document into chunks or reuses existing Chunk nodes from Neo4j. For resumption, it queries QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION (constants.py:801-808) to find the first chunk without an embedding, then processes from there. set_status_retry() (main.py:945-981) resets document counters and optionally deletes entities via QUERY_TO_DELETE_EXISTING_ENTITIES (constants.py:791-799) which only deletes entities not referenced by other documents.

Adding new documents. Each document is independent — uploading a new file creates a new :Document node and processes it end-to-end. There is no overall corpus index that requires rebuilding. claim_document_for_processing() (graphDB_dataAccess.py:334-360) uses an atomic WHERE d.status <> 'Processing' guard so concurrent requests don't duplicate work.

Deletion. delete_file_from_graph() (graphDB_dataAccess.py:362-428) removes a document and optionally its unique entities (those not referenced by other docs) or just the document+chunks. Orphaned __Community__ nodes are also cleaned up via query_to_delete_communities.

Caching of extraction results. There is no extraction-result cache at the LLM level; every call to get_graph_from_llm() sends chunk text to the LLM. However, chunk embeddings are not re-computed when resuming: create_chunk_embeddings() sets c.embedding on each Chunk node, and the resume logic (QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION) skips chunks that already have embeddings. The optional GCS_FILE_CACHE (main.py:56-60) caches uploaded files in Google Cloud Storage rather than locally, but this is file-level storage, not extraction caching.

How are LLM cost and latency controlled during indexing and query?

answered

Caching. There is no disk or database caching of LLM extraction outputs; every get_graph_from_llm() call sends chunk text to the LLM. At the file level, the GCS_FILE_CACHE option (main.py:56-60) stores uploaded files in GCS instead of a local temp directory, which is a storage cache, not an LLM cache.

Batching. Chunks are processed in configurable batches controlled by UPDATE_GRAPH_CHUNKS_PROCESSED environment variable (default 20, main.py:515). Within each batch, multiple chunks are combined via get_combined_chunks() (llm.py:158-182) into a single document with configurable chunks_to_combine — meaning up to UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine chunks are sent to the LLM in one extraction call, reducing the number of LLM invocations proportionally.

Model choice per stage. Different models can be assigned to different pipeline stages. The default community-creation model is a smaller/cheaper model (openai_gpt_5.4_mini, communities.py:17). The GRAPH_CLEANUP_MODEL env var (post_processing.py:157) controls which LLM merges node labels. The chat response also uses a model-per-deployment while the extraction model is configured per-source. For embeddings, the default is the free/small sentence-transformers/all-MiniLM-L6-v2 (384-dim) (common_fn.py:719).

Token budgets. When TRACK_USER_USAGE is enabled (main.py:504), the system enforces preconfigured daily (DAILY_TOKENS_LIMIT, default 250k) and monthly (MONTHLY_TOKENS_LIMIT, default 1M) token limits per user (common_fn.py:491-492). Before processing begins, track_token_usage() with operation_type="precheck" (common_fn.py:460-477) checks if the user has exceeded their limits and raises LLMGraphBuilderException if so. After each batch, actual token usage is persisted to a User node in a separate token-tracker Neo4j database.

Non-LLM shortcuts. The KNN SIMILAR relationship between chunks (graphDB_dataAccess.py:184-195) is computed purely via vector cosine similarity in Cypher (no LLM involved). Entity deduplication (graphDB_dataAccess.py:470-538) uses embedding similarity, substring matching, and Levenshtein distance — not an LLM. The local/global search modes that return raw chunk text or entity subgraphs without summarization avoid per-query LLM overhead for retrieval.

Embedding filter at query. The EmbeddingsFilter in the query pipeline (QA_integration.py:316-320) uses a similarity_threshold (default 0.10, constants.py:247) and TokenTextSplitter to pre-filter retrieved documents before the LLM sees them, reducing LLM context length and cost.

Editor's note. Correction: batching does not put UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine chunks into one LLM call. Each extraction call covers chunks_to_combine concatenated chunks, so a batch of 20 chunks costs 20 / chunks_to_combine LLM calls; the batch size only controls how often progress and embeddings are written.