# neo4j-labs/llm-graph-builder

> FastAPI and React app that turns files, URLs and transcripts into a Neo4j knowledge graph via LLMGraphTransformer, plus GraphRAG chat.

- Category: [Graph RAG](https://llms-technical-reviews.com/graph-rag/)
- Repository: https://github.com/neo4j-labs/llm-graph-builder (reviewed at commit `5ff7af3e9bb9226e1bbecd02f70f8d98697727a7`, 2026-07-30)
- Stars: 5274 · Language: Jupyter Notebook · License: Apache-2.0
- Canonical page: https://llms-technical-reviews.com/p/llm-graph-builder/

## Overview

LLM Graph Builder is Neo4j Labs' end-to-end application for turning unstructured content into a knowledge graph in Neo4j and then chatting with it. It is a full product, not a library. A React frontend lets you connect a Neo4j database, upload PDFs and documents, or point at S3, GCS, web pages, Wikipedia or YouTube. You pick an LLM and optional schema, watch extraction progress, inspect the graph, and ask questions in a chat panel with several retrieval modes. A FastAPI backend (`backend/score.py`) does the work.

Extraction is delegated to LangChain's `LLMGraphTransformer` (or Diffbot's transformer). The app's own contribution is the plumbing around it: a "lexical graph" of `Document` and `Chunk` nodes with embeddings and `NEXT_CHUNK`/`SIMILAR` links, `HAS_ENTITY` edges from chunks to extracted entities, resumable per-document processing, post-processing jobs (KNN similarity, full-text indexes, entity embeddings, LLM schema consolidation, GDS Leiden communities with LLM summaries), and a set of Cypher retrieval queries for chat.

Everything lives in Neo4j: graph, vectors, full-text indexes, chat history and, optionally, per-user token usage. The project fits teams that already run Neo4j and want a working GraphRAG demo or a starting point. It is less suited as an embeddable component, because the logic is spread across a 1,200-line FastAPI module, a large constants file of Cypher strings, and environment variables.

## Architecture

```mermaid
flowchart LR
  UI["React frontend"] --> API["FastAPI score.py"]
  API --> SRC["document_sources: file, S3, GCS, web, YouTube, Wikipedia"]
  SRC --> CH["CreateChunksofDocument (TokenTextSplitter)"]
  CH --> EMB["create_chunk_embeddings"]
  CH --> LLM["get_graph_from_llm: LLMGraphTransformer"]
  LLM --> SAVE["save graph docs + HAS_ENTITY"]
  EMB --> NEO["Neo4j: Document, Chunk, __Entity__"]
  SAVE --> NEO
  API --> PP["/post_processing jobs"]
  PP --> COM["GDS Leiden + community summaries"]
  COM --> NEO
  API --> QA["QA_RAG: chat modes"]
  QA --> NEO
```

| Component | Path | Role |
|---|---|---|
| API server | `backend/score.py` | FastAPI routes: `/upload`, `/url/scan`, `/extract`, `/post_processing`, `/chat_bot`, duplicates, schema, metrics |
| Orchestration | `backend/src/main.py` | Source loading, chunk-list building, batched `processing_source` / `processing_chunks` loop, retries |
| Sources | `backend/src/document_sources/` | Loaders for local files, S3, GCS, web pages, YouTube transcripts, Wikipedia |
| Chunking | `backend/src/create_chunks.py` | `TokenTextSplitter`, page numbers and YouTube timestamps |
| Extraction | `backend/src/llm.py` | `get_llm` from `LLM_MODEL_CONFIG_*` env vars; `LLMGraphTransformer` / Diffbot wrapper |
| Graph writes | `backend/src/make_relationships.py`, `graphDB_dataAccess.py` | Chunk nodes and embeddings, `HAS_ENTITY`, KNN `SIMILAR`, duplicates, deletes, indexes |
| Communities | `backend/src/communities.py` | GDS projection, Leiden, community levels, LLM summaries, community embeddings |
| Post-processing | `backend/src/post_processing.py` | Full-text and vector indexes, entity embeddings, LLM label consolidation |
| Chat | `backend/src/QA_integration.py`, `src/shared/constants.py` | Retrieval modes, Cypher retrieval queries, prompts, chat history |
| Frontend | `frontend/` | React 18 + Neo4j NDL/NVL components for upload, graph view and chat |

## How a request flows

Uploading a PDF and clicking "Generate Graph":

1. **Upload and register.** `/upload` stores the file in parts and creates a `Document` node. `/extract` then picks a loader by `source_type` ([score.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/score.py#L227-L260)).
2. **Chunk.** `processing_source` creates the chunk vector index and calls `get_chunkId_chunkDoc_list`. That function uses `CreateChunksofDocument.split_file_into_chunks` with LangChain's `TokenTextSplitter`, keeping page numbers or YouTube timestamps ([create_chunks.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/create_chunks.py#L29-L80)).
3. **Claim and batch.** The document is claimed with a status guard so it is not processed twice. Chunks are then processed in batches of `UPDATE_GRAPH_CHUNKS_PROCESSED` (default 20). Before each batch, the loop checks a cancel flag ([main.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L515-L536)).
4. **Embed and extract.** For each batch, `processing_chunks` embeds every chunk (one `embed_query` call per chunk) and merges it as `(:Chunk)-[:PART_OF]->(:Document)`. It then calls `get_graph_from_llm` ([main.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L637-L683), [make_relationships.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L35-L58)).
5. **LLM transform.** `get_graph_from_llm` concatenates every `chunks_to_combine` chunks into one LangChain `Document`. It parses `allowedNodes` and `allowedRelationship` (source, type, target triples), then runs `LLMGraphTransformer.aconvert_to_graph_documents`, which makes one LLM call per combined document. Models with structured output also extract a `description` property on nodes and relationships. Groq and non-tool models fall back to prompt-only mode ([llm.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L158-L182), [L195-L292](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L195-L292)).
6. **Write.** `save_graphDocuments_in_neo4j` writes graph documents through `Neo4jGraph.add_graph_documents(..., baseEntityLabel=True)`, so every entity also gets the `__Entity__` label, and retries on deadlocks. `merge_relationship_between_chunk_and_entites` then links each source chunk to its entities with `apoc.merge.node` on `{id}` plus `MERGE (c)-[:HAS_ENTITY]->(n)` ([make_relationships.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L12-L32)). Entity resolution is exact match on label plus `id` string.
7. **Post-process.** The frontend then calls `/post_processing` with a task list: `materialize_text_chunk_similarities` (KNN `SIMILAR` edges), hybrid and full-text indexes, entity embeddings (only with `ENTITY_EMBEDDING=true`), `graph_schema_consolidation`, and `enable_communities` ([score.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/score.py#L353-L413)).

## Key components

### Communities

`create_communities` drops existing communities and projects every non-Chunk, non-Document, non-Community node into GDS, weighting edges by count. It then runs `gds.leiden.write` with intermediate communities, `maxLevels=3` and `minCommunitySize=1` ([communities.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L11-L34), [L232-L247](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L232-L247)). Cypher creates `__Community__` nodes per level with `IN_COMMUNITY` and `PARENT_COMMUNITY` edges, plus rank and weight. `create_community_summaries` asks the LLM for a title and summary of each base community, then summarises parents from their children, in a thread pool ([communities.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L313-L372)). Summaries are embedded and indexed for the `global_vector` chat mode. Communities are rebuilt only when the job runs, not on each upload.

### Chat retrieval modes

`QA_RAG` loads the chat history from Neo4j. For `graph` mode it uses `GraphCypherQAChain` (text-to-Cypher). Every other mode goes through a `Neo4jVector` retriever configured from `CHAT_MODE_CONFIG_MAP` ([QA_integration.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/QA_integration.py#L710-L746), [constants.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L709-L780)):
- `vector` and `fulltext`: chunk vector search, optionally hybrid with the `keyword` full-text index.
- `graph_vector` and `graph_vector_fulltext` (default): chunk search, then a neighbourhood expansion from each chunk's top entities. Entities with no embedding, or with query similarity between 0.3 and 0.9, get one hop (20 paths). Entities above 0.9 get two hops (40 paths). Entities below 0.3 contribute only themselves ([constants.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L329-L370)).
- `entity_vector`: entity vector search, then each entity's chunks, communities and relationships.
- `global_vector`: hybrid search over community summaries.

Retrieved documents pass an `EmbeddingsFilter`, are cut to a per-model token limit, and fill a single answer prompt. Modes without a document filter refuse to run while documents are selected in the UI.

### Model configuration

`get_llm(model)` upper-cases the model key and reads `LLM_MODEL_CONFIG_<KEY>`, a comma-separated string of model name, key and sometimes base URL. It then dispatches on substrings (GEMINI, OPENAI, AZURE, ANTHROPIC, FIREWORKS, GROQ, BEDROCK, OLLAMA, DIFFBOT) to the matching LangChain chat class. Any other key is treated as an OpenAI-compatible endpoint given as `model,base_url,api_key` ([llm.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L23-L32), [L119-L140](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L119-L140)). Embeddings default to sentence-transformers `all-MiniLM-L6-v2`. OpenAI, Gemini/Vertex and Bedrock Titan are alternatives.

### Duplicate entities

`get_duplicate_nodes_list` finds same-label entity pairs by substring containment, edit distance under 3, or embedding cosine above 0.97. It only considers nodes that have an embedding. The UI shows the candidates, and the user approves a merge through `apoc.refactor.mergeNodes` ([graphDB_dataAccess.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L470-L538)). This is a manual cleanup step, not part of ingestion.

## Extending it

- **Models.** Add an `LLM_MODEL_CONFIG_<name>` env var and the name to the frontend's model list. A name without a known provider substring is treated as an OpenAI-compatible server, which covers vLLM, LiteLLM and similar setups without code changes.
- **Schema.** Pass `allowedNodes` and `allowedRelationship` per extraction, generate them from sample text (`/populate_graph_schema`), or add `additional_instructions`, which are sanitised and appended to the transformer prompt.
- **Sources.** Each loader in `document_sources/` returns LangChain `Document` pages. A new source needs a loader, a `main.py` entry point and a branch in `/extract`.
- **Retrieval.** Chat modes are data. Add a Cypher retrieval query and a `CHAT_MODE_CONFIG_MAP` entry to define a new one.

## Running it

- **Prerequisite:** a Neo4j 5.23+ database with APOC (Aura or self-hosted). GDS is required only for communities. `docker-compose.yml` starts only the backend and frontend, not Neo4j.
- `docker compose up` with a `.env` that sets the `LLM_MODEL_CONFIG_*` keys, the embedding provider and the frontend's model list. You can also run the backend with uvicorn and the frontend with Vite.
- Optional: `TRACK_USER_USAGE` with a separate token-tracker database enforces daily and monthly token limits per user. `AUTHENTICATION_REQUIRED` enforces login.

## Strengths and caveats

- **Strength: complete product.** Ingestion from many sources, progress tracking, cancel and retry, graph visualisation, schema tools and chat ship together. It is the fastest way in this category to see GraphRAG working on your own documents in Neo4j.
- **Strength: graph and vectors in one store.** Chunks, entities and communities carry embeddings and indexes in Neo4j, so hybrid Cypher retrieval needs no second system.
- **Strength: resumable processing.** Per-document status, chunk-position resume and entity-safe deletes make re-runs manageable.
- **Caveat: hidden chunk cap.** Unless the user's email ends in `@neo4j.com`, a document is cut to `MAX_TOKEN_CHUNK_SIZE / token_chunk_size` chunks (10,000 tokens' worth by default). The extra chunks are dropped with only a log line ([create_chunks.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/create_chunks.py#L43-L80)). Self-hosters must raise the env var.
- **Caveat: weak entity resolution.** Entities merge only on identical label plus `id`. Fuzzy deduplication is a manual UI step that depends on optional entity embeddings.
- **Caveat: no extraction cache.** Every run re-sends chunk text to the LLM, and chunk embeddings are computed one call per chunk.
- **Caveat: argument mix-up in the community job.** At this commit, `/post_processing` calls `create_communities(uri, user, password, database, email, embedding_provider, embedding_model)` positionally. The signature expects `model` in the sixth position, so the embedding provider name is passed as the summary LLM key ([score.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/score.py#L387-L390), [communities.py](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L519-L533)). Unless an `LLM_MODEL_CONFIG_` entry exists under that name, summary generation fails, and the error is only logged.
- **Caveat: configuration sprawl.** Behaviour depends on many environment variables, and prompts and retrieval live in a 900-line constants file. Customising it means editing the app, not configuring a library.

*Sources: code at 5ff7af3, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (12 pages), verified Q&A.*

## How neo4j-labs/llm-graph-builder answers the Graph RAG questions

### How is the knowledge graph extracted from documents? (answered)

**Chunking.** Source documents are split by LangChain's `TokenTextSplitter` (`create_chunks.py:42`) with configurable `token_chunk_size` and `chunk_overlap`. Non-Neo4j users are capped at `MAX_TOKEN_CHUNK_SIZE / token_chunk_size` chunks (`create_chunks.py:78-80`). Chunks receive SHA-1 content-derived IDs stored as `Chunk` nodes with `text`, `position`, `length`, `content_offset` properties (`make_relationships.py:60-146`).

**Entity and relation extraction.** Chunks are combined in groups of `chunks_to_combine` and sent to the LLM via `get_graph_from_llm()` (`llm.py:249`). This calls LangChain's `LLMGraphTransformer` (`llm.py:222-230`) which instructs the LLM to return entities (as nodes) and relationships (as edges) from the chunk text. When the LLM supports structured output (tool calling), `node_properties` and `relationship_properties` are set to `["description"]` and `ignore_tool_usage=False` (`llm.py:211-220`); for models that do not (e.g. Groq), property extraction is disabled and the raw text tool mode is used instead. Built-in `additional_instructions` (`constants.py:885-888`) tell the LLM to treat dates/numbers as properties, not separate nodes. For Diffbot, a dedicated `DiffbotGraphTransformer` is used (`llm.py:126-130`).

**Entity resolution / deduplication.** Entities are merged via `apoc.merge.node` using their `id` property (`make_relationships.py:29`), which prevents duplicate creation within a batch. An explicit duplicate-detection Cypher query (`graphDB_dataAccess.py:473-518`) uses embedding cosine similarity, substring containment, and Levenshtein distance (`apoc.text.distance`) to find near-duplicates, then merges them via `apoc.refactor.mergeNodes` (`graphDB_dataAccess.py:530-538`). A `graph_schema_consolidation` function (`post_processing.py:149-186`) uses an LLM to semantically merge similar node labels and relationship types.

**Schema / ontology.** An optional `allowedNodes` (comma-separated node labels) and `allowedRelationships` (triples of source, relation, target) constrain what the LLM extracts (`llm.py:257-276`). The schema can be pre-extracted from a sample text via `populate_graph_schema_from_text` (`main.py:931`), which uses structured-output LLM calls to produce triplet lists (`schema_extraction.py:61-87`).


Citations: [backend/src/create_chunks.py:29-82](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/create_chunks.py#L29-L82) · [backend/src/llm.py:195-292](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/llm.py#L195-L292) · [backend/src/make_relationships.py:12-32](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L12-L32) · [backend/src/shared/constants.py:884-889](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L884-L889) · [backend/src/graphDB_dataAccess.py:470-538](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L470-L538) · [backend/src/shared/schema_extraction.py:61-87](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/schema_extraction.py#L61-L87)

### Where and how is the graph stored? (answered)

**Graph database (Neo4j).** All graph data is stored in a Neo4j 5.23+ instance with APOC. The `langchain_neo4j.Neo4jGraph` wrapper is used for writes (`common_fn.py:281-295`) and a raw `neo4j.GraphDatabase.driver` for reads (`graph_query.py:26-37`).

**Node/edge schema.** The core labels are:
- `:Document {fileName, fileSize, fileType, fileSource, status, model, url, createdAt, nodeCount, relationshipCount, ...}` — one per uploaded file.
- `:Chunk {id, text, position, length, fileName, content_offset, page_number?, start_time?, end_time?, embedding}` — text fragments with vector embeddings.
- `:__Entity__` (base label via `baseEntityLabel=True`, `common_fn.py:285`) plus user-defined sub-labels (e.g. `:Person:__Entity__`, `:Organization:__Entity__`).
- `:__Community__ {id, level, summary, title, embedding, community_rank, weight}` — hierarchical community summaries (`communities.py`).

Relationships include `PART_OF` (Chunk→Document), `FIRST_CHUNK` / `NEXT_CHUNK` (Document→Chunk / Chunk→Chunk), `HAS_ENTITY` (Chunk→__Entity__), `SIMILAR` (Chunk↔Chunk, KNN-based), `IN_COMMUNITY` (__Entity__→__Community__), and `PARENT_COMMUNITY` (__Community__→__Community__ for hierarchy).

**How embeddings sit next to the graph.** Three Neo4j vector indexes exist (`graphDB_dataAccess.py:551-583`; `post_processing.py:41-45`): `vector` on `Chunk.embedding`, `entity_vector` on `__Entity__.embedding`, and `community_vector` on `__Community__.embedding`. All use cosine similarity. Embeddings are generated by a pluggable model (OpenAI, Gemini, Bedrock Titan, or sentence-transformers/all-MiniLM-L6-v2) and stored as array properties directly on the node, not in a separate vector store. A `keyword` full-text index on `Chunk.text` and a `community_keyword` full-text index on `__Community__.summary` support hybrid search. The KNN update (`graphDB_dataAccess.py:184-195`) places `SIMILAR` relationships between chunks whose embedding cosine similarity meets the `KNN_MIN_SCORE` threshold.


Citations: [backend/src/graphDB_dataAccess.py:540-585](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L540-L585) · [backend/src/post_processing.py:19-45](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/post_processing.py#L19-L45) · [backend/src/shared/constants.py:159-207](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L159-L207) · [backend/src/make_relationships.py:149-171](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/make_relationships.py#L149-L171) · [backend/src/graphDB_dataAccess.py:151-196](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L151-L196)

### Are communities, summaries or hierarchies built over the graph? (answered)

**Community detection.** Communities are built using the Neo4j Graph Data Science (GDS) library's Leiden algorithm (`communities.py:232-247`). `write_communities()` calls `gds.leiden.write()` on a named graph projection, with `includeIntermediateCommunities=True`, up to `MAX_COMMUNITY_LEVELS=3` hierarchical levels, and `minCommunitySize=1`. The projection is built over all non-Chunk/non-Document/non-__Community__ nodes and their relationships, with an undirected adjacency weight based on edge count (`communities.py:20-34`). The Leiden partition IDs are written as a `communities` array property on each `__Entity__` node.

**Hierarchical community creation.** `create_community_properties()` (`communities.py:468-499`) executes a multi-step Cypher pipeline:
1. A uniqueness constraint on `__Community__.id`.
2. `CREATE_COMMUNITY_LEVELS` creates `__Community__` nodes at level 0 (from lowest Leiden partition), then level 1 and 2 parent groups, with `PARENT_COMMUNITY` edges linking child→parent.
3. Ranks (distinct document count) and weights (chunk count) are computed per community via `CREATE_COMMUNITY_RANKS`/`CREATE_COMMUNITY_WEIGHTS` and their parent variants.

**Summaries.** `create_community_summaries()` (`communities.py:313-371`) generates LLM summaries for each level-0 community and then for parent communities (by summarizing child summaries). It uses a LangChain `ChatPromptTemplate | llm | StrOutputParser` pipeline, processing communities concurrently via `ThreadPoolExecutor` (max 10 workers). The LLM receives the community's subgraph (node IDs, types, descriptions and relationship triples) and returns a title and natural-language summary. Summaries are stored as `summary` and `title` properties on `__Community__` nodes via `STORE_COMMUNITY_SUMMARIES`.

**Embeddings and indexes** for communities are generated in `create_community_embeddings()` (`communities.py:374-403`) using the configured embedding model in batches of 100. Full-text (`community_keyword`) and vector (`community_vector`) indexes are also created for hybrid search.

**When they are computed.** Communities are a post-extraction step invoked via the `create_communities()` entry point (`communities.py:519-533`), which first clears existing communities, creates a new GDS graph projection, runs Leiden, then builds all properties and summaries. This is triggered manually — not automatically on each document ingestion.


Citations: [backend/src/communities.py:232-247](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L232-L247) · [backend/src/communities.py:37-84](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L37-L84) · [backend/src/communities.py:313-372](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L313-L372) · [backend/src/communities.py:468-500](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L468-L500) · [backend/src/communities.py:519-533](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L519-L533)

### How does query-time retrieval use the graph? (answered)

Query-time retrieval is handled by `QA_RAG()` (`QA_integration.py:710-748`), which supports multiple modes configured in `CHAT_MODE_CONFIG_MAP` (`constants.py:718-780`).

**Local vector mode** (`vector`, `fulltext`). Uses a `Neo4jVector` retriever on the `vector` index (Chunk.embedding). `VECTOR_SEARCH_QUERY` (`constants.py:302-326`) retrieves top-k matching chunks, their documents, and associated chunk details. For `fulltext`, a `keyword` full-text index is added for hybrid vector+keyword search.

**Graph-augmented modes** (`graph_vector`, `graph_vector_fulltext`). `VECTOR_GRAPH_SEARCH_QUERY` (`constants.py:329-507`) first retrieves top-k chunks, then for each chunk's entities, performs a 2-hop neighborhood traversal. Entity embedding similarity to the query vector determines traversal depth: entities with mid-range scores (0.3-0.9) get 1-hop neighbors (limit 20), entities with high scores (>0.9) get 2-hop neighbors (limit 40). This produces a combined context of chunk text plus entity/relationship subgraph, assembled into a structured text block with `Text Content:`, `Entities:` and `Relationships:` sections (`constants.py:482-498`).

**Entity vector mode** (`entity_vector`). `LOCAL_COMMUNITY_SEARCH_QUERY` (`constants.py:515-558`) retrieves top-k `__Entity__` nodes by vector similarity, then fetches their associated chunks (up to 3), communities (up to 3), internal relationships, and outside entity connections (up to 10). This gives a local neighborhood view around matching entities.

**Global community mode** (`global_vector`). `GLOBAL_VECTOR_SEARCH_QUERY` (`constants.py:681-695`) searches the `community_vector` index (__Community__.summary), combining full-text (via `community_keyword` index) and vector search. It returns community summaries as the context, providing a document-set overview rather than chunk-level detail.

**Graph (Cypher) mode** (`graph`). Uses `GraphCypherQAChain` (`QA_integration.py:562-583`) which generates a Cypher query from the user's natural language question, executes it against Neo4j, and then answers with the results via a QA LLM.

**Context assembly.** In all RAG modes, retrieved documents pass through `format_documents()` (`QA_integration.py:181-228`) which sorts by similarity score, truncates to a model-specific token cutoff (`CHAT_TOKEN_CUT_OFF`, `constants.py:249-253`), formats each as "Document start / source / content / Document end", and collects source metadata, entity IDs, and community IDs. The final context is passed to a `ChatPromptTemplate` with the `CHAT_SYSTEM_TEMPLATE` (`constants.py:256-297`). Default mode is `graph_vector_fulltext` (`CHAT_DEFAULT_MODE`, `constants.py:716`).


Citations: [backend/src/QA_integration.py:710-748](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/QA_integration.py#L710-L748) · [backend/src/shared/constants.py:302-508](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L302-L508) · [backend/src/shared/constants.py:514-608](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L514-L608) · [backend/src/shared/constants.py:680-705](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L680-L705) · [backend/src/shared/constants.py:718-780](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L718-L780) · [backend/src/QA_integration.py:181-228](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/QA_integration.py#L181-L228)

### How are updates and incremental indexing handled? (answered)

**Retry conditions for existing documents.** The system supports three retry modes per document (`constants.py:823-825`): `start_from_beginning` (reprocesses all chunks), `delete_entities_and_start_from_beginning` (deletes entity nodes unique to this document then reprocesses), and `start_from_last_processed_position` (resumes from the last chunk that lacks an embedding or has no `HAS_ENTITY` relationship). These are handled in `get_chunkId_chunkDoc_list()` (`main.py:685-744`), which either re-splits the document into chunks or reuses existing `Chunk` nodes from Neo4j. For resumption, it queries `QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION` (`constants.py:801-808`) to find the first chunk without an embedding, then processes from there. `set_status_retry()` (`main.py:945-981`) resets document counters and optionally deletes entities via `QUERY_TO_DELETE_EXISTING_ENTITIES` (`constants.py:791-799`) which only deletes entities not referenced by other documents.

**Adding new documents.** Each document is independent — uploading a new file creates a new `:Document` node and processes it end-to-end. There is no overall corpus index that requires rebuilding. `claim_document_for_processing()` (`graphDB_dataAccess.py:334-360`) uses an atomic `WHERE d.status <> 'Processing'` guard so concurrent requests don't duplicate work.

**Deletion.** `delete_file_from_graph()` (`graphDB_dataAccess.py:362-428`) removes a document and optionally its unique entities (those not referenced by other docs) or just the document+chunks. Orphaned `__Community__` nodes are also cleaned up via `query_to_delete_communities`.

**Caching of extraction results.** There is no extraction-result cache at the LLM level; every call to `get_graph_from_llm()` sends chunk text to the LLM. However, chunk embeddings are not re-computed when resuming: `create_chunk_embeddings()` sets `c.embedding` on each `Chunk` node, and the resume logic (`QUERY_TO_GET_LAST_PROCESSED_CHUNK_POSITION`) skips chunks that already have embeddings. The optional `GCS_FILE_CACHE` (`main.py:56-60`) caches uploaded files in Google Cloud Storage rather than locally, but this is file-level storage, not extraction caching.


Citations: [backend/src/main.py:685-744](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L685-L744) · [backend/src/main.py:945-981](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L945-L981) · [backend/src/shared/constants.py:783-825](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/constants.py#L783-L825) · [backend/src/graphDB_dataAccess.py:334-428](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L334-L428) · [backend/src/graphDB_dataAccess.py:362-428](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L362-L428)

### How are LLM cost and latency controlled during indexing and query? (answered)

**Caching.** There is no disk or database caching of LLM extraction outputs; every `get_graph_from_llm()` call sends chunk text to the LLM. At the file level, the `GCS_FILE_CACHE` option (`main.py:56-60`) stores uploaded files in GCS instead of a local temp directory, which is a storage cache, not an LLM cache.

**Batching.** Chunks are processed in configurable batches controlled by `UPDATE_GRAPH_CHUNKS_PROCESSED` environment variable (default 20, `main.py:515`). Within each batch, multiple chunks are combined via `get_combined_chunks()` (`llm.py:158-182`) into a single document with configurable `chunks_to_combine` — meaning up to `UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine` chunks are sent to the LLM in one extraction call, reducing the number of LLM invocations proportionally.

**Model choice per stage.** Different models can be assigned to different pipeline stages. The default community-creation model is a smaller/cheaper model (`openai_gpt_5.4_mini`, `communities.py:17`). The `GRAPH_CLEANUP_MODEL` env var (`post_processing.py:157`) controls which LLM merges node labels. The chat response also uses a model-per-deployment while the extraction model is configured per-source. For embeddings, the default is the free/small `sentence-transformers/all-MiniLM-L6-v2` (384-dim) (`common_fn.py:719`).

**Token budgets.** When `TRACK_USER_USAGE` is enabled (`main.py:504`), the system enforces preconfigured daily (`DAILY_TOKENS_LIMIT`, default 250k) and monthly (`MONTHLY_TOKENS_LIMIT`, default 1M) token limits per user (`common_fn.py:491-492`). Before processing begins, `track_token_usage()` with `operation_type="precheck"` (`common_fn.py:460-477`) checks if the user has exceeded their limits and raises `LLMGraphBuilderException` if so. After each batch, actual token usage is persisted to a `User` node in a separate token-tracker Neo4j database.

**Non-LLM shortcuts.** The KNN `SIMILAR` relationship between chunks (`graphDB_dataAccess.py:184-195`) is computed purely via vector cosine similarity in Cypher (no LLM involved). Entity deduplication (`graphDB_dataAccess.py:470-538`) uses embedding similarity, substring matching, and Levenshtein distance — not an LLM. The local/global search modes that return raw chunk text or entity subgraphs without summarization avoid per-query LLM overhead for retrieval.

**Embedding filter at query.** The `EmbeddingsFilter` in the query pipeline (`QA_integration.py:316-320`) uses a `similarity_threshold` (default 0.10, `constants.py:247`) and `TokenTextSplitter` to pre-filter retrieved documents before the LLM sees them, reducing LLM context length and cost.

> **Editor's note.** Correction: batching does not put `UPDATE_GRAPH_CHUNKS_PROCESSED × chunks_to_combine` chunks into one LLM call. Each extraction call covers `chunks_to_combine` concatenated chunks, so a batch of 20 chunks costs 20 / `chunks_to_combine` LLM calls; the batch size only controls how often progress and embeddings are written.

Citations: [backend/src/main.py:515-515](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/main.py#L515-L515) · [backend/src/communities.py:15-18](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/communities.py#L15-L18) · [backend/src/shared/common_fn.py:443-477](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/shared/common_fn.py#L443-L477) · [backend/src/QA_integration.py:302-325](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/QA_integration.py#L302-L325) · [backend/src/graphDB_dataAccess.py:182-196](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/graphDB_dataAccess.py#L182-L196) · [backend/src/post_processing.py:149-158](https://github.com/neo4j-labs/llm-graph-builder/blob/5ff7af3e9bb9226e1bbecd02f70f8d98697727a7/backend/src/post_processing.py#L149-L158)
