LLMs Technical Reviews
Home / Graph RAG / trustgraph

trustgraph-ai/trustgraph

Microservice Graph RAG platform that extracts RDF triples onto a message bus and answers by cross-encoder-filtered graph hops.

GitHub ↗★ 2.8kPythonApache-2.0commit ea19308 · 2026-10-01homepage ↗

Overview

TrustGraph is a self-hosted platform for building knowledge graphs from documents and querying them with Graph RAG, Document RAG, agents and structured queries. It is not a library you import. It is a set of several dozen small Python services (chunker, extractors, embedders, store writers, query services, RAG services, agent, gateway, librarian) that talk only through a pub/sub bus. Pulsar is the default; RabbitMQ and Kafka are alternatives (pubsub.py). A “flow” is a named wiring of these services, so one deployment can run several pipelines with different models and parameters side by side.

The data model is RDF. Every extractor emits Triple messages with IRI, literal and quoted-triple terms, plus named graphs for provenance and explainability. Entity names become deterministic IRIs, entity text is embedded for lookup, and every extracted subgraph is linked back to its chunk, page and document. At query time Graph RAG does not use community summaries. It grounds the question’s concepts in the entity vector index, then walks the graph for up to two hops, using a small cross-encoder to keep only the edges that matter.

It suits teams that want a full stack with an API gateway, IAM, a document library, metrics and many storage options, and are prepared to operate it. For a single-process Graph RAG experiment it is a lot of machinery.

Architecture

flowchart LR
  LIB["Librarian (docs + blobs)"] --> CH["Chunker"]
  CH --> DEF["kg-extract-definitions"]
  CH --> REL["kg-extract-relationships"]
  CH --> ONT["kg-extract-ontology"]
  CH --> AG["kg-extract-agent"]
  DEF --> TQ["Triples topic"]
  REL --> TQ
  ONT --> TQ
  DEF --> EC["Entity contexts"]
  ONT --> EC
  TQ --> TS["Triple store writer"]
  EC --> GE["graph-embeddings"]
  GE --> VS["Vector store writer"]
  GW["API gateway (REST + WebSocket)"] --> GR["graph-rag service"]
  GR --> VS
  GR --> TS
  GR --> RR["Reranker (flashrank)"]
  GR --> PR["Prompt service -> LLM"]
Component Path Role
Base framework trustgraph-base/trustgraph/base/ FlowProcessor, consumer/producer specs, clients, pub/sub backends
Chunking trustgraph-flow/trustgraph/chunking/ Recursive character splitter (2000/100 by default) and a token splitter
KG extractors trustgraph-flow/trustgraph/extract/kg/ definitions, relationships, ontology (OntoRAG), agent, topics, rows
Embeddings trustgraph-flow/trustgraph/embeddings/ Graph, document and row embedders over FastEmbed, Ollama or HF
Stores trustgraph-flow/trustgraph/storage/ Triples: Cassandra, Neo4j, Memgraph, FalkorDB. Vectors: Qdrant, Milvus, Pinecone. FTS5 keyword index
Retrieval trustgraph-flow/trustgraph/retrieval/ graph_rag, document_rag, NLP-to-GraphQL and structured query
Query services trustgraph-flow/trustgraph/query/ Triple pattern, SPARQL, GraphQL over rows, ontology-driven SPARQL/Cypher
Agents trustgraph-flow/trustgraph/agent/ ReAct agent, an orchestrator with plan/supervisor patterns, MCP tool bridge
Gateway, IAM, librarian trustgraph-flow/trustgraph/gateway/, trustgraph-flow/trustgraph/iam/, trustgraph-flow/trustgraph/librarian/ Public API, auth, document storage and processing triggers
Model adapters trustgraph-flow/trustgraph/model/text_completion/ OpenAI, Azure, Claude, Cohere, Mistral, vLLM, TGI, Ollama, LM Studio, llamafile; Bedrock and Vertex AI as separate packages

How a request flows

Ingestion:

  1. A document is added to the librarian, which stores it in a blob store and, on add_processing, pushes it into a flow (librarian.py). Decoders turn PDFs into text; the chunker splits on \n\n, \n, space, then characters (chunker.py).
  2. Each extractor in the flow consumes the same Chunk messages. kg-extract-definitions calls the extract-definitions prompt and emits rdfs:label and definition triples, plus two entity contexts per entity (its name and its definition) for embedding (definitions/extract.py). kg-extract-relationships does the same for subject-predicate-object triples, turning predicates into IRIs too.
  3. Entity IRIs come from to_uri: lower-case, spaces to hyphens, URL-quoted, under a fixed namespace (definitions/extract.py). The same name in two documents is the same node; there is no fuzzy or embedding-based resolution.
  4. Each batch of extracted triples gets a provenance subgraph in the urn:graph:source named graph, linking the triples to the chunk via tg:contains on quoted triples.
  5. Store writers persist triples per workspace and collection; graph-embeddings embeds entity contexts in one batched call and the vector writer stores them with the entity IRI and source chunk.

Graph RAG query (graph_rag.py):

  1. Concepts. The extract-concepts prompt splits the question into concept lines; if it returns nothing the raw question is used (graph_rag.py).
  2. Grounding. All concepts are embedded in one call and each searches the graph-embeddings store with entity_limit / n_concepts results (50 total by default); hits are deduplicated into seed IRIs (graph_rag.py).
  3. Hop and filter. For up to max_path_length (2) hops, the frontier is queried as subject, predicate and object (30 triples each). RDF/RDFS/OWL schema predicates are dropped, labels are resolved through a TTL cache, and each edge becomes short text showing only the new side of the edge. Up to 350 candidates are scored by the reranker against all concepts, and the top edge_limit (25) per hop survive; their endpoints form the next frontier (graph_rag.py).
  4. Sources. Selected edges are traced through tg:contains and prov:wasDerivedFrom to their source documents (graph_rag.py).
  5. Synthesis. Edges plus document metadata go to the kg-synthesis prompt, optionally streamed. Question, grounding, exploration, focus and synthesis steps are each emitted as explainability triples, with token counts.

Key components

Ontology-guided extraction (OntoRAG)

kg-extract-ontology splits each chunk into sentences and noun/verb phrases with NLTK, embeds them, and searches a FAISS inner-product index of the workspace’s ontology classes and properties for matches above 0.3. Small ontologies skip the search and are passed whole (ontology_selector.py). The selected subset goes into the extract-with-ontologies prompt, and results are checked against domain and range before conversion to triples. Entity IRIs include the ontology id and the class, so “Cornish pasty” as a Recipe and as a Food are different nodes (entity_normalizer.py). Ontologies arrive by config push and are reloaded per flow.

Document RAG

The document path supports vector, keyword (SQLite FTS5 BM25) and hybrid modes. Hybrid runs both and fuses them with weighted reciprocal rank fusion, degrading to vector-only if the keyword index fails (document_rag.py). Chunk text is fetched from object storage rather than kept in the vector store.

Storage backends

Cassandra keeps the full RDF model, including the named graph and term types (cassandra/write.py). The Neo4j writer maps triples to :Node, :Literal and :Rel with MERGE, scoped by workspace and collection, and does not store the graph field (neo4j/write.py). Provenance tracing and explainability therefore assume a store that keeps named graphs.

Knowledge cores

A knowledge core is a saved bundle of triples and graph embeddings. The knowledge service can list, export, import, delete and load cores into a flow, which is how a built graph is shipped between deployments without re-running extraction.

Extending it

  • New processors. Subclass FlowProcessor, declare ConsumerSpec/ProducerSpec/client specs, and the service joins any flow that wires its topics.
  • Prompts. Extraction, concept, synthesis and agent prompts are templates in configuration, editable per workspace.
  • Ontologies. Load an OWL-style ontology into config to switch on schema-guided extraction and ontology-driven SPARQL or Cypher answering.
  • Tools and agents. The agent services call Graph RAG, Document RAG, structured queries and MCP tools.
  • Backends. Pick triple, vector, LLM, embedding and pub/sub implementations per deployment; the bus contracts stay the same.

Running it

Deployment is configuration-generated: npx @trustgraph/config (or the hosted configuration UI) produces a Docker Compose or Kubernetes bundle for the chosen components. Container images are built from containers/ (base, flow, HF, OCR, Docling, MCP, Bedrock, Vertex AI). Expect a pub/sub broker, a triple store, a vector store, object storage for blobs, Prometheus metrics and an LLM endpoint. Clients use the gateway’s /api/v1 REST and WebSocket API, the Python API or the trustgraph-cli tools.

Strengths and caveats

  • Strength: provenance and explainability by design. Every extracted subgraph and every query step is itself stored as RDF, so answers can be traced to documents and retrieval decisions audited.
  • Strength: cheap, bounded graph retrieval. Only two LLM calls per Graph RAG query (concepts and synthesis); the hops are filtered by a MiniLM cross-encoder (ms-marco-MiniLM-L-12-v2 via flashrank by default), not by the LLM.
  • Strength: breadth. Four triple stores, three vector stores, hybrid document search, SPARQL, GraphQL over extracted rows, agents and multi-tenant workspaces in one platform.
  • Caveat: operational weight. Dozens of services, a broker and several databases. Debugging means following messages across queues.
  • Caveat: shallow entity resolution. Name-to-IRI hashing merges exact names and nothing else; aliases and spelling variants stay separate.
  • Caveat: no global or summary queries. No community detection or hierarchical summaries; retrieval is always local from grounded seeds.
  • Caveat: coarse deletion. Removing a document from the librarian deletes its blob and metadata, not its triples or embeddings (librarian.py); graph data is deleted per collection (neo4j/write.py).

Sources: code at ea19308, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the Graph RAG questions

Each answer was drafted by a code-reading agent at commit ea19308. Its citations were checked mechanically. Compare with the other graph rag →

How is the knowledge graph extracted from documents?

answered

Chunking. Documents are split by a recursive character splitter (chunking/recursive/chunker.py:76-77) that tries separators `

, , , then empty string (character fallback), with configurable chunk-size (2000) and chunk-overlap (100). ChunkingService (base/chunking_service.py:28-38`) registers these as parameters.

Multiple parallel extractors. After chunking, several processors consume Chunk messages concurrently:

  1. OntoRAG ontology-based extraction (extract/kg/ontology/extract.py): TextProcessor (text_processor.py:158-196) splits the chunk into sentences via NLTK and extracts noun/verb phrases using POS tagging. These segments are embedded and matched against pre-embedded ontology elements via OntologySelector (ontology_selector.py:107-166) using FAISS cosine similarity (threshold 0.3, top-k 10). Below 5 elements the selector is bypassed. Selected ontology subsets are passed to an LLM prompt (extract-with-ontologies) returning entities, relationships, and attributes in JSONL. TripleConverter (triple_converter.py:54-82) converts output to RDF triples, validating domain/range constraints and expanding URIs.

  2. Definitions extraction (extract/kg/definitions/extract.py:158-185): Calls prompt.extract_definitions(text=chunk) for entity-definition pairs.

  3. Relationships extraction (extract/kg/relationships/extract.py:110-117): Calls prompt.extract_relationships(text=chunk) for SPO triples.

  4. Agent-based extraction (extract/kg/agent/extract.py:178-215): Uses an Agent API with configurable templates.

Entity resolution. The OntoRAG pipeline uses EntityRegistry (entity_normalizer.py:113-165) mapping (entity_name, entity_type) to consistent normalized URIs per session. No cross-document deduplication exists.

Schema / ontology. Defined per-workspace via config push (extract/kg/ontology/extract.py:252-331). OntologyClass and OntologyProperty dataclasses (ontology_loader.py:15-85) capture OWL-like schema with labels, subclass-of, domain, range, inverse-of, and cardinality.

Editor's note. Correction: entity IRIs are deterministic (lower-cased, hyphenated name in the definitions/relationships extractors; ontology id + class + name in OntoRAG), so identical names from different documents merge into one node across the collection. What is missing is fuzzy or embedding-based resolution of aliases and variants, not cross-document deduplication.

Where and how is the graph stored?

answered

Graph database backends. Triples are stored in one of four backends, all consuming Triples messages:

  • Neo4j (storage/triples/neo4j/write.py): Nodes stored as :Node {uri, workspace, collection} for URI entities and :Literal {value, workspace, collection} for literals. Relationships as :Rel {uri, workspace, collection}. MERGE provides idempotent creation. Compound indexes on (workspace, collection, uri).
  • Memgraph (storage/triples/memgraph/write.py): Identical Cypher-based schema.
  • FalkorDB (storage/triples/falkordb/write.py): Same Cypher schema.
  • Cassandra (storage/triples/cassandra/write.py:168-191): EntityCentricKnowledgeGraph with insert(collection, s, p, o, g, otype, dtype, lang). Each workspace is a separate keyspace.

Embedding stores alongside the graph. Entity vectors are stored separately in:

  • Qdrant (storage/graph_embeddings/qdrant/write.py:100-133): Collections named t_{workspace}_{collection}_{dim}. Entity URI stored as payload, created lazily with cosine distance.
  • Milvus / Pinecone: Similar patterns.

Document/row embeddings go through storage/doc_embeddings/* and storage/row_embeddings/*. A keyword (BM25) index (storage/kw_index/fts5/) uses SQLite FTS5.

Document content is stored by the librarian service. Chunk provenance triples are emitted to GRAPH_SOURCE (chunking/recursive/chunker.py:203-223).

Isolation is per-workspace and per-collection. Deletion is collection-level only (storage/triples/neo4j/write.py:298-327).

Are communities, summaries or hierarchies built over the graph?

not applicable

This repository does not implement community detection, hierarchical community summaries, or any analogous graph grouping mechanism. There is no Leiden algorithm, Louvain, graph partitioning, or summarization of entity clusters.

Instead, TrustGraph relies on local graph traversal with cross-encoder reranking (retrieval/graph_rag/graph_rag.py:323-496) to navigate the knowledge graph at query time. From seed entities found via vector search, it performs up to max_path_length (2) iterative hops, retrieving edges adjacent to the frontier, scoring them with a cross-encoder reranker, and selecting the top edge_limit (25) edges.

The ontology provides a class hierarchy (extract/kg/ontology/ontology_loader.py:15-26) with subclass_of for domain/range validation, but this is an inheritance hierarchy, not community detection.

How does query-time retrieval use the graph?

answered

TrustGraph has three query modes:

Graph RAG (retrieval/graph_rag/): A five-stage pipeline:

  1. Concept extraction (graph_rag.py:156-178): LLM prompt (extract-concepts) decomposes the query into concept strings. Falls back to raw query if empty.

  2. Entity grounding (graph_rag.py:192-241): Each concept is embedded and used to query the graph embeddings vector store. Results are deduplicated into seed entity URIs.

  3. Iterative hop-and-filter (graph_rag.py:323-496): For up to max_path_length (2) hops, retrieves triples where frontier entities appear as subject, predicate, or object. Schema predicates (RDF/RDFS/OWL) filtered out. Labels resolved in batch with a per-request LRU cache. Each edge is rendered as text and scored against concepts by a cross-encoder reranker. Top edge_limit (25) edges survive, capped at max_reranker_input (350) candidates and max_reranker_text_length (240) chars.

  4. Source tracing (graph_rag.py:498-618): Edges traced to source documents via provenance chain (tg:contains -> prov:wasDerivedFrom).

  5. Synthesis (graph_rag.py:791-828): Edges and source metadata assembled into a kg-synthesis LLM prompt. Supports streaming. Explainability triples at every stage.

Document RAG (retrieval/document_rag/document_rag.py:216-530): Queries document embeddings. Supports vector, keyword (BM25), and hybrid modes (RRF fusion). Optional cross-encoder reranker with MMR diversity.

OntoRAG Query (query/ontology/query_service.py:144-218): Analyzes the question, matches ontology, generates SPARQL/Cypher, executes, and answers.

No local/global modes -- always local traversal.

Editor's note. Correction: edge_limit (25) is applied per hop, so the default two-hop traversal can pass up to 50 selected edges, plus traced source-document metadata, to the kg-synthesis prompt.

How are updates and incremental indexing handled?

insufficient evidence

This repository does not implement incremental indexing. Across the storage backends, the librarian, flow definitions and extraction processors:

  • No incremental document changes: Documents flow through a linear pipeline (document -> chunker -> extractors -> storage). Any new or changed document must be re-ingested through the entire pipeline. No change detection, fingerprinting, or differential update exists.

  • No per-document deletion: The only deletion mechanism is collection-level (delete_collection, e.g. storage/triples/neo4j/write.py:298-327), removing all data for a workspace+collection pair.

  • No extraction result cache: Extraction processors call the LLM on every chunk. Re-ingesting a document duplicates all previous extractions. The EntityRegistry (entity_normalizer.py:113-165) is per-request-scoped, not persistent.

  • No index staleness tracking: No document versioning or tracking which chunks/triples derive from which document version. Provenance triples are append-only.

Updates require full re-ingestion of the collection; deletion is collection-granularity only.

How are LLM cost and latency controlled during indexing and query?

answered

Caching. Graph RAG uses a per-request LRUCacheWithTTL (retrieval/graph_rag/graph_rag.py:94-134) for entity label lookups (max_size=5000, TTL=300s). OntoRAG has CacheManager (query/ontology/cache.py:463-610) with in-memory/file backends, TTL expiry, LRU eviction, and @cache_result decorator. Ontology loading cached per flow component (extract/kg/ontology/extract.py:237-243). IAM auth cached on gateway. SPARQL EXISTS cached (query/sparql/algebra.py:410-417).

Batching. Extraction processors batch output (triples_batch_size=50, entity_batch_size=5). OntologyEmbedder (ontology_embedder.py:149-170) batches at 50. Embeddings client accepts lists. Graph RAG does concurrent batch traversal (graph_rag.py:269-316).

Model choice per stage. LlmService (base/llm_service.py:81-91) reads a model flow parameter per request, configurable per pipeline stage. Supports API (OpenAI, Claude, Azure, Bedrock, Vertex AI) and local (vLLM, TGI, Ollama, LM Studio, llamafile) models.

Token budgets. Graph traversal caps at max_reranker_input=350 edges per hop, max_reranker_text_length=240 chars (retrieval/graph_rag/rag.py:40-43). Only edge_limit=25 edges reach synthesis. Document RAG limits doc_limit (retrieval/document_rag/rag.py:31-40). Response objects carry in_token/out_token/model.

Non-LLM shortcuts. Ontology selector bypassed when <5 elements (extract/kg/ontology/extract.py:161). FTS5 BM25 provides zero-LLM-cost sparse retrieval. OntoRAG answers SPARQL/Cypher queries structurally. Cross-encoder reranker is a small encoder model.

Editor's note. Correction: the 25-edge cap is per hop, not per query; with the default max_path_length of 2, synthesis can receive up to 50 edges plus source-document metadata edges.