semantica-agi/semantica
Large Python knowledge-graph toolkit with spaCy/LLM extraction, pluggable graph stores, and GraphRAG-style community and global search.
Overview
Semantica is a very large Python framework (about 420 modules and 220k lines under semantica/) for building and using knowledge graphs. Its own pitch is “context graphs” and “decision intelligence” for auditable agents. For a graph-RAG reader, the relevant part is a toolkit: ingestors, entity and relation extractors, a GraphBuilder, community detection, LLM community reports, and a ContextRetriever with local, global, drift and hybrid modes. Around it sit ontology generation, rule reasoning, provenance, a decision recorder, an MCP server, a FastAPI server and a web “Knowledge Explorer”.
It is a library of parts, not a fixed pipeline. Nothing calls community detection or report generation for you. A GraphRAG-style flow is something you assemble: build a graph, run CommunityHierarchyBuilder, run CommunitySummarizer.summarize_hierarchy, then pass the hierarchy and reports into ContextRetriever. The defaults lean local and free. Raw-text extraction uses spaCy NER and pattern-based triplets, and LLM extraction is opt-in per stage.
The graph itself is usually a plain Python dict, {"entities": [...], "relationships": [...]}. A Neo4j, FalkorDB, Apache AGE or Neptune store is optional. Who is it for? Python teams that want many knowledge-graph building blocks under one import, and accept glue work and uneven depth across modules. Teams that want a turnkey “index a folder, ask a question” GraphRAG will find more assembly here than in Microsoft GraphRAG or LightRAG.
Architecture
flowchart LR
SRC["Files / web / DB / streams"] --> ING["semantica.ingest + parse"]
ING --> GB["GraphBuilder.build"]
GB --> NER["NERExtractor (spaCy default)"]
GB --> TRI["Triplet / relation extractors"]
GB --> RES["EntityResolver (optional)"]
GB --> KG["KG dict: entities + relationships"]
KG --> GS["GraphStore: Neo4j / FalkorDB / AGE / Neptune"]
KG --> CH["CommunityHierarchyBuilder"]
CH --> CSUM["CommunitySummarizer (LLM reports)"]
KG --> CR["ContextRetriever"]
VS["Vector store"] --> CR
CSUM --> GR["GlobalGraphRetriever (map-reduce)"]
CSUM --> DR["DriftSearchEngine"]
GR --> CR
DR --> CR
| Component | Path | Role |
|---|---|---|
| Orchestrator | semantica/core/orchestrator.py |
Semantica.build_knowledge_base: ingest, run a pipeline, then build the graph and embeddings |
| Graph builder | semantica/kg/graph_builder.py |
Turns text, entity/relation objects or dicts into the KG dict; optional store persistence |
| Extraction | semantica/semantic_extract/ |
NERExtractor, RelationExtractor, TripletExtractor, LLM extraction, extraction cache |
| Communities | semantica/kg/community_detector.py, community_hierarchy.py |
Flat detection, and the multi-level Louvain/Leiden hierarchy |
| Reports | semantica/kg/community_summarizer.py |
Bottom-up CommunityReports with LLM or extractive fallback |
| Retrieval | semantica/context/context_retriever.py |
retrieve(mode=...): local, global, drift, hybrid |
| Global / DRIFT | semantica/context/global_retriever.py, drift_search.py |
Map-reduce over reports; facet-driven local expansion |
| Graph stores | semantica/graph_store/ |
GraphStore facade over Neo4j, FalkorDB, Apache AGE, Neptune |
| Vector stores | semantica/vector_store/ |
FAISS, Qdrant, Weaviate, Milvus, Pinecone, PgVector and others |
| LLM providers | semantica/llms/ |
OpenAI, Anthropic, Gemini, Groq, Ollama, DeepSeek, LiteLLM and more, with generate_typed |
| Services | semantica/server.py, semantica/mcp_server/, semantica/cli.py |
REST API + Explorer, stdio MCP server, CLI |
How a request flows
Building a graph from raw text with GraphBuilder().build(texts):
- Configure. The constructor builds an
EntityResolveronly whenmerge_entities=True, and aConflictDetectorwhenresolve_conflicts=True(the default) (graph_builder.py). - Extract. Each text goes to
_extract_from_text. NER defaults to"ml"(spaCy) and triplets to"pattern". Relation extraction runs only ifextract_relations=True. Each stage’s exceptions are logged and the build carries on (graph_builder.py). If spaCy or its model is missing,NERExtractorfalls back to regex patterns and, as a last resort, capitalised words (ner_extractor.py). - Resolve. With a resolver,
resolve_entitiesmerges duplicates and_remap_relationship_endpointsrewrites edges to the surviving ids. Synthetic endpoint entities are then dropped in favour of real ones with the same id. - Persist (optional). With a
graph_store, entities go toadd_nodesand relationships toadd_edgesassource_id/target_id/type/properties(graph_builder.py). - Check conflicts.
ConflictDetectorruns over the entities and logs how many it marked as resolved (graph_builder.py). The KG dict is returned.
Answering with ContextRetriever(knowledge_graph=kg, vector_store=vs, community_reports=reports, llm=llm).retrieve(q, mode=...):
- Dispatch.
retrieveroutesglobaltoretrieve_global,drifttoretrieve_drift, andhybridto a local call plus a global call merged by_rank_and_merge(context_retriever.py). - Local. Vector search, graph retrieval (with query-intent extraction and up to
max_expansion_hops, default 2) and memory search run one after another. A store that exposesquery()is queried directly; a KG dict is matched by embedding similarity and expanded (context_retriever.py). - Merge.
_rank_and_mergemin-max normalises scores per source and weights vector results by1 - hybrid_alphaand graph results byhybrid_alpha(context_retriever.py). - Global.
GlobalGraphRetriever.searchpicks a hierarchy level that fits the token budget, maps each candidate report to scored key points with parallel LLM calls, packs the best points into the reduce budget, and synthesises an answer with[Community id]citations (global_retriever.py).
Key components
Two community detectors
There are two code paths, and they are not equal. CommunityDetector.detect_communities_louvain actually calls NetworkX greedy_modularity_communities (Clauset-Newman-Moore, not Louvain). It also reads resolution from **options, so the named resolution argument is ignored (community_detector.py). Its detect_communities_leiden runs that same method and relabels the result “leiden” (community_detector.py).
The serious implementation is CommunityHierarchyBuilder. It runs real nx.louvain_communities with a seed, coarsens the graph, and repeats per level, with per-level resolutions (community_hierarchy.py). For Leiden it tries leidenalg on igraph, then cdlib, then a native Python Leiden (community_hierarchy.py). Use this builder for GraphRAG work.
Community reports
summarize_hierarchy walks levels from the bottom up, so each parent sees its children’s reports (community_summarizer.py). Each community gets a subgraph, centrality scores and a token-budgeted context. _call_llm then tries generate_typed against CommunityReportLLMSchema and works down a chain of weaker call styles. With no LLM it returns an extractive report (community_summarizer.py). Reports are cached by content hash.
Global and DRIFT search
GlobalGraphRetriever defaults to a 4,000-token context, a 600-token response budget and 4 map workers (global_retriever.py). DriftSearchEngine frames the query with the top reports, generates facets, and expands entities over an in-memory adjacency index with a drift threshold and a depth limit (drift_search.py).
Storage
GraphStore picks a backend class from a string: neo4j (default), falkordb, neptune or age (graph_store.py). The decision and agent side uses a separate in-memory ContextGraph (semantica/context/context_graph.py). That is what the MCP server loads and saves to SEMANTICA_KG_PATH (mcp_server/init.py). Graph-scope deletion exists as ContextGraph.purge_node, and ErasureCoordinator cascades an entity’s erasure to memory and vector stores (erasure.py). There is no “delete this document and everything extracted from it” operation in GraphBuilder.
Extending it
- Extractors. Choose
ner_method,relation_methodandtriplet_methodper build (ml,pattern,regex,huggingface,llm).NERExtractoralso accepts a list of methods withfallback,unionorconsensusmerging. - LLM providers. Classes in
semantica/llms/wrap provider SDKs and exposegenerateandgenerate_typed(prompt, schema)(openai.py). The summarizer and retrievers accept any object with those methods. - Registries and plugins. Most packages have a
registry.pyfor custom methods, andcore/plugin_registry.pyloads plugin bundles. - Stores. Add a graph backend next to
neo4j_store.pyand wire it in_initialize_store_backend. Vector stores follow the same pattern. - Agent surfaces. The MCP server exposes tools such as
extract_entities,add_entity,add_relationship,query_graph,record_decisionandfind_precedents. It has no community or global-search tool. LangChain, Agno and CrewAI integrations live underintegrations/.
Running it
- Install.
pip install semanticais a slim core (NetworkX, Pydantic, rdflib and similar). spaCy, local embeddings, FAISS, graph-database drivers and LLM SDKs are extras such assemantica[nlp-spacy],semantica[vectorstore-faiss],semantica[graph-neo4j]andsemantica[llm-all].semantica[all]installs everything. - Entry points.
semantica(CLI),semantica-server(FastAPI REST API plus Explorer SPA; protected routes needSEMANTICA_API_KEY),semantica-mcp(stdio MCP),semantica-workerandsemantica-explorer. - Services. None required for the in-memory path. Neo4j, FalkorDB, Postgres with AGE, Neptune and the vector databases are optional.
docker-compose.ymlis provided.
Strengths and caveats
- Strength: breadth. Ingestion, extraction, dedup, ontologies, reasoning, provenance, export formats and GraphRAG retrieval all sit in one package with consistent dict-shaped data.
- Strength: a faithful GraphRAG core. The hierarchy builder, bottom-up reports, token-budgeted map-reduce and DRIFT follow the published GraphRAG design closely, with typed Pydantic schemas and LLM-free fallbacks.
- Strength: works offline by default. Default extraction needs no API key, and each LLM stage is opt-in.
- Caveat: assembly required.
GraphBuilder, the hierarchy, the reports and the retriever are separate objects. Nothing keeps reports in sync when the graph changes. - Caveat: uneven depth and naming.
CommunityDetector’s “Louvain” and “Leiden” are both greedy modularity. Extraction errors are logged and swallowed. The conflict step thatbuild()runs by default only counts conflicts as “resolved”; the code comments that a real system would update the entity, and nothing is changed (conflict_detector.py). - Caveat: the default extraction is shallow. spaCy NER plus pattern triplets, or capitalised-word fallbacks without spaCy, give a far sparser graph than LLM extraction. Set
ner_method="llm"and friends for real GraphRAG quality. - Caveat: size. At 220k lines with many overlapping abstractions (
ContextGraph, the KG dict,GraphStore), finding the one code path that matters takes time. Global search falls back to citing every retained community when the model cites none.
Sources: code at 9a71df6, verified Q&A.
How it answers the Graph RAG questions
Each answer was drafted by a code-reading agent at commit 9a71df6. Its citations were checked mechanically. Compare with the other graph rag →
How is the knowledge graph extracted from documents?
answeredGraph construction flows through the GraphBuilder.build() method in semantica/kg/graph_builder.py. It processes sources that can be raw text, pre-extracted Entity/Relation objects, or dicts with entities/relationships keys. For raw text, the builder calls _extract_from_text which runs three extractors in sequence: NERExtractor (default method "ml", using spaCy; alternatives: "pattern", "regex", "huggingface", "llm"), RelationExtractor (default "pattern", alternatives up to "llm"), and TripletExtractor (default "pattern"). The NERExtractor (semantica/semantic_extract/ner_extractor.py) supports fallback chains — if the primary method returns nothing, it falls through pattern matching to a capitalized-word heuristic as a last resort. The GraphChunker module (semantica/split/kg_chunkers.py) provides EntityAwareChunker, RelationAwareChunker, GraphBasedChunker, OntologyAwareChunker, and HierarchicalChunker — which can preserve entity boundaries or triplet integrity during chunking. LLM enhancement is additive: LLMExtraction.enhance_entities() and enhance_relations() call an LLM provider with a structured prompt, merge results back by matching entity text or exact triples, and always preserve originals on failure (semantica/semantic_extract/llm_extraction.py:192-285). Entity resolution is done by EntityResolver (if merge_entities=True) using fuzzy/exact/ML strategies, and EntityMerger in semantica/deduplication/entity_merger.py detects duplicate groups via DuplicateDetector, then applies configurable strategies (keep_first, keep_most_complete, merge_all, etc.) with provenance tracking. After resolution, _remap_relationship_endpoints() rewrites relation source/target IDs to the surviving canonical entity ID (graph_builder.py:370-454). Conflict detection via ConflictDetector (optional, resolve_conflicts=True) runs after the graph is built and can auto-resolve conflicts. The output is a dict with entities, relationships, and metadata. Temporal edges are supported via add_temporal_edge(). No fixed schema/ontology is enforced at build time; entity types and relation predicates are free-form strings from the extractors, although the ontology/ module provides schema mapping post-hoc.
Where and how is the graph stored?
answeredThe graph is stored via a pluggable backend architecture. The core GraphStore class in semantica/graph_store/graph_store.py defines the interface (NodeManager, RelationshipManager, QueryEngine), and concrete implementations exist for Neo4j (Neo4jStore in neo4j_store.py, using Cypher queries via the neo4j driver), Apache AGE (ApacheAgeStore in age_store.py, wrapping openCypher through PostgreSQL psycopg2), FalkorDB (FalkorDBStore in falkordb_store.py), and Amazon Neptune. The builder's build() method optionally persists the graph to a store via self.graph_store.add_nodes(resolved_entities) and self.graph_store.add_edges(formatted_edges) (graph_builder.py:1012-1044). When no database backend is configured, the graph is returned as an in-memory Python dict {"entities": [...], "relationships": [...]} and held by the caller. Nodes are stored as property-graph entities with an id, optional labels (e.g. ["Person"]), and arbitrary properties. Edges carry source_id, target_id, type, and properties. The Neo4j driver (neo4j_store.py:96-100) normalizes Node/Relationship objects to plain dicts with underscored identity keys (_labels, _element_id, _type) so user properties are never shadowed. Embeddings sit alongside the graph in a separate vector store (semantica/vector_store/), supporting FAISS, Qdrant, Weaviate, Pinecone, Milvus, PgVector, and SQLiteVec. The ContextRetriever orchestrates hybrid retrieval: it queries the vector store via cosine similarity and the graph via semantic entity matching / BFS traversal, then merges results with configurable hybrid_alpha weighting — 0 = vector only, 1 = graph only, 0.5 = balanced (context_retriever.py:178). Embeddings for entities are generated via DecisionEmbeddingPipeline which uses the vector store's own embed method.
Are communities, summaries or hierarchies built over the graph?
answeredSemantica has a full community detection and hierarchical summarization pipeline. The CommunityDetector class (semantica/kg/community_detector.py) supports Louvain (via NetworkX greedy_modularity_communities with resolution parameter, and a basic fallback), Leiden (via optional igraph+leidenalg or cdlib, falling back to a native Python implementation with local-moving and refinement phases at lines 894-1034), overlapping k-clique communities, and label propagation (supports chunked processing for large graphs). The CommunityHierarchyBuilder (semantica/kg/community_hierarchy.py:553-1419) builds multi-level partitions by iteratively coarsening the graph via Louvain or Leiden at increasing resolution, then constructs a CommunityHierarchy — a tree of HierarchicalCommunity nodes keyed by SHA-256 content hashes with metrics (internal/external edges, density, conductance) and deterministic parent-child resolution. The CommunitySummarizer (semantica/kg/community_summarizer.py:547-2297) generates structured CommunityReport objects bottom-up via summarize_hierarchy(): for each community it extracts a subgraph, computes centrality scores (degree/betweenness/closeness/eigenvector/PageRank), packs context within a token budget (prioritizing child reports, edges, anchor entities, and source text), and calls an LLM with a Pydantic-typed schema (CommunityReportLLMSchema) to produce a title, summary, findings list, and impact rating. It has a 5-tier LLM unwrap strategy (_call_llm) and falls back to an extractive baseline when no LLM is configured. Reports are cached via SHA-256 content hashing with atomic disk persistence. The hierarchy is computed on demand when the user invokes community detection; it is not automatically triggered during ingestion or graph building.
CommunityDetector's Louvain is NetworkX greedy modularity (Clauset-Newman-Moore) with the resolution argument effectively fixed at 1.0, and its Leiden method just relabels that result. Real Louvain (nx.louvain_communities) and Leiden (leidenalg, cdlib, then a native fallback) are only in CommunityHierarchyBuilder in community_hierarchy.py.How does query-time retrieval use the graph?
answeredQuery-time retrieval supports three modes — local, global, and hybrid — selected via the mode parameter of ContextRetriever.retrieve() (context_retriever.py:210-295). Local mode runs three retrievers in parallel: (1) vector search via _retrieve_from_vector() which queries the vector store (FAISS/Qdrant/etc.) for semantically similar content; (2) graph retrieval via _retrieve_from_graph() which first extracts query intent (relationship types, entity types, keywords), then computes cosine similarity between the query embedding and every entity/relationship embedding — boosting scores when entity types or relation types match the intent — and finally performs BFS-based graph expansion (max_expansion_hops, default 2) to pull in related entities. All results are merged via _rank_and_merge() which normalizes scores per source, applies contextual boosting (graph results with more neighbors get up to 20% boost; results found in both vector and graph get 20% cross-source boost), and re-ranks using 70% original + 30% query-content similarity. Global mode delegates to GlobalGraphRetriever.search() (global_retriever.py:329+), which implements a Map-Reduce pattern over community reports: it dynamically selects a hierarchy level (promoting to coarser levels if token budget is exceeded), runs Map LLM calls concurrently via ThreadPoolExecutor (each extracting key points from a community report into a structured MapPointSchema with relevance scores), then synthesizes the Reduce response from top-scored points within a token budget, with automatic evidence/citation validation. Hybrid mode runs both local and global retrievers, concatenates results, and re-ranks. DRIFT mode (drift_search.py) combines global thematic framing with directed facet reasoning and local entity-hop graph traversal. The assembled context is returned as RetrievedContext objects with content, score, source, related_entities, and related_relationships — ready for LLM consumption.
How are updates and incremental indexing handled?
answeredIncremental updates are handled at multiple levels. The EntityMerger.incremental_merge() method (semantica/deduplication/entity_merger.py:388-473) merges new entities into an existing set by running DuplicateDetector.incremental_detect() between new and existing entities, then merging each candidate pair with configurable strategies — avoiding duplicate merges via processed-ID tracking. The GraphBuilder.build() method is designed to be re-runnable: it merges sources and can be called repeatedly on new documents, with merge_entities=True enabling deduplication against previously built entity sets. The builder's entity resolver (EntityResolver) can operate on the full accumulated set each time, and _remap_relationship_endpoints() rewrites relation endpoints after entity merges change canonical IDs (graph_builder.py:370-454). Ingestion is source-agnostic: FileIngestor, StreamIngestor (Kafka, RabbitMQ, Pulsar, Kinesis), WebIngestor, DBIngestor, and dozens of other ingestors in semantica/ingest/ each produce documents that feed into extraction. The Pipeline module (pipeline_builder.py, execution_engine.py) provides full orchestration with retry, parallelism, and resource scheduling. Version management via TemporalVersionManager (change_management/managers.py) supports snapshot creation, diffs between versions, and persistent storage (SQLite or in-memory). Deletion is not explicitly handled as a first-class operation — there is no built-in mechanism to remove a document's extracted entities/relations after they are added. The system is additive: incremental merges consolidate but do not remove stale data. Extraction caching (semantica/semantic_extract/cache.py) provides TTL-based caching for entity, relation, and triplet extraction results, with pluggable backends (in-memory LRU or persistent SQLite). The cache is keyed by text hash and method parameters, avoiding re-extraction from unchanged source texts across rebuilds.
ContextGraph.purge_node (tombstoned) and ErasureCoordinator.erase_entity, which cascades to agent memory and vector stores and returns a receipt.How are LLM cost and latency controlled during indexing and query?
answeredSemantica controls LLM cost and latency through several mechanisms. Extraction-level caching (semantica/semantic_extract/cache.py) stores entity, relation, and triplet extraction results keyed by text hash and method, with configurable TTL (default 3600s) and dual backends: in-memory LRU or persistent SQLite. The cache sits behind all extractor calls so repeated texts never reach the LLM. Batching is configurable in the default optimization config: batch_size: 10, max_tokens_per_batch: 2000, enable_batching: True (semantica/semantic_extract/config.py:58-71). The max_workers setting (default 8, auto-clamped to CPU count and capped at 32) controls parallelism in pipelines. Model choice per stage is fully configurable: the GraphBuilder allows separate selection of ner_method, relation_method, and triplet_method — each can be "ml" (spaCy, free), "pattern" (regex, free), or "llm" (paid). The default stack uses free local methods ("ml" for NER, "pattern" for relations/triplets), requiring no API keys (graph_builder.py:519-521). LLM enhancement is opt-in. The LLM provider layer (semantica/llms/ and semantica/semantic_extract/providers.py) supports 10+ providers including local/cheap options: Ollama for local open-source models, Groq for fast cheap inference, and LiteLLM for routing to the cheapest available model. Each provider exposes generate_typed() for structured output via Pydantic schemas, reducing the need for multiple retries. Community summarization has a max_tokens budget (default 4000) that limits prompt size, and centrality-based token budgeting (_pack_context) prioritizes the most central entities and highest-impact child reports within that budget (community_summarizer.py:1169-1728). The _extractive_fallback method produces deterministic reports with no LLM call when no model is configured. Global query uses similar token budgeting (max_context_tokens: 4000, response_token_budget: 600) with dynamic level promotion — if level-0 community reports exceed budget, it promotes to coarser levels where fewer but denser reports exist (global_retriever.py:428-534). The Map phase runs concurrently across community reports with a configurable max_workers (default 4).