SciPhi-AI/R2R
Python RAG server on Postgres/pgvector: hybrid and graph search, cited streaming answers, and RAG and research agents behind a REST API.
Overview
R2R (“RAG to Riches”) from SciPhi is a self-hosted RAG server. It is a FastAPI application with a /v3 REST API, Python and JavaScript SDKs, and a Next.js dashboard. It covers the whole loop: upload documents, parse and chunk them, embed them into PostgreSQL with pgvector, search with vector, full-text or hybrid retrieval plus an optional knowledge graph, and answer with an LLM that cites its sources. On top of plain rag it offers an agent endpoint with two modes. “rag” is a tool-using retrieval agent. “research” adds reasoning, critique and Python-execution tools.
The architecture is a provider registry. Every concern (auth, database, embeddings, LLM, ingestion, OCR, orchestration, file storage, email, scheduler) is an abstract provider. A TOML config selects the concrete one, and a factory builds it at startup. PostgreSQL does almost everything: chunks and vectors, a generated tsvector column for keyword search, documents, users, collections, conversations, prompts and graph tables. There is no separate vector database. LLM and embedding calls go through LiteLLM by default, so most providers work by changing a model string.
R2R suits teams that want a ready multi-user RAG backend with document management, collections-based access control and an API they can build a product on. The pinned commit is from November 2025.
Architecture
flowchart LR
C["SDK / dashboard / MCP"] --> API["FastAPI /v3 routers"]
API --> SVC["Services: ingestion, retrieval, graph, management"]
SVC --> ORCH["Orchestration: simple (inline) or Hatchet"]
ORCH --> ING["R2RIngestionProvider: parsers + splitter"]
ING --> OCR["Mistral OCR / VLM / Unstructured"]
ING --> EMB["Embedding provider (LiteLLM)"]
EMB --> PG["Postgres + pgvector"]
SVC --> PG
SVC --> LLM["LLM provider (LiteLLM)"]
SVC --> AG["Agents: RAG / research"]
AG --> LLM
PG --> CL["Graph clustering service (Leiden)"]
| Component | Path | Role |
|---|---|---|
| App and routers | py/core/main/app.py, py/core/main/api/v3/ |
Documents, chunks, collections, graphs, indices, prompts, retrieval, conversations, users, system |
| Config | py/r2r/r2r.toml, py/core/configs/*.toml, py/core/main/config.py |
Default config plus named variants (full, ollama, azure, …) |
| Provider factory | py/core/main/assembly/factory.py |
Builds each provider from config |
| Services | py/core/main/services/ |
IngestionService, RetrievalService, GraphService, ManagementService, AuthService |
| Orchestration | py/core/main/orchestration/{simple,hatchet}/ |
Ingestion and graph workflows, inline or as Hatchet tasks |
| Ingestion | py/core/providers/ingestion/, py/core/parsers/ |
Per-type parsers, text splitters, Unstructured alternative |
| Database | py/core/providers/database/ |
Postgres handlers for chunks, documents, graphs, filters, limits |
| Agents | py/core/agent/, py/core/base/agent/tools/ |
RAG and research agents, streaming, citations, tool registry |
| Clients | py/sdk/, js/, py/r2r/mcp.py |
Python and JS SDKs, a small MCP server |
How a request flows
Upload one PDF, then ask a question with POST /v3/retrieval/rag:
- Upload.
create_documentchecks the per-type size limit, stores the raw bytes with the file provider, and registers the document withingest_file_ingress. Withrun_with_orchestration, it callsorchestration.run_workflow("ingest-files"). If that fails, it falls back to the inline ingestor (documents_router.py). - Orchestrate. The default
simpleprovider has no worker.run_workflowsimply awaits the workflow function, so ingestion happens inside the HTTP request (simple.py). Thefullconfig switches to Hatchet for background workers. - Parse, summarize, embed, store.
ingest_filesmoves the document through PARSING, AUGMENTING (an LLM document summary unlessskip_document_summary), EMBEDDING, STORING and SUCCESS. It then assigns the document to the user’s default collection (ingestion_workflow.py). Parsing picks a parser fromDEFAULT_PARSERSby type, and chunking defaults to a recursive character splitter of 1,024 characters with 512 overlap (base.py, L160-L215). - Search.
RetrievalService.ragfirst callssearch, which dispatches onsearch_strategyto basic, HyDE or RAG-Fusion (retrieval_service.py)._vector_search_logicembeds the query, runs semantic, full-text or hybrid search onPostgresChunksHandler, then callsarerankand prefixes each chunk with its document title (L651-L730). - Fuse.
hybrid_searchruns the vector query and thewebsearch_to_tsqueryquery separately. It merges them in Python with weighted reciprocal rank fusion, usingsemantic_weight,full_text_weightandrrf_k(chunks.py). - Generate. The results go into a
SearchResultsCollectorand are formatted with short source ids. Thesystemandragprompts are loaded from the database and filled withqueryandcontext. Without streaming, oneaget_completioncall is made, and bracketed short ids in the answer becomeCitationobjects with full payloads (L985-L1080). Withstream: true, the same flow yields SSE events instead.
Key components
Ingestion and parsing
R2RIngestionProvider maps about 35 document types to parser classes: PDF, Office formats, email (EML/MSG), EPUB, HTML, Markdown, code files, images and MP3. PDFs have extra parsers (zerox VLM and ocr) configured in the default TOML. Images and audio go through the configured VLM and the Whisper model. The unstructured provider is a drop-in alternative for layout-aware partitioning. Only recursive and character splitting are implemented. basic and by_title raise NotImplementedError in the R2R provider. Optional chunk enrichment rewrites each chunk with neighbouring context, and it is off by default (r2r.toml).
Storage and indexes
One chunks table per project schema holds vec vector(N), an optional vec_binary bit(N) for INT1 quantization, text, metadata JSONB, a collection_ids UUID[] with a GIN index, and a generated English tsvector (chunks.py). HNSW or IVFFlat vector indexes are created explicitly through the indices API. The default embedding is openai/text-embedding-3-small at 512 dimensions via LiteLLM.
Reranking
Reranking is a method on the embedding provider. The LiteLLM provider supports only a HuggingFace TEI rerank endpoint, configured with rerank_model and rerank_url or HUGGINGFACE_API_BASE. Without them, arerank just truncates the results (litellm.py, L258-L288).
Knowledge graph
Graph extraction is an explicit step per document or collection. An LLM extracts entities and relationships, describes them and embeds the descriptions. Community detection is delegated to a separate clustering container (ragtoriches/cluster-prod) that runs Leiden. Communities are then summarized by LLM and become searchable next to chunks (graphs.py). automatic_extraction is true in the default config, but the simple workflow only logs that it is “not yet implemented” (ingestion_workflow.py).
Agents and streaming citations
AgentFactory picks a RAG or research agent, with streaming or XML-tool variants. R2RAgent.arun loops LLM calls and tool calls up to max_iterations, then summarizes if the loop never completed (base.py). The streaming agent emits message events. It scans the accumulated text for bracketed short ids, and sends a citation event with the source payload the first time each id appears (L440-L510). Default RAG tools are search_file_descriptions, search_file_knowledge and get_file_content. Web search (Serper), scraping (Firecrawl) and Tavily tools are opt-in.
Extending it
- Swap providers by config. Pick a TOML with
R2R_CONFIG_NAMEorR2R_CONFIG_PATH. Bundled variants cover Ollama, LM Studio, Azure, Gemini and the Hatchetfullstack. - Prompts. System, task, graph and enrichment prompts live in Postgres and are editable through
/v3/promptswithout a redeploy. - Custom agent tools.
ToolRegistryimports every Python file in auser_toolsdirectory, and the Docker image mountsdocker/user_tools. Add a tool name torag_toolsorresearch_toolsto enable it. - Search tuning per request.
search_settingstoggles semantic, full-text, hybrid and graph search, sets the HyDE or RAG-Fusion strategy, and applies Mongo-style metadata filters ($eq,$in,$overlap, …). - Clients. Python (
R2RClient), JavaScript (r2r-js), and an MCP server inpy/r2r/mcp.pythat wraps search and RAG for MCP clients.
Running it
- Light mode.
pip install r2randpython -m r2r.serve, with anOPENAI_API_KEYand a reachable Postgres with pgvector. Ingestion runs inline. - Docker.
docker/compose.yamlruns Postgres, MinIO, the graph-clustering service, R2R on port 7272 and the dashboard (compose.yaml).compose.full.yamladds Hatchet (engine, dashboard, RabbitMQ, its own Postgres) and an Unstructured service. - Defaults to change. The default config ships
require_authentication = falseand a well-known default admin password (r2r.toml). Set both before exposing the server.
Strengths and caveats
- Strength: one database. Vectors, keyword search, graph, users and conversations all sit in Postgres, so backup and multi-tenant filtering (
collection_ids,owner_id) stay simple. - Strength: complete API surface. Document versioning and status, collections, user management, conversation history, prompt management and streaming citations are already built. That is the boring 70% of a RAG product.
- Strength: flexible retrieval. Hybrid RRF, HyDE, RAG-Fusion, graph and community search, plus web search tools, all switchable per request.
- Caveat: inline ingestion by default. Without Hatchet, an upload blocks its HTTP request through parsing, an LLM summary and embedding. Large files or bursts need the heavier
fullstack. - Caveat: hybrid fusion in Python. Two separate SQL queries are fused in application code, and English-only
tsvectorkeyword search is weak for other languages. - Caveat: research agent’s Python tool.
python_executorwrites model-generated code to a temp file and runs it with the server’s interpreter in a subprocess, with a 10 s timeout and no sandbox (research.py). Keep it out ofresearch_toolson shared deployments. - Caveat: limited reranking. Only a TEI endpoint is supported, and it is off by default.
Sources: code at 9c5a94d, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit 9c5a94d. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredSupported formats. R2R supports 40+ document types via a parser registry at py/core/providers/ingestion/r2r/base.py:40-76. DEFAULT_PARSERS maps each DocumentType enum (py/shared/abstractions/document.py:18-93) to a dedicated parser class: BMPParser, CSVParser, DOCParser, DOCXParser, EMLParser, EPUBParser, HTMLParser, JSONParser, MDParser, MSGParser, ORGParser, BasicPDFParser (plus OCRPDFParser, VLMPDFParser, PDFParserUnstructured as extras), PPTParser, PPTXParser, RTFParser, TextParser, TSVParser, XLSParser, XLSXParser, ImageParser (GIF/JPEG/JPG/PNG/HEIC/SVG/TIFF), AudioParser (MP3), P7SParser, RSTParser, PythonParser, JSParser, TSParser, CSSParser.
OCR and layout parsing. The MistralOCRProvider (py/core/providers/ocr/mistral.py:13-53) wraps the Mistral OCR API for document image processing. For PDFs, alternative parsers include OCRPDFParser (local OCR via pdf2image + Mistral) and VLMPDFParser (zero-shot VLM-based). The UnstructuredIngestionProvider (py/core/providers/ingestion/unstructured/base.py) integrates with Unstructured's API or a local Unstructured service for layout-aware partitioning with options like pdf_infer_table_structure, hi_res_model_name, and ocr_languages. Table handling is further supported through XLSXParserAdvanced and CSVParserAdvanced.
Chunking strategy. The R2RIngestionProvider (py/core/providers/ingestion/r2r/base.py:160-215) builds a text splitter based on ChunkingStrategy (RECURSIVE, CHARACTER, BASIC, BY_TITLE). Default is RecursiveCharacterTextSplitter with chunk_size=1024 and chunk_overlap=512 (py/core/providers/ingestion/r2r/base.py:31-35). The recursive splitter (py/shared/utils/splitter/text.py:1219-1289) walks through `["
", "
", " ", ""] separator levels. The character splitter (py/shared/utils/splitter/text.py:620-646) uses a configurable separator. For OCR/VLM parsers, an optional vlm_ocr_one_page_per_chunkmode creates oneDocumentChunk` per page rather than re-splitting.
recursive (default, 1,024 characters with 512 overlap) and character splitting are implemented in the R2R provider; basic and by_title raise NotImplementedError.How are embeddings and indexes built and stored?
answeredEmbedding models. Three embedding providers: OpenAIEmbeddingProvider (py/core/providers/embeddings/openai.py:21-60) supporting text-embedding-ada-002, text-embedding-3-small, and text-embedding-3-large with configurable dimensions; LiteLLMEmbeddingProvider (py/core/providers/embeddings/litellm.py:25-65) routing through LiteLLM for any provider (Amazon, HuggingFace, etc.); and OllamaEmbeddingProvider for local models. EmbeddingConfig (py/core/base/providers/embedding.py:21-44) configures provider, base_model, base_dimension, batch_size, concurrent_request_limit, and vector quantization settings. Retry with exponential backoff is built in.
Vector store. All vectors are stored in PostgreSQL using pgvector. PostgresChunksHandler (py/core/providers/database/chunks.py:75-182) creates tables with a vec column (configurable dimension), optional vec_binary bit(N) column for INT1 quantization, text TEXT, and metadata JSONB. A full-text search column (fts tsvector) is auto-generated as to_tsvector('english', text). Index methods supported: IVFFlat (IndexArgsIVFFlat with n_lists) and HNSW (IndexArgsHNSW with m and ef_construction) via py/shared/abstractions/vector.py:27-107. Distance measures: cosine, L2, max-inner-product, L1, hamming, jaccard. Vector quantization types: FP32, FP16, INT1 (binary), SPARSE.
Hybrid / keyword (BM25). Full-text search uses PostgreSQL's websearch_to_tsquery('english', ...) with ts_rank ordering (py/core/providers/database/chunks.py:484-536). Hybrid search (py/core/providers/database/chunks.py:538-640) combines semantic and full-text via weighted Reciprocal Rank Fusion with configurable semantic_weight, full_text_weight, and rrf_k.
Metadata. Stored as JSONB alongside each vector entry. All document metadata (title, version, chunk_order, page_number, parser_generated) is carried through to the metadata column. Filtered at query time via the filter system in py/core/providers/database/filters.py.
How is retrieval performed?
answeredDense / sparse / hybrid search. The RetrievalService (py/core/main/services/retrieval_service.py:248-280) dispatches to three strategies: basic (vanilla semantic + optional graph), HyDE (LLM-generated hypothetical documents), and RAG Fusion (multi-query with RRF). Internally, _vector_search_logic (py/core/main/services/retrieval_service.py:651-729) routes to semantic_search, full_text_search, or hybrid_search on PostgresChunksHandler depending on SearchSettings flags. Semantic search (py/core/providers/database/chunks.py:327-482) computes cosine distance via vec <=> $1::vector(N) with a two-stage binary re-ranking path for INT1 quantization. Full-text search uses ts_rank(fts, websearch_to_tsquery('english', $1)). Hybrid search (py/core/providers/database/chunks.py:538-640) fuses both via weighted RRF.
Search strategies. HyDE (py/core/main/services/retrieval_service.py:554-614) generates N hypothetical documents via LLM, embeds each, runs parallel searches, and re-ranks results against the original query. RAG Fusion (py/core/main/services/retrieval_service.py:326-403) generates num_sub_queries alternative phrasings via LLM, runs each independently, then fuses via RRF and optionally re-ranks.
Reranking. The embedding provider's arerank() method is called after initial search. OpenAIEmbeddingProvider passes through (py/core/providers/embeddings/openai.py:227-243). LiteLLMEmbeddingProvider optionally uses a HuggingFace TEI reranking endpoint (py/core/providers/embeddings/litellm.py:47-59).
Graph search. _graph_search_logic (py/core/main/services/retrieval_service.py:732-794) searches entities, relationships, and communities via embedding similarity on their description vectors. Graph entities have their own vector tables with pgvector.
Filters. Rich filter system (py/core/providers/database/filters.py) supporting $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin, $like, $ilike, $overlap, $contains, $and, $or. Applied to the metadata JSONB column and top-level columns like document_id, owner_id, collection_ids.
How are answers generated and grounded?
answeredPrompt assembly. System prompts are stored in the database via PostgresPromptsHandler. The Agent base class (py/core/base/agent/agent.py:81-97) loads a static prompt (e.g., static_rag_agent) and a dynamic reasoning prompt at startup. The R2RAgent.arun() (py/core/agent/base.py:100-151) assembles messages from the conversation history, passes them to llm_provider.aget_completion(), and iterates up to max_iterations (default 10). GenerationConfig (py/shared/abstractions/llm.py) controls model, temperature, max_tokens, streaming, and extended thinking.
Citations / source attribution. During streaming (py/core/agent/base.py:351-638), the R2RStreamingAgent tracks a SearchResultsCollector and CitationTracker. As the LLM streams tokens, find_new_citation_spans() detects bracket references like [abc1234] in the text. Each citation emits an SSE citation_event with the short ID, span boundaries, and full source payload on first occurrence. On finish_reason=stop, consolidated citations with all spans and payloads are emitted in a final_answer_event (py/core/agent/base.py:599-633). Citations are stored in Message.metadata for persistence.
Streaming. The R2RStreamingAgent (py/core/agent/base.py:351-638) yields SSE-formatted events: thinking (model chain-of-thought), message (partial text tokens), tool_call (tool name/arguments), citation (source reference), and final_answer (complete answer with structured citations). The endpoint returns text/event-stream content type via FastAPI's StreamingResponse (py/core/main/api/v3/retrieval_router.py:408-427).
Agentic / multi-step. AgentFactory (py/core/main/services/retrieval_service.py:60-245) creates one of 10 agent types based on mode (RAG or research), streaming preference, and XML tool format. R2RRAGAgent and R2RResearchAgent support tool usage via ToolRegistry — tools include web_search, web_scrape, search_file_knowledge, get_file_content, reasoning, critique, python_executor. The research agent adds its own reasoning/critique/Python tools (py/core/agent/research.py:35-80).
How is quality evaluated or observed?
answeredR2R has no built-in evaluation framework — no scoring, no LLM judge, no integrated metrics like faithfulness or answer relevancy. The project ships with two observability mechanisms instead:
Sentry integration (
py/core/utils/sentry.py). Initialized at app startup viainit_sentry(), readsR2R_SENTRY_DSNfrom the environment. Supports configurable traces_sample_rate and profiles_sample_rate. This provides error tracking and performance tracing but no RAG-specific quality metrics.Structured logging (
py/core/utils/logging_config.py). Configures python-json-logger with HTTP status filters for uvicorn.access logs. Health endpoint requests are filtered out. There is no RAG-specific logging (no retrieval latency, generation quality, or user feedback capture).External evaluation guide. The project provides a cookbook at
docs/cookbooks/evals.mddemonstrating how to evaluate R2R outputs using the Ragas framework externally. It shows collecting R2R's RAG responses viaclient.retrieval.rag()and feeding them into Ragas for metrics like faithfulness, answer_relevancy, context_precision, and context_recall. This is an optional, bring-your-own-eval approach — Ragas is not a dependency and no integration code exists in the source tree.
There are no built-in evaluation endpoints, no A/B testing infrastructure, no LLM-as-judge calls, and no quality monitoring dashboards. Users must integrate external frameworks (Ragas, TruLens, LangFuse, etc.) themselves.
How is it deployed and operated?
answeredLibrary vs service. R2R is both a Python library (pip install r2r) and a FastAPI web service. The library exposes R2RClient and R2RAsyncClient (py/r2r/__init__.py:1-19) for programmatic access. The service runs via uvicorn on the core.main.app_entry:app FastAPI application (py/core/main/app_entry.py:109-139).
API. Versioned REST API under /v3/ with routers for search, RAG, agent conversations, ingestion, documents, collections, users, conversations, prompts, graphs, chunks, indices, and system health (py/core/main/app.py:14-23). The /retrieval/search, /retrieval/rag, and /retrieval/agent endpoints are the primary retrieval/RAG surface. Authentication is pluggable via AuthProvider (Supabase, JWT, Clerk). CORS is wide open (allow_origins=["*"]).
Configuration. Multiple toml config files (py/core/main/config.py:29-43) with environment variable overrides. Key env vars: R2R_POSTGRES_HOST/PORT/USER/PASSWORD/DBNAME, OPENAI_API_KEY, ANTHROPIC_API_KEY, MISTRAL_API_KEY, R2R_SENTRY_DSN, R2R_CONFIG_NAME, R2R_PROJECT_NAME.
Required infrastructure. PostgreSQL 15+ with pgvector extension is mandatory. Optional: Unstructured API (or self-hosted Unstructured service) for advanced document parsing; Mistral API for OCR; Redis for APScheduler jobs; Sentry DSN for error tracking.
Deployment. Docker image (py/Dockerfile) runs uvicorn on port 8000 (configurable via R2R_PORT). The Docker build uses multi-stage with Rust compilation for tokenizer dependencies. A CMD of uvicorn core.main.app_entry:app --host $R2R_HOST --port $R2R_PORT launches the service.
Scaling and multi-tenancy. No horizontal scaling infrastructure — the SimpleOrchestrationProvider (py/core/providers/orchestration/simple.py) is a stub with no background worker. A HatchetOrchestrationProvider exists for orchestration via an external Hatchet server. Multi-tenancy is handled via collection-based scoping: users and documents belong to collections, and all queries filter by collection_ids (py/core/providers/database/chunks.py:178-179 creates a GIN index on collection_ids). The database schema uses PostgreSQL schemas (one per project via R2R_PROJECT_NAME) for logical isolation.
simple orchestration is not a stub; it runs ingestion and graph workflows inline inside the HTTP request (Hatchet is used only by the full configs). Redis is not used: the scheduler is in-process APScheduler.