LLMs Technical Reviews
Home / RAG engines / quivr

The-Vibe-Company/quivr

Go engine that ingests content streams durably and serves hybrid search and alerts; parsing, embedding and ranking live in plugins.

GitHub ↗★ 40kGoMITcommit 63d76a9 · 2026-10-06homepage ↗

Overview

At the pinned commit, Quivr is “Quivr V2”: a Go service that ingests a continuous stream of content, makes it searchable, and lets you follow topics over time with saved queries and alerts. It is not the earlier Python “chat with your files” library that made the name popular. There is no LLM answer step anywhere in the engine. A search returns ranked segments with exact provenance (record, version, part, and Unicode code-point offsets), and producing an answer from them is left to the caller or to an agent over MCP.

The design takes durability seriously. Every write carries an idempotency key and is acknowledged only once it is committed to PostgreSQL. Corrections create new immutable record versions. The Weaviate index is treated as a rebuildable projection of canonical data held in PostgreSQL and S3-compatible storage. The engine also owns very little “AI”: since an internal refactor, it does not segment, embed or rank anything itself. Those jobs belong to pinned plugins that run as HTTP sidecars, and the API and worker refuse to start without an ingestion and a retrieval plugin.

The audience is teams that need a self-hosted retrieval backbone for news-like feeds (RSS, X lists, NewsML-G2, mailboxes, archives) more than a document chatbot. The README labels the API v0 and says it is at evaluation stage, not for production.

Architecture

flowchart LR
  C["Client / connector"] --> API["quivr api (REST v0, MCP)"]
  API --> PG["PostgreSQL: catalog, receipts, intents"]
  PG --> DISP["Dispatcher (polls intents)"]
  DISP --> TMP["Temporal"]
  TMP --> W["quivr worker (activities)"]
  W --> NORM["Normalizer plugins"]
  W --> ING["Ingestion plugin (core.ingest)"]
  ING --> TEI["TEI: multilingual-e5-small"]
  W --> S3["S3 / SeaweedFS artifacts"]
  W --> WV["Weaviate (BM25 + vectors)"]
  API --> RET["Retrieval plugin (core.retrieve)"]
  RET --> API
  API --> WV
  API --> MON["Monitoring: saved queries, matches, webhooks"]
Component Path Role
Binary entry cmd/quivr/main.go One binary: api, worker, migrate, plugin tooling and online commands such as mcp
Process setup internal/app/run.go Reads the QUIVR_CONFIG JSON, loads plugin pins, wires stores and services
HTTP API internal/transport/httpapi/ Strict OpenAPI dispatch for records, search, connectors, monitoring, admin
Content domain internal/content/ Commands, receipts, versions, manifests, segmentations, embedding artifacts
Orchestration internal/orchestration/temporal/ Intent dispatcher, ingestion workflow, backfill, rebuild and evaluation workflows
Processing internal/processing/ Baseline segmentation and indexing, then enrichment with vectors
Retrieval internal/retrieval/ Authorization, generation routing, and the multi-round plugin ranking protocol
Adapters internal/adapters/{postgres,weaviate,s3,tei,pluginhttp} Storage, search index, embedding server and plugin HTTP clients
Monitoring internal/monitoring/ Saved Query and Subscription versions, Matches and webhook delivery
Plugins plugins/ core-ingest, core-retrieve, jev-rerank, hosted-embed, pdf-text, newsml-g2, alerts, connectors
Plugin SDKs sdks/go, sdks/python Libraries for writing sidecar plugins
Demo UI quivr-search/ Node/Vite browser UI over the same API

How a request flows

Ingestion of one text record, then a hybrid search:

  1. Accept. POST /v0/records reaches handleIngestRecord, which validates the body against the OpenAPI schema and calls Submitter.Accept. It answers 202 with an ingestion receipt and a Location header (content.go). content.Service.Accept checks the content:accept permission and corpus scope, and refuses forged normalization provenance (content.go).
  2. Dispatch. An accepted receipt becomes a durable dispatch intent. A dispatcher polls ingestion intents every 200 ms, claims a batch and starts up to eight Temporal workflows in parallel. “Already started” counts as success, so a crash between start and completion is safe (dispatch.go).
  3. Baseline. materializeWorkflow runs process-token-windows-v2. If the version is a routed blob (a PDF, for example), it runs normalize-external and loops, at most three rounds (ingestion.go). The activity calls processing.Service.Run. It materializes the version, asks the ingestion plugin for segments only, and publishes them to Weaviate with Retrieval.Index. At this point the text is keyword-searchable (processing.go).
  4. Enrich. The workflow then runs enrich-e5, which calls Service.Enrich. That step gets vectors for the same segments from the plugin’s SegmentAndEmbed, stores them as embedding artifacts and projects them. A plugin that returns different segments with vectors than it did without is blocked as a derivation_conflict instead of silently overwriting (enrichment.go, plugin.go).
  5. Search. POST /v0/search takes one of 64 per-process search slots, or fails at once with Retry-After: 1 (search.go). retrieval.Service.Search defaults to hybrid mode and limit 10. It allows up to 16 corpora, authorizes each one and routes it to its projection generation. It refuses metadata or namespace filters that an older generation cannot honour, instead of dropping hits silently (retrieval.go).
  6. Rounds. rankProfile opens a retrieval session under the profile’s deadline and cost budget. Each round, the plugin either asks for candidates (a primitive plus a space) or returns a final ranking. The engine serves candidate requests against Weaviate, encodes the query with the space’s owner, and hydrates each hit from canonical storage (plugin.go, L516-L557). The Weaviate query is GraphQL bm25, nearVector or hybrid over title^2 and body (projection.go).
  7. Response. Each hit carries record, version, part, segment and generation ids, plus an excerpt with unicode_codepoint offsets and the ranking explanation.

Key components

Plugin contract

A plugin manifest declares up to five contributions: normalizer, subscription, connector, ingestion and retrieval (manifest.go). Plugins are separate processes that the engine calls over HTTP. Each answer is checked against JSON Schemas in contracts/. A retrieval round, for example, is one call.Invoke with a size cap and typed error envelopes (retrieval.go). Work is pinned to the “pipeline plan” it started on, so a plugin upgrade does not change an ingestion halfway through.

core.ingest (segmentation and embedding)

The default ingestion plugin cuts title and body parts into windows of at most 384 tokens with a 48-token overlap. Windows prefer paragraph, line and sentence ends. It embeds them with a pinned revision of intfloat/multilingual-e5-small served by TEI, using E5’s passage: and query: prefixes. It refuses inputs over 256 KiB, 64 parts or 256 segments (profile.json). hosted-embed is an optional alternative that calls OpenAI-compatible or Cohere embedding APIs and declares its own vector space.

Vector spaces, evaluation and promotion

Each ingestion plugin owns its vector space, and a space never changes owner. A deployment can index an evaluation space next to the served one on the same segments, compare them with evaluation_plugin/evaluation_space searches, and promote the new space later. The former one stays, so rollback is another promotion. This is more model-migration machinery than most RAG engines have.

core.retrieve and reranking

The default retrieval plugin makes one request per served space: bm25 for lexical, near_vector for semantic, or hybrid. Hybrid uses alpha 0.5 and relative-score fusion unless configured otherwise. It then merges by score, breaks ties by segment id and dedupes (main.go). Its deep profile currently ranks the same way as default, with a 3 s / 1 cent budget reserved (quivr-plugin.yaml). Real reranking comes from the optional jev.rerank plugin, which takes over deep and sends candidates to an external TypeSafe service (README).

Monitoring

internal/monitoring keeps immutable Saved Query and Subscription versions. Evaluation turns newly eligible record versions into unique Matches with pending webhook deliveries. Alert logic itself is pluggable: the first-party alerts plugin offers keyword, local vector “meaning” and Jev-checked “described” alerts.

MCP

quivr mcp exposes read and ingest tool profiles. The instructions tell agents to cite hits by record_id, version_id, part_key and code-point offsets, and to poll the receipt until the text is searchable (mcp_catalogue.go).

Extending it

  • Write a plugin. quivr plugin init scaffolds a Python normalizer, subscription or connector. dev runs it against fixtures with --watch, inspect validates the manifest, and test certifies every declared contribution (cli.go). Go plugins use sdks/go/quivrplugin.
  • New formats. Add a normalizer that maps a media type to a manifest of text parts. pdf-text (pypdf, one part per page, no OCR) is the worked example.
  • New embedding model. Pin another ingestion plugin as an evaluation space, compare, then promote.
  • New ranking. Pin a retrieval plugin and map a profile name to it. The engine enforces its round, request, candidate, latency and cost limits.
  • New sources. Connector plugins (RSS, X list, M365 mail, object-storage archives, push sources) feed the same idempotent record API.

Running it

  • Local. make dev starts PostgreSQL 17, Temporal, SeaweedFS, Weaviate and TEI with Docker Compose, runs migrations and launches the API and the worker. The first run downloads the E5 model (about 1 GB). It needs Go, Docker Compose v2, Python 3, Node.js 22+ and jq. make demo adds the browser UI.
  • Configuration. One JSON file named by QUIVR_CONFIG, decoded with unknown fields rejected. It needs at least a database_url, a 32-byte cursor_key and one scoped API key (run.go).
  • Commands. quivr api, quivr worker and quivr migrate come from the same binary. Plugin tooling runs without a config or stack (main.go).

Strengths and caveats

  • Strength: correctness under failure. Idempotent writes, an outbox-style dispatcher, Temporal retries with heartbeats, and searches that error instead of returning partial “success”. Few open-source RAG backends are this careful.
  • Strength: precise provenance. Every hit is rehydrated and re-authorized from canonical storage, with code-point offsets an LLM front end can cite exactly.
  • Strength: model migration built in. Owned vector spaces, evaluation spaces and promotion make changing embedding models a routine operation.
  • Caveat: not a chat-with-docs product. There is no generation, prompt assembly or conversation memory. You bring the LLM layer.
  • Caveat: heavy stack. Five infrastructure services plus plugin sidecars, even for a small corpus.
  • Caveat: thin document parsing. First-party normalizers cover PDF text (no OCR) and NewsML-G2. Office formats, tables and layout need your own normalizer.
  • Caveat: pre-production. v0 API, explicitly marked evaluation stage. Reranking needs a third-party paid service unless you write your own retrieval plugin.

Sources: code at 63d76a9, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the RAG engines questions

Each answer was drafted by a code-reading agent at commit 63d76a9. Its citations were checked mechanically. Compare with the other rag engines →

How are documents parsed and chunked?

answered

The engine does zero parsing or chunking itself — it delegates entirely to external plugins. Documents arrive as Command structs (content.go:42-80) containing a Manifest of typed Parts (text, blob, etc.). A normalizer plugin (normalization/normalization.go:1-14) runs first for routed binary blobs, converting them into the canonical Manifest/Part structure (the engine ships the NewsML-G2 normalizer for IPTC news formats as an example). The core.ingest plugin (plugins/core-ingest/README.md) handles standard text: it only processes title and body text Parts, cutting body Parts into windows of at most 384 tokens (controlled by profile.json), preferring paragraph/line/sentence boundaries, with 48 token overlap. A title with no body Part is one segment. Refusal limits: 256 KiB text, 2 titles, 64 Parts, or 256 windows (README.md:21-24). The tokenizer (tokenizer.py, tokenizer.json) is a helper Python process per plugin process, loaded once and kept for the plugin's lifetime (plugins/core-ingest/main.go:53-67). The IngestionPlugin.SegmentAndEmbed interface (processing/plugin.go:21) is the single entry point — the plugin returns segments with per-space vectors, and the engine stores the segmentation and embedding artifacts durably. The engine validates the output via the Contract Runner's rules (pluginhttp/ingestion.go:91). No OCR or layout analysis is present in the engine or first-party plugins.

How are embeddings and indexes built and stored?

answered

Embedding model: intfloat/multilingual-e5-small (384 dimensions, cosine distance), served through a TEI (Text Embeddings Inference) container (deploy/compose/compose.yaml:65-79). The TEI encoder (adapters/tei/encoder.go:26-65) validates 384-dim unit-norm vectors, prefixes documents with passage: and queries with query: (plugins/core-ingest/main.go:131). The engine's legacy E5 space is an older built-in encoder for pre-THE-777 Corpora (app/run.go:55). Vector store: Only Weaviate is supported (adapters/weaviate/projection.go). Vectors are stored as named vectors per generation: the legacy semantic_text_v1 for pre-named-space generations, or s_<hash> for plugin-owned spaces (projection.go:54-59). Distance metrics include cosine (default), dot, and L2-squared (projection.go:62-70). The index uses HNSW (projection.go:73). Hybrid/BM25: Weaviate provides built-in BM25 alongside vector search. The engine constructs GraphQL queries: bm25:{query} for lexical, nearVector:{vector} for semantic, and hybrid:{query, vector, alpha, fusionType} for hybrid (projection.go:552-558). Keyword fields default to title^2 and body, with an optional lexicalText property from plugins. Metadata: Projected alongside vectors into Weaviate as individual properties (adapters/weaviate/metadata.go). The Generation.MetadataProjected flag gates filter availability (content/searchable.go:58-68); filters apply as Weaviate where clauses (projection.go:480-502). Embeddings are stored as artifacts in SeaweedFS (S3) with SHA-256 content hashes (content/embeddings.go:43-61), carrying full provenance (version, segment, space, producer info).

Editor's note. Correction: TEI with multilingual-e5-small is the default via core.ingest, but the optional hosted.embed plugin embeds through OpenAI-compatible or Cohere v2 APIs in its own vector space, which can be evaluated and promoted.

How is retrieval performed?

answered

The engine itself does no ranking — it delegates to a retrieval plugin via a multi-round protocol. Three modes (retrieval/retrieval.go:209-211): lexical (BM25-only), semantic (vector-only), hybrid (both, the default). The Request struct carries Query, CorpusIDs, Mode, Profile, Limit (max 50), SourceNamespaces (max 50), Metadata filters, and an optional pre-encoded Vector (retrieval/retrieval.go:73-98). Search proceeds in rounds (retrieval/plugin.go:189-256): (1) the plugin is called with a SearchRequest containing the current session state; (2) the plugin responds with either requests for candidates (primitive+space combinations) or a final ranking; (3) the engine serves candidates by querying Weaviate with authorization, generation routing, source namespace, and metadata filters applied (weaviate/projection.go:470-564); (4) the plugin receives the candidates and can request more or return the ranking. The core.retrieve plugin (plugins/core-retrieve/main.go) runs in at most two rounds: round 1 requests candidates (single BM25 for lexical, near_vector per served space for semantic, hybrid per space with configurable alpha 0.5 and fusion type for hybrid), then final ranking merges by descending score, deduplicates per segment, and tie-breaks by segment ID. Query encoding is done by the ingestion plugin that owns each vector space (TEI for legacy E5, plugin's EmbedQuery for plugin-owned spaces). Reranking: Not built in; the deep profile exists for future reranking. Query rewriting/decomposition: Not implemented. Filters: Source namespace filter via SourceNamespaceProjected check (retrieval/retrieval.go:290), metadata filter via MetadataProjected check (line:284) and corpus.ResolveFilters validation.

Editor's note. Correction: reranking is not in the engine or in core.retrieve (whose deep profile ranks like default), but the optional first-party jev.rerank plugin serves deep and reranks candidates through the external TypeSafe API.

How are answers generated and grounded?

answered

Quivr does not implement answer generation. It is a retrieval engine, not an LLM orchestrator — there is no prompt assembly, no LLM call for answer synthesis, no citation formatting, and no streaming of generated text. Instead, it returns ranked segments with precise provenance suitable for an external generator. Every search hit includes record_id, version_id, part_key, segment_id, and an excerpt with exact Unicode code-point [start, end) offsets into the source Part text (transport/httpapi/search.go:75-77). This allows an external LLM frontend to cite sources with verifiable precision. The MCP interface (quivr mcp command) exposes these same capabilities to AI agents — the read profile provides search and read-record tools, and its instructions explicitly tell agents how to cite hits by record_id, version_id, part_key and offset ranges (online/mcp_catalogue.go:41-44). The ingest profile adds document ingestion. The monitoring/evaluation subsystem (monitoring/monitoring.go) supports Saved Queries and Subscriptions that evaluate new documents against persistent queries, producing Matches delivered via webhook — but this is alerting, not generation. For answer generation, users are expected to call Quivr's search API from a separate application that handles prompt construction, context assembly, LLM inference, and citation rendering.

How is quality evaluated or observed?

answered

No built-in retrieval quality evals (no RAGAS, nDCG, Recall@k, or LLM-as-judge). Quality assessment is operational/observational, not automated benchmarking. Observability has two tiers: (1) Per-process Prometheus-style metrics (observability/metrics.go:46-57) at GET /metrics — plugin calls by plugin/operation/outcome, search durations by mode/profile, matches created. (2) Rollup metrics (observability/rollup.go:16-36) aggregated in-memory and flushed every 5 seconds as time-bucketed rows per Organization — series for plugin_call, search, step, search_query (opt-in, 7-day retention), received documents, and connector_push. Latency histograms with 12 pre-defined bucket boundaries from 5ms to 30s. OpenTelemetry instrumentation is compiled in for traces and metrics, with OTLP gRPC/HTTP exporters (go.mod:14-23) and Temporal-OpenTelemetry integration. Monitoring subsystem (monitoring/monitoring.go): Saved Queries (persistent search definitions) paired with Subscriptions and evaluator plugins (keyword, vector similarity, or plugin-based) fire Matches via webhook when new documents match — this is alerting, not quality measurement. Ingestion evaluation (processing/evaluation.go:26-95) runs after served publication as a separate activity on dedicated capacity, projecting vectors for assessment spaces without affecting served search. Quality testing exists through certification and parity tests: golden test vectors for core-ingest (plugins/core-ingest/README.md:46-52), nightly vector parity verification, and controlled retrieval baselines for core-retrieve (plugins/core-retrieve/README.md:50-58) across all three search modes with 24 CC0 queries.

How is it deployed and operated?

answered

Service, not library: Deployed as a distributed service with multiple roles (api, worker, migrate) from a single Go binary (cmd/quivr/main.go). Infrastructure (deploy/compose/compose.yaml:1-90, deploy/railway/services.json): PostgreSQL 17 (metadata, routing, activity), Temporal (workflow orchestration via go.temporal.io/sdk), Weaviate standalone (vector+BM25 search index), TEI container (HuggingFace text-embeddings-inference, CPU-only), SeaweedFS (S3-compatible blob storage for artifacts). Two Go processes — api serves HTTP (REST/JSON on port 8080, admin UI, also serves the web frontend) and worker runs Temporal activities (ingestion pipeline, enrichment, backfill, rebuilds, evaluation). Web UI: Node.js/React app at quivr-search/. API: Documented via OpenAPI at contracts/http/v0/openapi.yaml. Also exposes an MCP interface (quivr mcp) for AI agents. Multi-tenancy: Via API keys mapped to corpus.Scope (Organization + action permissions + corpus list, app/run.go:70). Scalability: Concurrent search limited to 64 per API process (transport/httpapi/search.go:16). Worker capacity is elastic via Temporal task queue quivr-content-v0. The current Railway demo is single-node with 8 containers and no HA for Temporal dev server. Configuration: Single QUIVR_CONFIG JSON file with database URL, S3 credentials, TEI/Weaviate/Temporal addresses, TLS, telemetry, API keys/scopes, plugin pins, delivery/webhook settings, and observability toggles (app/run.go:50-90). Plugins: Run as sidecar processes (Go or Python) communicating over HTTP via Plugin API v0, pinned by deployment configuration. First-party plugins are built into the core Docker image (deploy/images/quivr.Dockerfile); third-party plugins use plugin.Dockerfile.