# How is quality evaluated or observed?

> RAG engines — a good answer covers: Built-in evals, metrics, tracing / observability hooks; if absent, say so.

Canonical page: https://llms-technical-reviews.com/rag/q/evaluation/

## Verdict

[LlamaIndex](/p/llama_index/) and [Haystack](/p/haystack/) are the only projects with answer-quality evaluators you can run. Among the applications, [Onyx](/p/onyx/) has the most usable evaluation and tracing.

**Evaluator libraries.** LlamaIndex ships faithfulness, relevancy, correctness and pairwise LLM judges. It also has `RetrieverEvaluator` (hit rate, MRR) and `BatchEvalRunner`, which runs evaluations concurrently. Haystack's evaluators are pipeline components: faithfulness, context relevance, MAP, MRR, NDCG, recall and semantic answer similarity. Its tracer interface is pluggable, and anonymous usage telemetry is on by default.

**Product-level evals and tracing.** Onyx runs chat datasets locally or in Braintrust and checks which tools were called. Every LLM flow is traced to Langfuse or Braintrust. It has no RAGAS-style metrics. [RAGFlow](/p/ragflow/) traces each chat to Langfuse (per-tenant keys) and has OpenTelemetry spans. We found its evaluation tables to be schema only: no service, route or UI runs them. [Quivr](/p/quivr/) exposes metrics, OpenTelemetry and per-organisation rollups. It also lets you index a new embedding model as an evaluation vector space next to the live one, which compares models but does not score answers.

**Runtime signals only.** [Kotaemon](/p/kotaemon/) has an LLM grade each retrieved chunk and computes a `qa_score` from logprobs. It warns when relevance falls below 0.3. We found no evaluation harness, but the evidence was incomplete. [AnythingLLM](/p/anything-llm/) records tokens, speed and cost for each message. [R2R](/p/r2r/) has Sentry and JSON logs, plus a Ragas cookbook that runs outside the server. [RAG-Anything](/p/rag-anything/)'s `MetricsCallback` counts documents and times each stage. [PageIndex](/p/pageindex/) turns Agents SDK tracing off. Its only quality checks run during indexing (`verify_toc`, tree cost before and after optimisation).

Pick: LlamaIndex or Haystack to measure retrieval and faithfulness in code.
Pick: Onyx or RAGFlow for production traces in Langfuse.
Pick: for every other project, plan to add an external evaluation framework such as Ragas.

## Per-project answers

### infiniflow/ragflow (answered)

**Database-backed evaluation framework:** RAGFlow has a built-in evaluation system with four entity types (`entity/evaluation.go`): `EvaluationDataset` (named groups of test cases with KB bindings), `EvaluationCase` (question + reference_answer + relevant_doc_ids/chunk_ids per case), `EvaluationRun` (ties a dataset to a dialog/chat config with a `MetricsSummary` JSON blob), and `EvaluationResult` (per-case generated_answer, retrieved_chunks, Metrics, execution_time, token_usage). Runs have a status lifecycle (PENDING → ...) and store a `ConfigSnapshot` of the chat configuration at run time.

**Metrics:** Individual `EvaluationResult` records carry a `Metrics` JSONMap that stores arbitrary evaluation metrics per case. The `EvaluationRun` aggregates these into a `MetricsSummary` JSONMap. The evaluation entity types include both `execution_time` and `token_usage` tracking for cost monitoring. This provides an offline benchmark framework rather than continuous production metrics — users create datasets of question-answer pairs, run the chat against them, and inspect results.

**Langfuse observability:** Beyond the built-in evaluation tables, RAGFlow integrates with Langfuse for production tracing (`service/langfuse.go:45-57`). Per-tenant Langfuse credentials are stored in `TenantLangfuse` and resolved at request time. The chat pipeline (phase 3, `chat_pipeline.go:311-326`) creates a Langfuse trace on each chat invocation when configured, posting trace metadata including stream mode, KB count, and session/user IDs. Observations are posted async via a background worker.

**No built-in metrics library:** The system does not include an implementation of standard RAG evaluation metrics like Faithfulness, Answer Relevancy, or Context Precision. The evaluation framework stores whatever metrics are computed externally or via custom scripts — it provides the storage infrastructure (datasets, cases, runs, results) and the mechanism to execute a dialog against a set of test questions, but the metric computation itself is not shipped in this Go codebase. What it stores in the JSONMap fields is up to the operator.

> **Editor's note.** Correction: the evaluation entities in `entity/evaluation.go` are schema only. No Go service, route or web UI creates or runs evaluations at this commit. Observability also includes OpenTelemetry spans in the harness (`internal/harness/graph/pregel/otel_telemetry.go`).

Citations: [internal/entity/evaluation.go:1-97](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/entity/evaluation.go#L1-L97) · [internal/service/langfuse.go:38-80](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/langfuse.go#L38-L80) · [internal/service/chat_pipeline.go:309-327](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/service/chat_pipeline.go#L309-L327)

### Mintplex-Labs/anything-llm (answered)

There is no built-in RAG evaluation framework — no LLM-as-judge metrics, no answer relevancy/faithfulness scoring, no ground-truth datasets, and no A/B comparison tool. The evaluation story is limited to operational observability:

**Telemetry.** An anonymous telemetry system (`Telemetry.sendTelemetry()`) reports high-level usage events (`server_boot`, `sent_chat`, `embed_sent_chat`, `workspace_created`) when `DISABLE_TELEMETRY` is not set. These include commit hash, LLM provider, embedder, vector-db type, and multi-user mode. In development mode telemetry is stubbed.

**Event logging.** The `EventLogs` model logs events (`sent_chat`, `embed_created`, `embed_updated`, etc.) to the Prisma SQLite database with user ID, timestamp, and metadata JSON. Events can be queried by type or user. These serve as an audit trail but provide no quality metrics.

**Per-request metrics.** The `LLMPerformanceMonitor` captures per-request metrics: `prompt_tokens`, `completion_tokens`, `total_tokens`, `duration` (seconds), `outputTps` (tokens per second). These are stored alongside each chat message and exposed via the API as `metrics`. Cost tracking (`addChatCostToMetrics`) enriches metrics with per-model pricing when pricing data is configured.

**Observability hooks.** The system uses `EmbeddingWorkerManager.emitProgress()` for real-time chunk-embedding progress and `abortConnectorOnClientDisconnect()` for signaling. HTTP logging is optional via middleware. There are no OpenTelemetry, Langfuse, or LangSmith integrations — nor any built-in dashboard for RAG quality metrics.


Citations: [server/utils/telemetry/index.js:1-33](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/telemetry/index.js#L1-L33) · [server/models/eventLogs.js:1-129](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/models/eventLogs.js#L1-L129) · [server/utils/helpers/chat/LLMPerformanceMonitor.js:1-20](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/chat/LLMPerformanceMonitor.js#L1-L20) · [server/utils/chats/stream.js:340-365](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/chats/stream.js#L340-L365) · [server/endpoints/chat.js:76-93](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/endpoints/chat.js#L76-L93)

### run-llama/llama_index (answered)

LlamaIndex has a substantial **built-in evaluation suite** in `llama_index.core.evaluation`. Evaluators are subclasses of `BaseEvaluator` and include:
- **`FaithfulnessEvaluator`** (`llama_index/core/evaluation/faithfulness.py:16-44`): LLM-judge that checks if each claim in the response is supported by the context, answering YES/NO with example few-shot prompts.
- **`RelevancyEvaluator`** / **`AnswerRelevancyEvaluator`**: measure how relevant the response is to the query and context.
- **`CorrectnessEvaluator`**: compares response against a reference answer.
- **`ContextRelevancyEvaluator`**: evaluates whether retrieved context is relevant to the query.
- **`SemanticSimilarityEvaluator`**: embedding-based similarity between response and reference.
- **`PairwiseComparisonEvaluator`**: A/B comparison of two responses.
- **`GuidelineEvaluator`**: checks responses against custom guidelines.

**Retrieval-specific evaluation**: `RetrieverEvaluator` evaluates retrieval quality with `HitRate` and `MRR` from `llama_index/core/evaluation/retrieval/metrics.py`.

**Batch evaluation**: `BatchEvalRunner` (`llama_index/core/evaluation/batch_runner.py:75-80`) runs multiple evaluators across queries in parallel with semaphore concurrency control and exponential-backoff retries.

**Observability/tracing**: Two-layer system — (1) `CallbackManager` (`llama_index/core/callbacks/base.py:28-80`) with `CBEventType` events (`CHUNKING`, `NODE_PARSING`, `EMBEDDING`, `LLM`, `QUERY`, `RETRIEVE`, `SYNTHESIZE`, etc.) and contextvar-based trace stacking. (2) The `instrumentation` package (`llama_index/core/instrumentation/__init__.py`) provides typed events (`QueryStartEvent`, `RetrievalEndEvent`, `SynthesizeStartEvent`, etc.) with `@dispatcher.span` decorators on key methods. `NullEventHandler` and `NullSpanHandler` are defaults; users register custom handlers for production monitoring.


Citations: [llama-index-core/llama_index/core/evaluation/faithfulness.py:16-44](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/faithfulness.py#L16-L44) · [llama-index-core/llama_index/core/evaluation/batch_runner.py:75-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/batch_runner.py#L75-L80) · [llama-index-core/llama_index/core/evaluation/__init__.py:1-44](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/evaluation/__init__.py#L1-L44) · [llama-index-core/llama_index/core/callbacks/base.py:28-80](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/callbacks/base.py#L28-L80) · [llama-index-core/llama_index/core/callbacks/schema.py:16-47](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/callbacks/schema.py#L16-L47)

### The-Vibe-Company/quivr (answered)

**No built-in retrieval quality evals** (no RAGAS, nDCG, Recall@k, or LLM-as-judge). Quality assessment is operational/observational, not automated benchmarking. **Observability** has two tiers: (1) Per-process Prometheus-style metrics (`observability/metrics.go:46-57`) at `GET /metrics` — plugin calls by plugin/operation/outcome, search durations by mode/profile, matches created. (2) Rollup metrics (`observability/rollup.go:16-36`) aggregated in-memory and flushed every 5 seconds as time-bucketed rows per Organization — series for plugin_call, search, step, search_query (opt-in, 7-day retention), received documents, and connector_push. Latency histograms with 12 pre-defined bucket boundaries from 5ms to 30s. **OpenTelemetry** instrumentation is compiled in for traces and metrics, with OTLP gRPC/HTTP exporters (go.mod:14-23) and Temporal-OpenTelemetry integration. **Monitoring subsystem** (monitoring/monitoring.go): Saved Queries (persistent search definitions) paired with Subscriptions and evaluator plugins (keyword, vector similarity, or plugin-based) fire Matches via webhook when new documents match — this is alerting, not quality measurement. **Ingestion evaluation** (processing/evaluation.go:26-95) runs after served publication as a separate activity on dedicated capacity, projecting vectors for assessment spaces without affecting served search. **Quality testing** exists through certification and parity tests: golden test vectors for core-ingest (plugins/core-ingest/README.md:46-52), nightly vector parity verification, and controlled retrieval baselines for core-retrieve (plugins/core-retrieve/README.md:50-58) across all three search modes with 24 CC0 queries.


Citations: [internal/observability/metrics.go:46-89](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/observability/metrics.go#L46-L89) · [internal/observability/rollup.go:1-80](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/observability/rollup.go#L1-L80) · [internal/observability/recorder.go:47-60](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/observability/recorder.go#L47-L60) · [internal/processing/evaluation.go:26-95](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/processing/evaluation.go#L26-L95) · [plugins/core-ingest/README.md:46-52](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/plugins/core-ingest/README.md#L46-L52)

### VectifyAI/PageIndex (answered)

**No built-in evaluation framework.** PageIndex has no built-in evaluation harness, no metrics (accuracy, precision, recall, F1, faithfulness), no online evaluation, no A/B testing infrastructure, and no tracing/observability hooks for LLM calls. There is no evaluation module, no eval datasets shipped with the repo, and no logging of retrieval or generation quality.

**What exists instead.** The project's evaluation is conducted through a **separate benchmark repository** at [PageIndex-OSS-Benchmark](https://github.com/VectifyAI/PageIndex-OSS-Benchmark) (`README.md:137-149`), which measures the quickstart setup on 62 lookup questions over 34 PDFs from MMLongBench-Doc-V2. Results (accuracy vs. cost per question) are reported externally, not produced by the SDK itself. FinanceBench results (98.7% accuracy) were achieved by the separate [Mafin2.5](https://github.com/VectifyAI/Mafin2.5-FinanceBench) system (`README.md:163-175`).

**TOC verification.** The standard pipeline does have a **TOC quality check** — it verifies a random sample of TOC entries by checking whether each section title actually appears on its assigned page (`page_index_classic.py:1066-1120`). If accuracy is <60%, it falls back to a more expensive extraction method (TOC-with-numbers → TOC-without-numbers → no-TOC). Items with incorrect page indices are fixed with up to 3 retries (`fix_incorrect_toc_with_retries`, line 1044-1060). This is a correctness gate during indexing, not an evaluation metric.

**Tree optimization metrics.** The `tree_optimize.py` module computes worst-case search cost (in pages) before and after merge/expand passes, reported in the `optimize` key of flash results (`flash/api.py:120-125`). This measures the efficiency of the tree structure for agent navigation, not answer quality.

**Observability.** The codebase uses Python's `logging` module throughout but defines no tracing, spans, or observability exports. The `disable_tracing=True` flag is set in the OpenAI Agents SDK `RunConfig` (`local_chat.py:486-489`), explicitly disabling tracing. No OpenTelemetry, LangSmith, or similar integrations exist.

**Cloud-only features.** The managed cloud chat endpoint returns usage statistics (token counts) and citation objects in responses (`cloud_api.py:340`). No evaluation-specific endpoints are exposed.


Citations: [pageindex/page_index_classic.py:1066-1120](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1066-L1120) · [pageindex/page_index_classic.py:1044-1060](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1044-L1060) · [pageindex/tree_optimize.py:1-54](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/tree_optimize.py#L1-L54) · [pageindex/local_chat.py:486-489](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L486-L489)

### onyx-dot-app/onyx (answered)

**Built-in eval framework:** Onyx has an eval system at `backend/onyx/evals/eval.py` (line 1+) with a module-level entry point that sends multi-turn or single-turn chat messages and evaluates the responses. The `EvalConfigurationOptions` (models.py, lines 86–98) configures the LLM, search permissions email, tool types to enable, dataset name, and Braintrust project/experiment name.

**Eval providers:** The framework supports two eval backends. **Braintrust** (`backend/onyx/evals/providers/braintrust.py`): runs evaluation tasks via the Braintrust SDK with a `tool_assertion_scorer` that scores whether expected tools were called, tracks pass/fail per turn, and reports `MultiTurnEvalResult` or `EvalToolResult`. It supports both local data and remote datasets from Braintrust. **Local** (`backend/onyx/evals/providers/local.py`): runs the task inline with a simple pass/fail assertion check. No persistent dashboard UI for local results.

**Assertions:** Per-turn `ToolAssertion` (models.py, lines 14–18) specifies which tools should be called and whether all must be invoked (`require_all`). The `EvalToolResult` records tools called, citations, timing, and assertion pass/fail.

**Timing:** `EvalTimings` (models.py, lines 21–29) captures `total_ms`, `llm_first_token_ms`, per-tool execution times, and stream processing time.

**Tracing/Observability:** All LLM calls are tagged with a standardized `LLMFlow` enum (`backend/onyx/tracing/flows.py`, lines 16–68). The tracing framework (`tracing/framework/create.py`) can send spans to **Langfuse** (`backend/onyx/tracing/langfuse_tracing_processor.py`) and **Braintrust** (`backend/onyx/tracing/braintrust_tracing_processor.py`). Every LLM invocation (chat, embedding, rerank, image generation, voice, intent classification) opens a generation span tagged with its flow. The indexing pipeline also has its own tracing (`INDEXING_PIPELINE_TRACE_NAME`).

**User usage tracking:** The `user_usage_processor.py` (tracing/processors/) tracks per-user token usage for monitoring and quota enforcement.

However, Onyx does not have built-in RAGAS-style evaluation metrics (faithfulness, answer relevancy, context precision) computed automatically. Quality evaluation relies on the Braintrust integration for structured LLM-as-judge scoring and manual assertion checks.


Citations: [backend/onyx/evals/eval.py:1-80](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/evals/eval.py#L1-L80) · [backend/onyx/evals/models.py:86-108](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/evals/models.py#L86-L108) · [backend/onyx/evals/providers/braintrust.py:1-160](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/evals/providers/braintrust.py#L1-L160) · [backend/onyx/tracing/flows.py:16-68](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/tracing/flows.py#L16-L68) · [backend/onyx/evals/models.py:14-29](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/evals/models.py#L14-L29)

### deepset-ai/haystack (answered)

**Built-in evaluators** – Haystack ships a set of component-based evaluators under `haystack/components/evaluators/`. The `FaithfulnessEvaluator` (`haystack/components/evaluators/faithfulness.py:53-269`) uses an LLM to split a generated answer into statements and check each against the provided contexts, producing a score of 0–1 representing the proportion of supported statements. The `ContextRelevanceEvaluator` similarly scores how relevant retrieved contexts are to a question. Other built-in metrics include `AnswerExactMatchEvaluator`, `DocumentMAPEvaluator` (Mean Average Precision), `DocumentMRREvaluator` (Mean Reciprocal Rank), `DocumentNDCGEvaluator` (Normalized Discounted Cumulative Gain), `DocumentRecallEvaluator`, and `SASEvaluator` (Semantic Answer Similarity). The `LLMEvaluator` (`haystack/components/evaluators/llm_evaluator.py:25-80`) is the base class for LLM-as-judge evaluations: user-defined instructions, input/output schema, and few-shot examples guide an LLM (default: OpenAI in JSON mode) to produce structured scores. All evaluators follow the Haystack `@component` pattern and can be plugged into Pipelines.

**Tracing / observability** – Haystack provides an abstract `Tracer` interface (`haystack/tracing/tracer.py:77-94`) with `trace()` context managers and `Span` objects for instrumenting operations. A built-in `LoggingTracer` (`haystack/tracing/logging_tracer.py:35-50`) logs operation names and tags with customizable ANSI coloring. Third-party tracers (e.g., OpenTelemetry) can be integrated by subclassing `Tracer` and calling `enable_tracing()`. Content-level tracing (queries, documents, answers) is opt-in via the `HAYSTACK_CONTENT_TRACING_ENABLED` environment variable.

**Telemetry** – Haystack uses PostHog for anonymous usage analytics (`haystack/telemetry/__init__.py`), tracking when pipelines are executed.

**What is absent** – There is no built-in evaluation result database or dashboard; the `EvalRunResult` dataclass (`haystack/evaluation/eval_run_result.py`) provides a lightweight container, but users are expected to aggregate and log results themselves.


Citations: [haystack/components/evaluators/llm_evaluator.py:25-80](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/evaluators/llm_evaluator.py#L25-L80) · [haystack/tracing/tracer.py:1-94](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/tracing/tracer.py#L1-L94) · [haystack/tracing/logging_tracer.py:35-50](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/tracing/logging_tracer.py#L35-L50)

### Cinnamon/kotaemon (insufficient evidence)

There is no built-in evaluation framework (no RAGAS, TruLens, DeepEval, or similar eval harness integrated). The system does provide runtime quality signals: (1) `LLMTrulensScoring` (`libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py`) scores retrieved documents on a 0–10 relevance scale per query; (2) `qa_score` is derived from the exponential average of LLM response log-probabilities (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, lines 274–277); (3) a context-relevance warning threshold (`CONTEXT_RELEVANT_WARNING_SCORE`) triggers a UI warning when the max LLM relevance score is below 0.3 (`libs/kotaemon/kotaemon/indices/qa/citation_qa.py`, line 36–38). There is no built-in logging/tracing system (e.g., LangSmith, Weights & Biases, MLflow) connected to the pipeline — observability would need to be added externally. I checked the Kotaemon library and the ktem app modules; no evaluation harness or metric tracker was found.


Citations: [libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py:96-182](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py#L96-L182) · [libs/kotaemon/kotaemon/indices/qa/citation_qa.py:274-277](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/qa/citation_qa.py#L274-L277) · [libs/kotaemon/kotaemon/indices/qa/citation_qa.py:36-38](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/qa/citation_qa.py#L36-L38)

### HKUDS/RAG-Anything (answered)

**Built-in evals.** RAG-Anything has no built-in evaluation framework (no RAGAS, no faithfulness/relevance metrics, no answer-grounding checks). The only quantitative measurement is the `MetricsCallback` (`callbacks.py:183-287`), which aggregates processing pipeline statistics: documents processed/failed, content blocks parsed, parse/insert/multimodal processing times, and query counts/times. It exposes a `summary()` method returning a human-readable report and a `reset()` method. This is strictly operational telemetry, not quality evaluation.

**Tracing / observability hooks.** The `ProcessingCallback` base class (`callbacks.py:61-180`) defines lifecycle hooks dispatched by `CallbackManager`: `on_parse_start/complete/error`, `on_text_insert_start/complete`, `on_multimodal_start/item_complete/complete`, `on_query_start/complete/error`, `on_document_complete/error`, and `on_batch_start/complete`. The `ProcessingEvent` dataclass captures timestamps, doc_ids, file paths, stages, and error details. Users can subclass `ProcessingCallback` to integrate with any observability system (e.g., a custom class could send events to OpenTelemetry, Datadog, or a log aggregator). The `CallbackManager` also supports optional internal event logging (`enable_event_log(true)`) and thread-safe callback registration (`callbacks.py:298-346`).

**LLM response caching.** Both text and multimodal queries can be cached via LightRAG's `llm_response_cache` (`query.py:297-380`), which avoids re-generating answers for identical inputs. Cache keys incorporate content fingerprints (SHA-256 of file content, MD5 of query parameters) to detect changes (`query.py:59-126`).

There are no eval benchmarks, no retrieval-quality metrics (hit rate, MRR, precision), and no generation quality metrics (faithfulness, helpfulness) in the codebase.


Citations: [raganything/callbacks.py:61-180](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/callbacks.py#L61-L180) · [raganything/callbacks.py:183-287](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/callbacks.py#L183-L287) · [raganything/callbacks.py:290-365](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/callbacks.py#L290-L365) · [raganything/query.py:59-126](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L59-L126) · [raganything/query.py:284-380](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/query.py#L284-L380)

### SciPhi-AI/R2R (answered)

R2R has **no built-in evaluation framework** — no scoring, no LLM judge, no integrated metrics like faithfulness or answer relevancy. The project ships with two observability mechanisms instead:

- **Sentry integration** (`py/core/utils/sentry.py`). Initialized at app startup via `init_sentry()`, reads `R2R_SENTRY_DSN` from the environment. Supports configurable traces_sample_rate and profiles_sample_rate. This provides error tracking and performance tracing but no RAG-specific quality metrics.

- **Structured logging** (`py/core/utils/logging_config.py`). Configures python-json-logger with HTTP status filters for uvicorn.access logs. Health endpoint requests are filtered out. There is no RAG-specific logging (no retrieval latency, generation quality, or user feedback capture).

- **External evaluation guide.** The project provides a cookbook at `docs/cookbooks/evals.md` demonstrating how to evaluate R2R outputs using the Ragas framework externally. It shows collecting R2R's RAG responses via `client.retrieval.rag()` and feeding them into Ragas for metrics like faithfulness, answer_relevancy, context_precision, and context_recall. This is an optional, bring-your-own-eval approach — Ragas is not a dependency and no integration code exists in the source tree.

There are no built-in evaluation endpoints, no A/B testing infrastructure, no LLM-as-judge calls, and no quality monitoring dashboards. Users must integrate external frameworks (Ragas, TruLens, LangFuse, etc.) themselves.


Citations: [py/core/utils/logging_config.py:1-50](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/utils/logging_config.py#L1-L50) · [docs/cookbooks/evals.md:1-50](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/docs/cookbooks/evals.md#L1-L50)
