LLMs Technical Reviews

How is quality evaluated or observed?

Built-in evals, metrics, tracing / observability hooks; if absent, say so.

Verdict

LlamaIndex and Haystack are the only projects with answer-quality evaluators you can run. Among the applications, Onyx has the most usable evaluation and tracing.

Evaluator libraries. LlamaIndex ships faithfulness, relevancy, correctness and pairwise LLM judges. It also has RetrieverEvaluator (hit rate, MRR) and BatchEvalRunner, which runs evaluations concurrently. Haystack’s evaluators are pipeline components: faithfulness, context relevance, MAP, MRR, NDCG, recall and semantic answer similarity. Its tracer interface is pluggable, and anonymous usage telemetry is on by default.

Product-level evals and tracing. Onyx runs chat datasets locally or in Braintrust and checks which tools were called. Every LLM flow is traced to Langfuse or Braintrust. It has no RAGAS-style metrics. RAGFlow traces each chat to Langfuse (per-tenant keys) and has OpenTelemetry spans. We found its evaluation tables to be schema only: no service, route or UI runs them. Quivr exposes metrics, OpenTelemetry and per-organisation rollups. It also lets you index a new embedding model as an evaluation vector space next to the live one, which compares models but does not score answers.

Runtime signals only. Kotaemon has an LLM grade each retrieved chunk and computes a qa_score from logprobs. It warns when relevance falls below 0.3. We found no evaluation harness, but the evidence was incomplete. AnythingLLM records tokens, speed and cost for each message. R2R has Sentry and JSON logs, plus a Ragas cookbook that runs outside the server. RAG-Anything’s MetricsCallback counts documents and times each stage. PageIndex turns Agents SDK tracing off. Its only quality checks run during indexing (verify_toc, tree cost before and after optimisation).

Pick: LlamaIndex or Haystack to measure retrieval and faithfulness in code. Pick: Onyx or RAGFlow for production traces in Langfuse. Pick: for every other project, plan to add an external evaluation framework such as Ragas.

Per-project answers

infiniflow/ragflow

answered

Database-backed evaluation framework: RAGFlow has a built-in evaluation system with four entity types (entity/evaluation.go): EvaluationDataset (named groups of test cases with KB bindings), EvaluationCase (question + reference_answer + relevant_doc_ids/chunk_ids per case), EvaluationRun (ties a dataset to a dialog/chat config with a MetricsSummary JSON blob), and EvaluationResult (per-case generated_answer, retrieved_chunks, Metrics, execution_time, token_usage). Runs have a status lifecycle (PENDING → ...) and store a ConfigSnapshot of the chat configuration at run time.

Metrics: Individual EvaluationResult records carry a Metrics JSONMap that stores arbitrary evaluation metrics per case. The EvaluationRun aggregates these into a MetricsSummary JSONMap. The evaluation entity types include both execution_time and token_usage tracking for cost monitoring. This provides an offline benchmark framework rather than continuous production metrics — users create datasets of question-answer pairs, run the chat against them, and inspect results.

Langfuse observability: Beyond the built-in evaluation tables, RAGFlow integrates with Langfuse for production tracing (service/langfuse.go:45-57). Per-tenant Langfuse credentials are stored in TenantLangfuse and resolved at request time. The chat pipeline (phase 3, chat_pipeline.go:311-326) creates a Langfuse trace on each chat invocation when configured, posting trace metadata including stream mode, KB count, and session/user IDs. Observations are posted async via a background worker.

No built-in metrics library: The system does not include an implementation of standard RAG evaluation metrics like Faithfulness, Answer Relevancy, or Context Precision. The evaluation framework stores whatever metrics are computed externally or via custom scripts — it provides the storage infrastructure (datasets, cases, runs, results) and the mechanism to execute a dialog against a set of test questions, but the metric computation itself is not shipped in this Go codebase. What it stores in the JSONMap fields is up to the operator.

Editor's note. Correction: the evaluation entities in entity/evaluation.go are schema only. No Go service, route or web UI creates or runs evaluations at this commit. Observability also includes OpenTelemetry spans in the harness (internal/harness/graph/pregel/otel_telemetry.go).

Mintplex-Labs/anything-llm

answered

There is no built-in RAG evaluation framework — no LLM-as-judge metrics, no answer relevancy/faithfulness scoring, no ground-truth datasets, and no A/B comparison tool. The evaluation story is limited to operational observability:

Telemetry. An anonymous telemetry system (Telemetry.sendTelemetry()) reports high-level usage events (server_boot, sent_chat, embed_sent_chat, workspace_created) when DISABLE_TELEMETRY is not set. These include commit hash, LLM provider, embedder, vector-db type, and multi-user mode. In development mode telemetry is stubbed.

Event logging. The EventLogs model logs events (sent_chat, embed_created, embed_updated, etc.) to the Prisma SQLite database with user ID, timestamp, and metadata JSON. Events can be queried by type or user. These serve as an audit trail but provide no quality metrics.

Per-request metrics. The LLMPerformanceMonitor captures per-request metrics: prompt_tokens, completion_tokens, total_tokens, duration (seconds), outputTps (tokens per second). These are stored alongside each chat message and exposed via the API as metrics. Cost tracking (addChatCostToMetrics) enriches metrics with per-model pricing when pricing data is configured.

Observability hooks. The system uses EmbeddingWorkerManager.emitProgress() for real-time chunk-embedding progress and abortConnectorOnClientDisconnect() for signaling. HTTP logging is optional via middleware. There are no OpenTelemetry, Langfuse, or LangSmith integrations — nor any built-in dashboard for RAG quality metrics.

run-llama/llama_index

answered

LlamaIndex has a substantial built-in evaluation suite in llama_index.core.evaluation. Evaluators are subclasses of BaseEvaluator and include:

  • FaithfulnessEvaluator (llama_index/core/evaluation/faithfulness.py:16-44): LLM-judge that checks if each claim in the response is supported by the context, answering YES/NO with example few-shot prompts.
  • RelevancyEvaluator / AnswerRelevancyEvaluator: measure how relevant the response is to the query and context.
  • CorrectnessEvaluator: compares response against a reference answer.
  • ContextRelevancyEvaluator: evaluates whether retrieved context is relevant to the query.
  • SemanticSimilarityEvaluator: embedding-based similarity between response and reference.
  • PairwiseComparisonEvaluator: A/B comparison of two responses.
  • GuidelineEvaluator: checks responses against custom guidelines.

Retrieval-specific evaluation: RetrieverEvaluator evaluates retrieval quality with HitRate and MRR from llama_index/core/evaluation/retrieval/metrics.py.

Batch evaluation: BatchEvalRunner (llama_index/core/evaluation/batch_runner.py:75-80) runs multiple evaluators across queries in parallel with semaphore concurrency control and exponential-backoff retries.

Observability/tracing: Two-layer system — (1) CallbackManager (llama_index/core/callbacks/base.py:28-80) with CBEventType events (CHUNKING, NODE_PARSING, EMBEDDING, LLM, QUERY, RETRIEVE, SYNTHESIZE, etc.) and contextvar-based trace stacking. (2) The instrumentation package (llama_index/core/instrumentation/__init__.py) provides typed events (QueryStartEvent, RetrievalEndEvent, SynthesizeStartEvent, etc.) with @dispatcher.span decorators on key methods. NullEventHandler and NullSpanHandler are defaults; users register custom handlers for production monitoring.

The-Vibe-Company/quivr

answered

No built-in retrieval quality evals (no RAGAS, nDCG, Recall@k, or LLM-as-judge). Quality assessment is operational/observational, not automated benchmarking. Observability has two tiers: (1) Per-process Prometheus-style metrics (observability/metrics.go:46-57) at GET /metrics — plugin calls by plugin/operation/outcome, search durations by mode/profile, matches created. (2) Rollup metrics (observability/rollup.go:16-36) aggregated in-memory and flushed every 5 seconds as time-bucketed rows per Organization — series for plugin_call, search, step, search_query (opt-in, 7-day retention), received documents, and connector_push. Latency histograms with 12 pre-defined bucket boundaries from 5ms to 30s. OpenTelemetry instrumentation is compiled in for traces and metrics, with OTLP gRPC/HTTP exporters (go.mod:14-23) and Temporal-OpenTelemetry integration. Monitoring subsystem (monitoring/monitoring.go): Saved Queries (persistent search definitions) paired with Subscriptions and evaluator plugins (keyword, vector similarity, or plugin-based) fire Matches via webhook when new documents match — this is alerting, not quality measurement. Ingestion evaluation (processing/evaluation.go:26-95) runs after served publication as a separate activity on dedicated capacity, projecting vectors for assessment spaces without affecting served search. Quality testing exists through certification and parity tests: golden test vectors for core-ingest (plugins/core-ingest/README.md:46-52), nightly vector parity verification, and controlled retrieval baselines for core-retrieve (plugins/core-retrieve/README.md:50-58) across all three search modes with 24 CC0 queries.

VectifyAI/PageIndex

answered

No built-in evaluation framework. PageIndex has no built-in evaluation harness, no metrics (accuracy, precision, recall, F1, faithfulness), no online evaluation, no A/B testing infrastructure, and no tracing/observability hooks for LLM calls. There is no evaluation module, no eval datasets shipped with the repo, and no logging of retrieval or generation quality.

What exists instead. The project's evaluation is conducted through a separate benchmark repository at PageIndex-OSS-Benchmark (README.md:137-149), which measures the quickstart setup on 62 lookup questions over 34 PDFs from MMLongBench-Doc-V2. Results (accuracy vs. cost per question) are reported externally, not produced by the SDK itself. FinanceBench results (98.7% accuracy) were achieved by the separate Mafin2.5 system (README.md:163-175).

TOC verification. The standard pipeline does have a TOC quality check — it verifies a random sample of TOC entries by checking whether each section title actually appears on its assigned page (page_index_classic.py:1066-1120). If accuracy is <60%, it falls back to a more expensive extraction method (TOC-with-numbers → TOC-without-numbers → no-TOC). Items with incorrect page indices are fixed with up to 3 retries (fix_incorrect_toc_with_retries, line 1044-1060). This is a correctness gate during indexing, not an evaluation metric.

Tree optimization metrics. The tree_optimize.py module computes worst-case search cost (in pages) before and after merge/expand passes, reported in the optimize key of flash results (flash/api.py:120-125). This measures the efficiency of the tree structure for agent navigation, not answer quality.

Observability. The codebase uses Python's logging module throughout but defines no tracing, spans, or observability exports. The disable_tracing=True flag is set in the OpenAI Agents SDK RunConfig (local_chat.py:486-489), explicitly disabling tracing. No OpenTelemetry, LangSmith, or similar integrations exist.

Cloud-only features. The managed cloud chat endpoint returns usage statistics (token counts) and citation objects in responses (cloud_api.py:340). No evaluation-specific endpoints are exposed.

onyx-dot-app/onyx

answered

Built-in eval framework: Onyx has an eval system at backend/onyx/evals/eval.py (line 1+) with a module-level entry point that sends multi-turn or single-turn chat messages and evaluates the responses. The EvalConfigurationOptions (models.py, lines 86–98) configures the LLM, search permissions email, tool types to enable, dataset name, and Braintrust project/experiment name.

Eval providers: The framework supports two eval backends. Braintrust (backend/onyx/evals/providers/braintrust.py): runs evaluation tasks via the Braintrust SDK with a tool_assertion_scorer that scores whether expected tools were called, tracks pass/fail per turn, and reports MultiTurnEvalResult or EvalToolResult. It supports both local data and remote datasets from Braintrust. Local (backend/onyx/evals/providers/local.py): runs the task inline with a simple pass/fail assertion check. No persistent dashboard UI for local results.

Assertions: Per-turn ToolAssertion (models.py, lines 14–18) specifies which tools should be called and whether all must be invoked (require_all). The EvalToolResult records tools called, citations, timing, and assertion pass/fail.

Timing: EvalTimings (models.py, lines 21–29) captures total_ms, llm_first_token_ms, per-tool execution times, and stream processing time.

Tracing/Observability: All LLM calls are tagged with a standardized LLMFlow enum (backend/onyx/tracing/flows.py, lines 16–68). The tracing framework (tracing/framework/create.py) can send spans to Langfuse (backend/onyx/tracing/langfuse_tracing_processor.py) and Braintrust (backend/onyx/tracing/braintrust_tracing_processor.py). Every LLM invocation (chat, embedding, rerank, image generation, voice, intent classification) opens a generation span tagged with its flow. The indexing pipeline also has its own tracing (INDEXING_PIPELINE_TRACE_NAME).

User usage tracking: The user_usage_processor.py (tracing/processors/) tracks per-user token usage for monitoring and quota enforcement.

However, Onyx does not have built-in RAGAS-style evaluation metrics (faithfulness, answer relevancy, context precision) computed automatically. Quality evaluation relies on the Braintrust integration for structured LLM-as-judge scoring and manual assertion checks.

deepset-ai/haystack

answered

Built-in evaluators – Haystack ships a set of component-based evaluators under haystack/components/evaluators/. The FaithfulnessEvaluator (haystack/components/evaluators/faithfulness.py:53-269) uses an LLM to split a generated answer into statements and check each against the provided contexts, producing a score of 0–1 representing the proportion of supported statements. The ContextRelevanceEvaluator similarly scores how relevant retrieved contexts are to a question. Other built-in metrics include AnswerExactMatchEvaluator, DocumentMAPEvaluator (Mean Average Precision), DocumentMRREvaluator (Mean Reciprocal Rank), DocumentNDCGEvaluator (Normalized Discounted Cumulative Gain), DocumentRecallEvaluator, and SASEvaluator (Semantic Answer Similarity). The LLMEvaluator (haystack/components/evaluators/llm_evaluator.py:25-80) is the base class for LLM-as-judge evaluations: user-defined instructions, input/output schema, and few-shot examples guide an LLM (default: OpenAI in JSON mode) to produce structured scores. All evaluators follow the Haystack @component pattern and can be plugged into Pipelines.

Tracing / observability – Haystack provides an abstract Tracer interface (haystack/tracing/tracer.py:77-94) with trace() context managers and Span objects for instrumenting operations. A built-in LoggingTracer (haystack/tracing/logging_tracer.py:35-50) logs operation names and tags with customizable ANSI coloring. Third-party tracers (e.g., OpenTelemetry) can be integrated by subclassing Tracer and calling enable_tracing(). Content-level tracing (queries, documents, answers) is opt-in via the HAYSTACK_CONTENT_TRACING_ENABLED environment variable.

Telemetry – Haystack uses PostHog for anonymous usage analytics (haystack/telemetry/__init__.py), tracking when pipelines are executed.

What is absent – There is no built-in evaluation result database or dashboard; the EvalRunResult dataclass (haystack/evaluation/eval_run_result.py) provides a lightweight container, but users are expected to aggregate and log results themselves.

HKUDS/RAG-Anything

answered

Built-in evals. RAG-Anything has no built-in evaluation framework (no RAGAS, no faithfulness/relevance metrics, no answer-grounding checks). The only quantitative measurement is the MetricsCallback (callbacks.py:183-287), which aggregates processing pipeline statistics: documents processed/failed, content blocks parsed, parse/insert/multimodal processing times, and query counts/times. It exposes a summary() method returning a human-readable report and a reset() method. This is strictly operational telemetry, not quality evaluation.

Tracing / observability hooks. The ProcessingCallback base class (callbacks.py:61-180) defines lifecycle hooks dispatched by CallbackManager: on_parse_start/complete/error, on_text_insert_start/complete, on_multimodal_start/item_complete/complete, on_query_start/complete/error, on_document_complete/error, and on_batch_start/complete. The ProcessingEvent dataclass captures timestamps, doc_ids, file paths, stages, and error details. Users can subclass ProcessingCallback to integrate with any observability system (e.g., a custom class could send events to OpenTelemetry, Datadog, or a log aggregator). The CallbackManager also supports optional internal event logging (enable_event_log(true)) and thread-safe callback registration (callbacks.py:298-346).

LLM response caching. Both text and multimodal queries can be cached via LightRAG's llm_response_cache (query.py:297-380), which avoids re-generating answers for identical inputs. Cache keys incorporate content fingerprints (SHA-256 of file content, MD5 of query parameters) to detect changes (query.py:59-126).

There are no eval benchmarks, no retrieval-quality metrics (hit rate, MRR, precision), and no generation quality metrics (faithfulness, helpfulness) in the codebase.

SciPhi-AI/R2R

answered

R2R has no built-in evaluation framework — no scoring, no LLM judge, no integrated metrics like faithfulness or answer relevancy. The project ships with two observability mechanisms instead:

  • Sentry integration (py/core/utils/sentry.py). Initialized at app startup via init_sentry(), reads R2R_SENTRY_DSN from the environment. Supports configurable traces_sample_rate and profiles_sample_rate. This provides error tracking and performance tracing but no RAG-specific quality metrics.

  • Structured logging (py/core/utils/logging_config.py). Configures python-json-logger with HTTP status filters for uvicorn.access logs. Health endpoint requests are filtered out. There is no RAG-specific logging (no retrieval latency, generation quality, or user feedback capture).

  • External evaluation guide. The project provides a cookbook at docs/cookbooks/evals.md demonstrating how to evaluate R2R outputs using the Ragas framework externally. It shows collecting R2R's RAG responses via client.retrieval.rag() and feeding them into Ragas for metrics like faithfulness, answer_relevancy, context_precision, and context_recall. This is an optional, bring-your-own-eval approach — Ragas is not a dependency and no integration code exists in the source tree.

There are no built-in evaluation endpoints, no A/B testing infrastructure, no LLM-as-judge calls, and no quality monitoring dashboards. Users must integrate external frameworks (Ragas, TruLens, LangFuse, etc.) themselves.

Cinnamon/kotaemon

insufficient evidence

There is no built-in evaluation framework (no RAGAS, TruLens, DeepEval, or similar eval harness integrated). The system does provide runtime quality signals: (1) LLMTrulensScoring (libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py) scores retrieved documents on a 0–10 relevance scale per query; (2) qa_score is derived from the exponential average of LLM response log-probabilities (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 274–277); (3) a context-relevance warning threshold (CONTEXT_RELEVANT_WARNING_SCORE) triggers a UI warning when the max LLM relevance score is below 0.3 (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, line 36–38). There is no built-in logging/tracing system (e.g., LangSmith, Weights & Biases, MLflow) connected to the pipeline — observability would need to be added externally. I checked the Kotaemon library and the ktem app modules; no evaluation harness or metric tracker was found.

← How are answers generated and grounded? · How is it deployed and operated? →