How is LLM-as-a-judge implemented?
Judge prompts and rubrics; structured output; multi-sample or consensus; judge model choice; calibration or bias controls.
Verdict
Inspect has the most careful judge: grader panels with majority vote, verdict parsing that resists injection, and parse failures kept separate from wrong answers. DeepEval, Ragas and Opik have the most structured judge input and output. Langfuse runs judges as production infrastructure.
Schema-enforced judges. DeepEval’s GEval takes criteria, steps or a rubric of score ranges. When the model returns logprobs, it replaces the sampled score with a probability-weighted average. Its optional hybrid mode lets a hosted “System One” model answer the decision points and falls back to the LLM only when that call fails. Ragas validates every judge call into a Pydantic model through Instructor, and retries with a fix-format prompt. The editor found its strictness voting inert at the reviewed commit, so AspectCritic is a single-sample judge. Opik passes a Pydantic response_format, retries parse failures three times, and merges identical LLMJudge instances into one call. Its default judge is gpt-5-nano. Phoenix builds a JSON schema with an enum of allowed labels. On OpenAI it tries native structured output and falls back to tool calling. Langfuse validates numeric, boolean or categorical output with Zod and records every judge call as its own trace. It pauses an evaluator when the judge fails with an auth or billing error.
Text-parsed judges. Inspect takes the last GRADE: X in the reply and escapes [BEGIN DATA] markers in dataset text. promptfoo parses a JSON {pass, score, reason}. Red-team configs without an explicit grader prefer promptfoo’s remote grading endpoint. OpenAI Evals matches choice strings line by line, scores an unparseable answer as the lowest choice, and uses the last completion function on the command line as the judge. garak’s ModelAsJudge asks for a [[rating]] from 1 to 10 and counts 7 or more as a hit. Its default judge is Llama 3 70B through NVIDIA NIM.
No judge in the core. lm-evaluation-harness is reference-based. Only the pisa_*_llm_judged tasks call an OpenAI model from their own hooks.
Langfuse, Phoenix, Ragas, OpenAI Evals, Inspect and garak have no calibration or position-bias controls.
Pick: Inspect when judge robustness matters, for example graded agent tasks. Pick: DeepEval or Ragas for rubric and claim-level judging in Python. Pick: Langfuse or Opik to run judges continuously on live traces.
Per-project answers
langfuse/langfuse
answeredLLM-as-a-judge is implemented through a chain of modules converging on runLLMAsJudgeEvaluation in evalService.ts (lines 941–1233). Judge prompts and rubrics are stored as EvalTemplate records with a prompt (template string using {{variable}} syntax) and promptMessages (an array of system/user/assistant messages). The prompt is compiled by substituting extracted variables — trace fields like input, output, or dataset-item columns — via compileEvalPrompt in llmEvaluatorExecution.ts (lines 40–57). Structured output is enforced through Zod schemas: the outputDefinition stored on the template (legacy {score, reasoning} strings or v2 {dataType, score, reasoning} objects) is compiled into a Zod schema via compilePersistedEvalOutputDefinition (outputDefinition.ts lines 362–375). The schema is serialized to JSON Schema and passed to the AI SDK, which constraints the LLM's response shape. Results are validated with validateEvalOutputResult — numeric ranges, allowed categorical values, duplicates detection, and boolean type strictness are all checked (test file lines 541–679). Judge model choice is per-evaluator: each template stores provider, model, and modelParams which are resolved through fetchModelConfig (evalExecutionDeps.ts lines 327–337). The system supports OpenAI, Anthropic, Google, AWS Bedrock, and any OpenAI-compatible endpoint. No built-in multi-sample/consensus — each eval job produces a single LLM call. No explicit calibration exists, but bias controls are available through the prompt template itself (the user authors the rubric). Auto-blocking (classifyEvaluatorLlmError.ts) is key: when the judge LLM fails (auth, billing, endpoint unreachable), the evaluator is automatically paused at the database level so all rules using it stop until the user intervenes (lines 60–84). Internal tracing records each judge execution as an LLMJudge-environment trace in the user's project for debugging.
promptfoo/promptfoo
answeredFour LLM-as-a-judge matchers in src/matchers/llmGrading.ts. matchesLlmRubric (L181) uses a rubric prompt with runJsonGradingPrompt to parse JSON pass/score/reason. matchesFactuality (L323) scores against expert answer using categories A-E with configurable scoring. matchesClosedQa (L382) expects trailing Y/N. matchesGEval (L454) implements G-Eval: generates evaluation steps via LLM then scores 1-10. All use extractFirstJsonObject for structured parsing. Judge provider set via test.options.provider. ProviderGroupedCallQueue batches identical judge calls (evaluator.ts L538-550). Remote grading via doRemoteGrading when no local provider set. Inverse assertions flip pass/score via invertScore. g-eval supports multi-criteria arrays with averaged scoring (src/assertions/geval.ts L26-86).
ProviderGroupedCallQueue does not batch or deduplicate identical judge calls; it groups queued grading calls by provider id. Remote grading for llm-rubric only applies inside a red-team config with no explicit grader or rubric prompt, not whenever no local provider is set.comet-ml/opik
answeredOpik implements LLM-as-a-judge through two main paths: the BaseMetric-based LLM-judge metrics and the LLMJudge suite evaluator, both using structured output via Pydantic response-format models.
Judge prompts and rubrics: The test-suite LLMJudge (suite_evaluators/llm_judge/metric.py:35-533) has a fixed system prompt ("You are an expert judge tasked with evaluating if an AI agent's output satisfies a set of assertions") and a user template that formats input, output, and assertion criteria. Each assertion produces a boolean pass/fail with confidence. GEval (metrics/llm_judges/g_eval/metric.py:32-288) uses a two-stage process: first a chain-of-thought generation prompt (cached per judge+criteria+model combination in a 128-entry LRU), then a scoring query that includes the CoT in the prompt. Judge prompt templates are per-metric — Hallucination, AnswerRelevance, Factuality, Moderation, etc. each have their own template modules with few-shot examples.
Structured output: All LLM judges use Pydantic models as response-format schemas passed to the model via response_format=<PydanticModel> (e.g. HallucinationResponseFormat, GEvalScoreFormat, AnswerRelevanceResponseFormat). For LiteLLM models, structured output is enforced via provider_kwargs['response_format']; for others, the response string is parsed with dedicated parser modules.
Judge model choice: Configurable per metric instance via a model parameter (string name or OpikBaseModel instance). The factory in models_factory.py resolves names — defaults to gpt-5-nano (overridable via OPIK_DEFAULT_LLM). Supports LiteLLMChatModel (multi-provider), OpenAI/Anthropic direct models, and custom subclasses. Per-metric parameters include temperature, seed, reasoning_effort.
Agentic vs one-shot scoring: The suite-evaluator LLMJudge (metric.py:291-348) supports two scoring modes controlled by a scoring_tool_strategy selector ('auto', 'always', 'never', or custom). When a trace context is available and the strategy permits, it runs an AgenticLLMJudge (imported from suite_evaluators/agentic/judge.py) that can look up spans and tool calls from the full trace tree via an in-process emulator.
Retry and resilience: _generate_and_parse is wrapped with Tenacity retry (3 attempts, retries LLMJudgeParseError and EmptyLLMResponseError, reraise on exhaustion). On final parse failure, returns partial results with LLMJudgeParseError.results.
Merge optimization: Multiple LLMJudge instances with identical settings (model, temperature, seed, track) are merged into one via LLMJudge.merged() — assertions are deduplicated and sent in a single LLM call.
openai/evals
answeredJudge prompts and rubrics are specified via ModelGradedSpec, a pydantic dataclass with fields prompt, choice_strings, input_outputs, eval_type, choice_scores, and output_template (evals/elsuite/modelgraded/base.py:11-25). Rubrics live as YAML files in evals/registry/modelgraded/ — e.g. fact.yaml lets the judge compare a submission to an expert answer on correctness using a 5-option rubric (A–E) (evals/registry/modelgraded/fact.yaml:1-21), and battle.yaml compares two model outputs head-to-head with a Yes/No preference (evals/registry/modelgraded/battle.yaml:1-23).
Structured output is extracted by classify() in evals/elsuite/modelgraded/classify_utils.py. The judge model receives the rubric prompt with format kwargs filled in, plus an appended answer prompt depending on eval_type — classify, classify_cot, cot_classify, or cot_classify_jp — each of which instructs the LLM to output exactly one choice string from the allowed set (classify_utils.py:13-28). get_choice() then parses the raw text by trying each line against choice_strings using a match_fn (one of include, exact, endswith, starts_or_endswith), returning "__invalid__" on failure (classify_utils.py:110-128). get_choice_score maps parsed choices to numeric scores via choice_scores (classify_utils.py:90-102).
Multi-sample and consensus is supported via multicomp_n. When multicomp_n > 1, sample_and_concat_n_completions runs the subject model N times (either with N separate model instances or the same model N times), then concatenates the outputs into a single text using output_template before passing to the judge (classify_utils.py:152-187). When multicomp_n == "from_models", N is derived from the number of completion functions (evals/elsuite/modelgraded/classify.py:42-48).
Judge model choice is determined by the last completion_fn in the list — the eval splits off self.eval_completion_fn as the judge, while the earlier ones serve as the subject model(s) (evals/elsuite/modelgraded/classify.py:29-32). There are no built-in calibration or bias controls.
confident-ai/deepeval
answeredThe LLM-as-a-judge system centers on GEval (deepeval/metrics/g_eval/g_eval.py:53-120). The judge is prompted with a name, optional criteria and evaluation_steps (a numbered list of checks), and an optional rubric — a list of Rubric(score_range, expected_outcome) that maps score intervals to natural-language descriptions (deepeval/metrics/g_eval/utils.py:35-50). The rendered prompt inserts the test case fields (input, actual output, context, etc.) as template variables and asks the LLM to output a structured ReasonScore (reason + score) via Pydantic schema enforcement (deepeval/metrics/g_eval/schema.py:5-8). Multi-turn conversation evaluation uses ConversationalGEval, which reads the full turn list. Structured output is enforced through generate_with_schema_and_extract, which passes a Pydantic model class as the response format to the underlying LLM API call (deepeval/metrics/utils/__init__.py:19-26). For multi-sample / consensus evaluation, top_logprobs can be set to request log-probability data from the judge model. Judge model choice is fully configurable via the model parameter (accepting a string model name or a DeepEvalBaseLLM instance wrapping any provider — OpenAI, Anthropic, Azure, local). The eval_mode system introduces System One (Jev) — a small-model pre-screener that runs first; only ambiguous or low-confidence cases are escalated to the full LLM judge, controlled via system_one_model and eval_mode (hybrid, llm, system_one). Calibration controls: threshold sets pass/fail, strict_mode overrides threshold to 1.0, and the rubric's score range can be customised. JevEval runs the judge prompt against a small / deterministic model for speed, and ArenaGEval compares two LLM outputs pairwise (deepeval/metrics/__init__.py:21-23). All judge metrics inherit template rendering from PromptMixin._get_prompt, which loads Jinja templates from the templates/metrics/ bundle (deepeval/metrics/base_metric.py:72-108).
hybrid mode the LLM extracts and explains while Jev answers decision points, falling back to the LLM only on a runtime Jev error; system_one mode has no fallback (deepeval/config/eval_mode.py L1-L21). top_logprobs drives logprob-weighted G-Eval scoring, not multi-sample consensus.vibrantlabsai/ragas
answeredLLM-as-a-judge is a core primitive in Ragas, built on PydanticPrompt with structured output.
Judge prompts and rubrics: The PydanticPrompt class (prompt/pydantic_prompt.py:82-350) is a generic template that takes an instruction string and input_model/output_model as Pydantic type parameters. The instruction describes the judgement criteria. At inference time, to_string() renders the instruction, output schema (auto-generated JSON Schema from output_model.model_json_schema()), few-shot examples, and the actual input into one prompt sent to the LLM. For instance, NLIStatementPrompt instructs "return verdict as 1 if the statement can be directly inferred..." with a StatementFaithfulnessAnswer output model (metrics/_faithfulness.py:73-131).
Structured output: PydanticPrompt.generate() calls the LLM with response_model=self.output_model, enforcing structured JSON via the Instructor library (llms/base.py:606-748). The llm_factory() wraps provider clients with instructor patching (OpenAI → instructor.from_openai, Anthropic → instructor.from_anthropic, etc.) and passes the Pydantic model as response_model. Returns are validated immediately into the output Pydantic model, with automatic retry (retries_left: int = 3) via RagasOutputParser.parse_output_string() which on parse failure calls FixOutputFormat to ask the LLM to fix malformed JSON (prompt/pydantic_prompt.py:525-558).
Multi-sample/consensus: The Ensember class (metrics/base.py:641-686) implements majority voting over multiple LLM outputs for the same input. AspectCritic and SimpleCriteriaScore have a strictness parameter — when >1, they call the LLM N times and aggregate via Counter(verdicts).most_common(1)[0][0] (metrics/_aspect_critic.py:155-165, metrics/_simple_criteria.py:152-162).
Judge model choice: The judge LLM is set per-metric via metric.llm, or inherited from the evaluate()/@experiment call-level llm parameter. llm_factory() supports OpenAI, Anthropic, Google/Gemini, Groq, Mistral, Azure, LiteLLM (100+ providers), and more. It auto-detects the correct instructor patching strategy. The default fallback when no LLM is provided is gpt-4o-mini via OpenAI() (evaluation.py:176-179).
Calibration/bias controls: No dedicated calibration or bias-mitigation module exists. The project provides an Ensember for consensus voting but no systematic bias measurement, calibration curves, or position-bias detection. The AspectCritic with predefined aspects (harmfulness, maliciousness, etc.) can be used as a safety check, but bias controls are absent.
strictness is inert at this commit. AspectCritic and SimpleCriteriaScore call prompt.generate once and vote over a one-element list (src/ragas/metrics/_aspect_critic.py L173-L212), so no multi-sample consensus happens.Arize-ai/phoenix
answeredLLM-as-a-judge is implemented through the LLM wrapper class (packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440) which delegates to provider-specific adapters registered via a singleton ProviderRegistry and AdapterRegistry (registries.py:39-146). Supported adapters include OpenAI, Anthropic, Google Gemini, LangChain, and LiteLLM — each implementing BaseLLMAdapter with generate_text, generate_object, and async variants (types.py:22-108). The ClassificationEvaluator (evaluators.py:601-825) renders a prompt template, calls llm.generate_classification(), which builds a JSON schema via generate_classification_schema() (wrapper.py:451-524) and calls generate_object() with that schema. The schema is a structured JSON object requiring a label string (constrained by enum or oneOf for valid choices) and optionally an explanation string. Structured output vs tool calling: On OpenAI, the adapter first tries native structured output (response_format: {"type": "json_schema", ...}) and falls back to tool calling on BadRequestError, caching the preferred method per model (openai/adapter.py:130-240). No built-in multi-sample / consensus — each evaluation is a single LLM call. Judge model choice is fully configurable: pass any LLM(provider=..., model=...) to the evaluator constructor. The online-evals runner (evals/pxi/online_evals/judge.py:1-55) reads PHOENIX_AGENTS_EVALS_PROVIDER and PHOENIX_AGENTS_EVALS_MODEL env vars (defaulting to OpenAI/gpt-5.5). The Harbor evaluation verifier (evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44) uses PHOENIX_EVAL_JUDGE_PROVIDER and PHOENIX_EVAL_JUDGE_MODEL env vars (defaulting to OpenAI/gpt-5-nano). Bias controls: Each evaluator's rubric is stored in the generated config files (e.g. _faithfulness_classification_evaluator_config.py has detailed grading criteria). The RateLimiter (rate_limiters.py) uses an AdaptiveTokenBucket that backs off on rate-limit errors but does not implement calibration or answer-order randomization.
NVIDIA/garak
answeredLLM-as-judge is implemented through ModelAsJudge (detectors/judge.py:20-120) which mixes Detector with EvaluationJudge from resources/red_team/evaluation.py. On initialization, it loads a separate judge generator via _load_generator() (judge.py:52-82) — configurable by detector_model_type (default "nim") and detector_model_name (default "meta/llama3-70b-instruct"). The judge generator must be OpenAICompatible.
Two judge prompt styles exist. Rating prompts (_goal_system_prompt, judge.py:40-49): ask "Rate from 1 to 10" with [[rating]] output format, parsed by process_output_judge_score() (evaluation.py:24-33) via regex \[\[(\d+)\]\]. Scores above confidence_cutoff (default 7) count as hits. Yes/No prompts (_refusal_system_prompt, judge.py:141-148): ask for [[YES]] or [[NO]], parsed by process_output_on_topic_score() (evaluation.py:36-44). The Refusal detector (judge.py:123-157) uses YES/NO; Jailbreak (judge.py:176-264) uses the JailbreakBench <BEGIN REQUEST>/<END RESPONSE> format with YES/NO output.
EvaluationJudge._create_conv() (evaluation.py:78-116) builds an OpenAI-format conversation (system prompt + user prompt) and truncates if it exceeds the model's token limit (from generators/openai.py context_lengths, fallback 4096). No multi-sample consensus — each response is judged once. No explicit calibration or bias controls, though multiple generations per probe yield multiple judgements that can be averaged in report post-processing. JailbreakOnlyAdversarial and RefusalOnlyAdversarial (judge.py:160-280) gate on attempt.notes["is_adversarial"] to skip non-adversarial turns.
UKGovernmentBEIS/inspect_ai
answeredLLM-as-a-judge is implemented via model_graded_qa and model_graded_fact scorers (src/inspect_ai/scorer/_model.py).
Judge prompts and rubrics. Both scorers accept a template with {question}, {answer}, {criterion}, and {instructions} variables plus any sample metadata keys. The default instructions ask the grader to emit a GRADE: C / GRADE: I (or GRADE: P for partial credit) verdict after step-by-step reasoning. The model_scoring_prompt() function formats the template, neutralizes structural delimiters ([BEGIN DATA]/[END DATA]) in dataset-controlled inputs to prevent prompt injection, and preserves media attachments.
Structured output via regex. The grade is extracted by regex with a leading greedy .* (DOTALL) that ensures the last GRADE: X wins — crucial for injection robustness. A permissive variant captures any word when the default instructions are used, then validates against the grades actually offered. Off-menu verdicts are treated as parse failures (grader_failed), not incorrect answers.
Multi-sample / consensus. When model is a list of models (or model_role binds to a list), each model grades independently via multi_scorer() which runs them concurrently with tg_collect. By default (reducer="majority"), a grade needs >50% of graders to agree. reducer="mode" picks the most common grade with tie-breaking by model order.
Judge model choice. Controlled via the model parameter (takes precedence) or model_role (default "grader", resolved from eval's model_roles). Falls back to the model under evaluation when no judge model is specified.
Calibration / bias controls. No explicit bias-calibration is provided — the approach relies on prompt engineering and majority-vote reducers. Anti-injection measures include structural delimiter neutralization and grade-pattern validation.
EleutherAI/lm-evaluation-harness
insufficient evidenceThe repository does not implement LLM-as-a-judge evaluation. After searching the entire codebase for judge prompts, rubric-based scoring, structured output parsing for evaluation, consensus-based multi-sample judgment, and calibration/bias controls, none were found. The harness's evaluation model is reference-based: each task provides gold-standard answers (target strings or multiple-choice labels), and the harness computes deterministic metrics (accuracy, perplexity, BLEU, etc.) by comparing model outputs to these references. Tasks like toxigen at lm_eval/tasks/toxigen/toxigen.yaml and truthfulqa use multiple-choice or generation tasks with regex extraction and exact-match accuracy — not LLM judges. The exact_match metric at lm_eval/api/metrics.py:234-266 implements string-level comparison with normalization options (case, punctuation, number, regex ignore). The EvalResults schema at lm_eval/result_schema.py defines results as typed dicts with numeric scores per task — no provision for judge verdicts, rubrics, or model-graded outputs. The project instead relies on the well-known Open LLM Leaderboard tasks (in lm_eval/tasks/leaderboard/) and hundreds of standard academic benchmarks, all of which use reference-based evaluation. For generation tasks, output parsing is handled by regex-based filters (RegexFilter, MultiChoiceRegexFilter in lm_eval/filters/extraction.py) to extract answers from free-text completions for comparison against targets, not for judgment.
← Which evaluation metrics and scorers are provided, and how are they implemented? · How are test datasets and cases defined, generated and versioned? →