# How is LLM-as-a-judge implemented?

> LLM evals and testing — a good answer covers: Judge prompts and rubrics; structured output; multi-sample or consensus; judge model choice; calibration or bias controls.

Canonical page: https://llms-technical-reviews.com/evals/q/llm-judge/

## Verdict

[Inspect](/p/inspect_ai/) has the most careful judge: grader panels with majority vote, verdict parsing that resists injection, and parse failures kept separate from wrong answers. [DeepEval](/p/deepeval/), [Ragas](/p/ragas/) and [Opik](/p/opik/) have the most structured judge input and output. [Langfuse](/p/langfuse/) runs judges as production infrastructure.

**Schema-enforced judges.** DeepEval's `GEval` takes criteria, steps or a rubric of score ranges. When the model returns logprobs, it replaces the sampled score with a probability-weighted average. Its optional `hybrid` mode lets a hosted "System One" model answer the decision points and falls back to the LLM only when that call fails. Ragas validates every judge call into a Pydantic model through Instructor, and retries with a fix-format prompt. The editor found its `strictness` voting inert at the reviewed commit, so `AspectCritic` is a single-sample judge. Opik passes a Pydantic `response_format`, retries parse failures three times, and merges identical `LLMJudge` instances into one call. Its default judge is `gpt-5-nano`. [Phoenix](/p/phoenix/) builds a JSON schema with an enum of allowed labels. On OpenAI it tries native structured output and falls back to tool calling. Langfuse validates numeric, boolean or categorical output with Zod and records every judge call as its own trace. It pauses an evaluator when the judge fails with an auth or billing error.

**Text-parsed judges.** Inspect takes the *last* `GRADE: X` in the reply and escapes `[BEGIN DATA]` markers in dataset text. [promptfoo](/p/promptfoo/) parses a JSON `{pass, score, reason}`. Red-team configs without an explicit grader prefer promptfoo's remote grading endpoint. [OpenAI Evals](/p/openai-evals/) matches choice strings line by line, scores an unparseable answer as the lowest choice, and uses the last completion function on the command line as the judge. [garak](/p/garak/)'s `ModelAsJudge` asks for a `[[rating]]` from 1 to 10 and counts 7 or more as a hit. Its default judge is Llama 3 70B through NVIDIA NIM.

**No judge in the core.** [lm-evaluation-harness](/p/lm-evaluation-harness/) is reference-based. Only the `pisa_*_llm_judged` tasks call an OpenAI model from their own hooks.

Langfuse, Phoenix, Ragas, OpenAI Evals, Inspect and garak have no calibration or position-bias controls.

Pick: Inspect when judge robustness matters, for example graded agent tasks.
Pick: DeepEval or Ragas for rubric and claim-level judging in Python.
Pick: Langfuse or Opik to run judges continuously on live traces.

## Per-project answers

### langfuse/langfuse (answered)

LLM-as-a-judge is implemented through a chain of modules converging on `runLLMAsJudgeEvaluation` in `evalService.ts` (lines 941–1233). **Judge prompts and rubrics** are stored as `EvalTemplate` records with a `prompt` (template string using `{{variable}}` syntax) and `promptMessages` (an array of system/user/assistant messages). The prompt is compiled by substituting extracted variables — trace fields like `input`, `output`, or dataset-item columns — via `compileEvalPrompt` in `llmEvaluatorExecution.ts` (lines 40–57). **Structured output** is enforced through Zod schemas: the `outputDefinition` stored on the template (legacy `{score, reasoning}` strings or v2 `{dataType, score, reasoning}` objects) is compiled into a Zod schema via `compilePersistedEvalOutputDefinition` (`outputDefinition.ts` lines 362–375). The schema is serialized to JSON Schema and passed to the AI SDK, which constraints the LLM's response shape. Results are validated with `validateEvalOutputResult` — numeric ranges, allowed categorical values, duplicates detection, and boolean type strictness are all checked (test file lines 541–679). **Judge model choice** is per-evaluator: each template stores `provider`, `model`, and `modelParams` which are resolved through `fetchModelConfig` (`evalExecutionDeps.ts` lines 327–337). The system supports OpenAI, Anthropic, Google, AWS Bedrock, and any OpenAI-compatible endpoint. **No built-in multi-sample/consensus** — each eval job produces a single LLM call. **No explicit calibration** exists, but **bias controls** are available through the prompt template itself (the user authors the rubric). **Auto-blocking** (`classifyEvaluatorLlmError.ts`) is key: when the judge LLM fails (auth, billing, endpoint unreachable), the evaluator is automatically paused at the database level so all rules using it stop until the user intervenes (lines 60–84). **Internal tracing** records each judge execution as an `LLMJudge`-environment trace in the user's project for debugging.


Citations: [worker/src/features/evaluation/evalService.ts:941-1233](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalService.ts#L941-L1233) · [packages/shared/src/server/evals/llmEvaluatorExecution.ts:40-87](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/llmEvaluatorExecution.ts#L40-L87) · [packages/shared/src/features/evals/outputDefinition.ts:286-427](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/features/evals/outputDefinition.ts#L286-L427) · [worker/src/features/evaluation/evalExecutionDeps.ts:245-325](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/evalExecutionDeps.ts#L245-L325) · [packages/shared/src/server/evals/classifyEvaluatorLlmError.ts:60-163](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/packages/shared/src/server/evals/classifyEvaluatorLlmError.ts#L60-L163) · [worker/src/features/evaluation/executeLLMAsJudgeEvaluation.test.ts:541-723](https://github.com/langfuse/langfuse/blob/1a21a4229c98dd0dbba2916631ca5fe3e05fd504/worker/src/features/evaluation/executeLLMAsJudgeEvaluation.test.ts#L541-L723)

### promptfoo/promptfoo (answered)

Four LLM-as-a-judge matchers in src/matchers/llmGrading.ts. matchesLlmRubric (L181) uses a rubric prompt with runJsonGradingPrompt to parse JSON pass/score/reason. matchesFactuality (L323) scores against expert answer using categories A-E with configurable scoring. matchesClosedQa (L382) expects trailing Y/N. matchesGEval (L454) implements G-Eval: generates evaluation steps via LLM then scores 1-10. All use extractFirstJsonObject for structured parsing. Judge provider set via test.options.provider. ProviderGroupedCallQueue batches identical judge calls (evaluator.ts L538-550). Remote grading via doRemoteGrading when no local provider set. Inverse assertions flip pass/score via invertScore. g-eval supports multi-criteria arrays with averaged scoring (src/assertions/geval.ts L26-86).

> **Editor's note.** Correction: `ProviderGroupedCallQueue` does not batch or deduplicate identical judge calls; it groups queued grading calls by provider id. Remote grading for `llm-rubric` only applies inside a red-team config with no explicit grader or rubric prompt, not whenever no local provider is set.

Citations: [src/matchers/llmGrading.ts:181-267](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L181-L267) · [src/matchers/llmGrading.ts:454-612](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L454-L612) · [src/matchers/llmGrading.ts:382-442](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/matchers/llmGrading.ts#L382-L442) · [src/evaluator.ts:538-550](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/evaluator.ts#L538-L550) · [src/assertions/geval.ts:9-116](https://github.com/promptfoo/promptfoo/blob/1df3eb8e0bacad7efc64bce2bd3a7df0094a1d61/src/assertions/geval.ts#L9-L116)

### comet-ml/opik (answered)

Opik implements LLM-as-a-judge through two main paths: the `BaseMetric`-based LLM-judge metrics and the `LLMJudge` suite evaluator, both using structured output via Pydantic response-format models.

**Judge prompts and rubrics**: The test-suite `LLMJudge` (`suite_evaluators/llm_judge/metric.py:35-533`) has a fixed system prompt ("You are an expert judge tasked with evaluating if an AI agent's output satisfies a set of assertions") and a user template that formats input, output, and assertion criteria. Each assertion produces a boolean pass/fail with confidence. `GEval` (`metrics/llm_judges/g_eval/metric.py:32-288`) uses a two-stage process: first a chain-of-thought generation prompt (cached per judge+criteria+model combination in a 128-entry LRU), then a scoring query that includes the CoT in the prompt. Judge prompt templates are per-metric — `Hallucination`, `AnswerRelevance`, `Factuality`, `Moderation`, etc. each have their own template modules with few-shot examples.

**Structured output**: All LLM judges use Pydantic models as response-format schemas passed to the model via `response_format=<PydanticModel>` (e.g. `HallucinationResponseFormat`, `GEvalScoreFormat`, `AnswerRelevanceResponseFormat`). For LiteLLM models, structured output is enforced via `provider_kwargs['response_format']`; for others, the response string is parsed with dedicated parser modules.

**Judge model choice**: Configurable per metric instance via a `model` parameter (string name or `OpikBaseModel` instance). The factory in `models_factory.py` resolves names — defaults to `gpt-5-nano` (overridable via `OPIK_DEFAULT_LLM`). Supports `LiteLLMChatModel` (multi-provider), `OpenAI`/`Anthropic` direct models, and custom subclasses. Per-metric parameters include `temperature`, `seed`, `reasoning_effort`.

**Agentic vs one-shot scoring**: The suite-evaluator `LLMJudge` (`metric.py:291-348`) supports two scoring modes controlled by a `scoring_tool_strategy` selector ('auto', 'always', 'never', or custom). When a trace context is available and the strategy permits, it runs an `AgenticLLMJudge` (imported from `suite_evaluators/agentic/judge.py`) that can look up spans and tool calls from the full trace tree via an in-process emulator.

**Retry and resilience**: `_generate_and_parse` is wrapped with Tenacity retry (3 attempts, retries `LLMJudgeParseError` and `EmptyLLMResponseError`, reraise on exhaustion). On final parse failure, returns partial results with `LLMJudgeParseError.results`.

**Merge optimization**: Multiple `LLMJudge` instances with identical settings (model, temperature, seed, track) are merged into one via `LLMJudge.merged()` — assertions are deduplicated and sent in a single LLM call.


Citations: [sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py:35-533](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py#L35-L533) · [sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py:32-288](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/metrics/llm_judges/g_eval/metric.py#L32-L288) · [sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py:350-365](https://github.com/comet-ml/opik/blob/f217a863ea3cbfcef7f9ab8217c543049bdff095/sdks/python/src/opik/evaluation/suite_evaluators/llm_judge/metric.py#L350-L365)

### openai/evals (answered)

**Judge prompts and rubrics** are specified via `ModelGradedSpec`, a pydantic dataclass with fields `prompt`, `choice_strings`, `input_outputs`, `eval_type`, `choice_scores`, and `output_template` (`evals/elsuite/modelgraded/base.py:11-25`). Rubrics live as YAML files in `evals/registry/modelgraded/` — e.g. `fact.yaml` lets the judge compare a submission to an expert answer on correctness using a 5-option rubric (A–E) (`evals/registry/modelgraded/fact.yaml:1-21`), and `battle.yaml` compares two model outputs head-to-head with a Yes/No preference (`evals/registry/modelgraded/battle.yaml:1-23`).

**Structured output** is extracted by `classify()` in `evals/elsuite/modelgraded/classify_utils.py`. The judge model receives the rubric prompt with format kwargs filled in, plus an appended answer prompt depending on `eval_type` — `classify`, `classify_cot`, `cot_classify`, or `cot_classify_jp` — each of which instructs the LLM to output exactly one choice string from the allowed set (`classify_utils.py:13-28`). `get_choice()` then parses the raw text by trying each line against `choice_strings` using a `match_fn` (one of `include`, `exact`, `endswith`, `starts_or_endswith`), returning `"__invalid__"` on failure (`classify_utils.py:110-128`). `get_choice_score` maps parsed choices to numeric scores via `choice_scores` (`classify_utils.py:90-102`).

**Multi-sample and consensus** is supported via `multicomp_n`. When `multicomp_n > 1`, `sample_and_concat_n_completions` runs the subject model N times (either with N separate model instances or the same model N times), then concatenates the outputs into a single text using `output_template` before passing to the judge (`classify_utils.py:152-187`). When `multicomp_n == "from_models"`, N is derived from the number of completion functions (`evals/elsuite/modelgraded/classify.py:42-48`).

**Judge model choice** is determined by the last `completion_fn` in the list — the eval splits off `self.eval_completion_fn` as the judge, while the earlier ones serve as the subject model(s) (`evals/elsuite/modelgraded/classify.py:29-32`). There are no built-in calibration or bias controls.


Citations: [evals/registry/modelgraded/fact.yaml:1-22](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/modelgraded/fact.yaml#L1-L22) · [evals/registry/modelgraded/battle.yaml:1-23](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/registry/modelgraded/battle.yaml#L1-L23) · [evals/elsuite/modelgraded/classify.py:29-52](https://github.com/openai/evals/blob/8eac7a7de5215c907fbddc30efdaf316913eccdd/evals/elsuite/modelgraded/classify.py#L29-L52)

### confident-ai/deepeval (answered)

The LLM-as-a-judge system centers on `GEval` (`deepeval/metrics/g_eval/g_eval.py:53-120`). The judge is prompted with a `name`, optional `criteria` and `evaluation_steps` (a numbered list of checks), and an optional `rubric` — a list of `Rubric(score_range, expected_outcome)` that maps score intervals to natural-language descriptions (`deepeval/metrics/g_eval/utils.py:35-50`). The rendered prompt inserts the test case fields (input, actual output, context, etc.) as template variables and asks the LLM to output a structured `ReasonScore` (reason + score) via Pydantic schema enforcement (`deepeval/metrics/g_eval/schema.py:5-8`). Multi-turn conversation evaluation uses `ConversationalGEval`, which reads the full turn list. **Structured output** is enforced through `generate_with_schema_and_extract`, which passes a Pydantic model class as the response format to the underlying LLM API call (`deepeval/metrics/utils/__init__.py:19-26`). For **multi-sample / consensus** evaluation, `top_logprobs` can be set to request log-probability data from the judge model. **Judge model choice** is fully configurable via the `model` parameter (accepting a string model name or a `DeepEvalBaseLLM` instance wrapping any provider — OpenAI, Anthropic, Azure, local). The `eval_mode` system introduces **System One (Jev)** — a small-model pre-screener that runs first; only ambiguous or low-confidence cases are escalated to the full LLM judge, controlled via `system_one_model` and `eval_mode` (`hybrid`, `llm`, `system_one`). **Calibration controls**: `threshold` sets pass/fail, `strict_mode` overrides threshold to 1.0, and the rubric's score range can be customised. `JevEval` runs the judge prompt against a small / deterministic model for speed, and `ArenaGEval` compares two LLM outputs pairwise (`deepeval/metrics/__init__.py:21-23`). All judge metrics inherit template rendering from `PromptMixin._get_prompt`, which loads Jinja templates from the `templates/metrics/` bundle (`deepeval/metrics/base_metric.py:72-108`).

> **Editor's note.** Correction: System One is not a pre-screener with confidence-based escalation. In `hybrid` mode the LLM extracts and explains while Jev answers decision points, falling back to the LLM only on a runtime Jev error; `system_one` mode has no fallback (deepeval/config/eval_mode.py L1-L21). `top_logprobs` drives logprob-weighted G-Eval scoring, not multi-sample consensus.

Citations: [deepeval/metrics/g_eval/g_eval.py:53-120](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/g_eval.py#L53-L120) · [deepeval/metrics/g_eval/utils.py:35-50](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/utils.py#L35-L50) · [deepeval/metrics/g_eval/schema.py:1-17](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/g_eval/schema.py#L1-L17) · [deepeval/metrics/base_metric.py:72-108](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/base_metric.py#L72-L108) · [deepeval/metrics/utils/__init__.py:1-26](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/utils/__init__.py#L1-L26) · [deepeval/metrics/__init__.py:21-23](https://github.com/confident-ai/deepeval/blob/ec9b9837a3bf4a41b7fc01f2de0bfeba4997c159/deepeval/metrics/__init__.py#L21-L23)

### vibrantlabsai/ragas (answered)

LLM-as-a-judge is a core primitive in Ragas, built on `PydanticPrompt` with structured output.

**Judge prompts and rubrics**: The `PydanticPrompt` class (prompt/pydantic_prompt.py:82-350) is a generic template that takes an `instruction` string and `input_model`/`output_model` as Pydantic type parameters. The instruction describes the judgement criteria. At inference time, `to_string()` renders the instruction, output schema (auto-generated JSON Schema from `output_model.model_json_schema()`), few-shot examples, and the actual input into one prompt sent to the LLM. For instance, `NLIStatementPrompt` instructs "return verdict as 1 if the statement can be directly inferred..." with a `StatementFaithfulnessAnswer` output model (metrics/_faithfulness.py:73-131).

**Structured output**: `PydanticPrompt.generate()` calls the LLM with `response_model=self.output_model`, enforcing structured JSON via the `Instructor` library (llms/base.py:606-748). The `llm_factory()` wraps provider clients with instructor patching (OpenAI → `instructor.from_openai`, Anthropic → `instructor.from_anthropic`, etc.) and passes the Pydantic model as `response_model`. Returns are validated immediately into the output Pydantic model, with automatic retry (`retries_left: int = 3`) via `RagasOutputParser.parse_output_string()` which on parse failure calls `FixOutputFormat` to ask the LLM to fix malformed JSON (prompt/pydantic_prompt.py:525-558).

**Multi-sample/consensus**: The `Ensember` class (metrics/base.py:641-686) implements majority voting over multiple LLM outputs for the same input. `AspectCritic` and `SimpleCriteriaScore` have a `strictness` parameter — when >1, they call the LLM N times and aggregate via `Counter(verdicts).most_common(1)[0][0]` (metrics/_aspect_critic.py:155-165, metrics/_simple_criteria.py:152-162).

**Judge model choice**: The judge LLM is set per-metric via `metric.llm`, or inherited from the `evaluate()`/`@experiment` call-level llm parameter. `llm_factory()` supports OpenAI, Anthropic, Google/Gemini, Groq, Mistral, Azure, LiteLLM (100+ providers), and more. It auto-detects the correct instructor patching strategy. The default fallback when no LLM is provided is `gpt-4o-mini` via `OpenAI()` (evaluation.py:176-179).

**Calibration/bias controls**: No dedicated calibration or bias-mitigation module exists. The project provides an `Ensember` for consensus voting but no systematic bias measurement, calibration curves, or position-bias detection. The `AspectCritic` with predefined aspects (harmfulness, maliciousness, etc.) can be used as a safety check, but bias controls are absent.

> **Editor's note.** Correction: `strictness` is inert at this commit. `AspectCritic` and `SimpleCriteriaScore` call `prompt.generate` once and vote over a one-element list (src/ragas/metrics/_aspect_critic.py L173-L212), so no multi-sample consensus happens.

Citations: [src/ragas/prompt/pydantic_prompt.py:82-350](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/prompt/pydantic_prompt.py#L82-L350) · [src/ragas/llms/base.py:606-748](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/llms/base.py#L606-L748) · [src/ragas/prompt/pydantic_prompt.py:525-558](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/prompt/pydantic_prompt.py#L525-L558) · [src/ragas/metrics/base.py:641-686](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/base.py#L641-L686) · [src/ragas/metrics/_aspect_critic.py:75-165](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/_aspect_critic.py#L75-L165) · [src/ragas/evaluation.py:170-180](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/evaluation.py#L170-L180)

### EleutherAI/lm-evaluation-harness (insufficient evidence)

The repository does **not** implement LLM-as-a-judge evaluation. After searching the entire codebase for judge prompts, rubric-based scoring, structured output parsing for evaluation, consensus-based multi-sample judgment, and calibration/bias controls, none were found. The harness's evaluation model is reference-based: each task provides gold-standard answers (target strings or multiple-choice labels), and the harness computes deterministic metrics (accuracy, perplexity, BLEU, etc.) by comparing model outputs to these references. Tasks like `toxigen` at `lm_eval/tasks/toxigen/toxigen.yaml` and `truthfulqa` use multiple-choice or generation tasks with regex extraction and exact-match accuracy — not LLM judges. The `exact_match` metric at `lm_eval/api/metrics.py:234-266` implements string-level comparison with normalization options (case, punctuation, number, regex ignore). The `EvalResults` schema at `lm_eval/result_schema.py` defines results as typed dicts with numeric scores per task — no provision for judge verdicts, rubrics, or model-graded outputs. The project instead relies on the well-known Open LLM Leaderboard tasks (in `lm_eval/tasks/leaderboard/`) and hundreds of standard academic benchmarks, all of which use reference-based evaluation. For generation tasks, output parsing is handled by regex-based filters (`RegexFilter`, `MultiChoiceRegexFilter` in `lm_eval/filters/extraction.py`) to extract answers from free-text completions for comparison against targets, not for judgment.

> **Editor's note.** Correction: the core harness has no judge abstraction, but some tasks do use one. The pisa_*_llm_judged tasks call an OpenAI chat model from their process_results hook (lm_eval/tasks/pisa/utils.py:85-111, 312-329) and map its one-token reply to acc.

Citations: [lm_eval/api/metrics.py:234-266](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/api/metrics.py#L234-L266) · [lm_eval/filters/extraction.py:15-63](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/filters/extraction.py#L15-L63) · [lm_eval/result_schema.py:1-108](https://github.com/EleutherAI/lm-evaluation-harness/blob/d6de81643928d653435c431bae19945d41d32520/lm_eval/result_schema.py#L1-L108)

### Arize-ai/phoenix (answered)

LLM-as-a-judge is implemented through the `LLM` wrapper class (`packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440`) which delegates to provider-specific **adapters** registered via a singleton `ProviderRegistry` and `AdapterRegistry` (`registries.py:39-146`). Supported adapters include OpenAI, Anthropic, Google Gemini, LangChain, and LiteLLM — each implementing `BaseLLMAdapter` with `generate_text`, `generate_object`, and async variants (`types.py:22-108`). The `ClassificationEvaluator` (`evaluators.py:601-825`) renders a prompt template, calls `llm.generate_classification()`, which builds a JSON schema via `generate_classification_schema()` (`wrapper.py:451-524`) and calls `generate_object()` with that schema. The schema is a structured JSON object requiring a `label` string (constrained by `enum` or `oneOf` for valid choices) and optionally an `explanation` string. **Structured output vs tool calling**: On OpenAI, the adapter first tries native structured output (`response_format: {"type": "json_schema", ...}`) and falls back to tool calling on `BadRequestError`, caching the preferred method per model (`openai/adapter.py:130-240`). **No built-in multi-sample / consensus** — each evaluation is a single LLM call. **Judge model choice** is fully configurable: pass any `LLM(provider=..., model=...)` to the evaluator constructor. The online-evals runner (`evals/pxi/online_evals/judge.py:1-55`) reads `PHOENIX_AGENTS_EVALS_PROVIDER` and `PHOENIX_AGENTS_EVALS_MODEL` env vars (defaulting to OpenAI/gpt-5.5). The Harbor evaluation verifier (`evals/harbor/verifiers/harbor_verifiers/llm_judge.py:1-44`) uses `PHOENIX_EVAL_JUDGE_PROVIDER` and `PHOENIX_EVAL_JUDGE_MODEL` env vars (defaulting to OpenAI/gpt-5-nano). **Bias controls**: Each evaluator's rubric is stored in the generated config files (e.g. `_faithfulness_classification_evaluator_config.py` has detailed grading criteria). The `RateLimiter` (`rate_limiters.py`) uses an `AdaptiveTokenBucket` that backs off on rate-limit errors but does not implement calibration or answer-order randomization.

> **Editor's note.** Correction: the PHOENIX_AGENTS_EVALS_* and PHOENIX_EVAL_JUDGE_* defaults belong to the internal evals/ scripts, not the product. In phoenix-evals the judge is whatever LLM(provider=..., model=...) the user passes.

Citations: [packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py:97-440](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/wrapper.py#L97-L440) · [packages/phoenix-evals/src/phoenix/evals/evaluators.py:601-825](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/evaluators.py#L601-L825) · [packages/phoenix-evals/src/phoenix/evals/llm/adapters/openai/adapter.py:130-240](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/adapters/openai/adapter.py#L130-L240) · [packages/phoenix-evals/src/phoenix/evals/llm/registries.py:39-146](https://github.com/Arize-ai/phoenix/blob/bbdce4f193ef85b8fcdad1df7109f18e93d44971/packages/phoenix-evals/src/phoenix/evals/llm/registries.py#L39-L146)

### NVIDIA/garak (answered)

LLM-as-judge is implemented through `ModelAsJudge` (`detectors/judge.py:20-120`) which mixes `Detector` with `EvaluationJudge` from `resources/red_team/evaluation.py`. On initialization, it loads a separate judge generator via `_load_generator()` (`judge.py:52-82`) — configurable by `detector_model_type` (default `"nim"`) and `detector_model_name` (default `"meta/llama3-70b-instruct"`). The judge generator must be `OpenAICompatible`.

Two judge prompt styles exist. **Rating prompts** (`_goal_system_prompt`, `judge.py:40-49`): ask "Rate from 1 to 10" with `[[rating]]` output format, parsed by `process_output_judge_score()` (`evaluation.py:24-33`) via regex `\[\[(\d+)\]\]`. Scores above `confidence_cutoff` (default 7) count as hits. **Yes/No prompts** (`_refusal_system_prompt`, `judge.py:141-148`): ask for `[[YES]]` or `[[NO]]`, parsed by `process_output_on_topic_score()` (`evaluation.py:36-44`). The `Refusal` detector (`judge.py:123-157`) uses YES/NO; `Jailbreak` (`judge.py:176-264`) uses the JailbreakBench `<BEGIN REQUEST>`/`<END RESPONSE>` format with YES/NO output.

`EvaluationJudge._create_conv()` (`evaluation.py:78-116`) builds an OpenAI-format conversation (system prompt + user prompt) and truncates if it exceeds the model's token limit (from `generators/openai.py` context_lengths, fallback 4096). No multi-sample consensus — each response is judged once. No explicit calibration or bias controls, though multiple `generations` per probe yield multiple judgements that can be averaged in report post-processing. `JailbreakOnlyAdversarial` and `RefusalOnlyAdversarial` (`judge.py:160-280`) gate on `attempt.notes["is_adversarial"]` to skip non-adversarial turns.


Citations: [garak/detectors/judge.py:20-120](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/detectors/judge.py#L20-L120) · [garak/detectors/judge.py:176-264](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/detectors/judge.py#L176-L264) · [garak/resources/red_team/evaluation.py:24-44](https://github.com/NVIDIA/garak/blob/bb30a7e79f4e78ef633a92295f142105a0e69941/garak/resources/red_team/evaluation.py#L24-L44)

### UKGovernmentBEIS/inspect_ai (answered)

LLM-as-a-judge is implemented via `model_graded_qa` and `model_graded_fact` scorers (`src/inspect_ai/scorer/_model.py`).

**Judge prompts and rubrics.** Both scorers accept a `template` with `{question}`, `{answer}`, `{criterion}`, and `{instructions}` variables plus any sample metadata keys. The default instructions ask the grader to emit a `GRADE: C` / `GRADE: I` (or `GRADE: P` for partial credit) verdict after step-by-step reasoning. The `model_scoring_prompt()` function formats the template, neutralizes structural delimiters (`[BEGIN DATA]`/`[END DATA]`) in dataset-controlled inputs to prevent prompt injection, and preserves media attachments.

**Structured output via regex.** The grade is extracted by regex with a leading greedy `.*` (DOTALL) that ensures the *last* `GRADE: X` wins — crucial for injection robustness. A permissive variant captures any word when the default instructions are used, then validates against the grades actually offered. Off-menu verdicts are treated as parse failures (`grader_failed`), not incorrect answers.

**Multi-sample / consensus.** When `model` is a list of models (or `model_role` binds to a list), each model grades independently via `multi_scorer()` which runs them concurrently with `tg_collect`. By default (`reducer="majority"`), a grade needs >50% of graders to agree. `reducer="mode"` picks the most common grade with tie-breaking by model order.

**Judge model choice.** Controlled via the `model` parameter (takes precedence) or `model_role` (default `"grader"`, resolved from eval's `model_roles`). Falls back to the model under evaluation when no judge model is specified.

**Calibration / bias controls.** No explicit bias-calibration is provided — the approach relies on prompt engineering and majority-vote reducers. Anti-injection measures include structural delimiter neutralization and grade-pattern validation.


Citations: [src/inspect_ai/scorer/_model.py:34-120](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L34-L120) · [src/inspect_ai/scorer/_model.py:246-362](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L246-L362) · [src/inspect_ai/scorer/_model.py:364-460](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_model.py#L364-L460) · [src/inspect_ai/scorer/_multi.py:19-57](https://github.com/UKGovernmentBEIS/inspect_ai/blob/aa20052a65b13516f1ee79d10ccceda00c205cc6/src/inspect_ai/scorer/_multi.py#L19-L57)
