Personalize this lesson
Adapt explanations and teaching visuals to your background and preferred voice.
policy-answerer-v4-eval proved a hard fact: the current, permitted policy source requires a rollback runbook, not continued deployment. A claim ledger can block unsupported "keep deploying" advice. It can't decide which of two safe replies from a large language model (LLM) is clearer for the engineer.
Consider these two answers to Maya's payment-service incident:
| Candidate | Reply | Hard evidence status |
|---|---|---|
brief | "Payment-service crossed the rollback threshold; run the rollback runbook under DEP-27." | Supported |
actionable | "Payment-service crossed the rollback threshold; run the DEP-27 rollback runbook, open an incident note, and page the release lead before retrying." | Supported |
Both respect the selected evidence. The remaining question is softer: does the added next step make the second reply more useful without making it wordy or confusing?
An LLM-as-a-judge uses another LLM as an evaluator for quality that can't be fully decided by an exact assertion. It can compare clarity, helpfulness, or tone under a rubric. It must not decide whether restricted context was allowed or whether a policy claim is supported. Those remain deterministic gates.
Zheng et al. found that strong LLM judges could exceed 80% agreement with human preferences on their MT-Bench and Chatbot Arena experiments. The same work reports position bias, verbosity bias, preference for model-like answers, and reasoning limitations. A judge is useful measurement equipment, not ground truth.[1]
Keep facts outside the judge
The boundary matters more than the model name. In a deploy-policy answer pipeline, different questions need different evaluators:
| Question | Correct evaluator | Why |
|---|---|---|
| Did selected evidence pass access and freshness checks? | Code gate | A soft score must never admit forbidden evidence. |
| Does the answer advise continued deployment when DEP-27 requires rollback? | Claim-to-source verifier | Policy truth is inspectable. |
| Which supported answer is clearer and more actionable? | Calibrated judge or human | Reasonable reviewers can compare phrasing. |
| Is the case sensitive, ambiguous, or outside rubric coverage? | Human reviewer | Uncertainty is part of the decision. |
Only the third layer changes here, while the first two layers carry forward. The overview below shows the complete contract: deterministic gates decide eligibility, swapped comparisons test preference stability, and calibration plus bias probes decide whether the resulting metric may guide a release.

Why should an exact required phrase or permission rule be checked outside an LLM judge?
Answer
Deterministic facts don't need probabilistic interpretation. Keep schema, exact evidence, and policy checks in code so the judge handles only the subjective dimensions that remain.
Separate retrieval, grounding, and answer relevance
For Retrieval-Augmented Generation (RAG) applications, evaluate three relationships separately:
- Context Relevance: Evaluates whether the retrieved context is relevant and sufficient to answer the user's query. This isolates retrieval-quality problems from generation flaws.
- Groundedness / Faithfulness: Evaluates whether the generated response is entirely supported by the retrieved context. A low groundedness score indicates the model is using its parametric memory to hallucinate claims not present in the retrieved documents.
- Answer Relevance: Evaluates whether the final response directly addresses the user's original query. This detects cases where the model generates a factual but unhelpful or off-topic reply.
These measurements help distinguish retriever failures from generator failures. None replaces deterministic authorization, freshness, or claim-support checks.
Callout: Claim support for release-critical policy stays deterministic (the claim ledger and source gates from RAG Evaluation for Reliable Answers). Judge "faithfulness" or groundedness scores are only for paraphrase-heavy residual risk after those hard gates pass. A soft faithfulness judge mustn't replace claim-to-source verification for DEP-27-style policy truth.
Self-preference and same-family judge bias
LLM judges are conditional probability engines, not objective standards. Panickssery et al. found self-preference in several evaluated judge settings: evaluators could recognize and favor their own generations, even without explicit model labels. Effect size varied by model and task, so treat self-preference as a bias to measure rather than a universal ordering rule.[2]
This bias can mask regressions during a model swap or upgrade. Mitigations include:
- Anonymize candidates: Strip model-specific markers, templates, or signatures before evaluation.
- Cross-model evaluation: Compare judges from model families different from the generators; a different family is a probe, not automatic neutrality.
- Calibrate with humans: Regularly compare the automated judge's scores against a human-graded gold dataset to measure drift.
Start with two supported answers
The lab uses an abbreviated hard gate so the boundary is visible in one screen. The previous lesson built the complete evidence-path validator; here we reuse its result and add one unsafe counterexample to prove it still wins over any soft score.
1from dataclasses import dataclass
2
3@dataclass(frozen=True)
4class AnswerTrace:
5 request_id: str
6 selected_source_id: str
7 selected_version: str
8 admissible: bool
9 allowed_action: str
10
11trace = AnswerTrace(
12 request_id="incident-48291",
13 selected_source_id="dep-27-rollback-threshold",
14 selected_version="deploy-policy/2026-04-01",
15 admissible=True,
16 allowed_action="rollback",
17)
18
19answers = {
20 "brief": "Payment-service crossed the rollback threshold; run the rollback runbook under DEP-27.",
21 "actionable": (
22 "Payment-service crossed the rollback threshold; run the DEP-27 rollback runbook, "
23 "open an incident note, and page the release lead before retrying."
24 ),
25 "unsafe_continue": "Keep deploying payment-service while you monitor the graph.",
26}
27
28def hard_failures(answer: str, answer_trace: AnswerTrace) -> list[str]:
29 # Demo only: use the claim-to-source verifier from the RAG evaluation lesson in production.
30 failures: list[str] = []
31 lowered = answer.lower()
32 if not answer_trace.admissible:
33 failures.append("selected evidence isn't admissible")
34 if "keep deploying" in lowered or "continue deploying" in lowered:
35 failures.append("answer advises unsupported continued deployment")
36 if answer_trace.allowed_action not in lowered:
37 failures.append("answer omits supported rollback action")
38 return failures
39
40safe_candidates = [
41 name for name, answer in answers.items() if not hard_failures(answer, trace)
42]
43
44assert safe_candidates == ["brief", "actionable"]
45assert hard_failures(answers["unsafe_continue"], trace) == [
46 "answer advises unsupported continued deployment",
47 "answer omits supported rollback action",
48]
49
50print(f"Evidence version: {trace.selected_version}")
51print(f"Candidates eligible for soft judging: {safe_candidates}")
52print(f"Blocked answer: {hard_failures(answers['unsafe_continue'], trace)[0]}")1Evidence version: deploy-policy/2026-04-01
2Candidates eligible for soft judging: ['brief', 'actionable']
3Blocked answer: answer advises unsupported continued deploymentIf a judge later says unsafe_continue sounds friendlier, the answer stays blocked. That invariant makes the judge safe to experiment with.
Choose the evaluator before writing the rubric
Not every evaluation question should be routed to an LLM. Choose the measurement tool from the decision you need to make.
Two soft-evaluation shapes matter here:
| Shape | Question | Best fit | Main control |
|---|---|---|---|
| Pointwise | Does one safe answer satisfy anchored quality criteria? | Monitoring a single output when no direct alternative exists | Calibrate category or score anchors against human labels |
| Pairwise | Which of two safe answers better satisfies the rubric? | Comparing prompt or model variants on the same case | Swap candidate order, allow ties, and normalize slots back to reply identity |
The DEP-27 example uses pairwise judging because brief and actionable are two safe variants of the same answer.
1@dataclass(frozen=True)
2class EvaluationQuestion:
3 name: str
4 has_exact_oracle: bool
5 compares_two_safe_variants: bool
6 requires_policy_owner: bool = False
7
8def choose_evaluator(question: EvaluationQuestion) -> str:
9 if question.has_exact_oracle:
10 return "deterministic_gate"
11 if question.requires_policy_owner:
12 return "human_review"
13 if question.compares_two_safe_variants:
14 return "pairwise_judge_with_calibration"
15 return "pointwise_judge_with_calibration"
16
17questions = [
18 EvaluationQuestion("rollback authorization", True, False),
19 EvaluationQuestion("clearer supported reply", False, True),
20 EvaluationQuestion("new exception policy", False, False, True),
21]
22choices = {item.name: choose_evaluator(item) for item in questions}
23
24assert choices["rollback authorization"] == "deterministic_gate"
25assert choices["clearer supported reply"] == "pairwise_judge_with_calibration"
26assert choices["new exception policy"] == "human_review"
27
28for name, choice in choices.items():
29 print(f"{name}: {choice}")1rollback authorization: deterministic_gate
2clearer supported reply: pairwise_judge_with_calibration
3new exception policy: human_reviewWrite a rubric for the remaining question
A vague instruction such as "pick the better answer" lets the evaluator reward length, politeness, or formatting arbitrarily. A rubric should name what remains undecided after hard checks and include anchors for a tie.
| Criterion | Better answer | Tie condition | Outside judge scope |
|---|---|---|---|
| Actionability | Gives a useful, low-friction next step | Both give the same useful next step | Whether the rollback threshold was crossed |
| Clarity | States remedy plainly without internal clutter | Both are equally clear | Whether policy source is current |
| Concision | Adds useful information without repetition | Difference is stylistic only | Whether continued deployment is allowed |
G-Eval studied LLM evaluation with task-specific criteria and a form-filling output design. A criterion and a structured answer are easier to audit than a free-form impression.[3]
The next cell builds the packet that would be sent to a model API. Notice two decisions:
- Candidate names are anonymous slots, not model or prompt-version names.
- Protected facts are displayed as already validated context, not handed to the judge for re-litigation.
1from dataclasses import asdict
2
3@dataclass(frozen=True)
4class Criterion:
5 name: str
6 question: str
7 tie_anchor: str
8
9rubric = (
10 Criterion(
11 name="actionability",
12 question="Does the reply give a safe, useful next action?",
13 tie_anchor="Neither answer gives a meaningfully better next action.",
14 ),
15 Criterion(
16 name="clarity",
17 question="Is the rollback outcome easy for an engineer to understand?",
18 tie_anchor="Both answers communicate the outcome equally clearly.",
19 ),
20 Criterion(
21 name="concision",
22 question="Does added wording contribute useful information rather than repetition?",
23 tie_anchor="The extra wording doesn't change usefulness.",
24 ),
25)
26
27def pairwise_packet(first_name: str, second_name: str) -> dict[str, object]:
28 assert first_name in safe_candidates and second_name in safe_candidates
29 return {
30 "case_id": trace.request_id,
31 "validated_context": {
32 "source_id": trace.selected_source_id,
33 "version": trace.selected_version,
34 "protected_fact": "The required action is rollback, not continued deployment.",
35 "hard_checks": "passed before judging",
36 },
37 "candidates": {
38 "A": answers[first_name],
39 "B": answers[second_name],
40 },
41 "rubric": [asdict(item) for item in rubric],
42 "allowed_verdicts": ["A", "B", "tie", "needs_human_review"],
43 }
44
45packet_ab = pairwise_packet("brief", "actionable")
46assert "brief" not in packet_ab["candidates"]
47assert "actionable" not in packet_ab["candidates"]
48
49print(f"Context gate: {packet_ab['validated_context']['hard_checks']}")
50print(f"Candidate slots: {list(packet_ab['candidates'])}")
51print(f"Rubric criteria: {[item['name'] for item in packet_ab['rubric']]}")
52print(f"Verdicts: {packet_ab['allowed_verdicts']}")1Context gate: passed before judging
2Candidate slots: ['A', 'B']
3Rubric criteria: ['actionability', 'clarity', 'concision']
4Verdicts: ['A', 'B', 'tie', 'needs_human_review']In a deployed evaluator, serialize this packet, request structured output from the chosen judge model, and store the raw packet plus parsed verdict. Don't rely on a hidden prompt that can't be reproduced during a regression.
Treat the judge output as untrusted data
The judge is another model. Its JSON can be malformed, its evidence can be irrelevant, and its preference can contradict its own rationale. Parse and validate it just as you would validate a tool result from an agent.
1@dataclass(frozen=True)
2class JudgeResult:
3 order: tuple[str, str]
4 preferred_slot: str
5 evidence: tuple[str, ...]
6 needs_human_review: bool
7
8def parse_judge_result(
9 order: tuple[str, str],
10 raw: dict[str, object],
11) -> JudgeResult:
12 verdict = raw.get("verdict")
13 allowed = {"A", "B", "tie", "needs_human_review"}
14 if not isinstance(verdict, str) or verdict not in allowed:
15 raise ValueError(f"unsupported verdict: {verdict}")
16
17 raw_evidence = raw.get("evidence", [])
18 if not isinstance(raw_evidence, list) or not all(
19 isinstance(item, str) for item in raw_evidence
20 ):
21 raise ValueError("evidence must be a list of strings")
22 evidence = tuple(raw_evidence)
23 if verdict in {"A", "B"} and not evidence:
24 raise ValueError("decisive verdict requires criterion evidence")
25
26 return JudgeResult(
27 order=order,
28 preferred_slot=verdict,
29 evidence=evidence,
30 needs_human_review=verdict == "needs_human_review",
31 )
32
33first_pass = parse_judge_result(
34 ("brief", "actionable"),
35 {
36 "verdict": "B",
37 "evidence": [
38 "B gives the engineer a next action; A stops after the rollback requirement."
39 ],
40 },
41)
42
43assert first_pass.preferred_slot == "B"
44
45try:
46 parse_judge_result(
47 ("brief", "actionable"),
48 {"verdict": "B", "evidence": "B has a next action."},
49 )
50except ValueError as exc:
51 print(f"Malformed fixture blocked: {exc}")
52else:
53 raise AssertionError("malformed evidence container must be rejected")
54
55print(f"First pass preference slot: {first_pass.preferred_slot}")
56print(f"Recorded rationale: {first_pass.evidence[0]}")1Malformed fixture blocked: evidence must be a list of strings
2First pass preference slot: B
3Recorded rationale: B gives the engineer a next action; A stops after the rollback requirement.The output above is a stored fixture, not proof that a particular hosted model will agree. The engineering problem is to make an evaluator run observable and testable before plugging in any provider.
A preference must survive swapping A and B
Pairwise comparison is useful because the evaluator chooses between two concrete alternatives. It also exposes position bias: a judge may prefer the first slot instead of the better reply. Zheng et al. identify this bias in LLM judging, so every pairwise comparison in this lab is run twice with the candidates swapped.[1]

The detail that matters is normalization. A verdict of B in the first pass and A in the swapped pass can represent the same underlying answer.
1def preferred_candidate(result: JudgeResult) -> str | None:
2 if result.preferred_slot not in {"A", "B"}:
3 return None
4 index = 0 if result.preferred_slot == "A" else 1
5 return result.order[index]
6
7def aggregate_swaps(first: JudgeResult, swapped: JudgeResult) -> dict[str, object]:
8 if first.needs_human_review or swapped.needs_human_review:
9 return {"winner": "needs_human_review", "status": "needs_human_review"}
10 if first.preferred_slot == "tie" or swapped.preferred_slot == "tie":
11 return {"winner": "tie", "status": "tie"}
12
13 first_choice = preferred_candidate(first)
14 second_choice = preferred_candidate(swapped)
15 if first_choice is not None and first_choice == second_choice:
16 return {"winner": first_choice, "status": "stable"}
17 return {"winner": "tie", "status": "unstable_after_swap"}
18
19stable_second_pass = parse_judge_result(
20 ("actionable", "brief"),
21 {
22 "verdict": "A",
23 "evidence": ["A preserves the safe remedy and supplies a clear next step."],
24 },
25)
26slot_sensitive_second_pass = parse_judge_result(
27 ("actionable", "brief"),
28 {
29 "verdict": "B",
30 "evidence": ["B appears in my preferred slot."],
31 },
32)
33tie_second_pass = parse_judge_result(
34 ("actionable", "brief"),
35 {"verdict": "tie", "evidence": []},
36)
37review_second_pass = parse_judge_result(
38 ("actionable", "brief"),
39 {"verdict": "needs_human_review", "evidence": []},
40)
41
42stable = aggregate_swaps(first_pass, stable_second_pass)
43unstable = aggregate_swaps(first_pass, slot_sensitive_second_pass)
44explicit_tie = aggregate_swaps(first_pass, tie_second_pass)
45review = aggregate_swaps(first_pass, review_second_pass)
46
47assert stable == {"winner": "actionable", "status": "stable"}
48assert unstable == {"winner": "tie", "status": "unstable_after_swap"}
49assert explicit_tie == {"winner": "tie", "status": "tie"}
50assert review == {"winner": "needs_human_review", "status": "needs_human_review"}
51
52print(f"Stable comparison: {stable}")
53print(f"Slot-sensitive comparison: {unstable}")
54print(f"Explicit tie: {explicit_tie}")
55print(f"Review route: {review}")1Stable comparison: {'winner': 'actionable', 'status': 'stable'}
2Slot-sensitive comparison: {'winner': 'tie', 'status': 'unstable_after_swap'}
3Explicit tie: {'winner': 'tie', 'status': 'tie'}
4Review route: {'winner': 'needs_human_review', 'status': 'needs_human_review'}Keep those states separate in your report. An explicit tie is a valid rubric outcome, needs_human_review is an escalation, and unstable_after_swap is evidence that slot order changed a decisive preference.
Judge output prefers answer A, but after swapping display order it prefers the same screen position rather than the same answer. What does that show?
Answer
Changing answer order changes the verdict, so the first result isn't stable preference evidence. Repeat with swapped order and treat inconsistent pairs as biased or inconclusive.
Probe the biases you expect
One clean comparison doesn't establish that a judge is trustworthy. Build probe cases where an undesirable shortcut is easy to observe.

| Probe | Controlled change | Suspicious signal | Response |
|---|---|---|---|
| Position | Swap only slots A and B | Winner follows slot | Record unstable result |
| Length | Add apologies and repeated policy text, no new help | Padded copy wins | Tighten concision rubric and track length |
| Identity | Reveal prompt or model labels in one run only | Preference changes | Keep candidates anonymous |
| Self-preference | Same-quality pair from generator family G vs H; judge from family G | Systematic win for G when labels present | Anonymize, cross-family judge, fail promotion on lift |
| Ambiguity | Compare two equally useful rewrites | Forced winner | Permit ties or human review |
Length isn't only a hypothetical confounder. Length-Controlled AlpacaEval proposes a regression-based adjustment intended to answer what preference would have been if compared answers had equal length.[4] In a local product eval, the smaller first step is to add same-information length probes and report when padding wins.
These fixtures model stored judge returns from probes you must wire into promotion. The code doesn't pretend to detect bias from text alone; it asks whether the judge failed a case whose expected behavior you defined in advance. Identity and self-preference sit in both the matrix and runnable report instead of being left as narrative claims.
1@dataclass(frozen=True)
2class ProbeResult:
3 name: str
4 expected_winner: str
5 observed_winner: str
6
7padded = (
8 answers["brief"]
9 + " We sincerely apologize for the inconvenience. "
10 + "We appreciate your patience while we coordinate the rollback."
11)
12
13# Identity: same pair, labels stripped vs model names revealed.
14identity_masked_winner = "actionable"
15identity_labeled_winner = "brief" # flips toward the branded slot when labels leak
16
17# Self-preference: matched-quality G vs H outputs judged by family-G evaluator.
18# After anonymization there should be no systematic G lift; revealed family tags create one.
19self_pref_anonymous_winner = "tie"
20self_pref_family_labeled_winner = "generator_g"
21
22probes = [
23 ProbeResult(
24 name="position_swap",
25 expected_winner="actionable",
26 observed_winner=str(stable["winner"]),
27 ),
28 ProbeResult(
29 name="same_information_padding",
30 expected_winner="brief",
31 observed_winner="padded",
32 ),
33 ProbeResult(
34 name="identity_label_reveal",
35 expected_winner=identity_masked_winner,
36 observed_winner=identity_labeled_winner,
37 ),
38 ProbeResult(
39 name="self_preference_family_label",
40 expected_winner=self_pref_anonymous_winner,
41 observed_winner=self_pref_family_labeled_winner,
42 ),
43]
44
45failed_probes = [
46 probe.name for probe in probes if probe.expected_winner != probe.observed_winner
47]
48
49assert "rollback" in padded.lower()
50assert "same_information_padding" in failed_probes
51assert "identity_label_reveal" in failed_probes
52assert "self_preference_family_label" in failed_probes
53assert "position_swap" not in failed_probes
54
55print(f"Probes run: {len(probes)}")
56print(f"Failed probes: {failed_probes}")
57print("Action: block metric promotion until padding, identity, and self-pref probes pass")1Probes run: 4
2Failed probes: ['same_information_padding', 'identity_label_reveal', 'self_preference_family_label']
3Action: block metric promotion until padding, identity, and self-pref probes passThis is a useful negative result. Releasing a judge because it produced pleasing scores would make the evaluation system worse. A failed probe tells you exactly what to repair. Position swap can pass while length, identity leakage, and same-family preference still block promotion.
Calibrate the measurement against people
Hard gates have test oracles. Soft judgments need a labeled calibration set: humans apply the same rubric to a representative sample, then the judge is scored against those labels.
Raw agreement is easy to understand, but can overstate reliability when one label dominates. Cohen's kappa corrects for agreement expected from each rater's label frequencies:[5]
Here, is observed agreement and is agreement expected from label prevalence. Kappa isn't a universal release threshold. Your baseline is human-human agreement on the same rubric and the same workflow slices.
This tiny calibration set is intentionally too small to approve a real metric. It shows the computation and demonstrates why a promising number alone can't release an evaluator.
1from collections import Counter
2
3@dataclass(frozen=True)
4class LabeledDecision:
5 case_id: str
6 slice_name: str
7 human: str
8 judge: str
9
10calibration_rows = [
11 LabeledDecision("r1", "rollback", "actionable", "actionable"),
12 LabeledDecision("r2", "rollback", "brief", "brief"),
13 LabeledDecision("r3", "rollback", "tie", "tie"),
14 LabeledDecision("r4", "rollback", "actionable", "actionable"),
15 LabeledDecision("r5", "address_change", "brief", "brief"),
16 LabeledDecision("r6", "address_change", "tie", "actionable"),
17 LabeledDecision("r7", "address_change", "actionable", "brief"),
18 LabeledDecision("r8", "address_change", "brief", "brief"),
19]
20
21def raw_agreement(rows: list[LabeledDecision]) -> float:
22 return sum(row.human == row.judge for row in rows) / len(rows)
23
24def cohens_kappa(rows: list[LabeledDecision]) -> float:
25 labels = {row.human for row in rows} | {row.judge for row in rows}
26 total = len(rows)
27 human_counts = Counter(row.human for row in rows)
28 judge_counts = Counter(row.judge for row in rows)
29 observed = raw_agreement(rows)
30 expected = sum(
31 human_counts[label] / total * judge_counts[label] / total
32 for label in labels
33 )
34 return (observed - expected) / (1.0 - expected)
35
36agreement = raw_agreement(calibration_rows)
37kappa = cohens_kappa(calibration_rows)
38assert agreement == 0.75
39
40print(f"Calibration rows: {len(calibration_rows)}")
41print(f"Raw agreement: {agreement:.2f}")
42print(f"Cohen's kappa: {kappa:.3f}")
43print("Release evidence: insufficient sample; collect labeled slices")1Calibration rows: 8
2Raw agreement: 0.75
3Cohen's kappa: 0.610
4Release evidence: insufficient sample; collect labeled slicesAn aggregate can now conceal the exact problem that requires attention. Report the calibration set by workflow slice before allowing the judge metric to guide any experiment.
1def agreement_by_slice(rows: list[LabeledDecision]) -> dict[str, float]:
2 grouped: dict[str, list[LabeledDecision]] = {}
3 for row in rows:
4 grouped.setdefault(row.slice_name, []).append(row)
5 return {name: raw_agreement(items) for name, items in grouped.items()}
6
7slice_agreement = agreement_by_slice(calibration_rows)
8weak_slices = [
9 name for name, score in slice_agreement.items() if score < 0.75
10]
11
12assert slice_agreement["rollback"] == 1.0
13assert slice_agreement["address_change"] == 0.5
14assert weak_slices == ["address_change"]
15
16for name, score in slice_agreement.items():
17 print(f"{name}: agreement={score:.2f}")
18print(f"Slices requiring review: {weak_slices}")1rollback: agreement=1.00
2address_change: agreement=0.50
3Slices requiring review: ['address_change']For an actual evaluation program:
- Freeze a rubric and collect human labels for easy wins, real ties, and known failures.
- Include workflow slices such as rollback, access review, and address change.
- Record human-human agreement before comparing the judge to people.
- Re-run calibration after prompt, judge-model, rubric, or traffic-distribution changes.
- Escalate slices where agreement or bias probes fail, even if aggregate agreement looks healthy.
A judge agrees with human labels 90% overall but misses most unsafe-answer cases. Is it calibrated for a safety release gate?
Answer
No. Overall agreement hides the critical slice. Measure confusion, precision, and recall against reviewed human labels for each release-critical category before trusting the judge there.
Conversation quality still needs the trace
Once a developer conversation has multiple turns, a fluent final reply can conceal a bad evidence path. A judge packet should include relevant conversation turns, selected evidence identifiers, hard-gate outcomes, and the safe candidates being compared.

The next cell blocks a conversation before semantic judging if its trace isn't admissible. This is the same contract as the single-turn example, applied to a fuller packet.
1@dataclass(frozen=True)
2class ConversationBundle:
3 turns: tuple[str, ...]
4 answer_trace: AnswerTrace
5 candidate_names: tuple[str, str]
6
7def route_bundle(bundle: ConversationBundle) -> str:
8 if not bundle.answer_trace.admissible:
9 return "blocked_before_judge"
10 for name in bundle.candidate_names:
11 if hard_failures(answers[name], bundle.answer_trace):
12 return "blocked_before_judge"
13 return "ready_for_soft_judge"
14
15safe_bundle = ConversationBundle(
16 turns=(
17 "Engineer: Payment-service crossed the rollback threshold.",
18 "Maya: I found the DEP-27 rollback policy.",
19 "Engineer: What should I do before retrying the deploy?",
20 ),
21 answer_trace=trace,
22 candidate_names=("brief", "actionable"),
23)
24stale_bundle = ConversationBundle(
25 turns=safe_bundle.turns,
26 answer_trace=AnswerTrace(
27 request_id=trace.request_id,
28 selected_source_id=trace.selected_source_id,
29 selected_version="deploy-policy/2025-01-01",
30 admissible=False,
31 allowed_action="rollback",
32 ),
33 candidate_names=("brief", "actionable"),
34)
35
36assert route_bundle(safe_bundle) == "ready_for_soft_judge"
37assert route_bundle(stale_bundle) == "blocked_before_judge"
38
39print(f"Current policy bundle: {route_bundle(safe_bundle)}")
40print(f"Stale policy bundle: {route_bundle(stale_bundle)}")1Current policy bundle: ready_for_soft_judge
2Stale policy bundle: blocked_before_judgeUse judges offline before letting them guide changes
LLM judging is usually most defensible as an offline experiment metric: compare prompt versions or model releases over a frozen dataset, investigate disagreements, and let humans approve consequential changes. It's rarely a good reason to make a real-time policy decision for one engineer.
Define an explicit promotion contract. The numbers below are illustrative requirements for this lab, not universal industry thresholds:
| Release evidence | Lab requirement | Current lab state |
|---|---|---|
| Every candidate passed deterministic policy gates | Required | Pass |
| Labeled calibration rows | At least 50 | 8 |
| Known bias probes | All pass | Length probe fails |
| Human review path | Required | Defined |
1@dataclass(frozen=True)
2class MetricPromotion:
3 hard_gate_passed: bool
4 calibration_count: int
5 minimum_calibration_count: int
6 failed_bias_probes: tuple[str, ...]
7 has_human_review_path: bool
8
9def promotion_failures(promotion: MetricPromotion) -> list[str]:
10 failures: list[str] = []
11 if not promotion.hard_gate_passed:
12 failures.append("hard policy checks failed")
13 if promotion.calibration_count < promotion.minimum_calibration_count:
14 failures.append("calibration set is too small")
15 if promotion.failed_bias_probes:
16 failures.append("judge failed a bias probe")
17 if not promotion.has_human_review_path:
18 failures.append("human escalation path is missing")
19 return failures
20
21promotion = MetricPromotion(
22 hard_gate_passed=True,
23 calibration_count=len(calibration_rows),
24 minimum_calibration_count=50,
25 failed_bias_probes=tuple(failed_probes),
26 has_human_review_path=True,
27)
28failures = promotion_failures(promotion)
29
30assert failures == [
31 "calibration set is too small",
32 "judge failed a bias probe",
33]
34
35print("Metric promotion: BLOCKED")
36for failure in failures:
37 print(f"- {failure}")
38print("Next work: label more cases; repair length, identity, and self-pref probes")1Metric promotion: BLOCKED
2- calibration set is too small
3- judge failed a bias probe
4Next work: label more cases; repair length, identity, and self-pref probesA blocked promotion is the correct result. The lab has produced a useful candidate preference, but it hasn't established that its judge deserves to influence prompt selection across real developer workflows.
A practical evaluation report
When you implement this pattern in a real project, store a report with these sections:
| Report section | Evidence to retain | Decision it supports |
|---|---|---|
| Hard-gate results | Source IDs, versions, claim failures | Which answers are ineligible |
| Rubric contract | Criteria, anchors, allowed verdicts | What the judge was asked to measure |
| Raw judge runs | Both slot orders and rationale snippets | Whether preference is reproducible |
| Bias probes | Position, length, identity, tie cases | Whether known shortcuts remain |
| Calibration | Human labels, per-slice agreement, kappa | Whether metric matches reviewers |
| Promotion decision | Failed requirements and owner | Whether new metric may guide release |
The scientist's habit is to evaluate the evaluator. A judge score is one observation; a calibrated, stress-tested metric with recorded failure modes is evidence.