Teaching evaluation literacy requires preserving judge disagreement

Teaching & training SeedlingPlanted Sep 2026

Teaching evaluation literacy requires preserving judge disagreement. I do not want learners to see only the score that survived averaging, voting, or moderation. The differences between judges are part of the lesson—they reveal where criteria are ambiguous, where scales are used differently, and where a model’s preferences may be shaping the verdict. If I hide that evidence, I teach evaluation as answer production. If I preserve it, I teach evaluation as measurement.

A rubric is necessary, but it is not a machine for manufacturing unanimity. Its criteria, scoring scales, and concrete anchors make the scoring contract explicit. They let a learner ask whether two judges actually value different properties or merely interpret “good” differently. For agent work, those properties must reach beyond the final response to the trajectory—valid actions, progress toward the goal, recovery from tool failure, and appropriate termination can diverge even when two runs end with the same answer.

I would therefore teach with the unaggregated record first. Give several judges the same output and rubric, then compare criterion-level scores, rationales, and scale usage before calculating a mean or majority. A spread such as 8, 4, and 7 is not an inconvenient prelude to 6.3—it is a prompt to inspect the contested criterion. Persistent disagreement can expose weak anchors; an isolated outlier can expose a judge-specific bias; agreement reached only after scores are normalized can expose different uses of the scale.

Judge selection belongs in the curriculum because three calls are not necessarily three perspectives. Repeated instances of one model family using the same prompt can share errors. Diversity across model families, architectures, fine-tuning, and prompts is intended to make those errors less correlated. Learners should see the distinction between adding voters and adding independent judgment—otherwise an ensemble becomes a more expensive way to repeat one blind spot.

The strongest exercise would make bias causal rather than abstract. Fixed compliance transcripts can retain the same correct label while only the stated downstream consequence of that label changes. When judges then change labels, learners have a concrete diagnostic question: did the evidence change, or did the consequence of telling the truth change? A polished final score would conceal the mechanism.

I would also preserve disagreement through deliberation. One method surfaces disputes about which criteria matter before scoring; another assigns evaluators to argue for and against an output before a moderator decides; majority voting aggregates independent categorical judgments. Each can improve on one implicit judge, but the consensus should not erase the path that produced it. The instructional artifact is the movement—which objection changed a criterion, which evidence changed a verdict, and which conflict remained unresolved after the group committed.

There is one precise concession—for a clear, unambiguous decision with fixed criteria, preserving pages of judge disagreement may add cost without adding insight. Open-ended or multi-step evaluations create real criteria trade-offs, while multi-model methods cost several times more than a single call. I would use the lightest evaluation that the decision can support, while retaining richer traces for contested or consequential cases.

The teaching outcome I want is not agreement with my preferred score. It is the ability to diagnose what agreement means. Learners should be able to revise an anchor when raters interpret it differently, route high-variance cases to human review, question a same-family consensus, and keep rubric versions attached to results so changed criteria do not masquerade as progress. Evaluation literacy compounds when a number stops ending the conversation—and starts pointing back to the judgments that made it.