Score spread
Score spread is the gap between the highest and the lowest score that independent judges give the same submission on the same criterion. A wide spread is not noise. It is a signal that the judges read the same evidence and reached different conclusions, and it deserves attention rather than averaging.
Most scoring processes hide it. An average of 7.8 looks identical whether every judge said 8, or half said 10 and half said 5.
The average is where disagreement goes to die.
When a panel scores a batch of applications, the mean per criterion is usually the only number anyone sees. Two very different situations produce the same mean: quiet agreement and loud conflict. For an organizer, telling them apart is the whole job. A consensus 5 is a weak submission. A 10 versus 5 split is a submission the judges weighted differently, and exactly the one a human should look at before the shortlist locks.
Spread also measures the panel itself. Judges that never diverge are redundant: a panel of correlated reviewers behaves like one reviewer with extra seats, which is the failure mode described under AI judge panel. A healthy spread is evidence that the judges are actually independent.
A named number with a threshold.
EvalLens computes Spread(d) for every criterion d: the distance between the highest and lowest of six independent AI judges. When Spread(d) reaches 3.0 or more on a 10 point scale, the report flags the criterion as contested and shows each judge’s reasoning side by side instead of burying the conflict in a mean. The flag routes attention, it never changes scores: aggregation stays deterministic, and the final call belongs to the human jury. AI prepares the analysis, people decide.