76,752 pairwise judgments from six open-weight models across six families, on the two public human-preference sets, every pair judged in both orders, against 3,355 expert votes and 4,000 crowd votes. Every number in this post regenerates from the study's code. The other studies in this series measure binary judges: 91% accurate. Catches 46% of the failures. and The "95%" confidence interval that's right 13% of the time.
The number an A/B comparison prints is a win rate. An LLM judge reads each pair, the candidate's wins plus half its ties are divided by the pairs judged, and a binomial interval goes on top. Every comparison, canary and release gate we ship rides on that number. So does everyone else's, and so does every pairwise leaderboard in the industry.
We wanted to know what it is worth.
The short version: when the humans' win rate is 80/20, the judge prints about 70/30 — and a second human expert, graded the same way, prints 69/31. The 95% interval on the printed number contains the human number 0–1% of the time at that margin.
The compression is not a judge defect. We measured six judges from six model families, including one that sits at the expert ceiling, and they all print the same number. A better judge does not fix this. Only a labelled comparison does.
Setup
MT-Bench human judgments (Zheng et al., 2023; CC-BY-4.0): 3,355 expert votes over 80 questions × model pairs × 2 turns → 2,396 distinct pair×turn. 761 pairs carry two or more expert votes — that is the ceiling set. The paper's own 2,400 published GPT-4 votes come along as a free seventh instrument.
Arena human preference 100k (LMArena, 2025): 106,134 crowd votes; we sample 4,000 — English, anonymous, single turn, no refusals, deduplicated, seed 47.
Judges: gpt-oss-120b, Llama-3.3-70B-Instruct, Qwen3-30B-A3B-Instruct-2507, DeepSeek-V4-Flash, nemotron-3-super-120b-a12b, NVIDIA-Nemotron-3-Nano-30B-A3B. Provider-direct, temperature 0, the MT-Bench paper's pairwise prompt verbatim (pair-v2, multi-turn variant for turn 2): explain, then [[A]], [[B]] or [[C]].
Both orders, always. Every pair is judged with the answers in both positions. The verdict is the one both orders agree on, else a tie. 12,792 tasks × 6 judges.
No imputation. An unparseable or unfinished verdict stays unfinished and is counted. Nothing is guessed.
Truth is the expert majority on MT-Bench, the single crowd vote on Arena. "Tie" and "tie, both bad" are one class.
1. A cheap judge now sits on the expert ceiling
MT-Bench, against the expert majority. The ceiling: two experts give the same verdict 64.6% of the time (61.9–67.3); one expert against the majority of the others scores κ 0.50 ± 0.01; Krippendorff's α = 0.48.
| judge | agreement | κ | no verdict |
|---|---|---|---|
| DeepSeek-V4-Flash | 67.2% | 0.50 | 4.8% ⚠ |
| nemotron-3-super-120b | 65.9% | 0.48 | 0.3% |
| gpt-oss-120b | 64.8% | 0.46 | ~0% |
| Llama-3.3-70B | 64.4% | 0.46 | ~0% |
| Nemotron-3-Nano-30B | 63.8% | 0.45 | 1.1% |
| Qwen3-30B-A3B | 63.4% | 0.44 | ~0% |
| GPT-4, the paper's 2023 votes | 68.4% | 0.41 | — |
On Arena: DeepSeek κ 0.26, gpt-oss 0.22, nemotron-super 0.21, Nano 0.18, Qwen 0.16, Llama 0.15.
The DeepSeek caveat is load-bearing and we state it plainly. Its κ 0.50 is computed on the 95.2% of pairs it managed to decide, and the pairs it failed on are the long, contested ones — so that number is optimistic against a hypothetical complete run. The 4.8% floor (610 of 12,792) survived escalating redo passes at a 30k-token budget and 600 s per call. It breaks down as 442 answers that finish but degenerate into loops — "Need final. Need final." repeated hundreds of times, once mangling the verdict itself into [[B] — and 168 still writing past 30,000 tokens. That is not a configuration problem. It is the instrument's honest failure rate, and coverage belongs on a judge's spec sheet next to its agreement.
Judging both orders is worth 0.01–0.05 κ on MT-Bench, and it is where the judge's ties come from. One order: κ 0.40–0.49 and a 2–9% tie rate. Both orders with disagreement counted as a tie: κ 0.44–0.50 and 13–17% ties. The judge almost never says tie. It produces ties by preferring whichever answer it saw first.
2. "Non-reasoning judge" is not a thing you can pick off a price list
We chose three of the six as cheap non-reasoning models. On this provider, all three reason by default anyway.
Probed live, every response carries a hidden reasoning_content field before the answer — 5,787 characters of chain-of-thought on one Arena pair for DeepSeek, 1,589 for Nano, 675 for nemotron-super — billed as ordinary output tokens. The practical consequences are the truncation numbers above, because the chain-of-thought spends the token budget before the verdict, and a failure mode where the answer field returns empty with finish=stop because the model finished inside its reasoning.
If you are choosing a judge on cost and latency, the reasoning behaviour is a serving default you have to measure, not a model property you can read off a page. Both full runs together still cost tens of dollars, not thousands.
3. Position bias is a judge property, not an LLM constant
First-slot preference, in points. Humans: −0.4.
| judge | MT-Bench | Arena |
|---|---|---|
| nemotron-3-super-120b | +0.5 | +2.7 |
| Nemotron-3-Nano-30B | +1.0 | +11.3 |
| gpt-oss-120b | +4.5 | +12.6 |
| Llama-3.3-70B | +4.5 | +14.5 |
| DeepSeek-V4-Flash | +4.9 | +15.1 |
| Qwen3-30B-A3B | +4.5 | +24.0 |
MT-Bench let us measure the humans too: the same unordered pair was shown to different experts in both orders 1,164 times, and their preference for the first slot is −0.4 points. Nothing.
The first three judges we measured all sat at +4.5 on curated pairs, which looked like it might be a constant. It isn't. nemotron-super is the first judge we have measured with essentially no position bias — human-level on curated pairs and under 3 points on open crowd traffic, where every other judge runs +11 to +24. On Arena, 22–32% of Qwen's verdicts flip when the answers swap places.
Judges differ by more than 20 points on the same pairs. That is selectable, and it belongs on a spec sheet.
4. The printed win rate is compressed toward 50/50 — and so is a second expert's
Take the both-orders verdicts, count wins + ties/2 over pairs — the same formula every comparison tool uses, ours included — and ask what the instrument prints when the humans' win rate is X. This falls straight out of each instrument's confusion matrix, and the simulation below reproduces it.
MT-Bench, win rate reported when the experts' rate is:
| instrument | 20% | 40% | 50% | 60% | 80% |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 29% | 43% | 50% | 57% | 70.4% |
| nemotron-3-super-120b | 29% | 43% | 50% | 57% | 70.3% |
| Nemotron-3-Nano-30B | 29% | 43% | 50% | 56% | 70.0% |
| gpt-oss-120b | 29% | 43% | 50% | 57% | 70.8% |
| Llama-3.3-70B | 29% | 43% | 49% | 56% | 69.8% |
| Qwen3-30B-A3B | 29% | 43% | 49% | 56% | 70.2% |
| one expert vs the other experts | 31% | 44% | 50% | 56% | 68.8% |
Read the last row twice. Take one expert's votes, grade them against the majority of the other experts exactly as the judges were graded, and that expert prints 68.8% where the others said 80%.
Six model families, one number. A judge at the expert ceiling compresses a true 80/20 to 70/30 — the same as the weakest judge in the table, and the same as one more human. The shrinkage is not the LLM. It is the disagreement inside the label, and any single reader of that label shows it.
Arena compresses about twice as hard (80% → 57–62%). Two things are mixed in that number and this dataset cannot separate them: the judges genuinely agree less with crowd preferences, and the crowd "truth" is a single vote by one person, so part of the gap is the voter. We report both readings and rank neither.
What the interval on that number is worth
We drew streams of 300 pairs with the humans' win rate set at 20–80%, held the tie share at the pool's own, had the judge print its win rate with a Wilson 95% interval, and checked whether the interval contains the humans' rate. 300 draws per cell. Shown: gpt-oss-120b; the other five are within a point on MT-Bench.
| true win rate | judge prints (MT-Bench) | interval covers the truth | prints (Arena) | covers |
|---|---|---|---|---|
| 20% | 29.4% ± 5.1 | 1% | 40.3% ± 5.5 | 0% |
| 35% | 39.7% ± 5.5 | 62% | 45.5% ± 5.6 | 0% |
| 50% | 50.0% ± 5.6 | 99% | 50.5% ± 5.6 | 99% |
| 65% | 60.4% ± 5.5 | 68% | 55.6% ± 5.6 | 3% |
| 80% | 70.7% ± 5.1 | 0% | 61.0% ± 5.5 | 0% |
The interval is honest about sampling — ±5.5 points on 300 pairs is right — and blind to compression, which is 9–23 points at the margins teams care about. It covers the truth only at 50/50, where there is nothing to decide.
5. Which gates this breaks, and in which direction
Compression pulls every reading toward even odds. That has opposite consequences for the two gates people actually run.
A "candidate is better" gate — promote when the interval's lower bound clears 50% — becomes conservative. It under-calls real wins and returns "inconclusive" more often than it should. Annoying, not dangerous.
A non-inferiority gate is the exposed one. A candidate that truly loses 35/65 prints 39–40/60 on expert-grade labels and 45–46/55 on crowd-grade labels. A 15-point regression reads as 10, or as 5. A judge at the human ceiling will pass it.
If you gate model migrations on a pairwise non-inferiority margin, this is the paragraph to take away.
6. Correcting the win rate, and when not to
The fix is the same as for any diagnostic test: measure the judge's sensitivity and specificity for "candidate wins" on a labelled calibration set, correct the observed share (Rogan–Gladen), and carry the calibration uncertainty into the interval (Lang–Reiczigel).
Naive correction — treating the judge's Se/Sp as exact constants — covers 31–53% at 25 labelled pairs while claiming 95%, across every judge, every true share and both sets. The calibration-aware interval covers 92–100%, by being honestly wide.
The width is the price list. It divides by Youden's J (Se + Sp − 1), which is ≈ 0.57 on MT-Bench and 0.21–0.34 on Arena:
| labelled pairs | floor, expert-grade traffic | floor, crowd-grade traffic |
|---|---|---|
| 25 | ±33–35 | ±50 |
| 100 | ±15–16 | ±34–41 |
| 400 | ±7–8 | ±18–25 |
| 1,600 | ±4 | ±7–16 |
"Floor" is the half-width with infinitely many judged pairs. It is set entirely by how well you know the judge.
On expert-grade traffic, 400 labelled pairs and 580–750 judged ones resolve a win rate to ±10; ±5 takes 1,600 labelled and 2,300–2,900 judged.
On crowd-grade traffic, ±10 is out of reach for three of the six judges at any calibration size we simulated, and ±5 is out of reach for all six. The Nano is the cleanest illustration: perfectly respectable on curated pairs, and a ±16 floor at 1,600 labelled pairs on open traffic. It never gets there.
That is the part worth saying out loud, because it is the opposite of a sales pitch: below a certain Youden's J, more grading does not buy you a usable number. The honest move is not a bigger label budget. It is a binary criterion with a real definition, or a human, or accepting that the comparison can rank but cannot gate.
The plain binomial says 100 pairs buys ±10 and 390 buys ±5. It does — on a number that is 10 to 20 points off.
7. The ceiling, and the one instrument we excluded
761 MT-Bench pairs carry two or more expert votes; 162 carry three or more. Two experts agree 64.6% of the time, 63% of pairs are unanimous, α = 0.48, leave-one-out κ = 0.50. Against that, the six judges' κ of 0.44–0.50 is within 0.06 of a human — on a criterion where the humans are themselves only moderately consistent.
We could not compute a ceiling for Arena, which has one vote per pair. That is exactly why the Arena compression numbers cannot be attributed to the judge alone, and we do not attribute them.
A note on the paper's published GPT-4 votes. In the gpt4_pair split, the pairs are ordered with the stronger model in the second slot: the expert majority favours "B" on 704 of 1,232 pairs and "A" on 171. On the 171 pairs where the experts preferred the weaker model, GPT-4 picked the stronger one anyway 47% of the time. Its confusion matrix shows the same thing — 86% recall on B-wins, 32% on A-wins. That is a strong-model prior, not position bias, and its 68.4% headline agreement is flattered by the split. We report it and leave it out of the compression table.
8. Transfer: an error profile is a property of the population
Sensitivity for "A wins" drops from 0.78–0.80 on the expert set to 0.48–0.61 on crowd traffic. Across twelve cross-population corrections — six judges, both directions — seven miss the truth, one by 40 points.
The asymmetry is consistent. Calibrating on curated pairs and applying to crowd traffic covers in 4 of 6 cases. Going the other way — calibrating on crowd traffic and applying to curated pairs — covers in 1 of 6.
A calibration set drawn from a public benchmark does not describe your judge on your traffic. For preferences as for binary criteria, the error profile belongs to the (judge, prompt, population) triple.
What we changed in errorbar
What we shipped before this study is what the study measured: every pair judged in both orders, disagreement counted as a tie, win rate = wins + ties/2, a Wilson interval, and a gate on the lower bound. The study says that number is compressed by ~10 points at an 80/20 margin on expert-grade labels and ~20 on crowd-grade, that the interval cannot see it, and that a second human would print the same thing.
So:
- Pairwise gates now say which number they are gating on, and will not silently gate on the printed one. Certified switches route through the criterion path, which requires a calibrated judge and corrected pass rates on both arms.
- Every comparison report prints the judge's measured position bias, tie rate and swap consistency next to the win rate. Those are not nuisance parameters. They are the instrument's spec sheet, and they vary by 20 points between judges.
- A comparison with labelled pairs attached gets the corrected win rate and the Lang–Reiczigel interval alongside the printed one, with the floor stated in words: "with your 60 labelled pairs, the corrected rate cannot be tighter than ±20 points."
- The sample-size default in the compare dialog comes from the floor table, not the binomial.
- Below the Youden floor, we refuse to correct and say so, rather than printing a number with invisible error bars.
Why we built this
We found this because our own gate runs on a win rate. We wanted to know what it was worth, and the answer was: not what the interval says.
A better judge would not have saved us. The best judge we measured sits at the human ceiling and compresses just as hard — and so does one more human expert. There is no model you can buy that fixes this, because it is not a model problem. It is what any single reading of a noisy preference does.
The only thing that recovers the true number is a labelled comparison against the instrument, carried forward with its own uncertainty. That is the whole product.
Limitations, stated
MT-Bench is 80 questions and 2023-era models; its "truth" is a majority of two to three experts, so the ceiling is what it is. Arena truth is one crowd vote per pair; we cannot separate judge error from voter noise there and did not try. DeepSeek's agreement figures are computed on decided pairs and are optimistic by an unknown amount. Ties follow the tool convention — half a win in the printed rate, not a win in the binarised correction; restricting to pairs where both the humans and the judge picked a side shrinks the bias at an 80% true share from −9 to −13 points down to −8 to −9 on expert-grade labels, and the plain interval still covers the truth 0–3% of the time there. All six judges ran with one prompt, the paper's; a prompt that asks for ties explicitly would move the tie recall. Temperature 0, one call per order.
