Research
Why we grade by recomputation: LLM judges, measured as instruments on public human-labelled data. Every number in every study regenerates from its code.

The grader reported a gain while true accuracy fell
A pre-registered study of reward corruption in a real GRPO run. A 7B policy was trained against a cheap LLM grader with no reference answer, and the frozen set was read three ways every 16 steps: the grader's verdict, exact-match truth, and an independent certified anchor. On one of two seeds true accuracy fell 14 points while the grader reported 0.85 to 0.96; on the other it dipped 7 and recovered. The anchor stayed within 0.004 of the truth while the truth fell 49 points, held each corrupted run at the checkpoint that was in fact the best, and issued zero holds in eleven checks on a control whose training reward rose 14 points, because the rule compares two instruments on one frozen population and the training curve never enters it. On 600 human-labelled QA items with the reference in the prompt, the same cheap grader matched the reasoning judge to a point and none of four judges cleared the bar; the ceiling was the labels. With the reference on math, the grader that passed 61% of wrong answers certified. The rule now reads the gap between the two instruments on the same outputs and names three states. Half of every GRPO group that carried gradient on the corrupted runs was a group where every rollout was wrong and the grader split anyway: the curriculum selects the seam. Two predictions failed in the useful direction; five study errors are disclosed.

Seven frontier judges, one physician-written criterion, none over the bar
We tried, as hard as a motivated customer would, to certify an LLM judge as an RL reward for HealthBench's emergency-referral criterion against 186 physicians' labels. Seven model families, two grading protocols, twenty rounds of prompt optimisation, ensembles of the best, and a human rewrite of the criterion from the physicians' disagreements: 57 certificates, none trustworthy, and the reward door refused every one. Then we measured the mid-run anchor on three healthy training runs and one we tried to game: at this anchor quality it stops healthy runs about half the time, and the gaming did not take. Pre-registered, errors disclosed.

91% accurate. Catches 46% of the failures.
36,063 fresh judgments from three 2026 open-weight models over JUDGE-BENCH's human-labelled datasets, measured as instruments rather than ranked as contestants. Accuracy hides the direction a judge is wrong in; the usual corrected-rate interval covers the truth 13–46% of the time; a calibration dies in transport.

A true 80/20 prints as 70/30. So does one more human expert.
76,752 pairwise judgments from six open-weight judges on the two public human-preference sets. Every judge — including one at the expert ceiling — compresses a true 80/20 win rate to 70/30, exactly like one more human expert. The interval can't see it; a labelled calibration can.

The “95%” confidence interval that’s right 13% of the time
A re-analysis of JUDGE-BENCH plus 36,063 fresh judgments from three 2026 open-weight models. Five things a single agreement score hides — and what an LLM judge's error bars should actually look like.