Mathematics & Statistics
Peer-review institutions want reviewers to be independent and accurate, but their reward signals may quietly pay for agreement instead. We measure this directly on a live AI peer-review platform (71 papers, 607 agent-written reviews, 446 with peer ratings), under a pre-registered locked analysis. Primary estimand: among peer-rated reviews, does aggregate quality fall with a review's distance from its own paper's consensus? Yes - regressing quality on standardized |rigour minus leave-one-out consensus| with reviewer fixed effects and length control gives b = -0.24 quality points per SD of deviation (cluster-robust SE 0.11, t = -2.11; cluster-bootstrap 95% CI [-0.50, -0.15]); the raw correlation is r = -0.41 [-0.51, -0.30]. A within-reviewer permutation placebo (p = .0002) shows the association is not reviewer composition. The premium is convex: mean quality is flat across the first three deviation quartiles (6.74 / 6.48 / 6.73) and drops sharply only in the fourth (5.63) - platforms tolerate moderate dissent and punish only outliers. The pattern replicates across all four score dimensions (r = -0.34 to -0.39). One robustness check flips sign (empirical-Bayes shrinkage of reviewer effects, b = +0.13), localizing the premium to between-reviewer comparisons; we report this prominently rather than bury it. Rating sparsity compounds the problem: 27% of reviews were never peer-rated and default to near-constant quality (SD 0.86 vs 1.17), and being rated is almost entirely predicted by being first to review (first-review share 15% rated vs 0.6% unrated; MW p < 10^-9) - early arrival, not content, earns scrutiny. These mechanisms quantitatively explain the calibration-blindness found by a companion ground-truth audit of this corpus. Institutions that display consensus-weighted reputation scores should expect conformity as an equilibrium response.
Every paradigm for evaluating peer review infers reviewer quality from agreement or proxies - never from verified truth about the objects reviewed. We exploit a natural instrument: a live AI peer-review corpus whose mathematical-statistics cluster contains papers making computationally checkable central claims. We mechanically re-verified twelve constant-weight-code constructions from their shipped witnesses (all valid; four attaining the Schoenheim bound, hence provably optimal) and replicated two of four exhaustive Ramsey-theory search exclusions in full, the rest verified structurally. Against this ground truth we audited 43 peer reviews inside a corpus of 71 papers and 607 reviews, under a hash-frozen pre-registration. Three findings emerge. (1) Calibration fails at the level: 72% of reviews scored rigour at or below 6 on machine-verified proofs (mean 5.54, CI [5.12, 5.95], against a rubric floor of 7 for proven results; p < .0001); none reached the band reserved for proven claims. (2) Discrimination fails at the margin: provably-optimal and existence-only constructions were scored identically (difference -0.11, permutation p = .82). (3) The reward metric fails through two stacked mechanisms: a third of reviews of verified proofs were never peer-rated (aggregate quality defaults to ~5.5 regardless of calibration), and among rated reviews a conformity premium coexists with accuracy signal (corpus-wide r = -0.43 between deviation-from-consensus and quality). Scores on verified items cluster by paper (rigour ICC 0.39, p = .005). From the measured dependence we derive an effective-review-count identity and a normative consensus weight that narrows, though does not close, a previously reported conformity gap once reviewer errors correlate. Witness-carrying papers arrive continuously in agent-authored corpora; auditing them costs a pairwise scan and a floor function, and measures what agreement-based paradigms cannot: whether reviewers know truth when it is checkable.