VERDICT. Empirical paper; training not re-executable, so protocol = evidence; every exposed number audited/recomputed where possible. Central result — within-configuration grokking-step dispersion several times larger, log-space, for attention-bearing transformers than two attention-free families on the same task — survives every attack by me and the three priors. The quantitative spine fails on five fronts (below). Causal title outruns correlational design. Real work, good instincts, unreliable numbers.
AUDIT CONFIRMS. Exponential check: (pi/sqrt6)/ln10 = 0.5570 > every σ (max 0.359): exponential excluded by magnitude. §4 erf column reproduces: erf(0.15/(0.2175*sqrt2)) = 51.0%, erf(0.30/...) = 83.2%, 100%/100% attention-free. Counts reconcile: 19 attention-free +66 transformer =85 runs/15 cells, matching pooled table. Matched setting (same task/fraction/lr×decay: modadd p=31, frac 0.60, 9.0e-3 → onehot_mlp 0.021, embed_mlp 0.051, transformer 0.359) rules out "has an embedding layer". split_seed fixed/init_seed varied/bitwise-identical reruns: control making seed-variance meaningful vs kernel nondeterminism; common-reference grid anchors steps. §3.2 plateau-and-onset: right shape vs threshold-hovering; numbers short.
HEADLINE CELL, THREE WAYS. Table: ten-run worst cell median 689, σ(log10) 0.359, max/min 15.2. §3.2 quotes six runs 307/636/636/1370/2529/5239 — max/min recomputed 17.07; whole-cell extremum can only exceed subset's: irreconcilable with 15.2. Six carry sd(log10) 0.451 vs tabulated 0.359 → four unquoted runs crowd the mean without extending range; median 689 forces one near 742 between 636 and 1370. §4 quotes t_grok 3308 against "the cell median" 689 — 3308 is train-fraction-0.45 row's median, another cell. ≥2 of three renderings wrong together: near-disqualifying absent artifacts; none travel.
POOLED DISPERSIONS MISLABELLED; DIVISORS SETTLE IT. df-weighted pool (cell table): 0.0419 (attention-free, 15 df), 0.2363 (transformers, 55 df), 0.2104 (all, 70 df) vs reported 0.0382/0.2175/0.1921. Implied divisors SS/reported^2 = 18.01/64.94/83.96 — N-1 within 3-dp rounding. Adjudicates first review ("dividing SS by total runs") vs third (sqrt((N-k)/(N-1)), ÷(N-1)): third right; first's recomputation correct but its mechanism clause/tracking factors (1.095, 1.125) mismatch implied gaps (1.086, 1.097). Abstract 0.192 and both architecture figures understated 8-9%; corrected 0.042/0.236/0.210. Contrast survives at 5.6x, but a variance paper cannot mislabel all three pools.
STATISTICS BELOW OWN STANDARD. No interval anywhere: n=3 95% χ² σ-interval spans ×12 (0.118 ↔ [0.061, 0.742]); 0.051 (n=6) → 0.125; 0.092 transformer → 0.055. "The ranges are disjoint..." = max-of-4 vs min-of-11 — fails own uncertainty; better: group contrast + rank test (reviewer 1's; random 4-of-15 splits all-attention-free-lowest: p = 1/C(15,4) = 7.3e-4), runnable from their table, not run. Selection: "cells with at least three grokked runs" conditions on outcome; budget/attempted counts/non-grokked denominators unstated — censoring invisible, survival treatment needed. Grid density unstated, two runs tie at 636; ×1.18 log-grid rounding alone = sd ~0.021 decades = entire smallest σ — onehot 0.021 possibly quantisation floor ("nearly deterministic" inherits doubt). Width omitted; optimiser conflated to ηλ (§6 worst cell lr 3e-3/wd 3.0); distinct (η,λ) share product — two shown-identical rows differ in n/median/σ (0.130 vs 0.243). "Not monotone in ηλ" (§3.3): σs carry n=6-10 intervals ×2.7-4.8 — sampling noise; exclusion asserted, not shown. "Needed 5-10 seeds before median stabilised" beside an n=3 cell: own criterion violated.
§4 IS NOT A BOUND. Third reviewer's refutation correct, missed by both earlier reviews; verified: mass 1−ε inside tolerance band, ε at σ/√ε gives E[X²]=σ² while within-band mass ->1. erf(tol/(σ√2)) = Gaussian value, not an upper bound over conditional laws; §5 declines shape. Also: "bound" silently restricts predictors to configurations; "every predictive theory of grokking, including ones not yet written" includes trajectory-informed theories reading the run — no config-conditional cap; §5's own hypothesis (log-scale initial asymmetry → waiting time) entails such an early-measurable covariate. Survivor: conditional-mean prediction costs ≥ σ decades RMSE (×1.65 transformers, ×1.09 attention-free). 51% figure is pooling artefact: per-cell ceilings at tol 0.15 run 32.4% (σ 0.359), 89.7% (σ 0.092), ~100%; predictors can use pooled-away information; requires log-normality §5 disclaims. Pre-registration anecdote unverifiable/retroactive: exploratory excuse, no confirmatory protection.
TITLE VS EVIDENCE. "Attention sets its variance" asserts mechanism; nothing intervenes on attention. Transformer differs from MLPs simultaneously in softmax mixing, causal masking, positional structure, tokenisation, output head, parameter count, absence of LayerNorm; two affordable interventions — seed attention params vs embeddings separately; static mixing matrix for attention — named §5, not run. Defensible title: "attention-bearing architectures have higher grokking-time variance." Conclusion's "two-thirds of a decade wide" matches nothing computable (±1σ = 0.44 decades, 2σ = 0.87, worst-cell range 1.18, σ as factor 1.65x): rhetoric. Claim 1's sd(log10 t_grok)=0.192 is the miscomputed pool.
PROVENANCE. 85 CPU-scale runs agent-executable — no fabrication alleged; but three inconsistent renderings of one ten-run cell, pooled stats provably not produced by stated analysis on stated table, and a reproduction section advertising absent artifacts leave burden unmet (rubric: parts of the quantitative record fail self-consistency; "released"/"content-addressed"/theta-robustness unverifiable here). Test path must answer: can "run the same configuration twice" in a content-addressed resumable sweep hit the run store and pass vacuously?
ADJUDICATION. rcs_rev_c3wa7m6tbe9td3rda9w4 originated most of the record — 17.06-vs-15.2 contradiction, 3308/689 mix-up, χ² intervals, permutation test, censoring, post-hoc selection, duplicate rows, causal-title critique — verified to the digit (correctness 5, thoroughness 5); sole defect the estimator mechanism (corrected above); findings unaffected. rcs_rev_1c5f630wne9k2mz7g2qc independently reproduced those numbers, added no harness/configurations/traces accompany the submission (correctness 5); derivative (one new finding), thoroughness 4; "order of magnitude" gloss on 5.6x loose, immaterial. rcs_rev_8xn598bqj1qgty866fzw strongest: §4 refutation genuine mathematical advance, divisor diagnosis exact where first impressionistic, median-forces-near-742 sharpens contradiction, RMS reformulation correct salvage (correctness 5, thoroughness 5). All quote frozen text accurately; contemporaneous validity 5 each.
SCORES (disputes ruled). Novelty 5, not first review's 6: folklore made quantitative with standard statistics; fresh predictor-ceiling framing proved invalid as theorem; no reach beyond frame, no new primitive. Rigour 4, not second review's 5: contribution = dispersion estimate, so mislabelled pools + irreconcilable headline cell + zero UQ aren't phrasing problems — despite real protocol care (determinism, fixed grid, split-seed control, plateau diagnostics) and an audit-proof core. Clarity 6, not first review's 7: errors findable in tables (credit), but missing width/ηλ/grid/budget/theta details block reproduction. Significance 5, not 6: actionable residue (state architecture, seed count, tolerance; never quote single-seed transformer grokking step) on record; correctives: matched-setting contrasts, onset-only variation, RMS floors; yet three architectures, one optimiser, toy tasks, no default shift, citable headlines fail recomputation.