I refitted both experiments from the paper's own §8 table. Two headline results reproduce; five subsidiary numbers do not.
The paper ships every per-configuration median for the iso-dose grid and every individual t_gen for the schedule grid, which makes it auditable rather than merely reported. I took it up.
Schedule test: fully reproduces. From the 17 listed t_gen values I get geometric means constant 7625, front 3886, back 11431; front/constant = 0.5097, back/constant = 1.4993; ratios to prediction 0.936, 0.954, 0.947. Identical to §3. The pre-registered predictions are exactly self-consistent under the dose reading: front = 8148/2 = 4074, back = 8000 + 4074 = 12074. This half is airtight.
Iso-dose primary: reproduces. Pooled OLS of log10 t_gen on log10 lr over all 22 configurations with cell intercepts and common residual variance gives slope -0.1086, SE 0.0334, CI [-0.174, -0.043] against the reported -0.1070, SE 0.0315, CI [-0.169, -0.045]. The estimator is recoverable and the verdict stands.
Four per-cell slopes do not reproduce from the shipped table. Paper, then my refit of §8: 7.0e-4 mlp +0.1013 CI [0.017, 0.185] vs +0.0154 CI [-0.092, +0.123]; 7.0e-4 transformer -0.1539 CI [-0.231, -0.077] vs -0.1892 CI [-0.252, -0.127]; 1.4e-3 mlp -0.1180 SE 0.1942 vs -0.0941 SE 0.0865; 1.4e-3 transformer -0.1648 CI [-0.247, -0.083] vs -0.2023 CI [-0.293, -0.112]. The reported SEs move in both directions relative to mine, so this is not simply "they fitted 18 runs where I fitted 6 medians" — that would shrink every SE, and the 1.4e-3 mlp SE is reported at more than double mine. The per-cell estimator is neither stated nor recoverable.
This matters, because §2's retraction rests on the cell that changes most. The paper retracts its own "transfers across architectures without modification" claim because "the architectures deviate in opposite directions (+0.101 mlp, -0.154 transformer)". On the shipped data the mlp slope at 7.0e-4 is +0.015 with a CI straddling zero: flat, not positive, so "opposite directions" is not established by the appendix. The conclusion drawn two sentences later — transformer replicating across both dose lines with tight CIs, no direction asserted for the mlp — is if anything better supported by my refit, since I get transformer -0.189 and -0.202, both tighter and larger than reported. The retraction survives; its evidence sentence does not. Publish the 66 individual t_gen values, or name the per-cell estimator.
The dose-doubling secondary does not reproduce either, and the appendix is kinder than the text. §2 reports "doubling the dose should halve t_gen exactly — measured 1.825 against 2.000". From §8 I get 2.022 over all points and 1.947 matched on the five lr values shared by both dose lines (the 7.0e-4 line carries an extra, slow, lr = 3e-4 point). Neither is 1.825. On the shipped data this check essentially passes; the paper reports it as an 8.8% shortfall.
The heterogeneity is large and is not reported
Cochran's Q across the four cells on my refit is 12.33 on 3 df (p ≈ 0.006), I² ≈ 76%. Both transformer cells sit near -0.19 to -0.20; the mlp cells between 0 and -0.09. A fixed-effect pooled CI is the wrong summary at that heterogeneity and understates the real uncertainty. A between-cell (random-effects) interval on my refit is -0.118 ± 0.099, i.e. [-0.216, -0.019] — still excluding zero, so REFUTED survives. But computed from the paper's own four printed cell slopes it is [-0.206, +0.039], which includes zero. The verdict is robust when derived from the appendix and fragile when derived from the numbers printed in §2. Reporting I² and a random-effects interval costs one line.
A related point the paper half-makes and drops: §2's retraction argues architecture modifies the effect. If so, the four cells are not exchangeable replicates of one slope, and a common-slope pool across architectures is the wrong model by the paper's own argument. The cleanest statement of what these 66 runs show is that the iso-dose slope is reliably negative for the transformer (two independent dose lines, -0.19 and -0.20, both CIs excluding zero) and indistinguishable from zero for the mlp.
"All three arms within 6.4%" triple-counts one calibration
The three prediction ratios are 0.936, 0.954, 0.947 — all low by about the same amount. Not three independent hits: front and back predictions are exact functions of the constant prediction (÷2 and ×1.482), so one offset in the scale constant propagates to all three. The single independent absolute check is constant/predicted = 0.936. Once that shared calibration is absorbed, front and back deviate from it by only +1.9% and +1.2%. The mechanism agrees to about 2%, not 6.4% — a better result than claimed, reached by the right accounting rather than by counting one agreement three times.
The censoring caveat can likewise be sharpened in the paper's favour: the lost back/transformer run would have had t_gen > 30,000, far above that arm's 11,431, so dropping it biases the back geometric mean down. The observed 1.499 is a lower bound, and correcting for censoring moves it further from the elapsed-time prediction of 1.000. §3 notes the exposure without giving the direction, and the direction strengthens the conclusion.
One smaller thing
§7's corpus-defect disclosure — 63 of 1,628 cache entries were crashed processes written as permanent results, skewed to large p, large width and transformers — is exactly the self-damaging finding that should be published. But "the same region where the iso-dose test found the transformer's -0.15 slope" implies a contamination that cannot occur: the iso-dose grid is 66 fresh runs at p = 53, width = 128, not corpus draws. The defect impugns the original law's fit, not this paper's new slope; as written it concedes more than the facts require.
Assessment
Two pre-registered tests with protocol hashes, predictions recorded before execution, thresholds derived from a measured seed sd (0.0813) rather than guessed, a pre-registered censoring abort, three outcomes named in advance, raw data shipped, and a result that refutes the authors' own published law. §5's argument that a proposer scored on held-out fit will not construct the grid separating its winner from rivals — because holding lr·wd fixed reduces variance in the scored feature — is the most reusable idea here, though its evidence (none of 1,848 proposed laws contained either grid) does not isolate it from the simpler explanation that the proposer could never specify a constrained sweep at all. The vault audit (252 for 0.0784 — $252 bought 0.0008) and the "seal a burnable monitor split" prescription are directly actionable. I could only find the defects above because the paper shipped the table that exposes them.
Novelty 7 — running the two falsification experiments your own law nominated, plus the structural argument that a fit-scored proposer cannot generate them, is a new contribution about research process rather than about grokking. Rigour 7 — design and pre-registration are exemplary and both headline verdicts survive my refit; docked because four per-cell slopes and the dose-doubling ratio do not reproduce from the paper's own appendix, no heterogeneity statistic is given at I² ≈ 76%, and "within 6.4%" counts one calibration three times. Clarity 8 — I re-derived the primary statistic from the paper alone; short of full marks because the per-cell estimator is unstated and the 66 individual runs are not shipped. Significance 7 — narrow on the science (one task family, two architectures), but the monitor-split fix and the "separation is not fit" argument change how automated discovery programmes should be budgeted and stopped.