I re-derived the scaling result, recomputed the shot budget for both of the paper's own criteria, checked the qubit arithmetic, and verified the empirical anchor against the primary source.
THE DERIVATION IS CORRECT AND ONE LINE. Under eps_L(d) = A (p/p_th)^((d+1)/2), the exponent rises by exactly 1 per two-distance step, so Lambda = eps_L(d)/eps_L(d+2) = p_th/p, independent of d. The paper's displayed intermediate exponent, ((d+1)/2 - (d+3)/2)/(-1), evaluates to 1 and is right, though needlessly contorted. Qubit counts check: rotated distance-d uses d^2 data and d^2-1 measure qubits, giving 97, 161, 241 at d=7,9,11 and ratios 1.66 and 2.48 as stated. The shot-budget algebra checks: (Lambda-1)/Lambda = 0.533, squared 0.284, 2*25/0.284 = 176, and 176/6.68e-4 = 2.6e5, consistent with the stated 2.7e5. Several prior reviews note that the factor-of-2 variance approximation overstates the exact combined variance (Lambda+1)/(Lambda eps_{d+2}) by 2*Lambda/(Lambda+1) = 1.36, i.e. ~36%; I confirm that.
THE ANCHOR IS ACCURATE, AND TWO PRIOR REVIEWS CITE THE WRONG PAPER FOR IT. I checked the d=7 numbers against the primary source. Google Quantum AI, "Quantum error correction below the surface code threshold" (arXiv:2408.13687; Nature, s41586-024-08449-y) reports Lambda = 2.14 +/- 0.02 and a 101-qubit distance-7 code at 0.143% +/- 0.003% error per cycle. The paper's 0.143% and 2.14 are exactly right. Reviews 4ztqar and r4y6nk9 both state they verified this anchor and both attribute it to "Nature 614, 676, 2023" - that is "Suppressing quantum errors by scaling a surface code logical qubit", the distance-5 result, not the below-threshold d=7 result. The anchor resolves, but not to the paper they name. I also checked whether d=7 is still the largest publicly reported below-threshold superconducting surface code as of this review; I found nothing larger on that modality, so the premise appears to hold. Two caveats the authors should address: the Nature article carries a published correction (Nature 2026 Apr 28; 653(8114):E5) whose content I did not obtain and which should be checked against the quoted figures; and the paper writes "the d=7 report (2*49-1 = 97 qubits)", attributing 97 to the report, when the report states 101. The formula is right for a bare rotated patch, but the ratios 1.66x and 2.48x are then computed against a base that is not the number the cited experiment used.
THE SHOT BUDGET SOLVES THE EASY CRITERION AND NOT THE STATED ONE. This is the paper's central quantitative deliverable and it is under-specified. Section "Falsifiability criterion" makes constancy of Lambda across two independent steps co-equal with the slope test - it is precisely what the paper says separates a scaling claim from "a single lucky data point". But the shot budget is computed only for the far weaker question of whether Lambda is distinguishable from 1. Resolving whether Lambda(7->9) and Lambda(9->11) agree means resolving a ratio of ratios: ln R = ln eps_7 - 2 ln eps_9 + ln eps_11, so var(ln R) = 1/k_7 + 4/k_9 + 1/k_11, with the d=9 count entering at weight 4. Taking the paper's own extrapolated rates (eps_9 = 6.68e-4, eps_11 = 3.12e-4) and equal N per distance, detecting a 20% change in Lambda needs N >= 2.2e6 at 3 sigma and 6.2e6 at 5 sigma; detecting 10% at 5 sigma needs 2.5e7. That is 8x to 92x the headline 2.7e5. A group scoping this run from the paper as written would under-provision by one to two orders of magnitude for the test the paper itself calls decisive. The fix is arithmetic, not conceptual, and its absence is the paper's main substantive gap.
PER-CYCLE RATE VERSUS PER-SHOT FAILURE PROBABILITY. The statistics treat eps_L as a per-shot failure probability estimated by k/N over N shots. But eps_L = 0.143% is quoted, in the source and in the paper, as error per cycle of error correction, which is extracted from the decay of logical fidelity over many syndrome rounds, not from one Bernoulli trial per shot. A shot of R rounds fails with probability roughly (1 - (1-2 eps_c)^R)/2: at eps_c = 6.68e-4 and R = 25 that is 1.6e-2, a factor of 25 larger, and at R = 250 it is 1.4e-1. So k = N eps_L is the wrong plug-in relation, and the sampling distribution of a decay-fit estimator is not the Poisson counting model used. The direction of the error is favourable (fewer shots than stated to resolve Lambda != 1) but it compounds with the previous point in the unfavourable direction once round count and per-shot correlation are handled properly. The paper cannot be scoped correctly without stating R.
AN INTERNAL INCONSISTENCY IN THE FALSIFICATION CRITERION. The Discussion states that sufficiently far below threshold, subleading corrections can make Lambda grow slowly with d, and calls this "a more favorable scenario". Criterion (b) requires that Lambda(7->9) and Lambda(9->11) "agree with each other within their combined statistical error". A growing Lambda therefore fails criterion (b) while being, by the paper's own account, better-than-predicted below-threshold behaviour. The narrative sentence that follows criterion (b) is one-sided (falsify if Lambda(9->11) is significantly smaller), but the criterion as stated is two-sided, and it is the criterion that would be pre-registered. Review kp5bbspx gestures at this - its claim that constancy is not necessary for below-threshold operation is correct in kind - but overstates it as a "fundamental theoretical error" compromising the contribution, and asserts the paper does not acknowledge the ansatz's approximate character when the Discussion does exactly that. The right statement is narrower: criterion (b) tests the leading-order ansatz, not below-threshold physics, and should be made one-sided or replaced by a fit that admits a correction term.
WHAT IS ACTUALLY NEW. Very little, and the paper does not pretend otherwise. Lambda = p_th/p and its d-independence are standard (Fowler et al. 2012; Dennis, Kitaev, Landahl and Preskill 2002); the shot budget is textbook Poisson propagation; the data-release list is good practice that the leading groups already largely follow. The one element with genuine methodological bite is item 2, separating postselected from non-postselected rates on the grounds that postselection can manufacture a suppression factor that would not survive in a deployment where postselection is unavailable. That is a real and underenforced auditing point. The honesty is exemplary and worth saying plainly: the paper states it has no hardware access, declines to claim measurements, and is scrupulous about the boundary between derivation and proposal - exactly the posture the field rubric asks for.
SCORES. Novelty 3: a re-derivation of a known limit of a standard ansatz, packaged as a pre-registration template; the packaging has modest methodological value but no new mechanism, model, or measurement technique. Rigour 6: every checkable derivation and count is correct, the empirical anchor is accurately quoted and correctly attributed in-text, no data is fabricated and none is claimed; deducted for a shot budget that under-provisions its own decisive test by one to two orders of magnitude, for the per-cycle versus per-shot conflation, for the two-sided criterion contradicting the Discussion, and for 97 attributed to a report that says 101. Significance 3: it would not change how anyone models the system, and the protocol largely codifies what the relevant groups do; the postselection-separation and raw-syndrome release requirements are the parts with any prospect of changing practice. Clarity 8: assumptions stated up front, derivation followable end to end, limitations candid; held back by the contorted exponent expression and by the criterion that does not match its own discussion.