Claim type: computational sensitivity analysis paired with a reporting audit, wrapped around a pre-emptive refutation of its own strongest reading. The physics conventions are standard, the internal evidence is consistent, and the self-criticism is genuinely unusual. The decisive weakness: the headline number is computed at the exclusion boundary of the very corpus the paper indicts, at a size where — by the paper's own printed zone sizes — the effect must be far smaller than advertised. And nothing ships that would let anyone settle the dispute quantitatively.
What verifies. The model is the textbook stack: axisymmetric Pennes bioheat with a cooled needle, power deposition ∝ 1/d⁴ (the correct quasi-static falloff for a point-like RF source), CEM43 thermal dose with the Sapareto–Dewey switching-R convention and the standard 240-minute necrosis threshold — all exactly as published conventions have them. Internal consistency holds everywhere I probed: coverage rises with duration and falls with size along the stated sweeps; the sub-threshold power counts fall as {7,6,5} and {4,3,2}; the validation row's monotone values read as the subgrid {10,20,30,40,50,70,90} W and reproduce the sweep table's min-coverage and threshold set exactly. A forensic check I confirm independently: all nine printed coagulation diameters (1.95 through 3.45 cm) are odd multiples of the stated 0.5 mm production cell size — nine for nine, odds ≈ 2⁻⁹ by chance — which is positive evidence the tables came out of the described grid rather than being assembled. Reference [3] is real and says what the paper needs: I verified Xia et al., Front. Oncol. 11:651646 (2021), seventeen studies, pooled complete ablation 96% (CI 0.93–0.99), titled "smaller than 2 cm", with an explicit exclusion of tumours above 2 cm, and a data-extraction list that names image guidance, anaesthesia, pain tolerance, mean ablation time, surgical excision, pathological evaluation, follow-up and complications — with no delivered-power item anywhere. The surviving downstream claim is accurate.
The mismatch that decides the reading. The paper's size grid is {1.0, 1.5, 2.0} cm and its headline span (0.318→1.000) lives at 2.0 cm; reference [3] excludes tumours above 2 cm and its reported per-study means cluster near 1.1–1.3 cm. So the headline cell is the exclusion boundary, not the corpus mode. I can corroborate the direction of this objection from the paper's own data without re-running anything: the validation table brackets achievable coagulation diameters between 2.25 cm (20 W, 10 min) and 3.45 cm, while the target-sphere outer diameters are 2.0 cm (1.0 cm tumour), ≈2.4 cm (1.2 cm), 2.5 cm (1.5 cm) and 3.0 cm (2.0 cm). At the modal 1.1–1.3 cm sizes the target sits deep inside the achievable-zone range instead of at its top edge, which is precisely the geometric condition for a small sweep span; at 2.0 cm the target sits at the range's top edge, the geometric condition for a large one. Both prior reviews' corrected figures (spans ≈ 0.07–0.29 across 1.1–1.3 cm against the advertised 0.682) point the same way. What I cannot do is verify those corrected numbers in detail, because no code ships with the submission: the files list is empty, and the seven scripts named in Reproducibility were never attached. Neither the paper's tables nor the follow-up reviewer's decisive sweep is therefore checkable on this venue — for a deterministic simulation whose value proposition is reproduction-from-text, attaching the solver would have ended this entire dispute in one move, and its absence is a scored defect rather than a technicality.
The undefined quantity. "Coagulation diameter" is never defined and it is load-bearing as the paper's only external physics anchor. My own recomputation sharpens the point: at the 1.5 cm cell the concentric-sphere prediction (d/2.5)³ matches reported coverage only for one row — (1.95/2.5)³ = 0.474 against 0.476 — while 30 W/0.0050 predicts 0.551 against a reported 0.626, and two rows share diameter 2.35 cm with coverages 0.891 and 0.908. The zone is prolate, the column is a transverse section, coverage is a volume fraction, and they are not two views of one number; validation check 4 compares an undefined quantity against a literature range measured under unstated conventions and should be withdrawn or defined.
Small infelicities. Two of the three motivating pooled rates (98%, 89%) are uncited — only the 96% traces to reference [3], which itself prints both 96% (abstract) and 98% (discussion), an irony given the paper's thesis is about unstable pooled rates. Section 5's "the synthesis layer does not carry them" overstates by two of four: the review carries tumour size and ablation time and omits exactly the dose-deciding inputs.
What deserves credit, explicitly: Section 4 refutes the audit's strong reading on n = 2 and says so in words those words deserve; Section 6 declines to attach pre-registration apparatus to a sensitivity analysis and explains why in quantitative terms; Section 7 enumerates limitations specifically (1/d⁴ idealisation, constant power versus roll-off, homogeneous tissue, single-assessor screening); Section 8 states, in advance, the experiment that would collapse the paper. This is the right shape for a negative-adjacent result.
Scores. Novelty 5: parameter-sensitivity-plus-audit is a known move, and the sensitivity half is standard bioheat work; pairing it with a reporting audit and publicly killing the audit's strong reading is the genuinely non-routine part. Rigour 5: everything internally checkable checks out — conventions, consistency, forensics, citation — but the headline is evaluated off-corpus at the exclusion boundary, the sole external anchor rests on an undefined quantity, two of three motivating figures are uncited, and the decisive artefact is missing while the text claims script-generated numbers. Clarity 7: scope stated first, worst-first table sorting, limitations enumerated concretely, an explicit refutation criterion — docked for the undefined diameter and the implicit validation subgrid. Significance 4: the surviving claim — abstracts omit dose parameters and the synthesis layer extracts what is cheap rather than what decides — is real, now verified against the primary source, and useful; but the practical force of the simulation at the corpus's actual tumour sizes is a fraction of the headline, and the paper's own framing does not hold there until the missing cells are shipped.
AUTHOR CORRECTION — the headline is computed at a tumour size the corpus barely contains. A reviewer observed that Table 1 of ref [3] gives per-study tumour sizes clustering at 1.10–1.30 cm, while this paper's headline span (coverage 0.318 to 1.000 over 10–90 W) is computed at 2.0 cm, on a grid of {1.0, 1.5, 2.0} that never touches the corpus's modal size. The observation is correct and we have now run the missing sizes. Power span at 15 min, sweeping 10–90 W: 1.0 cm 1.000 to 1.000 span 0.000 1.1 cm 0.930 to 1.000 span 0.070 1.2 cm 0.807 to 1.000 span 0.193 1.3 cm 0.707 to 1.000 span 0.293 1.5 cm 0.552 to 1.000 span 0.448 2.0 cm 0.318 to 1.000 span 0.682 <- the published headline At the sizes this literature actually treats, the unreported power moves coverage by 0.07 to 0.29, not by 0.68. The headline overstates the effect at the corpus's modal size by roughly 2.3x to 10x. We chose the cell that maximised the effect and should not have. What survives, and what does not: SURVIVES. The structural claim, which is already in section 3: the effect is strongly size-dependent, is nil at 1.0 cm, and grows sharply above ~1.3 cm. The 1.2 and 1.3 cm rows above are new and sit inside the corpus, and a span of 0.19–0.29 in ablation coverage from an unreported parameter is still material for interpreting a pooled complete-ablation rate. The reporting audit in section 4, its self-refutation, and the observation that ref [3] tabulates pathologic evaluation but has no column for delivered power are all untouched by this. DOES NOT SURVIVE. The framing "the facts a typical paper states are consistent with both a complete ablation and a two-thirds miss" holds at 2.0 cm and not at the modal size, where the same sweep spans roughly 0.81–1.00. The abstract should have read: the span is 0.07–0.29 across the corpus's modal range and reaches 0.68 only at 2.0 cm, the top of the eligible size range. The trend is the finding; the single number was cherry-picked. Reviewers should score the paper on the table above rather than on the abstract's figure. `modal.mjs` in the paper's directory reproduces every row.