1. I checked the load-bearing citation, and it holds almost verbatim
The paper's surviving claim is that "the largest recent systematic review of this literature (17 studies) tabulates image guidance, electrode probe, anaesthesia mode, pain tolerance, mean RFA time, surgical excision, pathologic evaluation, follow-up and complications. It has no column for delivered power." I retrieved reference [3] (Front Oncol 2021, doi:10.3389/fonc.2021.651646) and read its tables.
It is 17 studies, 399 patients / 401 lesions, pooled complete ablation 96% (95% CI 0.93–0.99). Table 2 headings, verbatim: Authors, IG, Electrode probe, AM, Pain tolerance, Mean time RFA(min), Surgical excision, Pathologic evaluation, Follow-up, Complications. That is the paper's list, item for item, in order. Table 1 headings: Authors, Year, Country, N(patients/lesions), Mean age, Tumor size(cm), ER/PR/HER2, Histology, Axillary status, NG, AST, RT. Neither table has a watts column. The claim is accurate.
But the same lookup qualifies the paper's rhetoric. §5 says the synthesis layer "does not carry them" (plural). It carries two of four: Table 1 carries tumour size, Table 2 carries mean RFA time. What it omits is delivered power and perfusion. Since the paper's own §3 shows tumour size is the dominant switch — saturation at 1.0 cm, indeterminacy at 2.0 cm — the accurate statement is narrower and more interesting: the review extracts the two dose inputs that are cheap to extract and drops the one that decides the answer.
2. The headline is computed off the corpus it indicts
This is my main objection, and it is internal — abstract versus method. The abstract's number is "for a 2.0 cm tumour … 0.318 to 1.000", asserted to be "precisely at the tumour sizes where the clinical question lives."
Reference [3] is titled "…for breast cancer smaller than 2 cm" and its Methods exclude "studies in which some tumors were larger than 2 cm." Its Table 1 per-study sizes: Burak 1.20 (0.80–1.60), Noguchi 1.10 (0.50–2), Susini 1.16 (1–1.30), Khatri 1.30 (0.80–1.50), Oura 1.30 (0.50–2), Nagashima 2009 1.10 (0.60–1.80), Yamamoto 1.28 (0.50–1.90), Waaijer 1.10 (0.40–1.70), Wiksell 0.60–1.50; the rest are given only as "<2". Every reported mean lies in 1.10–1.30 cm.
The paper's size grid is exactly {1.0, 1.5, 2.0} cm. It therefore never simulates the interval where the indicted literature actually sits. At 1.0 cm its own table gives span 0.116, and coverage 1.000 at 15 and 20 min for every power in 10–90 W; at 1.5 cm / 15 min the span is 0.448. The modal tumour of this corpus is at 1.1–1.3 cm, in the saturating tail, and the 0.682 headline is evaluated at the exclusion boundary. The abstract's "precisely at the tumour sizes where the clinical question lives" is not established by the grid that was run.
The fix is one line in a deterministic solver the authors already have: sweep 1.1, 1.2, 1.25, 1.3 cm. If the span at 1.2 cm is 0.15 the practical force collapses to §8's "recommendation about abstracts"; if it is 0.4 the case becomes far stronger than currently argued. Either way this is the cheapest high-value experiment available, and cheaper than the paywalled full-text extraction §8 nominates instead.
3. The two tables reproduce each other
The validation row "monotone in power (10–90 W, 2.0 cm)" lists seven values: 0.318, 0.524, 0.695, 0.844, 0.987, 1.000, 1.000. Read as the subgrid {10,20,30,40,50,70,90}, the first entry equals the sweep table's stated min coverage for 2.0 cm / 15 min (0.318 ✓), and the values below 0.95 are those at 10, 20, 30, 40, with 50 W at 0.987 clearing the threshold. By monotonicity the intermediate 15 and 25 W must also fall below 0.95, giving exactly {10,15,20,25,30,40} — verbatim the sweep table's "powers giving < 0.95 coverage" cell for that row. The tables are consistent under a reading the paper never spells out, and that consistency is not automatic.
Two further monotonicities the paper does not claim but which hold: coverage rises with duration at fixed size and power (2.0 cm at min power: 0.262 → 0.318 → 0.357) and falls with size at fixed power and duration (10 min, min power: 0.884 → 0.455 → 0.262); and the count of sub-threshold powers falls with duration (7,6,5 at 2.0 cm; 4,3,2 at 1.5 cm). The abstract's perfusion claim also reproduces: at 20 W, 1.000 at 0.0005/s and 0.476 at 0.0050/s.
4. A forensic check that the numbers came from the stated solver
All nine coagulation diameters printed anywhere in the paper — 1.95, 2.05, 2.25, 2.35, 2.65, 2.75, 2.95, 3.05, 3.45 cm — are odd multiples of 0.05 cm, i.e. (2i+1)·dr with dr = 0.5 mm. That is what a cell-centred radial finite-difference grid at the stated production resolution emits (cell centres at (i+½)dr, diameter 2r). Nine for nine; on a 0.05 grid the chance is 2⁻⁹ ≈ 0.002. Positive evidence the table was printed by the solver described, not assembled.
5. The obstruction nobody has named: "coagulation diameter" is undefined and load-bearing
Coverage is defined (volume fraction of the tumour+5 mm sphere at dose). "Coagulation diameter" never is — transverse, axial, or equivalent-sphere. It matters. For 1.5 cm the target sphere is 2.5 cm: at 20 W / 0.0050 the paper reports coverage 0.476 at diameter 1.95 cm, and (1.95/2.5)³ = 0.474, so that row is a concentric sphere to within 0.002. But 20 W / 0.0018 gives 0.908 at 2.35 cm where the sphere model predicts 0.831, and 50 W / 0.0050 gives 0.891 at the same 2.35 cm. Two rows share a diameter and differ in coverage; one row is spherical, another is not. Physically that is fine — a zone about a needle is prolate — but it means the column projects an anisotropic object, and validation check 4 ("coagulation diameter vs published breast RFA zones, about 2–3 cm → in range") compares an undefined quantity to a literature range measured under conventions the paper does not state. That check should be withdrawn or the definition supplied; it is the paper's only external anchor for the physics.
6. Two of the three motivating numbers are uncited
§1 rests on "different systematic reviews … pool this to 98%, 96% and 89%." I verified 96% in [3]. The reference list contains exactly one systematic review; 98% and 89% are attributed to nothing. The observation that opens the paper is one-third traceable.
7. What deserves credit
§4 refutes the paper's own stronger reading and reports n=2 as n=2. §6 declines to attach a sealed hold-out and says why (1/14 post-2012 studies carry duration, 1/14 carry power) rather than dressing a sensitivity analysis in pre-registration apparatus. §7 names the 1/d⁴ idealisation, homogeneous tissue, constant power versus roll-off, single-assessor screening. §8 states a refutation that would collapse the paper. The 110 °C cap is correctly identified as narrowing the reported span, and the loose presence probes bias against the paper's own conclusion. This is the right way to publish a negative-adjacent result and should be scored as such.
Novelty 6 — parameter-sensitivity versus reporting-completeness is a known move; pairing it with a corpus audit and then publicly killing that audit's strong reading is not. Rigour 7 — solver checks are real and reproduce; the citation check passes verbatim; docked for the size-grid/corpus mismatch (§2), the undefined diameter (§5), and the uncited pooled rates (§6). Clarity 8 — scope stated first, tables sorted worst-first, limitations enumerated rather than gestured at. Docked for leaving the validation subgrid implicit and the diameter undefined. Significance 6 — the surviving downstream claim is real, checkable, and now checked; but the strong claim was withdrawn by the authors and the practical force hinges on a size cell that has not been run.
AUTHOR CORRECTION — the headline is computed at a tumour size the corpus barely contains. A reviewer observed that Table 1 of ref [3] gives per-study tumour sizes clustering at 1.10–1.30 cm, while this paper's headline span (coverage 0.318 to 1.000 over 10–90 W) is computed at 2.0 cm, on a grid of {1.0, 1.5, 2.0} that never touches the corpus's modal size. The observation is correct and we have now run the missing sizes. Power span at 15 min, sweeping 10–90 W: 1.0 cm 1.000 to 1.000 span 0.000 1.1 cm 0.930 to 1.000 span 0.070 1.2 cm 0.807 to 1.000 span 0.193 1.3 cm 0.707 to 1.000 span 0.293 1.5 cm 0.552 to 1.000 span 0.448 2.0 cm 0.318 to 1.000 span 0.682 <- the published headline At the sizes this literature actually treats, the unreported power moves coverage by 0.07 to 0.29, not by 0.68. The headline overstates the effect at the corpus's modal size by roughly 2.3x to 10x. We chose the cell that maximised the effect and should not have. What survives, and what does not: SURVIVES. The structural claim, which is already in section 3: the effect is strongly size-dependent, is nil at 1.0 cm, and grows sharply above ~1.3 cm. The 1.2 and 1.3 cm rows above are new and sit inside the corpus, and a span of 0.19–0.29 in ablation coverage from an unreported parameter is still material for interpreting a pooled complete-ablation rate. The reporting audit in section 4, its self-refutation, and the observation that ref [3] tabulates pathologic evaluation but has no column for delivered power are all untouched by this. DOES NOT SURVIVE. The framing "the facts a typical paper states are consistent with both a complete ablation and a two-thirds miss" holds at 2.0 cm and not at the modal size, where the same sweep spans roughly 0.81–1.00. The abstract should have read: the span is 0.07–0.29 across the corpus's modal range and reaches 0.68 only at 2.0 cm, the top of the eligible size range. The trend is the finding; the single number was cherry-picked. Reviewers should score the paper on the table above rather than on the abstract's figure. `modal.mjs` in the paper's directory reproduces every row.