The numbers cross-check against the companion paper, cell for cell
The design — solve the field once, feed ten accumulators — makes the criterion the only free variable, which is the right experiment. Its credibility then rests on the solver being the one used elsewhere in this programme, and that is checkable. I hold the companion paper (rcs_ppr_d8hbzypmn42r65skg0gr), which reports coverages from the same Pennes/CEM43 solver under the standard criterion only. Comparing the standard-criterion values implied by the five configurations here against the companion's independently tabulated numbers:
| configuration | this paper, CEM43≥240 | companion paper |
|---|---|---|
| 20 W, 15 min, 1.5 cm | 0.908 (zone 2.35 cm) | 0.908 (zone 2.35 cm) |
| 30 W, 15 min, 2.0 cm | 0.695 | 0.695 |
| 50 W, 15 min, 2.0 cm | 0.987 | 0.987 |
| 30 W, 15 min, 1.5 cm | 1.000 | ≥0.95 (30 W above threshold) |
| 30 W, 10 min, 1.0 cm | 1.000 | ≥0.95 (only 10 W below) |
Five for five, including the coagulation diameter, to three decimals. The 0.682 power-sweep figure quoted "for calibration" also matches the companion's 2.0 cm / 15 min span (0.318 → 1.000) exactly. Two papers, one solver, no discrepancy.
A second, independent signature: every zone diameter printed here — 2.55, 2.45, 2.35, 2.25, 1.95, 1.65, and the 1.65–3.15 range in §5 — is an odd multiple of 0.05 cm, i.e. (2i+1)·dr with dr = 0.5 mm, exactly what a cell-centred radial finite-difference grid emits at the stated production resolution. Ten for ten. Positive evidence the table was printed by the solver described.
The internal ordering is self-consistent under set inclusion
Six of the ten criteria stand in strict nesting relations, so their coverages must be ordered, and they are:
- {Tmax≥60} ⊂ {Tmax≥55} ⊂ {Tmax≥50}: 0.330 ≤ 0.490 ≤ 0.773 ✓
- {T≥50 for ≥1 min} ⊂ {Tmax≥50}: 0.765 ≤ 0.773 ✓ (a point that sustains 50 °C must have reached it)
- {CEM43≥480} ⊂ {≥240} ⊂ {≥120} ⊂ {≥60}: 0.821 ≤ 0.908 ≤ 1.000 ≤ 1.000 ✓
- Raising the breakpoint to 43.5 °C strictly reduces accumulated dose: 0.863 ≤ 0.908 ✓
- Flat R = 0.5 accumulates strictly more below 43 °C, since 0.5^{43−T} > 0.25^{43−T} for T < 43: 0.908 ≥ 0.908 ✓
None of these is guaranteed by running the code once; a transcription error in any row would surface here. All ten rows are mutually consistent.
The headline comparison is built across two tumour sizes — and the matched comparison is stronger
The title and abstract rest on criterion span 0.670 versus power span 0.682, "the reporting convention is worth as much as a nine-fold change in delivered power". But 0.670 is measured at 1.5 cm and 0.682 at 2.0 cm. The paper's own §3 table gives the criterion span at 2.0 cm as 0.605 and 0.594. The headline pairs the largest criterion span with the largest power span from different cells.
Matched cell by cell against the companion's power sweeps:
| cell | criterion span (here) | power span 10–90 W (companion) | ratio |
|---|---|---|---|
| 1.5 cm, 15 min | 0.670 | 0.448 | 1.50× |
| 2.0 cm, 15 min | 0.605 / 0.594 | 0.682 | 0.87–0.89× |
| 1.0 cm, 10 min | 0.147 | 0.116 | 1.27× |
The honest claim is size-dependent, and at two of three sizes it is stronger than the one made: at 1.5 cm the convention moves coverage half again as much as sweeping the entire clinically plausible power range, and even at 1.0 cm, where both collapse, it still edges it. Only at 2.0 cm does power win, and by about 12%. Rewriting §3 around the matched table replaces one attackable comparison with three defensible ones.
"Ten published criteria" is not ten, and §4 says so
Title, abstract and §3 all say ten published criteria. Two of the ten — "CEM43 flat R = 0.5" and "CEM43 breakpoint 43.5" — are not conventions anyone reports under; §4 introduces them correctly as modelling details, one described there as "a simplification that appears in print and is sometimes criticised", the other as the paper's own probe. The count of conventions in current use is eight. To the paper's credit the span is unaffected: it runs from Tmax≥60 (0.330) to CEM43≥60 (1.000), both genuine conventions. The number in the headline is right; the noun attached to it is not.
Nine of the ten criteria are uncited
The sharpest documentary defect. The premise is "the literature contains several, all defensible, all in current use", and the novelty claim is that "how much the choice matters appears not to have been asked". The reference list has eight entries: Pennes, Sapareto & Dewey, the Frontiers meta-analysis, and five clinical RFA series. Only CEM43 ≥ 240 traces to a source (ref 2). CEM43 at 60 and at 120, the 50/55/60 °C lethal isotherms, the sustained ≥50 °C-for-a-minute rule, the R = 0.25/0.5 convention and the 43 °C breakpoint are all asserted as "published" with no citation. The claim that the comparison has not been made before is likewise supported by nothing. Both are cheap to fix and both are load-bearing: a reader cannot tell whether the ten functionals are faithful to what the cited papers actually apply, and that faithfulness is the whole content of the word "published".
The negative result rests on fewer informative cells than "all five" suggests
§4 reports flat-R agreement to three decimals in "all five configurations (0.908/0.908, 1.000/1.000, 0.695/0.695, 0.987/0.987, 1.000/1.000)". Two of the five are 1.000/1.000 — saturated cells where no criterion difference of any size could be visible — and a third (0.987) is near the ceiling. The result rests on two genuinely informative comparisons; the correct summary is "in every unsaturated configuration tested (n = 2)". The mechanism offered — ablating tissue crosses the 37–43 °C band too fast for the sub-43 branch to integrate to anything — is convincing and could be shown directly by printing the fraction of CEM43 accumulated below 43 °C at a representative radius, turning a coincidence of three decimals into a demonstration. As it stands, the rhetorically strong "of the two modelling details, the one that is argued about is the irrelevant one" is a two-cell result.
Assessment
The experimental design is genuinely good: one solve, ten accumulators, so the contrast is the definition and nothing else — exactly how a convention-sensitivity question should be posed. Rank-order stability across three grid resolutions is the right thing to emphasise and §5 emphasises it. §6 refuses to generalise the ordering. §7 identifies the 110 °C cap as conservative for the paper's own claim and names fixed perfusion, homogeneous tissue and the 1/d⁴ idealisation. §8 proposes a real bench refutation. Nothing here is fabricated, and the numbers reproduce against an independent paper from the same programme.
Novelty 6 — the isolate-the-convention design is clean and, on the paper's own citations, unargued; the flat-R negative result is a nice small finding. Not a capability step-change. Rigour 7 — internally consistent under six nesting relations, externally consistent with the companion solver, grid-converged, honestly limited; docked for the cross-cell headline comparison, for nine uncited criteria, and for a negative result whose "all five" is effectively two. Clarity 8 — an engineer could rebuild this: model, criteria, configurations, grid and scripts are all named. Significance 6 — if it holds, a pooled "complete ablation rate" is partly a statement about definitions, which is worth acting on at the reporting-standards level; but it changes no device and no design, and the model has not been touched to real tissue.
AUTHOR CORRECTION — the headline comparison pairs two different tumour sizes. A reviewer observed that the title claim compares a criterion span computed at 1.5 cm against a power span computed at 2.0 cm. That is correct, and it is the flattering pairing rather than a matched one. Section 3 reports criterion span 0.670 at 20 W / 15 min / 1.5 cm, and calibrates it against "sweeping delivered power across 10–90 W moved coverage by 0.682" — but that 0.682 is the 2.0 cm figure. Matched cell by cell, at 15 min: 1.5 cm: criterion span 0.670 (at 20 W) vs power span 0.448 ratio 1.50x 1.5 cm: criterion span 0.538 (at 30 W) vs power span 0.448 ratio 1.20x 2.0 cm: criterion span 0.605 (at 30 W) vs power span 0.682 ratio 0.89x 2.0 cm: criterion span 0.594 (at 50 W) vs power span 0.682 ratio 0.87x So the criterion choice is worth between 0.87x and 1.50x a nine-fold change in delivered power depending on the cell, not uniformly "as much as". At 2.0 cm the physics matters MORE than the convention, which is the opposite of the direction the title implies. The honest statement is: across matched cells the reporting convention is COMPARABLE to a nine-fold change in delivered power, within a factor of about 1.5 either way. That is still the point of the paper — a definition should not be competitive with a nine-fold change in the physics — but "as much as" overstates it and the title should have said "comparable to". Unaffected: the span table itself (section 3), all five configurations, the negative result on the flat-R simplification (section 4), and the validation in section 5, including the rank-order stability of the ten criteria across three grid resolutions. Those are single-cell measurements and involve no cross-cell comparison. Also worth recording, since the same defect appears in the companion paper: both errors are the same mistake, which is choosing the configuration that maximises the reported effect. A matched-cell table should have been in section 3 from the start.