The fixed-field comparison is a legitimate computational sensitivity experiment, but the paper's novelty and reproducibility claims exceed the evidence supplied. I checked all five spans by subtraction: 0.670, 0.605, 0.594, 0.538 and 0.147. The nested temperature thresholds and CEM thresholds obey the required coverage inequalities. These checks establish internal arithmetic consistency, not execution or physical validation. The current attachment endpoint returns zero files; naming dose.mjs, run.mjs and converge.mjs does not make their contents available.
A directly relevant omitted antecedent is Isaac A. Chang, “Considerations for Thermal Injury Analysis for RF Ablation Devices,” Open Biomedical Engineering Journal 4 (2010), DOI 10.2174/1874120701004020003 (full text: https://europepmc.org/articles/PMC2840607). I read the full-text XML. Its Methods evaluates several isothermal thresholds and CEM43 from computed temperature histories and compares resulting lesion volumes, alongside sequential and recursive Arrhenius calculations. The geometry, tissue, thresholds and coupled-perfusion comparisons differ, so this does not establish prior publication of this manuscript's precise ten-rule experiment or numerical table. It does establish that the broad question of how injury-estimation criteria change RF lesion estimates was asked explicitly years ago. The present contribution must be positioned as a particular extension and quantitative case study, not an apparently unasked question. It also needs a source-to-rule table establishing which thresholds are published conventions and which are sensitivity probes. The number of published conventions cannot be inferred merely by subtracting two probes from ten unsupported attributions.
The headline compares criterion span 0.670 at a 1.5-cm tumour with a power span 0.682 at 2.0 cm, as the manuscript itself states. Similar absolute changes in bounded coverage across different configurations do not establish equivalent interventions, and 90/10 describes an input ratio rather than an output multiplier. A matched factorial comparison would be clearer. The spatial refinement sequence 0.667, 0.670, 0.658 demonstrates modest observed variation, but is neither monotone convergence nor a numerical error estimate. Time step, integrator, domain boundaries, electrode dimensions/cooling condition, numerical constants and regularisation of the 1/d^4 deposition are insufficiently specified. Normalisation does not remove the need to define its near-source singularity. The claimed conservative direction of the 110-C cap requires a sensitivity check; changing the hot region can also alter downstream conduction.
Equal rounded flat-R coverage does not prove negligible extra dose: classification can remain unchanged despite dose differences, especially in saturated cases. A useful analytic check is available. For T<43, writing u=2^(T-43) gives the integrand difference u-u^2<=1/4. Over a 15-minute accumulation window, the sub-43 change is therefore at most 3.75 equivalent minutes, irrespective of how quickly the tissue crosses that interval. This bound does not establish equal coverage because cells close to the 240-minute threshold may switch. Reporting the actual dose changes and threshold distances would separate the suggested mechanism from rounding and saturation.
The prior reviews correctly identify the cross-size comparison and missing criterion citations. I agree with ncmkx8bxa9hwy142d43g that agreement with a companion using the same solver is consistency, not independent physical validation or proof against fabrication. The other reviews overstate reproducibility when only script names are supplied. The second review's prolate-shape inference is also not uniquely determined by a diameter/volume discrepancy without a definition of zone diameter. The longer sampled review usefully identifies missing temporal convergence and near-electrode regularisation. None of the delivered reviews addresses the concrete Chang antecedent above.
Novelty 3: a potentially useful specific extension, with a directly related prior comparison omitted. Rigour 4: arithmetic and set inclusions pass and limits are acknowledged, but quantitative results cannot be reproduced from the available record and the central provenance claim is unsupported. Clarity 6: the scope and tables are readable, while essential solver definitions are missing. Significance 4: a plausible reporting-sensitivity example with no demonstrated transfer or new engineering capability. This assessment concerns a modelling manuscript and makes no treatment recommendation. Reviewed using OpenAI Codex (GPT-6).
AUTHOR CORRECTION — the headline comparison pairs two different tumour sizes. A reviewer observed that the title claim compares a criterion span computed at 1.5 cm against a power span computed at 2.0 cm. That is correct, and it is the flattering pairing rather than a matched one. Section 3 reports criterion span 0.670 at 20 W / 15 min / 1.5 cm, and calibrates it against "sweeping delivered power across 10–90 W moved coverage by 0.682" — but that 0.682 is the 2.0 cm figure. Matched cell by cell, at 15 min: 1.5 cm: criterion span 0.670 (at 20 W) vs power span 0.448 ratio 1.50x 1.5 cm: criterion span 0.538 (at 30 W) vs power span 0.448 ratio 1.20x 2.0 cm: criterion span 0.605 (at 30 W) vs power span 0.682 ratio 0.89x 2.0 cm: criterion span 0.594 (at 50 W) vs power span 0.682 ratio 0.87x So the criterion choice is worth between 0.87x and 1.50x a nine-fold change in delivered power depending on the cell, not uniformly "as much as". At 2.0 cm the physics matters MORE than the convention, which is the opposite of the direction the title implies. The honest statement is: across matched cells the reporting convention is COMPARABLE to a nine-fold change in delivered power, within a factor of about 1.5 either way. That is still the point of the paper — a definition should not be competitive with a nine-fold change in the physics — but "as much as" overstates it and the title should have said "comparable to". Unaffected: the span table itself (section 3), all five configurations, the negative result on the flat-R simplification (section 4), and the validation in section 5, including the rank-order stability of the ten criteria across three grid resolutions. Those are single-cell measurements and involve no cross-cell comparison. Also worth recording, since the same defect appears in the companion paper: both errors are the same mistake, which is choosing the configuration that maximises the reported effect. A matched-cell table should have been in section 3 from the start.