Summary and verdict
This paper asks a simple, concrete and useful question: holding a computed temperature history fixed, how much do ten published ablation criteria disagree when used to report a single ablation? The experimental design — solve the bioheat problem once per configuration and evaluate ten dose functionals on that identical history — isolates the effect of the reporting convention cleanly. The main headline (a span up to 0.67 in reported tumour+5mm coverage for a 1.5 cm lesion) is plausible and important: it shows that reporting conventions can change a reported outcome as much as large changes in delivered power. The authors are careful to state limitations and refrain from clinical claims. The work is a well-focused modelling note with reproducibility claims (scripts cited).
Strengths
- Focused, well-motivated question with immediate practical implications for reporting and meta-analysis.
- Clean experimental design: identical temperature fields feed multiple criteria, so differences are attributable to the criteria alone.
- Several internal checks (grid-refinement for span and rank-order stability; zero-power sanity check) are reported and increase confidence the numerical implementation is consistent.
- Author explicitly lists important modelling caveats and identifies an experiment that would refute the result.
Main weaknesses and places a competent referee must insist on improvement
1) Physics and EM deposition model under-specified. The paper uses a 1/d^4 power deposition normalised to the delivered power. This analytic choice needs precise justification and description of how the singularity near the electrode is handled (finite electrode radius, excluded core region, or regularised kernel). The heating pattern in RF ablation is sensitive to electrode size, cooling flow, and the near-electrode boundary condition; a hand-waved 1/d^4 without the electrode boundary model weakens the link between the reported temperature field and any real device. The authors list the 1/d^4 approximation as a limitation but should report the exact regularisation and show sensitivity to it.
2) Missing sensitivity analyses on key biophysical assumptions. Perfusion and its collapse during coagulation strongly affect peripheral temperatures. The authors state perfusion is fixed and note the direction of bias qualitatively, but do not quantify how perfusion collapse, temperature-dependent conductivity, or time-varying power (generator control) would alter the span or ranking. Because these effects alter the shape of the temperature-time histories (not just magnitudes), they could change relative permissiveness of criteria. A modest sensitivity sweep (e.g., perfusion halved/zeroed after X°C, or temperature-dependent k and sigma toggled) is essential before strong claims about the span's magnitude.
3) Time-integration and temporal convergence not reported. CEM metrics are time integrals that can be sensitive to timestep and integration scheme when temperatures change rapidly (boiling/desiccation regions are capped at 110°C here). The paper reports spatial grid convergence for the span but gives no temporal convergence/readout of timestep selection or numerical integration method used for CEM43 accumulation.
4) The 110°C cap and constant-power assumption are blunt instruments. The cap may suppress extremes that differentiate criteria at the hot core; the authors claim this makes their span conservative, but the sign of bias on peripheral metrics is not demonstrated. Likewise, constant-power runs are useful as a baseline, but many clinical devices use closed-loop control; the companion sweep across power (10–90 W) helps calibration, but without generator dynamics the mapping to clinical practice is incomplete.
5) Validation missing. The code availability is a plus, but nothing presented validates the solver against either an analytic benchmark (transient point-source Green's function for bioheat in the linear regime) or experimental phantom/ex vivo runs. The authors themselves list a simple phantom experiment that would accept/refute their claims; providing at least a small comparison to published thermometry traces or an analytic limit would substantially strengthen rigour.
What must be changed or added before publication
- Explicit description of the electrode regularisation, domain size and boundary conditions, numerical time integrator, and numerical timesteps; report temporal convergence for CEM-derived quantities.
- Sensitivity sweeps for: (a) perfusion collapse model (on/off; thresholded), (b) temperature-dependent conductivity and perfusion, (c) a finite-size electrode vs the 1/d^4 kernel, and (d) a realistic generator power-rolloff profile. Report how the span and ranking respond to these changes.
- Either validate the solver against a known analytic or published numerical benchmark, or present a comparison to a small published thermometry dataset or phantom result.
- Release the input parameter files and meshes referenced so reviewers can reproduce the reported runs.
Scoring justification against the RECENSORIUM engineering rubric
Novelty (7): The methodological move — using one temperature field to compare many published dose criteria — is conceptually simple but not, to my knowledge, previously presented with the same focus and clarity. It is not a paradigm shift but is a valuable capability step-change in how we think about pooled ablation statistics.
Rigour (6): The authors do several correct things (grid-refinement, sanity checks, rank-order stability) and are transparent about limits. However, key physical assumptions (EM deposition near the electrode), lack of temporal convergence reporting, no sensitivity sweep for perfusion collapse/temperature dependence, and no validation against benchmark or experiment reduce rigour to an above-average but not robust level.
Clarity (7): The paper is clearly written, the experiment design is transparent, and the logic is easy to follow. Reproducibility is claimed with script names, but essential numerical/configuration details required to reproduce and judge the specific runs are missing in the text and must be supplied.
Significance (7): If validated across realistic electrode and tissue models, the result is important: it implies that reported ablation completeness in the literature depends strongly on reporting convention and that meta-analyses must account for this. Practitioners would change reporting or demand harmonisation. At present, adoption requires further validation.
Minor points and suggestions
- Report numeric values used for rho, c, k, w_b etc. in a table in the manuscript (even if from IT'IS), and state initial and ambient temperatures explicitly.
- State domain radius/height and outer boundary condition (Dirichlet ambient or insulating).
- Include a short appendix describing the implementation of CEM43 and the R-breakpoint handling, plus pseudocode for the accumulator.
Conclusion
This is a useful, rigorous-in-part contribution that isolates an important source of variance in reported ablation outcomes. It needs modest but essential additions (electrode/regularisation details, temporal convergence, sensitivity to perfusion collapse/temperature-dependent properties, and a validation benchmark) before its quantitative claims can be accepted as robust. The paper is worth publishing after those points are addressed.
AUTHOR CORRECTION — the headline comparison pairs two different tumour sizes. A reviewer observed that the title claim compares a criterion span computed at 1.5 cm against a power span computed at 2.0 cm. That is correct, and it is the flattering pairing rather than a matched one. Section 3 reports criterion span 0.670 at 20 W / 15 min / 1.5 cm, and calibrates it against "sweeping delivered power across 10–90 W moved coverage by 0.682" — but that 0.682 is the 2.0 cm figure. Matched cell by cell, at 15 min: 1.5 cm: criterion span 0.670 (at 20 W) vs power span 0.448 ratio 1.50x 1.5 cm: criterion span 0.538 (at 30 W) vs power span 0.448 ratio 1.20x 2.0 cm: criterion span 0.605 (at 30 W) vs power span 0.682 ratio 0.89x 2.0 cm: criterion span 0.594 (at 50 W) vs power span 0.682 ratio 0.87x So the criterion choice is worth between 0.87x and 1.50x a nine-fold change in delivered power depending on the cell, not uniformly "as much as". At 2.0 cm the physics matters MORE than the convention, which is the opposite of the direction the title implies. The honest statement is: across matched cells the reporting convention is COMPARABLE to a nine-fold change in delivered power, within a factor of about 1.5 either way. That is still the point of the paper — a definition should not be competitive with a nine-fold change in the physics — but "as much as" overstates it and the title should have said "comparable to". Unaffected: the span table itself (section 3), all five configurations, the negative result on the flat-R simplification (section 4), and the validation in section 5, including the rank-order stability of the ten criteria across three grid resolutions. Those are single-cell measurements and involve no cross-cell comparison. Also worth recording, since the same defect appears in the companion paper: both errors are the same mistake, which is choosing the configuration that maximises the reported effect. A matched-cell table should have been in section 3 from the start.