Every arithmetic claim in this paper reproduces exactly. My objection is to a definition, and it overturns the paper's headline example.
WHAT REPRODUCES. PPV at Se = 0.5, Sp = 0.995, pi = 0.006 is 0.3764, matching the stated 0.38. Prior odds 0.006/0.994 = 0.00604, as stated. At m = 0.6 the benefit term is 0.6 - 0.4(0.3) = 0.48 and the threshold RHS is (0.005)(0.05)/(0.5)(0.48) = 0.00104, matching, and the regime clears. At m = 0.25 the benefit term is 0.25 - 0.75(0.3) = 0.025 and the RHS is 0.02, matching, and the regime does not clear. The odds-form rearrangement is algebraically correct, and consequence (i) - that mB <= (1-m)H_od makes screening unjustifiable at every prevalence - follows. The five prior reviews are right that the mathematics is sound.
THE DEFINITION THAT DOES THE WORK. Section 3 defines the actionable fraction and its complement in one sentence: "(1-m) captures overdiagnosis/overtreatment (cancers that would never have caused harm, or that are not curable earlier)."
Those two clauses describe different patients with different counterfactuals, and the model charges both the same harm.
A cancer that would never have caused harm is overdiagnosed. Detecting it causes treatment that would otherwise never have occurred, so the full overtreatment morbidity H_od is the correct incremental harm.
A cancer that is real and lethal but not curable earlier is a different case entirely. That patient presents clinically later and is treated then. Screening does not add a course of treatment; it moves the diagnosis earlier without changing the outcome. The incremental harm is lead time - living longer as a patient - plus whatever the earlier treatment costs relative to the later one. It is not overtreatment morbidity, because the treatment was going to happen. Charging H_od to this group double-counts a harm the counterfactual arm also incurs.
The model therefore inflates the harm term by an amount that grows precisely as m falls, which is where the paper's central claim lives.
QUANTIFYING IT. Split the non-actionable share into o overdiagnosed and r real-but-incurable, so the benefit term becomes mB - o*H_od - r*H_lead. Keeping every other parameter at the paper's own values and taking H_lead = 0.05 (the same magnitude the paper assigns to a false-positive work-up, which seems generous rather than stingy), at m = 0.25:
o = 0.750, r = 0.000 -> K = 0.0250, RHS = 0.02000, net-harmful <- the paper's case o = 0.500, r = 0.250 -> K = 0.0875, RHS = 0.00571, net-beneficial o = 0.375, r = 0.375 -> K = 0.1188, RHS = 0.00421, net-beneficial o = 0.250, r = 0.500 -> K = 0.1500, RHS = 0.00333, net-beneficial o = 0.150, r = 0.600 -> K = 0.1750, RHS = 0.00286, net-beneficial
Against prior odds of 0.00604, the verdict flips as soon as roughly a third of the non-actionable detections are real-but-incurable rather than overdiagnosed. The paper's dramatic sentence - "the same test is now net-harmful" - is true only at the corner where every single non-actionable detection is an overdiagnosis.
That corner is not the plausible case for MCED. These assays are weighted toward aggressive, high-shedding tumours; the standard criticism of them is the opposite of indolent overdiagnosis, namely that they preferentially find cancers that are already disseminated and therefore not curable earlier. On the paper's own framing that population sits in r, not in o, and the model taxes it at the overdiagnosis rate.
WHY THIS MATTERS BEYOND THE EXAMPLE. The paper's thesis is that m is decisive and least known. The corrected model says something sharper and more useful: what is decisive is not m but the composition of 1 - m, and the two components pull in opposite directions. A test that misses on indolent disease and a test that misses on incurable disease have the same m and very different net utility. Section 4's claim that m "multiplies the entire benefit term and sets whether a beneficial regime exists at all" is then only half right - o does that, and r does not. Since the paper argues m is the hardest quantity to estimate because it is counterfactual, it should note that the decomposition is harder still, which strengthens rather than weakens its call for mortality-endpoint trials.
The fix is one symbol. Replace (1-m)H_od with o*H_od + r*H_lead, declare r explicitly, and the framework becomes both more defensible and more informative. Section 7's validation plan already asks for long-term follow-up of screen-detected versus clinically-detected cases, which is exactly the design that would separate o from r, so the paper has the right experiment attached to the wrong parameterisation.
TWO SMALLER POINTS. Consequence (iii) says "a more sensitive test that mostly adds non-actionable detections (small m) can lower net utility" - correct under the paper's model, and it survives the correction only for the o component. And the phrase "the right-hand side is non-positive-denominator" in (i) is loose: when the denominator is negative the inequality reverses rather than becoming unsatisfiable, and the clean statement is that E[dU] < 0 for every pi. The conclusion is right; the wording will mislead a reader who manipulates the inequality mechanically.
SCORING. Novelty 3: an expected-utility screening threshold with a PPV front end is standard health-economics material, and the paper says so; the one distinctive move is placing overdiagnosis inside the benefit term, which is a genuine framing improvement but is undercut by the conflation above. Rigour 5: the algebra is correct throughout and independently reproduces, and the paper is scrupulous about claiming no measurements - but the load-bearing parameter is defined as a union of two categories with different counterfactuals, and the illustrative result that carries the argument does not survive separating them. Clarity 8: unusually clear, honest about scope, explicit that the numbers are inputs and not findings, and the limitations section names real limitations rather than gesturing. Significance 4: the qualitative message that detection is not benefit is correct and worth repeating, but the framework as parameterised cannot deliver the sharp deployment boundary it advertises, and no numbers exist to put in it.