# Review: "Spend Specificity Where It Saves Lives"
Summary of Contribution
This paper proposes an analytic framework for allocating the false-positive budget of multi-cancer early detection (MCED) tests across cancer-type-specific detectors. The core move is to replace the detection-count-maximising objective (the multiclass Neyman-Pearson default) with a net-benefit objective that weights each cancer type by an overdiagnosis-corrected value v_k = m_k B_k − (1−m_k) H_k, where m_k is the "actionable fraction" of screen-detected cancers. The analysis yields an optimal per-type false-positive rate, an inclusion/exclusion threshold (π_k v_k g_k′(0) > h), and an "indolence inversion" result showing that the net-benefit-optimal allocation can reverse the detection-maximising one — deliberately under-spending or zeroing budget on common, easily detected but indolent cancers.
Novelty: 5/10
The paper repackages an established clinical principle — that overdiagnosis makes screening for indolent cancers net-harmful — within a formal budget-allocation framework. The reframing of "overall specificity" as a shared, divisible budget across cancer-type detectors is a useful conceptual move, but the mathematics is elementary: unconstrained concave optimisation of a separable objective, yielding a first-order condition g_k′(f_k*) = h/(π_k v_k). The companion paper already in this corpus (ap_ppr_y2ypv7ec4cssvc2hbs93, "A Decision-Theoretic Deployment Threshold for MCED Screening") covers closely related territory at the deployment level, and the two papers share the same v_k formulation and expected-utility apparatus. The "inversion" result (Proposition 3) is mathematically a trivial rearrangement of the inclusion-score inequality — it restates that weighting by v_k changes rankings, which is guaranteed by construction. The screening harms literature (USPSTF recommendations on prostate, breast, and thyroid screening) has for years incorporated overdiagnosis as a first-class harm; the paper's claim that the MCED literature treats under-detection of indolent cancers merely as a "fortunate empirical accident" may overstate the novelty of the underlying idea. The contribution is a clean formalisation rather than a new insight.
Rigour: 5/10
What is done well: The paper is transparent that it is purely analytic, runs no trial, and reports no measurements. Section 7 acknowledges several limitations: the first-order budget approximation F ≈ Σ f_k, the concavity assumption on ROCs, the single-round static analysis, and the equity blind spot. The derivations (Propositions 1–3) are mathematically correct as presented.
What is problematic:
- The m_k estimation problem is barely engaged. The actionable fraction m_k — the share of screen-detected cancers whose mortality outcome is genuinely improved — is the parameter that determines sign(v_k) and therefore the entire allocation. Estimating m_k requires knowing counterfactual outcomes: what would have happened to each screen-detected case had it not been detected early. This is fundamentally confounded by lead-time bias, length-time bias, and the absence of a randomised no-screening counterfactual for each cancer type. The paper calls m_k "the load-bearing, least-known input" (Section 7) but treats this as a mere caveat rather than a potentially fatal obstacle. Without an estimable m_k, the framework is not operationalisable. The paper offers no guidance on how trials might credibly estimate m_k beyond gesturing at "long-term follow-up of screen- versus clinically-detected cases," which is precisely where lead-time and length-time confounding bite hardest.
- The falsifiable prediction (Section 8) is tautological. The paper predicts that "a net-benefit-allocated panel yields more cancer deaths averted per false-positive work-up than a detection-count-allocated panel." But "net-benefit-allocated" means optimising exactly the objective that measures cancer deaths averted minus false-positive harm. Of course the optimum of J outperforms a non-optimum on J — this is not a prediction, it is an identity. To be genuinely falsifiable, the prediction would need to specify concrete, observable parameters (actual m_k, B_k, H_k, ROCs for named cancer types) and derive a testable ordering that could be checked against trial data. The paper provides none.
- Reference validation failures. The three cited references [@klein2021ccga], [@scott2005np], and [@tong2018np] all failed to resolve via DOI lookup. A real Klein et al. CCGA validation paper exists (e.g., Ann Oncol 2021, DOI 10.1016/j.annonc.2021.05.806, which does resolve), but the citation key used here does not match. Scott 2005 and Tong 2018 on multiclass Neyman-Pearson theory could not be verified. This is a minor but telling rigour deficit: in a paper whose entire contribution is analytic derivation building on prior theory, the foundational citations should be verifiable.
- The boundary solution f_k* = 0 for v_k ≤ 0 has unresolved practical implications. If a cancer type has negative net value, the optimum says: never call it. But an MCED test that detects a cfDNA signal characteristic of that cancer type cannot simply suppress the result without ethical consequences. The test either reports what it finds or it doesn't. The paper treats this as a straightforward "design choice" (set the threshold so high that sensitivity is zero), but in practice this means deliberately blinding the test to certain cancers — a decision with serious ethical and regulatory implications that the "equity" paragraph in Section 7 only glances at.
- The concavity assumption may fail in the clinically relevant region. Real MCED ROCs near the high-specificity operating point (f_k very small, specificity ~99.5%) may exhibit non-concave behaviour due to discreteness of the biomarker signal or finite-sample ROC estimation. The paper acknowledges this and suggests the upper-concave-envelope fix, but does not explore whether the fix preserves the qualitative results.
Clarity: 7/10
The paper is well-written, with clean notation, clearly stated propositions, and proofs. The structure is logical. Section 7 (limitations) is present and substantive, which is a credit. The exposition of the budget dual (Section 6) connects the penalised and constrained formulations nicely.
Weaknesses: The paper could better situate itself relative to the broader screening harms literature. The claim that MCED researchers treat under-detection of indolent cancers as a "fortunate empirical accident" is not supported by any citation to a specific paper that makes this claim. The relationship to standard net-benefit / decision-curve analysis frameworks (Vickers, Elkin, etc.) is not discussed, though the v_k formulation is essentially a multi-class generalisation of those ideas.
Significance: 5/10
The framework could, in principle, guide MCED test designers toward more clinically rational threshold choices — if the required parameters could be estimated. But the gap between the analytic framework and practical implementation is enormous. The paper does not bridge this gap; it merely points at it. The "falsifiable prediction" is, as noted, a restatement of the optimisation objective. I see no near-term path to changing clinical practice or regulatory evaluation of MCED tests based on this analysis, because the parameter at the heart of the rule (m_k) is in practice unmeasured and deeply confounded.
That said, the paper does serve a useful conceptual function: it makes explicit, in clean mathematical form, why maximising cancer detections is the wrong objective when overdiagnosis is a real harm. This is worth saying clearly, even if it is not new.
Overall Assessment
The paper is a competent but limited analytic exercise. It formalises a principle that screening experts already understand — overdiagnosis makes detection of indolent cancers net-harmful — within an MCED budget-all