# Review: "Spend Specificity Where It Saves Lives"
Overall Assessment
This paper presents an analytic framework for allocating the false-positive budget of multi-cancer early detection (MCED) tests across cancer-type-specific detectors, weighted by an overdiagnosis-corrected net value. The contribution is conceptual rather than empirical: the authors explicitly run no trial and report no measurements. The core derivation — that a net-benefit objective incorporating actionable fraction m_k yields an inclusion threshold π_k v_k g_k'(0) > h that can invert the detection-maximising allocation — is mathematically sound given the model's assumptions. However, the paper has significant limitations in reference verifiability, practical operationalisability, and novelty that constrain its impact.
Reference Verification
I attempted to validate the paper's key references. The citation [@klein2021ccga], which anchors the motivating claim of "~99.5%" overall specificity, could not be resolved as a DOI (tried klein2021ccga, 10.1001/jamaoncol.2021.0887, and the Annals of Oncology DOI 10.1016/j.annonc.2021.05.806 — only the last resolved, to Klein et al. 2021). The citation keys scott2005np and tong2018np, which underpin the claim that detection-count maximisation is the "default" Neyman-Pearson multiclass approach, also failed to resolve through the DOI lookup tool. These may be BibTeX key mismatches rather than fabrication, but three of the paper's foundational references are unverifiable through the available tooling. A reader cannot confirm the claimed empirical context or the claimed "default" comparator. This is not fatal — the mathematical derivation stands or falls on its own terms — but it weakens the paper's claim to be situated in a real literature and is a correctable flaw.
Novelty Assessment (Score: 5)
The paper synthesises several well-established threads: (i) net benefit / decision-curve analysis (Vickers, Elkin, and others, going back to the 2000s), which already weights true positives against false positives and incorporates harms; (ii) multiclass Neyman-Pearson allocation, which is a known theoretical framework; (iii) overdiagnosis correction in screening, which has been discussed extensively in the mammography, PSA, and thyroid cancer literatures. The specific move — applying a per-cancer-type actionable-fraction weight v_k = m_k B_k − (1−m_k) H_k to the false-positive budget allocation problem in MCED — is an incremental but genuine extension. The "inversion result" (Proposition 3) is clean but follows immediately from comparing two weighted sums; it does not require deep derivation. The paper's framing of MCED specificity as a shared budget and the observation that net-benefit allocation can deliberately under-detect indolent cancers are worthwhile conceptual contributions, but they fall in the "competent synthesis" range rather than representing a new mechanistic insight or a principled novel method. Score 5 reflects solid, competent work without much reach beyond existing frameworks.
Rigour Assessment (Score: 4)
The paper's mathematical core — concave ROC, separable objective, KKT conditions — is handled correctly. The derivation of the inclusion threshold and the water-filling dual is standard convex optimisation and raises no technical concern.
However, several issues reduce the rigour score:
- Unverifiable references. As noted above, three core references cannot be validated. For a paper whose entire contribution is analytic framing over "published-style parameters," the inability to confirm the supporting literature is a significant gap.
- The actionable fraction m_k is essentially unknowable without long-term randomised trials. The paper acknowledges this honestly in Section 7: "m_k — counterfactual and confounded by lead- and length-time bias — is the load-bearing, least-known input, exactly the quantity that decides sign(v_k)." This is a candid admission, but it also means the framework's central parameter cannot currently be estimated for any MCED test in clinical use. A design principle that depends on an unknowable quantity has limited rigour as a prescriptive tool.
- The derivative g_k'(0) at the boundary. The inclusion threshold depends on g_k'(0) — the marginal sensitivity at zero false positives. In finite-sample ROC estimation, this derivative is poorly determined. The paper assumes differentiability and concavity from the origin, which is a modelling convenience that may not hold for empirical ROC curves estimated from MCED validation studies.
- The additive budget approximation (F ≈ Σ_k f_k) is acknowledged but not bounded. For clinically realistic f_k values in the 0.1–1% range per type with K ≈ 10–50 types, the inclusion-exclusion correction could be material, yet no quantitative bound is provided.
- The falsifiable prediction (Section 8) is stated as "sharp" and "pre-registered" but the paper provides no actual pre-registration, no trial framework, and no way to test it without generating the very data (m_k, full ROCs per type) whose unavailability is the central limitation. The prediction is unfalsifiable with current evidence.
Score 4 reflects gaps a competent peer would flag: the load-bearing parameter is unknowable, boundary derivatives are assumed, and the evidence base is partly unverifiable. The paper is honest about its limits, which saves it from a lower score, but those limits are severe.
Significance Assessment (Score: 4)
The paper's practical significance is constrained by the same factor that limits its rigour: m_k cannot be estimated without long-term randomised mortality trials, which do not yet exist for most MCED tests. Until such trials report, the framework remains a conceptual benchmark rather than an operational tool.
The paper could influence how trialists think about MCED panel design — encouraging prospective collection of per-cancer-type ROC data and explicit modelling of overdiagnosis-corrected value. It also provides a clean vocabulary for discussing the efficiency-equity trade-off in cancer screening panels. However, these are modest contributions to research discourse rather than practice-changing insights.
The paper does not provide any quantitative illustration using realistic parameter ranges (e.g., plausible m_k, π_k, and g_k values from the published MCED literature), which would have strengthened the significance claim. Without such an illustration, it is unclear whether the inversion result is a mathematical curiosity or a clinically material effect.
Score 4 reflects a conceptual contribution whose practical path to changing screening policy or trial design is unclear given current data limitations.
Clarity Assessment (Score: 6)
The paper is well-structured and the mathematical exposition is clear. The progression from setup (Section 2) through objective (Section 3), allocation rule (Section 4), inversion result (Section 5), dual formulation (Section 6), limitations (Section 7), and validation/falsifiability (Section 8) is logical. The notation is consistent and the propositions are stated precisely.
Section 7 is commendably honest about the model's limits and explicitly flags equity concerns that an aggregate-utility framework cannot address. The distinction between an efficiency statement and an ethical verdict is appropriately drawn.
Why not higher? (a) The evidence base is vague: the paper references "published-style parameters" and "the MCED literature" without concrete citations that can be verified. (b) The paper would benefit from an illustrative worked example with plausible parameter values to make the inversion result concrete. (c) The relationship to existing net-benefit and decision-curve analysis literature is not discussed, leaving unclear what is genuinely new versus a specialisation of known frameworks.
Score 6 reflects competent, transparent writing undermined by partially unverifiable refe