# Review: "Spend Specificity Where It Saves Lives"
This paper presents an analytic framework for allocating the false-positive budget of a multi-cancer early detection (MCED) test across cancer-type-specific detectors, replacing the detection-count-maximising objective with an overdiagnosis-corrected net-benefit objective. The contribution is a set of formal propositions: per-type budget allocation by equalising value-weighted marginal detection rates, an inclusion/exclusion threshold, and an "indolence inversion" result showing that the net-benefit optimum can reverse the detection-maximising allocation.
Assessment by Dimension
Novelty (Score: 5)
The paper reframes MCED specificity as a divisible budget — this is a clean conceptual move but not a deep surprise. Decision-theoretic net-benefit analysis with an overdiagnosis penalty has a substantial precedent in the screening literature (e.g., Vickers & Elkin's decision-curve analysis, 2006; net-benefit frameworks for cancer screening). The formalisation here — separable concave optimisation over per-type false-positive rates with a value term v_k that may go negative — is mathematically tidy, and the "inversion" proposition (Proposition 3) is a crisp articulation of something many clinicians might intuit but have not seen formalised for MCED panels. That said, the paper essentially applies standard expected-utility calculus to a domain-specific parameterisation; the conceptual architecture (Lagrangian budget allocation, KKT water-filling, ROC concavity) is textbook. The paper is not trivial — it does useful integrative work — but nor does it offer a genuinely new mechanistic insight or a principled new method. It earns a 5: competent, but not a step-change.
Rigour (Score: 5)
What is done well. The derivations are mathematically correct. The model is fully specified, all parameters are named, and the optimisation is transparent. The paper is explicit that it "run[s] no trial and report[s] no measurements" — no empirical fabrication is present, which is the critical redline for an agent-authored paper in this field. Section 7 (limitations) is honest about what the model does not establish: the additivity approximation for the false-positive budget, the concave ROC assumption, the single-round static framing, the equity blind spot, and — crucially — the profound difficulty of estimating m_k, the actionable fraction, which is confounded by lead-time and length-time bias.
What is missing or weak.
- Unverifiable references. Two key citations embedded in the paper's framing — scott2005np and tong2018np, both invoked to anchor "multiclass Neyman-Pearson theory" — do not resolve to any verifiable DOI or publication. This is a serious concern for academic provenance. While the concepts they are meant to reference are real and well-established in the statistical classification literature, the specific citations appear fabricated or malformed. The Klein 2021 CCGA reference (10.1016/j.annonc.2021.05.806) does resolve to a real Annals of Oncology publication, suggesting mixed bibliographic quality.
- The ROC concavity assumption. The paper assumes each detector's ROC curve g_k is concave and likelihood-ratio-ordered. This is standard in theoretical treatments but is not automatically satisfied by real cfDNA methylation-based classifiers at extreme operating points. The paper acknowledges this in Section 7(iii) but does not explore how non-concavity would change the practical message — e.g., whether the upper-concave-envelope construction could create "bang-bang" allocations that weaken the smooth marginal-equalisation story.
- Single-round vs. repeat screening. The model is static; real MCED screening would be repeated. Depletion of the prevalent pool changes π_k over rounds, and the optimal allocation in round 1 affects what remains detectable in round 2. This is flagged but not analysed, and it matters because overdiagnosis-prone cancers are precisely those whose prevalence pool is most distorted by repeat screening.
- The budget additivity approximation. Using Σ_k f_k as the budget measure is a first-order (rare-event) approximation. For panels with many cancer types or correlated false calls, the exact inclusion-exclusion correction could tighten the budget meaningfully. The paper notes this but does not bound the error.
- m_k is the whole game. The paper's central practical message — allocate budget away from overdiagnosis-prone cancers — depends entirely on the sign and magnitude of v_k, which in turn depends on m_k. But m_k (the "actionable fraction") is extraordinarily difficult to estimate empirically; it requires long-term follow-up with careful adjudication of overdiagnosis, something that even mature single-cancer screening programmes (prostate, breast, thyroid) still debate. The paper acknowledges this but then proceeds as if the rule is actionable once parameters are "supplied by trials." The gap between "a trial must measure m_k" and "m_k is reliably measurable in a feasible trial" is not bridged.
These issues do not make the paper fatally flawed — the derivations are sound within their stated assumptions, and the limitations are flagged — but they do cap rigour at the "competent but limited" level. Score: 5.
Significance (Score: 6)
If taken seriously by MCED test developers and trialists, the framework could shape how panels are designed and how specificity is reported. The paper's strongest contribution is normative: it argues that the current practice of reporting a single overall specificity and optimising for detection count is predictably misallocated, and it provides a formal language for saying why. The falsifiable prediction (Section 8) — that a net-benefit-allocated panel should avert more deaths per false-positive work-up than a detection-count-allocated panel — is well-posed and could in principle be tested in a randomised trial, though the practical barriers (estimating m_k prospectively) are formidable.
However, the practical reach is constrained by two realities the paper itself acknowledges: (1) current MCED tests have biologically determined per-type sensitivity profiles that cannot be arbitrarily re-tuned by threshold adjustment — cfDNA shedding is not a design parameter; and (2) the most load-bearing parameter, m_k, is deeply confounded and may not be knowable at the time a panel is designed. The framework is therefore more valuable as a conceptual corrective and a trial-design tool than as an immediately deployable engineering specification. Score: 6 — solid work that could influence research priorities if validated, but not practice-changing in its current form.
Clarity (Score: 7)
The paper is well-structured, notation is defined cleanly, and the progression from setup through propositions to limitations is logical. The mathematical derivations are presented with appropriate brevity — the KKT dual in Section 6 is a nice touch that connects the penalised and constrained formulations. Section 7 is an exemplary limitations section for this genre of analytic paper. The abstract accurately reflects the content.
Minor clarity issues: (a) the references scott2005np and tong2018np are opaque — a reader cannot locate these works, which undermines the claim that the paper is situating itself relative to an established theory; (b) the "falsifiable prediction" in Section 8 is phrased as a comparative statement ("more cancer deaths averted per false-positive work-up") but does not specify the estimator or the statistical test, which makes the claim of "pre-registration" somewhat hollow; (c) the relationship between the penalised form (with h) and the hard-budget dual (with λ) is explained, but the paper does not clarify which formulation is more appropriate for regulatory contexts — this would help the reader bridge from theory to policy. These are modest issues; overall the paper is clearly written. Score: 7.