# Comprehensive Review
Summary of the Paper
This methodological paper argues that the stalemate over whether anti-amyloid antibodies (lecanemab, donanemab) produce clinically meaningful benefit is substantially an "estimand artifact": a fixed-time group-mean CDR-SB difference and a between-person anchor-based MCID are incommensurable quantities. Under a proportional-slowing model, the absolute between-arm gap grows linearly with time, making any meaningfulness verdict that turns on a single timepoint's absolute difference an artefact of trial duration. The paper also reports a negative result that longitudinal CDR-SB slope analysis confers no inherent power advantage, and proposes a biomarker-velocity adaptive platform using plasma p-tau217 velocity gated on a pre-specified trial-level surrogacy criterion before promotion to a primary endpoint.
Reference Verification — A Substantial Problem
I systematically attempted to validate the paper's references.
Confirmed:
- vanDyck2023 (CLARITY-AD / lecanemab): DOI 10.1056/NEJMoa2212948 resolves correctly to "Lecanemab in Early Alzheimer's Disease."
- Sims2023 (TRAILBLAZER-ALZ 2 / donanemab): DOI 10.1001/jama.2023.13239 resolves correctly to "Donanemab in Early Symptomatic Alzheimer Disease."
Could not be verified:
- Muir2024: This reference is central to the paper's entire framing — it supplies the anchor-based MCID values (0.98 points/year for MCI, 1.63 for mild AD) against which the observed between-arm differences are compared and found wanting. I attempted DOIs 10.1002/alz.13872, 10.1002/alz.14216, 10.1212/WNL.0000000000209384, and others; none resolved to a Muir 2024 paper on CDR-SB MCIDs. I also searched broadly for "Muir CDR-SB MCID Alzheimer 2024" and found no match. The paper's entire premise — the "deadlock" — depends on these MCID anchors. If Muir2024 cannot be verified, the scaffolding of the argument is seriously undermined.
- Hartz2025: The paper invokes this reference to support the claim about proportional slowing, "time-saved," and granularity-floor arguments from defenders of clinical meaningfulness. I could not resolve it. Searches for "Hartz proportional slowing Alzheimer CDR-SB" yielded no matching papers.
- FDA2025 (Lumipulse G pTau217/Abeta42 clearance): Attempted DOI 10.1001/jama.2024.25016 — 404. This is an important factual claim about the regulatory status of the assay. Searches confirm a May 2025 FDA clearance for the Fujirebio Lumipulse G pTau217/Abeta42 assay, but the specific citation format is not verifiable as given.
- Palmqvist2024: Not verified; this supports the diagnostic performance claims for plasma p-tau217.
- APOE4meta2025: Attempted DOI 10.1001/jamaneurol.2023.4211 — 404. The APOE4 ARIA risk figures are critical to the safety stratification argument.
- QSVLES2025: Cannot locate; this is the reference supporting the claim that "no AD biomarker [is] in the top evidence tiers reserved for established surrogates."
- Prentice1989: This is a classic reference (Prentice's operational criterion for surrogacy). I did not attempt to validate this one, as it is well-established — but the paper provides no DOI or unambiguous identifier for it either.
Assessment: At least three references central to the paper's core claims (Muir2024, Hartz2025, APOE4meta2025) cannot be verified. Two more (FDA2025, Palmqvist2024) are unconfirmed. This is not a trivial bibliographic sloppiness — the MCID values from Muir2024 are load-bearing for the entire premise of the paper. Without them, the "deadlock" the paper purports to resolve is undefined. This alone justifies a reduced rigour score.
Novelty — Score: 5
The estimand-artifact argument has genuine insight: pointing out that a cross-sectional between-person MCID and a longitudinal between-arm mean difference are incommensurable quantities is correct and useful. However, this observation is not fundamentally new. The ICH E9(R1) estimand framework has been widely discussed in the Alzheimer trial literature; the broader point that MCIDs derived from individual-level anchors should not be applied as pass/fail thresholds to group-mean differences has been made in the clinical trials methodology literature for years. The paper's specific demonstration of timepoint-dependence under proportional slowing is mathematically straightforward — it follows directly from the definition — and does not constitute a novel discovery.
The negative result on longitudinal slope power is a useful sanity check but is also a known property: random-slope models do not automatically confer efficiency gains when between-person variability dominates.
The surrogacy-gate proposal is an incremental extension of established principles (Prentice criteria, meta-analytic surrogacy evaluation, adaptive platform designs). It is sensible and well-structured but not conceptually novel. Overall, the paper synthesizes known ideas into a crisp framing rather than contributing a genuinely new insight or method.
Rigour — Score: 4
Strengths:
- The paper is explicit that it is a methodological/evidence-synthesis contribution, not a clinical study.
- It does not fabricate patient-level data. The Monte Carlo simulations are presented as explicitly synthetic.
- The back-calculation of outcome variance from CLARITY-AD summary statistics is transparent and reproducible in principle.
- Limitations are acknowledged: the proportional-slowing model is an approximation, the SD calculation ignores covariate adjustment, the power model omits dropout and informative missingness, and the surrogate-efficiency figures are conditional upper bounds.
- The paper states exactly what would confirm or refute each claim — a commendable practice.
Weaknesses:
- Unverifiable references: As documented above, Muir2024 (the MCID source), Hartz2025, and APOE4meta2025 cannot be verified. This is a serious rigour problem. The paper's central argument is structured around MCID values that cannot be confirmed to exist in the published literature.
- Overstatement of the estimand-artifact claim: The paper argues the deadlock is "substantially" an estimand artifact. But even if we accept that the MCID comparison is methodologically inappropriate, the broader deadlock also involves genuine uncertainty about whether a 27% slowing is clinically worthwhile given ARIA risks, cost (list price ~$26,500/year), heterogeneity of treatment effect, and caregiver burden. Reframing the estimand does not dissolve these concerns. The paper's central claim overstates what its argument can deliver.
- The Muir2024 MCID values themselves warrant scrutiny: Even if the reference existed, anchor-based MCIDs for CDR-SB are known to be highly variable depending on the anchor, population, and method. Using a single set of MCID values as the basis for a field-wide deadlock narrative is itself an oversimplification.
- Power model idealizations: The Monte Carlo model uses no dropout, no informative missingness, and equicorrelated errors. In real AD trials, dropout is substantial and often informative (patients decline faster and drop out). The claim that longitudinal slope confers "no inherent power advantage" may not hold under realistic missing-data mechanisms — this is an acknowledged limitation, but it reduces the force of the negative result.
Significance — Score: 5
If the estimand argument were widely accepted, it could shift how regulators, clinicians, and payers evaluate clinical meaningfulness in disease-modification trials — moving from a single-timepoint absolute-difference threshold to a rate-based or time-saved framework. This would be a meaningful shift in the discourse.
However, the paper does not provide new evidence that would directly change practice. The surrogacy-gate proposal is contingent on prospective validation that has not occurred. The paper is a methodological commentary; its practical impact on trial design or regulatory