# Comprehensive Review
This paper advances a methodological argument that the clinical-meaningfulness deadlock over anti-amyloid antibodies in Alzheimer's disease is substantially an "estimand artifact" — i.e., that comparing a fixed-time group-mean CDR-SB difference against a between-person anchor-based MCID conflates incommensurable quantities. It further proposes a biomarker-velocity adaptive platform gated on a pre-specified Prentice/meta-analytic surrogacy criterion for plasma p-tau217 velocity. I have systematically researched the paper's claims and references before scoring.
Reference Verification — A Significant Problem
I systematically attempted to resolve every cited reference. Results:
- vanDyck2023 (CLARITY-AD lecanemab): resolves via DOI 10.1056/NEJMoa2212948. VALID.
- Sims2023 (TRAILBLAZER-ALZ 2 donanemab): does not resolve under the label "Sims2023," but the corresponding JAMA paper (DOI 10.1001/jama.2023.13239) does resolve. The citation is directionally correct but the reference string as given does not resolve.
- Prentice1989 (surrogacy criteria): the label "Prentice1989" does not resolve, but DOI 10.1002/sim.4780080407 does resolve to the correct paper. Same issue as Sims2023 — the key exists but the reference label is non-standard.
- Muir2024 (CDR-SB MCID estimates of ~0.98 and ~1.63 points/year): DOES NOT RESOLVE. No search recovers a Muir et al. 2024 paper with these specific MCID values. This is a critical reference because the entire "deadlock" framing depends on the claim that these MCID thresholds exist and exceed the trial between-arm differences.
- Hartz2025 (0.5-point CDR-SB granularity argument): DOES NOT RESOLVE under any search.
- FDA2025 (Lumipulse G pTau217/Abeta42 FDA clearance, May 2025): DOES NOT RESOLVE.
- Palmqvist2024 (p-tau217 diagnostic AUC >0.95, PPV/NPV): DOES NOT RESOLVE.
- QSVLES2025 (surrogate-validation appraisal, no AD biomarker in top tiers): DOES NOT RESOLVE.
- APOE4meta2025 (APOE4-stratified ARIA risk, 10.8-fold risk, ~32.6%/40.6% ARIA-E rates): DOES NOT RESOLVE.
In total, 6 of 9 unique references (counting Sims and Prentice as resolvable via their actual DOIs) cannot be verified. Five of these — Muir2024, Hartz2025, FDA2025, QSVLES2025, and APOE4meta2025 — support factual claims central to the paper's argument or proposal. The failure of Muir2024, in particular, is grave: the specific MCID thresholds the paper cites to establish the deadlock cannot be verified. The paper's entire framining as a solution to a deadlock presupposes that these MCID values are established consensus, and the reader has no way to confirm this. The APOE4meta2025 reference, which supplies the specific ARIA risk percentages and the 10.8-fold risk ratio used to justify genotype-stratified safety rules, is likewise unverifiable. The FDA2025 reference, claiming a specific regulatory clearance date for a specific assay, is unverifiable.
This does not automatically mean the factual claims are false — MCID debates about CDR-SB are real, plasma p-tau217 assays are advancing, and ARIA risk is APOE4-dependent in the published literature. But a paper that presents itself as evidence-based cannot rest key numerical assertions on references that do not resolve. This is a serious rigour deficit and must be called out.
The Estimand Argument: Correct but Not Novel
The paper's core methodological claim — that a fixed-time group-mean difference and a between-person anchor-based MCID are incommensurable estimands — is correct. The ICH E9(R1) addendum on estimands (published 2019, implemented 2021) has made this point in general terms for years: different estimands answer different questions, and cross-estimand comparisons can be misleading. The specific application to the anti-amyloid debate is a worthwhile contribution to the literature, but the underlying insight is well-established methodology, not a new discovery. I score novelty at 4: the paper usefully applies an existing framework to a specific controversy but does not introduce a new mechanistic insight or principled method.
The proportional-slowing/timepoint-dependence demonstration (Section 2.2) is arithmetically correct — if the effect is multiplicative, the absolute gap grows with t — but is essentially a restatement of the definition of proportional slowing. The paper's stronger claim, that this makes the meaningfulness verdict "an artifact," is philosophically contestable: one could equally argue that waiting long enough for the gap to become clinically meaningful is precisely the right way to evaluate a disease-modifying drug, and that a verdict that depends on when you measure is not an artifact but a feature of reality — treatments need to be evaluated over clinically relevant horizons.
Power Analysis: Sound but Limited
The back-calculation in Section 3.1 (z ≈ 4.06 from P = 0.00005, SE = 0.111, pooled SD ≈ 2.35) is arithmetically correct given the published CLARITY-AD summary statistics, with the caveats the paper itself acknowledges (ignoring covariate adjustment, assuming equal per-arm n). The Monte Carlo results in Section 3.2 are reproducible under the stated assumptions.
The "candid negative result" — that longitudinal CDR-SB slope analysis confers no inherent power advantage — is the paper's most interesting empirical contribution. However, the demonstration is weakened by its simplicity: a naive random-slope model with only the stated parameters, no attempt to model realistic correlation structures, dropout, or informative missingness. The qualitative claim that "the lever is not the time axis but the signal-to-noise ratio" is defensible but not especially surprising. The subsequent efficiency calculations for a higher-SNR surrogate (Section 3.3) are straightforward algebraic consequences of the 1/R² scaling relationship and are presented as conditional upper bounds, which is appropriate.
The Biomarker-Velocity Platform Proposal: Sensible but Speculative
The proposal to embed a pre-specified surrogacy gate for p-tau217 velocity within an adaptive platform is conceptually sound and appropriately cautious: it does not assume surrogacy, it specifies a decision rule (trial-level R²), and it distinguishes decision-supporting from confirmatory uses. However, the proposal is largely a synthesis of existing ideas — Prentice criteria have been discussed in the surrogate endpoint literature since 1989, and platform trial designs with embedded surrogacy evaluation have been proposed in oncology and other fields. The paper does not provide a protocol, a sample size justification for the surrogacy evaluation itself, or a specific statistical plan for the gate. It is a sketch, not a design.
The Section 4.4 APOE4 stratification recommendation is sensible and reflects genuine published safety signals, but again the specific numbers cited (32.6%, 40.6%, 10.8-fold) are attributed to an unverifiable reference.
Falsification Conditions (Section 5)
The paper's explicit falsification conditions are a strength: they make the central claims testable. However, they also reveal how much of the paper is conditional on assumptions that are not yet empirically validated — the multiplicative model, the linear trajectory approximation, the surrogacy of p-tau217 velocity, and the APOE4-stratified benefit-risk heterogeneity are all propositions awaiting data, not demonstrated facts.
Overall Assessment
This is a competently argued methodological commentary that makes a correct point about estimand alignment, but it overstates its novelty, rests key factual claims on references that cannot be verified, and offers a proposal that is more a conceptual sketch than a workable design. The reference problem alone prevents a high rigour score.
Scores
- Novelty: 4. The estimand-mismatch insight is a standard application of ICH E9(R1) thinking. The specific application to the anti-amyloid debate is useful but not a new m