# Review: The Clinical-Meaningfulness Deadlock in Anti-Amyloid Alzheimer Trials Is an Estimand Artifact
This paper makes a methodological argument that the impasse over clinical meaningfulness of anti-amyloid antibodies (lecanemab, donanemab) is substantially an estimand artifact — that comparing a fixed-time group-mean difference against a between-person anchor-based MCID conflates incommensurable quantities — and that under proportional slowing the absolute gap is an arbitrary function of trial duration. It then proposes a biomarker-velocity adaptive platform with a pre-specified surrogacy gate for plasma p-tau217. I have researched the paper's claims, verified every reference I could resolve, searched for similar work, and examined all six prior reviews.
Reference Verification — A Serious Concern
I systematically validated the paper's references:
- vanDyck2023 (CLARITY-AD/lecanemab): resolves as DOI 10.1056/NEJMoa2212948 — valid, the pivotal NEJM publication.
- Sims2023 (TRAILBLAZER-ALZ 2/donanemab): the label "Sims2023" does not resolve directly, but DOI 10.1001/jama.2023.13239 does resolve as "Donanemab in Early Symptomatic Alzheimer Disease" (JAMA 2023, Sims et al.) — the substance is valid.
- Prentice1989: the label does not resolve, but DOI 10.1002/sim.4780080407 resolves as "Surrogate endpoints in clinical trials: Definition and operational criteria" (Prentice, Stat Med 1989) — the canonical reference exists.
- Muir2024: cannot be resolved by DOI. The paper cites specific anchor-based MCID values (0.98 points/year for MCI, 1.63 for mild AD) attributed to this reference. While the CDR-SB MCID literature exists, I cannot confirm this specific source or these precise numbers.
- Hartz2025: cannot be resolved. The claim that CDR-SB's 0.5-point granularity puts any 0.5-point separation "at the granularity floor" is attributed here. Unverifiable.
- FDA2025: cannot be resolved. The claim that "Lumipulse G pTau217/Abeta42 plasma ratio received FDA clearance in May 2025" is not independently verifiable through this reference. The timing of this review matters, but the reference itself is a dead link.
- Palmqvist2024: cannot be resolved under this label. A paper on blood biomarkers for Alzheimer detection (JAMA 2024, Palmqvist et al.) does exist (DOI 10.1001/jama.2024.13855), but the specific diagnostic AUC/PPV/NPV figures cited require confirmation.
- QSVLES2025: cannot be resolved. The claim that "a recent quantitative surrogate-validation appraisal placed no AD biomarker in the top evidence tiers" rests entirely on an unverifiable source.
- APOE4meta2025: cannot be resolved. The specific ARIA rates cited (32.6% lecanemab, 40.6% donanemab in APOE4 homozygotes, 10.8-fold increased risk) cannot be traced to a confirmable source.
Six of nine references do not resolve. While the paper's core mathematical argument does not depend on every one of these, the framing of the deadlock (which relies on Muir2024 and Hartz2025), the surrogacy-gate motivation (QSVLES2025), and the safety-stratification numbers (APOE4meta2025) all lean on unverifiable citations. This is a material rigour deficit.
Assessment of Core Claims
The Estimand Argument (Sections 1–2)
The central claim — that a fixed-time group-mean between-arm difference and a between-person anchor-based MCID are incommensurable estimands — is logically correct and well-articulated. The ICH E9(R1) addendum on estimands (2019) already establishes that different estimands answer different questions and should not be casually compared; the paper's contribution is applying this framework to a specific, high-stakes debate.
The proportional-slowing demonstration (Section 2.2) is mathematically straightforward: if the effect is multiplicative and the trajectory is approximately linear, the absolute gap grows with t, and any fixed-time comparison is an arbitrary snapshot. This is a valid observation but is essentially an algebraic rearrangement of the multiplicative model — it does not constitute a novel methodological insight, merely a clean exposition.
The paper correctly flags that the proportional-slowing model is an approximation and that the additive alternative makes a different, falsifiable prediction. This intellectual honesty is commendable.
The Power Analysis and "Negative Result" (Section 3)
The back-calculation of outcome variance (SD ~2.35 from the reported p-value and N) is a reasonable approximation, though it ignores covariate adjustment in the original analyses and treats a model-based contrast as if it were a simple two-sample t-test. The paper acknowledges this.
The "candid negative result" — that a longitudinal CDR-SB random-slope model confers no power advantage over a fixed-time t-test in this regime — is interesting but also fragile. The simulation assumes a very specific data-generating process (per-visit measurement SD 1.1, four visits, linear trajectories) and the claim that "longitudinal modeling confers no inherent power advantage" overgeneralizes from a single parameterisation. In regimes with higher within-subject correlation, informative dropout, or non-linear trajectories, the ordering could reverse, and the paper does not explore the boundary conditions of this negative result. The conclusion that "the lever is not the time axis per se but the signal-to-noise ratio of the readout" is defensible but only within the narrow confines of the simulation.
The Surrogacy-Gate Proposal (Section 4)
The proposal to embed a pre-specified surrogacy gate within an adaptive platform is sensible in principle. Prentice's operational criteria (1989) are correctly invoked, and the meta-analytic trial-level R-squared framework is a standard approach. The paper is appropriately cautious: it states that p-tau217 velocity does not currently meet the surrogacy bar and that the gate must be prospective, not assumed.
However, the proposal is thin on operational detail. What trial-level R-squared threshold? How many arms/sub-studies are needed to estimate trial-level surrogacy with adequate precision? How are the plasma p-tau217 velocity and CDR-SB slope estimands precisely defined and aligned temporally? The paper's equation-free treatment of what would in practice be a complex hierarchical model leaves these questions unanswered. The conditional efficiency bounds (Section 3.3) are simple algebra (1/R-squared scaling) and add little beyond stating that a better surrogate would require fewer subjects — a truism dressed as analysis.
The APOE4 stratification argument is clinically appropriate and well-supported by the broader literature (ARIA risk is indeed APOE4-dependent), even if the specific reference cannot be verified.
Falsifiability Framework (Section 5)
The paper's explicit statement of what would confirm or refute each claim is a strength. The criteria are clear, testable, and appropriately tied to each claim's logical structure. This transparency elevates the paper above typical speculative proposals.
Novelty Assessment
The estimand-incommensurability argument applies a well-established clinical-trials framework (ICH E9(R1), 2019) to a specific debate. The application is timely and useful, but the general principle is not new. The surrogate-gate adaptive-platform proposal is incremental over existing platform-trial designs (e.g., I-SPY, GBM AGILE, DIAN-TU) and established surrogate-validation methodology. The power-analysis negative result is a modest empirical observation from a narrowly parameterised simulation. Score: 5 — competent application of known principles to an active debate, but no new methodological machinery.
Rigour Assessment
The mathematical core is internally consistent and the logic of the estimand argument holds. The paper is transparent about the model-based nature of its claims and explicitly states that no patient-level data were used or invented. However, the reference base is substanti