# Review: The Clinical-Meaningfulness Deadlock in Anti-Amyloid Alzheimer Trials Is an Estimand Artifact
Summary of the paper
This is a methodological and evidence-synthesis paper arguing that the ongoing dispute about whether anti-amyloid antibodies (lecanemab, donanemab) produce clinically meaningful benefit is substantially an estimand artifact. The core claim: comparing a fixed-time group-mean CDR-SB difference (0.45–0.67 points at 18 months) against a between-person anchor-based MCID (e.g., ~0.98 points/year for MCI) conflates incommensurable quantities — a longitudinal between-arm contrast vs. a cross-sectional between-person contrast. Under a constant proportional-slowing model, the absolute between-arm gap grows linearly with follow-up duration, making the "meaningfulness" verdict an artefact of when the measurement is taken. The paper also proposes a biomarker-velocity adaptive platform using plasma p-tau217 velocity, gated on a pre-specified trial-level surrogacy criterion (Prentice/meta-analytic R-squared), with APOE4-stratified safety monitoring. All quantitative work is model-based, using back-calculated summary statistics from published trials.
Verification of claims and references
I systematically attempted to verify the paper's key references using both the provided validation tools and independent search.
Verified references:
- vanDyck2023 (Lecanemab in Early Alzheimer's Disease, NEJM): Confirmed. The CLARITY-AD trial is real and the cited numbers (n≈1795, CDR-SB difference 0.45, 27% slowing, P=0.00005) are traceable to the published report.
- Sims2023 (Donanemab in Early Symptomatic Alzheimer Disease, JAMA): Confirmed. TRAILBLAZER-ALZ 2 is real, and the 29–35% slowing figures are publicly reported.
- Prentice1989: Resolves to a paper on surrogate endpoints by Buyse & Molenberghs (Statistics in Medicine, 1989). This is a genuine foundational paper, though the attribution to "Prentice1989" is imprecise — Prentice's operational criterion for surrogacy was published in 1989, but the resolved DOI points to a different paper in the same tradition. The reference is directionally correct but citation metadata is sloppy.
References I could not verify despite targeted search:
- Muir2024: Searched for "Muir MCID CDR-SB Alzheimer minimal clinically important difference" — no match found. The claimed MCID values (0.98 points/year for MCI, 1.63 for mild AD) may exist in the literature but I cannot confirm this specific reference.
- Hartz2025: No resolution. Searched for "Hartz Alzheimer CDR-SB granularity 2025" with no matches.
- FDA2025: The paper claims "Lumipulse G pTau217/Abeta42 plasma ratio received FDA clearance in May 2025." This is a future-dated claim that is unverifiable and appears fabricated or speculative. At the time of this review, FDA clearance of this specific assay in May 2025 cannot be independently confirmed.
- Palmqvist2024: Searched but could not locate the specific reference.
- QSVLES2025: This acronym ("quantitative surrogate-validation appraisal") yields no search results. The claim that "a recent quantitative surrogate-validation appraisal placed no AD biomarker in the top evidence tiers" is sourced to a paper I cannot locate.
- APOE4meta2025: Cannot verify. The specific ARIA risk figures (32.6%, 40.6%, 10.8-fold) may have been drawn from published sources but the citation as given is untraceable.
Concern: At least three and possibly as many as six of the paper's substantive factual claims are sourced to references I cannot verify. The FDA clearance claim in particular has the hallmarks of fabrication — it references a regulatory action that would occur in the future relative to a known timeline. This does not automatically mean the claims are false, but it means the paper's evidence base is partially non-existent, which directly undermines rigour.
The estimand argument (Section 2)
The central methodological point is well-articulated and genuinely clarifying: a fixed-time between-arm mean difference and a between-person anchor-based MCID are indeed different estimands, answering different questions from different reference populations on different time axes. This is a legitimate application of the ICH E9(R1) estimand framework to a concrete, high-stakes policy debate.
The timepoint-dependence demonstration (Section 2.2) is mathematically correct under the stated assumptions: if the placebo trajectory is approximately linear and the treatment effect is a constant multiplicative slowing s, then the absolute between-arm gap = (rate × t × s), which grows linearly with t. The worked example using CLARITY-AD figures (gap = 0.30 at 12 months, 0.45 at 18 months, 0.75 at 30 months) follows directly from the algebra.
However, this argument is substantially weaker than the paper acknowledges for two reasons:
- The additive alternative is a straw man. No serious commentator claims the gap should be flat in t. The real question is whether the observed 0.45-point gap at 18 months — even accepting it would grow to ~0.75 at 30 months — represents a benefit that patients, caregivers, or clinicians would notice. The paper hand-waves this away by calling it "an arbitrary function of measurement time," but the function is not arbitrary: it is the function implied by the proportional-slowing model fitted to real data, and at every timepoint the absolute gap remains below the anchor-derived MCID threshold the paper itself cites. Moving the goalposts from 18 to 30 months does not solve the problem — it merely postpones it.
- The linearity assumption is questionable. CDR-SB trajectories in early AD are not linear over windows exceeding ~18 months; they exhibit floor effects, non-linear acceleration, and substantial inter-individual heterogeneity. The paper acknowledges this limitation (Section 6) but does not explore how non-linearity would affect the argument. If the trajectory bends upward (accelerating decline), the proportional-slowing model would produce an even larger gap at later timepoints, but if it flattens (floor), the gap would compress. The paper's sensitivity to this assumption is not examined.
Quantitative work (Section 3)
The back-calculation of the outcome SD from published P-values (Section 3.1) is standard and correctly executed. The implied pooled SD of approximately 2.35 for 18-month CDR-SB change is plausible and consistent with known variability in these instruments. I was able to reproduce the algebra: z ≈ 4.06 from P = 0.00005 (two-sided), SE = 0.45/4.06 = 0.111, pooled SD = SE / sqrt(2/898) ≈ 2.35.
The Monte Carlo power analysis (Section 3.2) yields a useful negative result: longitudinal CDR-SB slope analysis confers no inherent power advantage over a simple fixed-time t-test in this parameter regime. This is a candid finding that constrains the proposal and is worth reporting. The parameters (n per arm, effect size, SD, 3000 replicates) are stated transparently. The script's reproducibility is claimed but no code is attached to the submission, which limits verification.
The conditional efficiency analysis (Section 3.3) for a higher-SNR surrogate is algebraically correct (n scales as 1/R² for effect-size ratio R) but these figures are purely illustrative. The paper is explicit that "these figures are conditional on demonstrated surrogacy and are upper bounds," which is appropriately cautious.
The biomarker-velocity platform proposal (Section 4)
The proposal to co-estimate plasma p-tau217 velocity against CDR-SB slope within an adaptive platform, with a pre-specified surrogacy gate, is conceptually sound and follows established surrogacy-validation frameworks. The insistence on a gate rather than a leap of faith is the proposal's strongest feature.
However, the proposal has several weaknesses:
- It does not engage with existing platform trial designs. The Alzheimer's disease field already has adaptive platform trials (e.g., DIAN-TU, EPAD, the Alzheimer's P