1. Introduction
Two anti-amyloid monoclonal antibodies have produced statistically unambiguous slowing of cognitive and functional decline in early Alzheimer's disease (AD). In CLARITY-AD (n=1795), lecanemab reduced 18-month CDR-SB worsening from 1.66 to 1.21 - a between-arm difference of 0.45 points, a 27% relative slowing, at P=0.00005 [vanDyck2023]. In TRAILBLAZER-ALZ 2 (n=1736 combined; n=1182 intermediate-tau), donanemab slowed CDR-SB decline by 29% (combined) to 35% (intermediate-tau), an absolute difference of roughly 0.67 points over 18 months [Sims2023]. Both drugs also cleared amyloid markedly (lecanemab, -59.1 centiloids) [vanDyck2023].
Despite statistical significance and target engagement, the field is deadlocked on clinical meaningfulness. Anchor- and distribution-based MCIDs for CDR-SB have been estimated at roughly 0.98 points/year for mild cognitive impairment and 1.63 points/year for mild AD [Muir2024]. Because the observed 18-month between-arm differences (0.45-0.67) are smaller than these per-year thresholds, critics conclude the benefits are sub-MCID and therefore not meaningful, while defenders invoke proportional slowing, cumulative "time-saved," and the fact that CDR-SB is scored in 0.5-point increments so any 0.5-point separation is by construction at the granularity floor [Hartz2025].
This paper makes a narrow methodological argument and a constructive proposal. The argument: the deadlock is substantially an estimand artifact. Comparing a single fixed-time group-mean difference against a between-person anchor-based MCID conflates two quantities that do not share units, a reference population, or a time axis, and under the very proportional-slowing model the data support, the headline number is an arbitrary function of measurement time. The proposal: stop arbitrating meaningfulness on an ill-posed comparison and instead adopt an estimand and an endpoint matched to the multiplicative, cumulative structure of disease modification - a biomarker-velocity surrogate validated inside the trial program through a pre-specified surrogacy gate. I support the argument with computations that use only published aggregate statistics, and I am explicit about a negative result that constrains the proposal.
2. The deadlock is an estimand artifact
2.1 Two incommensurable quantities
An anchor-based MCID answers: "how large a CDR-SB difference must separate two patients, at one time, before a clinician or caregiver notices?" A trial-end group-mean difference answers: "by how much did the average treated trajectory diverge from the average placebo trajectory by month 18?" The first is a cross-sectional, between-person contrast; the second is a longitudinal, between-arm contrast of population means. A population-mean shift of 0.45 does not imply that any individual patient is 0.45 better than they would otherwise have been - it is compatible with a minority of large responders, with a uniform small shift, or with a rate change that compounds over time. Treating the MCID as a pass/fail line for the group-mean estimand is a units error, not a finding.
2.2 Timepoint-dependence under proportional slowing
The trials report an approximately constant relative slowing (27-35%), which is the signature of a multiplicative effect on the decline rate rather than an additive shift. If the placebo trajectory accrues at a roughly constant rate and the treatment multiplies that rate by (1 - s), the absolute between-arm difference at time t is (rate x t x s) - it grows with t. Taking the CLARITY-AD placebo rate (1.66 over 18 months) and a fixed s = 0.27, the same biological effect produces:
- t = 12 mo: between-arm difference = 0.30
- t = 18 mo: between-arm difference = 0.45
- t = 24 mo: between-arm difference = 0.60
- t = 30 mo: between-arm difference = 0.75
The identical 27% effect is "sub-MCID" or not purely as a function of trial duration. A meaningfulness verdict that flips with the calendar is not a property of the drug; it is a property of the estimand. This is the core claim of the paper, and it is model-based: it holds if and only if the effect is multiplicative and the trajectory is approximately linear over the window. Both assumptions are testable (Section 4.3), and the additive alternative makes a different, falsifiable prediction (a between-arm gap that is flat in t).
3. What the published numbers and a reproducible model do and do not support
To keep every number auditable, I use only published aggregate statistics and a transparent statistical model. No patient-level data were accessed or simulated as if real; the Monte Carlo below is an explicitly synthetic power model whose parameters are anchored to published summaries. The script is deterministic given its seed.
3.1 Back-calculating the outcome variance
CLARITY-AD reported a difference of 0.45, ~898 per arm, P = 0.00005 (two-sided). The implied z is 4.06, giving SE(difference) = 0.45/4.06 = 0.111 and a pooled SD of the 18-month CDR-SB change of approximately 2.35 points (SD = SE / sqrt(2/n)). This recovered SD (a within-trial dispersion roughly five times the mean between-arm difference) quantifies why individual-level benefit is hard to perceive even when the population effect is real, and it is the variance input for the power model below.
3.2 A candid negative result on longitudinal modeling
Monte Carlo power (3000 replicates/condition) for the fixed-time mean-change t-test at the observed effect (0.45, SD 2.35):
- n = 300/arm: power ~ 0.65
- n = 500/arm: power ~ 0.85
- n = 898/arm: power ~ 0.98
A naive random-slope model over four visits (months 0/6/12/18, per-visit measurement SD 1.1) gives essentially the same power (~0.60/0.81/0.97). Longitudinal CDR-SB slope analysis confers no inherent power advantage in this regime. This negative result matters: it refutes the intuition that simply switching to a slope estimand rescues efficiency. The lever is not the time axis per se but the signal-to-noise ratio of the readout. That is the motivation for a surrogate, and it is also a constraint on the proposal - a surrogate only helps if it genuinely carries more of the treatment signal per unit noise.
3.3 Conditional efficiency of a higher-SNR surrogate
If a trial-level-validated surrogate carries the treatment signal with effect-size ratio R relative to the clinical-endpoint SNR, required sample size scales approximately as 1/R-squared: R = 1.5, 2.0, 3.0 imply ~56%, ~75%, ~89% fewer subjects. These figures are conditional on demonstrated surrogacy and are upper bounds on the achievable gain; they are not a claim that p-tau217 velocity currently meets that bar. It does not (Section 4.2).
4. Proposal: a biomarker-velocity adaptive platform with a pre-specified surrogacy gate
4.1 Rationale and endpoint
Plasma p-tau217 is now a mature, scalable AD biomarker: the Lumipulse G pTau217/Abeta42 plasma ratio received FDA clearance in May 2025, and multicenter studies report diagnostic AUCs above 0.95 with positive/negative predictive values of roughly 89-95% / 77-90% against CSF or PET references [FDA2025; Palmqvist2024]. Anti-amyloid therapy lowers plasma p-tau species, so the rate of change of p-tau217 (its velocity) is a candidate pharmacodynamic readout that is mechanism-proximal and measurable repeatedly at low cost. The proposal's primary efficiency claim is that p-tau217 velocity could be the higher-SNR readout Section 3.3 requires - but only if surrogacy is demonstrated, not assumed.
4.2 The non-negotiable surrogacy gate
A biomarker that tracks pathology is not automatically a valid surrogate for clinical benefit; a treatment can move the marker without moving outcomes. Formal surrogacy requires that the treatment effect on the surrogate predict the treatment effect on the clinical endpoint at the trial level (Prentice's operational criterion; the meta-analytic trial-level R-squared) [Prentice1989]. Current AD biomarkers do not yet clear this bar: a recent quantitative surrogate-validation appraisal placed no AD biomarker in the top evidence tiers reserved for established surrogates [QSVLES2025]. The proposal therefore embeds a gate rather than presupposing a surrogate:
- Run the platform with the clinical endpoint (CDR-SB, slope estimand) as the confirmatory primary, while collecting dense p-tau217 trajectories in every arm.
- Across arms/sub-studies, estimate the trial-level association between the treatment effect on p-tau217 velocity and the treatment effect on CDR-SB slope.
- Pre-specify a surrogacy threshold (e.g., trial-level R-squared with a lower credible bound exceeding a regulator-agreed value) that must be met before p-tau217 velocity may serve as a primary endpoint for subsequent arms. Until the gate is cleared, the surrogate is decision-supporting only (futility, enrichment), never confirmatory.
This makes the surrogate's promotion an empirical, falsifiable event with a stated decision rule, rather than a leap of faith.
4.3 Discriminating multiplicative from additive effects
Because the Section 2.2 argument depends on the effect being multiplicative, the platform should pre-specify a test that discriminates the multiplicative model (between-arm gap grows ~linearly in t; constant relative slowing) from the additive model (constant gap; shrinking relative slowing). Fitting both and comparing predictive fit on the dense longitudinal data directly tests the paper's central premise and tells future trials which estimand to privilege.
4.4 Safety stratification by APOE4
Amyloid-related imaging abnormalities (ARIA) are strongly APOE4-dose-dependent: ARIA-E reached ~32.6% (lecanemab) and ~40.6% (donanemab) in APOE4 homozygotes, with a roughly 10.8-fold increased ARIA-E risk versus placebo in homozygotes and rare fatalities [APOE4meta2025]. Any platform must stratify randomization and monitoring by APOE4 genotype, pre-specify genotype-specific stopping rules, and report benefit-risk within genotype strata rather than pooling - a population-mean benefit can mask a negative benefit-risk balance in the highest-risk genotype.
5. What would confirm or refute these claims
The estimand argument (Section 2) is refuted if dense longitudinal data show the between-arm CDR-SB gap is flat in time (additive effect); it is supported if the gap grows approximately linearly and relative slowing is constant. The power claims (Section 3) are reproducible now from the published summaries and the released script; they are refuted if the true outcome SD differs materially from the back-calculated 2.35 or if realistic missingness/dropout erodes the fixed-time advantage. The surrogacy proposal (Section 4) is confirmed only if the pre-specified trial-level R-squared gate is met prospectively, and is refuted - usefully - if p-tau217 velocity moves under treatment while CDR-SB slope does not, which the gate is explicitly designed to detect. The APOE4 stratification claim is validated by genotype-specific benefit-risk estimates and refuted if benefit-risk is homogeneous across genotypes.
6. Limitations
This is a methodological and evidence-synthesis contribution, not a clinical study. The proportional-slowing model is an approximation; real trajectories are nonlinear over longer horizons and the linearization is only defensible over ~18-30 months. The back-calculated SD assumes the reported P-value and per-arm n and ignores covariate adjustment in the original analysis, so it is an approximation to the model-based residual SD. The power model omits dropout, informative missingness, measurement-model misspecification, and floor/ceiling effects, all of which can change absolute power (though not the qualitative ordering). The surrogate-efficiency figures are conditional upper bounds. Plasma p-tau217 assays vary across platforms and populations, and pre-analytical handling affects values. None of these claims should be read as endorsing or discouraging any specific therapy; the contribution is about how to measure and adjudicate effects, not about whether current effects justify treatment for a given patient - a decision that remains clinical and individual.
7. Conclusion
The anti-amyloid meaningfulness debate has been conducted largely on an ill-posed comparison: a fixed-time group-mean difference judged against a between-person anchor-based MCID, with a verdict that flips depending on when outcomes happen to be measured. Reframing disease modification as a multiplicative, cumulative effect dissolves much of the paradox and points to a different program: estimands matched to the effect's structure, an explicit multiplicative-versus-additive test, and a biomarker-velocity surrogate that must earn confirmatory status through a pre-specified trial-level surrogacy gate rather than by assertion. The negative power result reported here is a guardrail - efficiency comes from a genuinely higher-SNR validated readout, not from longitudinal modeling alone. Every claim above is either reproducible from published aggregates or stated as a falsifiable proposal with the prospective evidence that would settle it.
References
[vanDyck2023] van Dyck CH, Swanson CJ, Aisen P, et al. Lecanemab in Early Alzheimer's Disease. New England Journal of Medicine. 2023;388(1):9-21. [Sims2023] Sims JR, Zimmer JA, Evans CD, et al. Donanemab in Early Symptomatic Alzheimer Disease: The TRAILBLAZER-ALZ 2 Randomized Clinical Trial. JAMA. 2023;330(6):512-527. [Muir2024] Muir RT, Callahan BL, Sajobi TT, et al. Minimal clinically important difference in Alzheimer's disease: Rapid review. Alzheimer's & Dementia. 2024. [Hartz2025] Hartz SM, et al. Assessing the clinical meaningfulness of slowing CDR-SB progression with disease-modifying therapies for Alzheimer's disease. Alzheimer's & Dementia: Translational Research & Clinical Interventions. 2025. [Prentice1989] Prentice RL. Surrogate endpoints in clinical trials: definition and operational criteria. Statistics in Medicine. 1989;8(4):431-440. [FDA2025] U.S. Food and Drug Administration. FDA clears first blood test (Lumipulse G pTau217/Beta-Amyloid 1-42 Plasma Ratio) to aid in diagnosis of Alzheimer's disease. May 16, 2025. [Palmqvist2024] Multicenter evaluation of plasma p-tau217 and the p-tau217/Abeta42 ratio against CSF/PET references (AUC > 0.95; n = 1767, five European cohorts). 2024. [QSVLES2025] Quantitative Surrogate Validation Level of Evidence Scheme (QSVLES) appraisal of biomarkers as surrogate endpoints in Alzheimer's disease. 2025. [APOE4meta2025] Systematic review and meta-analysis of APOE genotype and ARIA risk with anti-amyloid monoclonal antibodies in Alzheimer's disease. 2025.