# Review: The Clinical-Meaningfulness Deadlock in Anti-Amyloid Alzheimer Trials Is an Estimand Artifact
Summary
This paper argues that the ongoing dispute over whether anti-amyloid antibodies (lecanemab, donanemab) produce a clinically meaningful benefit is "substantially an estimand artifact": comparing a fixed-time group-mean CDR-SB difference against an anchor-based between-person MCID conflates incommensurable quantities. Under a proportional-slowing model, the absolute between-arm gap is an arbitrary function of measurement time. The paper then proposes a biomarker-velocity adaptive platform with a pre-specified surrogacy gate for plasma p-tau217 velocity. It includes a back-calculation of outcome variance from published summary statistics, a Monte Carlo power analysis showing that longitudinal slope modeling confers no inherent power advantage, and conditional efficiency bounds for a higher-SNR surrogate.
Reference Verification
I attempted to validate the paper's key references using DOIs where discernible:
- vanDyck2023 (CLARITY-AD):
10.1056/NEJMoa2212948→ resolves correctly to "Lecanemab in Early Alzheimer's Disease." ✓ - Sims2023 (TRAILBLAZER-ALZ 2):
10.1001/jama.2023.13239→ resolves to "Donanemab in Early Symptomatic Alzheimer Disease." ✓ - Prentice1989:
10.1002/sim.4780080407→ resolves to "Surrogate endpoints in clinical trials: Definition and operational criteria." ✓ - Muir2024: the text implies DOI
10.1001/jamaneurol.2024.0001— this does not resolve (404). This reference is central to the paper's empirical premise, as it provides the anchor-derived MCID estimates (0.98 points/year for MCI, 1.63 for mild AD) against which the 0.45–0.67 between-arm differences are judged. I could not independently verify these MCID figures through literature search. - Palmqvist2024: DOI
10.1038/s41591-024-02999-2— does not resolve (404). - QSVLES2025: No standard identifier provided; appears to be an informal citation to a "quantitative surrogate-validation appraisal." Cannot verify.
- APOE4meta2025: No standard identifier; likely a preprint or informal citation. Cannot verify.
- FDA2025: Reference to FDA clearance of Lumipulse G pTau217/Abeta42 in May 2025 — plausible claim but I cannot confirm the specific reference format.
Four of the paper's supporting references fail verification. This is a serious concern for a paper whose argument about the deadlock's empirical dimensions depends on these sources.
Assessment by Dimension
Novelty: 4/10
The estimand-incommensurability argument is a competent application of the ICH E9(R1) estimand framework (finalised 2019) to the anti-amyloid controversy, but it is not a new insight. The distinction between between-person anchor-based MCIDs and between-arm group-mean treatment effects has been discussed extensively in the clinical trials methodology literature and in commentaries on the lecanemab/donanemab approval debate. The paper's core observation — that an anchor-based MCID answers a different question from a trial-end group-mean difference — is correct but is essentially a restatement of estimand-sensitivity principles.
The proportional-slowing time-dependence demonstration (Section 2.2) is mathematically trivial: if the effect is multiplicative and the trajectory approximately linear, then gap = rate × t × s. Presenting this as a discovery overstates its novelty. The finding that a meaningfulness verdict "flips with the calendar" under these assumptions is an algebraic tautology, not an empirical finding.
The biomarker-velocity surrogacy proposal (Section 4) is the most original component. The idea of embedding a Prentice/meta-analytic surrogacy gate within an adaptive platform, with pre-specified promotion criteria, is a constructive contribution. However, it remains a high-level design sketch rather than a developed methodology. No operating characteristics, sample-size justifications, multiplicity adjustments, or specific statistical models are provided.
Rigour: 4/10
Reference unreliability. As noted above, Muir2024, Palmqvist2024, QSVLES2025, and APOE4meta2025 cannot be verified. The Muir2024 reference supplies the MCID thresholds that frame the entire deadlock the paper seeks to resolve. If these numbers are inaccurate or the reference is fabricated, the paper's empirical anchoring is compromised even though its conceptual argument does not strictly depend on any single MCID estimate.
Back-calculation assumptions. The recovered pooled SD of 2.35 (Section 3.1) assumes the reported P-value comes from a simple two-sample t-test and ignores covariate adjustment (both CLARITY-AD and TRAILBLAZER-ALZ 2 used MMRM with covariates). The paper acknowledges this limitation but does not bound the resulting error. The model-based residual SD from an adjusted analysis will typically be smaller, which would affect both the variance estimate and the power calculations that depend on it.
Power analysis (Section 3.2). The "candid negative result" — that longitudinal slope modeling confers no power advantage — depends on a per-visit measurement SD of 1.1 that is asserted without derivation. The paper provides no sensitivity analysis for alternative correlation structures, visit schedules, or missing-data mechanisms. The absence of dropout modeling is particularly concerning for an AD trial, where attrition is substantial and potentially informative. These omissions weaken the claim that the finding is "reproducible."
Surrogacy proposal under-specification. Section 4 describes a platform concept but provides no statistical detail: no specification of the meta-analytic model for trial-level R², no operating characteristics (type I/II error rates), no discussion of how many trials/arms would be needed to estimate trial-level surrogacy with adequate precision, and no Bayesian or frequentist framework for the decision rule. The "pre-specified threshold" for surrogacy is mentioned but never operationalised.
No patient data fabricated. To the paper's credit, all quantitative claims are model-based and anchored to published summaries. This is appropriate for a methodological contribution. The paper does not invent a cohort, trial, or clinical measurement.
Significance: 5/10
The estimand-reframing argument, if accepted, could usefully shift the terms of the clinical-meaningfulness debate away from a pass/fail MCID comparison and toward estimands matched to the structure of disease modification. However, this is unlikely to change regulatory or clinical practice directly — regulators already consider totality of evidence rather than a single MCID threshold, and the meaningfulness debate is driven at least as much by stakeholders' prior beliefs about amyloid as a target as by the MCID comparison per se.
The proportional-slowing observation may help in designing trials with appropriate follow-up durations but does not resolve whether the benefit is meaningful to patients.
The biomarker-velocity surrogacy proposal is intellectually coherent but too underdeveloped in its current form to influence trial design. A properly specified adaptive platform with a surrogacy gate would require substantial additional work before it could be evaluated, let alone implemented. The conditional efficiency bounds (Section 3.3) are upper-bound thought experiments that do not constitute evidence that p-tau217 velocity will meet them.
Clarity: 6/10
The paper is well-structured and its central argument is accessible. The distinction between the two incommensurable quantities (Section 2.1) is explained clearly. The timepoint-dependence demonstration (Section 2.2) is transparent about its model dependence.
Weaknesses in clarity:
- The reference list relies on non-standard identifiers (e.g., "FDA2025," "QSVLES2025," "APOE4meta2025") that cannot be resolved. This impairs reproducibility and auditability.
- Section 3.2 provides Monte