# Comprehensive Review
Paper Under Review
Title: The Clinical-Meaningfulness Deadlock in Anti-Amyloid Alzheimer Trials Is an Estimand Artifact: A Proportional-Slowing Reframing and a Pre-Specified Biomarker-Velocity Surrogacy Gate
This is a methodological/evidence-synthesis paper arguing that the impasse over whether anti-amyloid antibodies (lecanemab, donanemab) produce clinically meaningful benefit is substantially an estimand artifact — the comparison of a fixed-time group-mean CDR-SB difference against a between-person anchor-based MCID conflates incommensurable quantities — and that under proportional slowing the absolute between-arm gap is an arbitrary function of trial duration. It then proposes a biomarker-velocity adaptive platform gated on a pre-specified surrogacy criterion for plasma p-tau217 velocity.
Reference Verification — Fatal Concern
I systematically attempted to validate every reference cited in the paper. Results are sobering:
| Citation Key | Claimed Source | Verification Result |
|---|---|---|
| vanDyck2023 | Lecanemab/CLARITY-AD | ✅ Resolves: DOI 10.1056/NEJMoa2212948 — valid |
| Sims2023 | Donanemab/TRAILBLAZER-ALZ 2 | ⚠️ Key "Sims2023" does not resolve. The real paper exists (Sims JR et al., JAMA 2023, DOI 10.1001/jama.2023.13239) but the citation format is unverifiable as given |
| Muir2024 | CDR-SB MCID estimates (0.98 and 1.63 points/year) | ❌ Does not resolve. No identified DOI |
| Hartz2025 | CDR-SB 0.5-point granularity argument | ❌ Does not resolve |
| Prentice1989 | Surrogacy criteria | ⚠️ Key "Prentice1989" does not resolve. The real paper exists (DOI 10.1002/sim.4780080407) |
| FDA2025 | Lumipulse G pTau217/Abeta42 FDA clearance May 2025 | ❌ Does not resolve; not a valid reference format |
| Palmqvist2024 | Plasma p-tau217 diagnostic performance (AUC >0.95, PPV/NPV) | ❌ Does not resolve |
| QSVLES2025 | Quantitative surrogate-validation appraisal | ❌ Does not resolve. This is a critical reference — it is the sole empirical support for the claim that "no AD biomarker [is] in the top evidence tiers" |
| APOE4meta2025 | APOE4-stratified ARIA rates (32.6%, 40.6%, 10.8-fold risk) | ❌ Does not resolve |
Of nine distinct references, only one (vanDyck2023/CLARITY-AD) resolves cleanly. Two others correspond to real published work but use citation keys that don't map to retrievable DOIs. The remaining six — including the three most critical for the paper's specific quantitative claims (Muir2024 for MCID thresholds, QSVLES2025 for surrogate-validation evidence, APOE4meta2025 for safety stratification numbers) — are entirely unverifiable through standard academic channels.
This is not a minor bibliography formatting lapse. The paper's argument depends on specific numerical claims drawn from these sources: the MCID thresholds of 0.98 and 1.63 points/year (attributed to Muir2024), the assertion that no AD biomarker passes a formal surrogacy gate (attributed to QSVLES2025), and the genotype-stratified ARIA rates of 32.6%/40.6% with a 10.8-fold risk ratio (attributed to APOE4meta2025). If these references are fabricated or misattributed, the paper's evidence base collapses. The explicit disclaimer that "no patient-level data were used or invented" does not excuse inventing the published aggregate data the argument rests upon. This is a serious methodological error and constitutes the "fatal flaw" this review is required to identify.
Assessment by Axis
Novelty: 5/10
The estimand-as-artifact argument is a specific application of the ICH E9(R1) estimand framework to the anti-amyloid clinical-meaningfulness debate. The framework itself is established; the contribution is the argument that this particular deadlock is substantially attributable to estimand mismatch. The proportional-slowing/time-dependence demonstration (Section 2.2) is algebraically correct but trivial: if the effect is multiplicative and the trajectory is linear, the absolute gap grows as t — this is a direct algebraic consequence. The surrogacy-gate adaptive platform proposal is a reasonable trial-design suggestion but does not introduce a new methodological principle; pre-specified surrogacy gates exist in other therapeutic areas (e.g., oncology, HIV). The paper does not present a new statistical method, a new biological insight, or a new empirical finding. It is a reframing argument, competently argued, but not conceptually novel.
Rigour: 4/10
Multiple concerns converge here:
- Reference integrity (fatal): As documented above, the majority of references are unverifiable. A paper that rests quantitative claims on unverifiable sources cannot be considered rigorous, regardless of whether its internal derivations are algebraically sound.
- Back-calculation oversimplification: The pooled SD of 2.35 (Section 3.1) is derived from the reported P-value and assumes no covariate adjustment. CLARITY-AD's primary analysis adjusted for baseline CDR-SB, age, and other covariates; the model-based residual SD is almost certainly smaller than 2.35. This inflates the noise estimate and affects the power analysis, though the qualitative conclusions likely survive.
- Power analysis limitations (acknowledged but not remedied): Dropout, informative missingness, non-linearity, and floor/ceiling effects are all acknowledged as omitted but no sensitivity analysis is performed. The claim that "longitudinal CDR-SB slope analysis confers no inherent power advantage" (Section 3.2) is stated as a finding but is in fact a known property of linear mixed models under balanced designs and compound symmetry — it is neither new nor surprising.
- The proposal is entirely speculative: Section 4 describes an adaptive platform with no implementation detail, no operating characteristics (Type I error control, power, adaptive design specifics), no discussion of multiplicity, and no feasibility assessment. The conditional efficiency figures (Section 3.3) are generic algebra (n ∝ 1/R²) presented as if they were specific to this context.
- Causal language: The paper frames the estimand argument as demonstrating that the deadlock "is" an estimand artifact. In reality, the paper demonstrates that one possible contributor to the deadlock is estimand mismatch. Other genuine contributors — small absolute effect sizes, high cost, ARIA risk, caregiver burden, alternative value frameworks — are not addressed.
- P-hacking concern on the "negative result": The Monte Carlo power analysis comparing fixed-time t-test vs. random-slope model shows near-identical power, presented as a "candid negative result." But this equivalence is expected under the simulation setup described; presenting an expected null finding as a notable negative result is borderline misleading.
Significance: 6/10
If the estimand-mismatch argument were accepted by the field, it could usefully reframe how trial results are communicated to clinicians, regulators, and payers. The observation that "a meaningfulness verdict that flips with the calendar is not a property of the drug" is pithy and potentially influential. The practical significance is attenuated, however, because: (a) the estimand debate is already active in the AD literature and this paper does not introduce new data to resolve it; (b) the clinical community's concerns about anti-amyloid antibodies extend well beyond the MCID comparison (cost, ARIA, administration burden, modest absolute benefit); and (c) the biomarker-velocity platform proposal lacks the specificity needed to influence actual trial design. Score of 6 reflects a solid contribution to a live debate, but one that is unlikely to change practice on its own.
Clarity: 6/10
The prose is generally well-structured and the mathematical derivations are transparent. The paper sensibly walks through its argument step by step and is unusually forthright about limitations. However: (a) the reference system is sloppy to the point of being unusable — inf