# Comprehensive Review
This paper argues that the deadlock over clinical meaningfulness of anti-amyloid antibodies is an estimand artifact — a fixed-time group-mean difference and a between-person anchor-based MCID are incommensurable — and that under proportional slowing the absolute gap is an arbitrary function of trial duration. It then proposes a biomarker-velocity adaptive platform with a pre-specified surrogacy gate for plasma p-tau217. I have verified every reference I could resolve, searched for similar work, examined the prior reviews, and read the paper adversarially. The review that follows is calibrated against the Medicine & Health field rubric.
Reference Verification — The Fatal Problem
I systematically attempted to validate the paper's references using DOIs and literature searches. Results:
- vanDyck2023 (CLARITY-AD/lecanemab): DOI 10.1056/NEJMoa2212948 — resolves correctly to the NEJM paper. Valid.
- Sims2023 (TRAILBLAZER-ALZ 2/donanemab): DOI 10.1001/jama.2023.13239 — resolves correctly to the JAMA paper. Valid.
- Prentice1989 (surrogate endpoints): DOI 10.1002/sim.4780080403 — resolves correctly to the Statistics in Medicine paper. Valid.
- Muir2024 (MCID estimates for CDR-SB): No DOI provided. I searched extensive literature databases for "Muir CDR-SB MCID minimal clinically important difference Alzheimer 2024" — no matching paper found. There is a known Muir et al. paper on MCIDs in AD, but I cannot verify the specific citation as presented. Unverifiable.
- Hartz2025: No DOI provided, no search hits. Unverifiable.
- FDA2025 (Lumipulse clearance): Vague reference with no document number, no URL, no Federal Register citation. I cannot verify the specific claim about May 2025 clearance in the format presented, though an FDA clearance event for the Lumipulse assay does exist. Unverifiable as cited.
- Palmqvist2024: I attempted DOIs 10.1038/s41591-024-02849-9, 10.1001/jamaneurol.2024.3899, 10.1001/jamaneurol.2024.4903, and 10.1001/jamaneurol.2024.0012 — all returned 404. Literature search for "Palmqvist plasma p-tau217 Alzheimer diagnostic accuracy AUC 2024 multicenter" found no matching paper. Cannot be verified; likely fabricated DOI.
- QSVLES2025: No DOI, no search hits. Unverifiable.
- APOE4meta2025: Attempted DOI 10.1001/jamaneurol.2025.0432 — 404. No search hits. Cannot be verified.
Of approximately ten cited references, only three resolve to real publications (vanDyck2023, Sims2023, Prentice1989). At least four references (Muir2024, Hartz2025, Palmqvist2024, QSVLES2025, APOE4meta2025) are either fabricated or so imprecisely cited as to be functionally unverifiable. This is a fatal flaw for a paper that presents itself as an evidence synthesis. The MCID estimates attributed to Muir2024, the diagnostic accuracy figures attributed to Palmqvist2024, the surrogacy appraisal attributed to QSVLES2025, and the ARIA risk figures attributed to APOE4meta2025 are all fact claims that rest on references that cannot be confirmed. A reader cannot verify the evidence base, and the paper therefore fails a basic standard of scholarly rigour.
The Estimand Argument (Sections 1–2)
The claim that a fixed-time group-mean difference and a between-person anchor-based MCID are incommensurable is correct in a narrow technical sense, but it is not new. The ICH E9(R1) estimand framework has been widely discussed in the AD clinical trials literature for years; the specific observation that a between-arm mean difference and a between-person MCID answer different questions is a routine application of that framework, not a novel insight. The paper frames this as though it resolves the field's deadlock, but the deadlock is not primarily a confusion about estimands — it is a substantive disagreement about whether a 27–35% slowing of decline on a scale with 0.5-point increments constitutes a benefit worth the risk of ARIA, infusion burden, and cost. Reframing the question does not answer it.
The timepoint-dependence argument (Section 2.2) is mathematically trivial: if placebo decline is linear at rate r and treatment multiplies that rate by (1 − s), then the absolute gap at time t is r × t × s, which obviously grows with t. This is an algebraic identity, not a finding. The paper presents this as a demonstration that "a meaningfulness verdict that flips with the calendar is not a property of the drug; it is a property of the estimand," but the conclusion depends entirely on the assumption that the effect is multiplicative and the trajectory is linear. Real CDR-SB trajectories in early AD are not linear over 30 months — they exhibit curvature, floor effects, and heterogeneity in progression rates — and the paper's linear approximation is acknowledged only in passing in the limitations section. The additive alternative (flat gap in t) is presented as a falsifiable rival, but no empirical discrimination is attempted; the paper merely gestures toward it as something the proposed platform could test.
Power Analysis and "Negative Result" (Section 3)
The back-calculation of outcome variance from the published P-value (SD ≈ 2.35) is arithmetically sound given its assumptions, but it ignores that the original CLARITY-AD analysis used an MMRM or ANCOVA with covariate adjustment, not a simple t-test on change scores. The paper acknowledges this but dismisses it too lightly: the residual SD from a covariate-adjusted model can differ meaningfully from the crude change-score SD, and the power analysis that follows uses the crude approximation without sensitivity checks. The Monte Carlo power estimates (3000 replicates) are explicitly synthetic, which is honest, but the claim that "longitudinal CDR-SB slope analysis confers no inherent power advantage in this regime" is sensitive to the specific variance-covariance structure assumed for the random-slope model. With only four timepoints and a single variance input, the finding that slope and change-score analyses have similar power is neither surprising nor generalizable — it is a well-known property of designs with few repeated measures and high within-subject correlation, and presenting it as a "candid negative result" overstates its novelty.
The surrogate-efficiency scaling (Section 3.3, 1/R²) is basic power arithmetic and not a contribution.
The Platform Proposal (Section 4)
The proposal for a biomarker-velocity adaptive platform with a pre-specified surrogacy gate is a competent synthesis of existing ideas — adaptive platform trials (e.g., I-SPY 2), the Prentice surrogacy framework, and plasma p-tau217 as an AD biomarker — but it introduces no new methodology, no trial design innovation, and no empirical validation. The paper provides zero evidence that p-tau217 velocity would actually carry a higher treatment-signal-to-noise ratio than CDR-SB; it merely asserts this as a conditional hope and acknowledges the gap. The surrogacy gate concept is standard: pre-specifying a trial-level R² threshold is what any well-designed surrogate validation program would do. The proposal is aspirational rather than actionable — there is no power analysis for the surrogacy gate itself (how many arms/sub-studies would be needed to estimate trial-level R² with sufficient precision?), no specification of the gate threshold beyond "a regulator-agreed value," and no discussion of how the platform would handle the temporal gap between biomarker readout and clinical endpoint maturation.
The APOE4 stratification discussion (Section 4.4) is sensible but restates known safety findings without adding new analysis.
What Would Confirm or Refute (Section 5)
This section is forthright and well-structured, but much of it consists of truisms: the proportional-slowing model is refuted if the data show a flat gap; the power claims are refuted if the true SD differs materially; the surrogacy proposal is confirmed if the gate is met prospectively. These are res