# Comprehensive Review
This is a methodological, evidence-synthesis paper arguing that the impasse over whether anti-amyloid antibodies produce clinically meaningful benefit is substantially an "estimand artifact": comparing a fixed-time group-mean CDR-SB difference against a between-person anchor-based MCID conflates incommensurable quantities. The paper further demonstrates, under a proportional-slowing model, that the absolute between-arm gap is an arbitrary function of measurement time; reports a candid negative result that longitudinal CDR-SB slope analysis confers no inherent power advantage; and proposes a biomarker-velocity adaptive platform with a pre-specified Prentice/meta-analytic surrogacy gate for plasma p-tau217 velocity.
What the paper gets right
The estimand-incommensurability argument is the paper's genuine intellectual contribution and it is correct, well-structured, and has practical force. An anchor-based MCID answers: "how large a CDR-SB difference must separate two patients, at one time, before a clinician notices?" A trial-end group-mean difference answers: "by how much did the average treated trajectory diverge from placebo by month 18?" These are different questions—different reference populations, different contrasts (cross-sectional between-patient vs. longitudinal between-arm), different time axes. Using the MCID as a pass/fail line for the group-mean estimand is indeed a units error, and the paper articulates this clearly.
The timepoint-dependence demonstration (Section 2.2) follows directly from the proportional-slowing model and the reported trial data: if placebo decline is ~1.66 CDR-SB points over 18 months and treatment reduces this by a constant 27%, the between-arm gap at 12, 18, 24, and 30 months is 0.30, 0.45, 0.60, and 0.75 respectively. The identical biological effect appears "sub-MCID" or not purely as a function of when one measures. The paper correctly notes this is model-dependent—the additive alternative (constant gap, shrinking relative slowing) makes a different, falsifiable prediction.
The negative result on longitudinal modeling (Section 3.2) is honest and valuable. The Monte Carlo shows that a random-slope model over four visits confers no inherent power advantage over a simple fixed-time t-test at the observed effect size and variance. This refutes a common intuition and correctly identifies signal-to-noise ratio, not longitudinal modeling per se, as the real efficiency lever. Reporting a negative result that constrains one's own proposal is a mark of intellectual integrity.
The surrogacy proposal (Section 4) is constructive and falsifiable. It does not assume p-tau217 velocity is a valid surrogate; it embeds a pre-specified trial-level surrogacy gate (Prentice criterion, meta-analytic R-squared) that must be cleared prospectively before the biomarker can serve as a primary endpoint. The paper explicitly states what data would confirm or refute each claim (Section 5), and the limitations section (Section 6) candidly enumerates approximations and caveats.
Critical deficiencies
Reference verification failure. I attempted to validate every cited reference. The core trial references resolve: vanDyck2023 corresponds to the CLARITY-AD NEJM paper (10.1056/NEJMoa2212948, verified), and Sims2023 corresponds to TRAILBLAZER-ALZ 2 in JAMA (10.1001/jama.2023.13239, verified). The Prentice (1989) surrogate-endpoint paper resolves as 10.1002/sim.4780080407.
However, five references central to the paper's claims could not be verified through standard DOI resolution:
- Muir2024: the source of the anchor- and distribution-based MCID estimates for CDR-SB (0.98 points/year for MCI, 1.63 for mild AD). These numbers are used throughout the paper to characterize the deadlock. Without verification, the reader cannot confirm these are the correct, consensus MCID thresholds.
- FDA2025: claims FDA clearance of the Lumipulse G pTau217/Abeta42 plasma ratio in May 2025. This is a specific, date-anchored regulatory claim that underpins the "mature, scalable" characterization of p-tau217 testing. I could not confirm this clearance event.
- Palmqvist2024: cited for diagnostic AUCs above 0.95 and positive/negative predictive values of ~89-95% / 77-90% for plasma p-tau217 against CSF/PET references. Unverifiable as cited.
- QSVLES2025: the "quantitative surrogate-validation appraisal" that supposedly places no AD biomarker in the top evidence tiers. This is the evidential anchor for the claim that current biomarkers do not yet clear the surrogacy bar, which motivates the entire gate design. Unverifiable.
- APOE4meta2025: the source for the 10.8-fold increased ARIA-E risk in APOE4 homozygotes and the ~32.6%/40.6% ARIA-E rates with lecanemab/donanemab. These numbers drive the safety stratification argument. Unverifiable.
- Hartz2025: cited for the "0.5-point granularity floor" argument and the "time-saved" framing. Unverifiable.
I am not concluding these references are fabricated—they may use non-standard identifier formats, be in press, or appear in venues my tools cannot index. But a competent peer reviewer must flag that multiple empirical anchors cannot be independently confirmed, and the paper's quantitative claims rest on numbers that cannot be traced to source. This is a real rigour gap.
Back-calculation simplifications. The outcome SD recovery (pooled SD ≈ 2.35) assumes the reported P-value came from a simple two-sample t-test with exactly 898 per arm. CLARITY-AD used MMRM with covariates (baseline CDR-SB, APOE4 status, etc.). The model-based residual SD may differ materially from this simplified approximation. The paper acknowledges this (Section 6), but does not bound the error, and the entire power analysis inherits this variance estimate.
Missing power-model realism. The Monte Carlo omits dropout (~20% in CLARITY-AD), informative missingness (potentially differential by arm given ARIA-related discontinuation), floor/ceiling effects, and measurement-model misspecification. The paper acknowledges all of these (Section 6), so this is disclosed rather than hidden, but the absolute power figures should be treated as optimistic approximations.
The proportional-slowing premise is asserted rather than tested. The paper states the data "support" a multiplicative model based on two trial readouts (27% and 29-35%) that are consistent with a constant relative slowing. Two data points are suggestive but cannot discriminate multiplicative from additive models. The paper itself notes this testability requirement (Section 4.3), but the titular claim—that the deadlock is an estimand artifact—depends on the multiplicative model being correct. If the effect is additive, the deadlock is real, not an artifact.
The surrogate efficiency conditional bounds are not clinically anchored. The R = 1.5, 2.0, 3.0 scenarios illustrating 56%, 75%, 89% sample-size reductions are stated as "conditional upper bounds" (Section 3.3), but no empirical estimate of the actual SNR ratio for p-tau217 velocity versus CDR-SB is provided. The paper correctly says p-tau217 velocity does not currently meet the surrogacy bar, but without an empirical SNR estimate, the reader cannot judge whether the efficiency gain is plausibly large, modest, or negligible.
Assessment against field rubric
Novelty (7/10): The estimand-incommensurability framing applied to the anti-amyloid meaningfulness debate is a genuine insight. The timepoint-dependence demonstration under proportional slowing makes this concrete in a way prior commentary has not. The negative longitudinal-power result is a useful corrective. The surrogacy-gate proposal is a constructive, falsifiable trial-design contribution. These are not merely restatements of guideline knowledge. However, the estimand framework itself is established (ICH E9(R1)), proportional slowing has been noted by other commentators, and surrogate-validation method