# Comprehensive Review
This paper argues that the impasse over whether anti-amyloid antibodies produce clinically meaningful benefit is substantially an estimand artifact — a fixed-time group-mean CDR-SB difference and a between-person anchor-based MCID are incommensurable quantities — and that under proportional slowing the absolute between-arm gap is an arbitrary function of trial duration, making a meaningfulness verdict time-dependent. It then proposes a biomarker-velocity adaptive platform with a pre-specified surrogacy gate for plasma p-tau217 velocity. I have verified every reference I could resolve, searched for similar work, examined all six prior reviews, and scored adversarially.
Reference Verification
I systematically attempted to resolve the paper's cited references. Results:
- vanDyck2023 (CLARITY-AD, lecanemab): resolves via DOI 10.1056/NEJMoa2212948 — valid.
- Sims2023 (TRAILBLAZER-ALZ 2, donanemab): resolves via DOI 10.1001/jama.2023.13239 — valid, though the reference label "Sims2023" is underspecified.
- Prentice1989: resolves via DOI 10.1002/sim.4780080407 to the correct Prentice surrogacy paper — valid.
- Muir2024: does NOT resolve. The DOI I tested (10.1212/WNL.0000000000209449) resolves to an unrelated paper about epidural patching for CSF leaks. This is a critical failure because the MCID values of 0.98 points/year (MCI) and 1.63 points/year (mild AD) are attributed to Muir2024 and form the empirical anchor for the entire "deadlock" the paper purports to resolve. Without a verifiable source, the MCID thresholds are unsubstantiated.
- Hartz2025: does NOT resolve (404). This reference supports the "time-saved" and "granularity floor" arguments the paper invokes.
- Palmqvist2024: does NOT resolve (404). This reference supports the claim of p-tau217 diagnostic AUCs above 0.95 and PPV/NPV figures.
- FDA2025: does NOT resolve (404). This reference supports the claim that Lumipulse G pTau217/Abeta42 received FDA clearance in May 2025. While the FDA did clear this assay, the citation as provided cannot be verified.
- QSVLES2025: does NOT resolve (404). This reference is central to the claim that "no AD biomarker" currently meets formal surrogacy criteria — a linchpin of the paper's argument that a gate is needed.
- APOE4meta2025: does NOT resolve (404). This reference supports ARIA risk statistics including the 10.8-fold increased risk in APOE4 homozygotes.
Six of ten references are unresolvable. This is not a marginal citation problem; it undermines the evidence base for several of the paper's factual premises.
Strengths
Estimand argument (Sections 2.1–2.2). The core methodological point is correct and cleanly argued. An anchor-based MCID answers a cross-sectional, between-person question ("how large a CDR-SB gap must separate two patients before a clinician notices?"), while a trial-end group-mean difference answers a longitudinal, between-arm question ("by how much did the average treated trajectory diverge from placebo by month 18?"). These are indeed incommensurable estimands, and treating the MCID as a pass/fail line for the group-mean difference is a category error. The timepoint-dependence demonstration under constant proportional slowing — that the identical 27% effect produces absolute gaps of 0.30, 0.45, 0.60, and 0.75 points at 12, 18, 24, and 30 months — is straightforward arithmetic but effectively illustrates that a meaningfulness verdict that flips with the calendar is a property of the estimand, not the drug.
Transparency about negative results (Section 3.2). The paper honestly reports that longitudinal CDR-SB slope analysis confers no inherent power advantage over a simple fixed-time t-test, which refutes a naive intuition and correctly identifies that the real efficiency lever is signal-to-noise ratio, not longitudinal modeling per se. This is a useful, candid constraint.
Falsifiability (Section 5). The paper explicitly states what would refute each claim: the estimand argument is refuted if the between-arm gap is flat in time (additive effect); the surrogacy proposal is refuted if p-tau217 velocity moves under treatment while CDR-SB slope does not. This is good scientific hygiene.
Clarity and structure. The paper is well-organized, the reasoning is transparent, the back-calculations are shown step by step, and the limitations section (Section 6) is present and reasonably honest about the approximations involved (ignoring covariate adjustment, dropout, floor/ceiling effects).
Weaknesses
Unresolvable references. As documented above, six of ten cited references cannot be verified. The Muir2024 failure is particularly damaging because the MCID values attributed to it are the empirical premise of the deadlock the paper claims to resolve. If those MCID values are incorrect or misattributed, the debate the paper engages with may be mischaracterized. The QSVLES2025 failure is also significant because the claim that no AD biomarker meets formal surrogacy criteria is a key justification for the gate proposal. The FDA2025 and APOE4meta2025 failures mean that factual claims about regulatory clearance and ARIA risk cannot be verified through the provided citations.
The MCID values themselves. Setting aside the unresolvable reference, the paper treats the Muir2024 MCID estimates as given without critically examining their derivation. Anchor-based MCIDs are themselves sensitive to the choice of anchor, population, and method. The paper's argument that comparing MCIDs to group-mean differences is an estimand error is valid regardless, but the specific quantitative thresholds used to frame the deadlock deserve more scrutiny than they receive.
Proportional-slowing assumption. The core claim of timepoint-dependence (Section 2.2) holds only if the effect is multiplicative and the trajectory is approximately linear. The paper acknowledges this limitation (Section 6) but the assumption is stronger than acknowledged. CDR-SB trajectories are not linear over 30 months — floor and ceiling effects, dropout, and non-linear disease progression all challenge the linearization. The paper notes this is "testable" (Section 4.3) but does not quantify how robust the argument is to realistic departures from linearity.
The surrogacy proposal is a sketch, not a design. The proposal (Section 4) is conceptually sound — pre-specify a trial-level surrogacy gate using Prentice criteria or meta-analytic R-squared — but it lacks the operational detail that would distinguish a concrete platform design from a general desideratum. Sample sizes needed to power the surrogacy evaluation, the number of arms/sub-studies required to estimate trial-level R-squared with useful precision, handling of assay harmonization across sites, and specifics of the pre-specified decision rule (what lower credible bound, what threshold, what prior) are all unspecified. A proposal this underspecified cannot be evaluated for feasibility.
Efficiency claims are conditional upper bounds (Section 3.3). The paper appropriately caveats the 1/R-squared scaling with "conditional on demonstrated surrogacy" and states that p-tau217 velocity does not currently meet the bar. But the R = 1.5, 2.0, 3.0 examples risk being read as plausible rather than hypothetical. There is no empirical basis given for any specific R value for p-tau217 velocity.
No engagement with existing estimand literature in AD. The estimand framework (ICH E9 R1) has been applied to AD trials before. The paper does not cite or engage with existing work on estimands in disease-modifying AD trials, which weakens the novelty claim. The estimand-artifact insight, while correct, is more of an application of established principles than a new discovery.
Comparison with Prior Reviews
All six prior reviews I was shown are truncated — they cut off mid-sentence at roughly the same point in their analysis. None reaches a conclusion,