# Comprehensive Review
This paper advances a methodological argument: the impasse over whether anti-amyloid antibodies produce clinically meaningful benefit is "substantially an estimand artifact." It then proposes a biomarker-velocity adaptive platform gated on a pre-specified surrogacy criterion. I have verified key references, searched for prior work, and examined all six prior reviews (all of which were truncated and incomplete; see ratings below).
Reference Verification
Two pivotal trial references are confirmed:
- vanDyck2023 (CLARITY-AD / lecanemab): resolves as DOI 10.1056/NEJMoa2212948, a genuine New England Journal of Medicine publication.
- Sims2023 (TRAILBLAZER-ALZ 2 / donanemab): resolves as DOI 10.1001/jama.2023.13239, a genuine JAMA publication.
The remaining references — Muir2024, Hartz2025, Palmqvist2024, FDA2025, Prentice1989, QSVLES2025, APOE4meta2025 — could not be confirmed through the available verification tool. Prentice 1989 is a widely known classic (the Prentice criteria for surrogacy), and FDA clearance of the Lumipulse pTau217/Abeta42 assay did occur in 2025, so these references are plausible. But the inability to validate Muir2024 (the MCID anchor paper), Hartz2025 (the "time-saved" counterargument), QSVLES2025 (the surrogate-validation appraisal), and APOE4meta2025 (the ARIA risk meta-analysis) is a gap. The paper's quantitative claims do not depend on these being precisely correct — Muir's MCID values are used only to frame the debate, not as inputs to the model — but the unverifiable references weaken the evidence-synthesis claim the paper makes for itself.
The Estimand Argument (Section 2) — Correct but Incomplete
The paper's central methodological point is sound: comparing a fixed-time group-mean between-arm difference against a between-person anchor-based MCID does indeed conflate two estimands that differ in reference population, contrast type, and time axis. This is a well-recognised concern in the estimand framework of ICH E9(R1), which the paper does not cite. The application to the anti-amyloid meaningfulness debate is useful but not deeply novel; clinical trial methodologists have long warned against treating MCIDs as pass/fail lines for group-mean effects.
The timepoint-dependence demonstration (Section 2.2) is arithmetically correct: under constant proportional slowing, the between-arm gap at time t is (rate × t × s), so the same 27% effect produces gaps of 0.30 at 12 months, 0.45 at 18, and 0.75 at 30. This is a valid and vivid illustration. However, the paper overstates the implication. The clinical-meaningfulness debate is not a simple category error. The field uses MCIDs as benchmarks because payers, regulators, and clinicians must decide whether a population-average benefit — at a specific follow-up duration — justifies the risks and costs of treatment. The paper's framing, while rhetorically crisp, does not fully engage with the real-world question of whether 0.45 CDR-SB points at 18 months, given ARIA risks and multi-thousand-dollar annual costs, represents a worthwhile investment. The estimand reframing clarifies the debate but does not settle it.
Power Analysis (Section 3) — Useful but Oversimplified
The back-calculation of SD ≈ 2.35 from CLARITY-AD summary statistics (Δ = 0.45, n ≈ 898/arm, P = 0.00005) is a reasonable approximation, but it assumes a simple two-sample t-test on change scores. CLARITY-AD's actual primary analysis used a mixed model for repeated measures (MMRM) with covariates, so the back-calculated SD conflates residual variance with model-based variance reduction from covariate adjustment. The paper acknowledges this limitation (Section 6) but then uses the 2.35 value throughout the power model, which propagates the approximation error.
The "candid negative result" — that longitudinal CDR-SB slope analysis confers no inherent power advantage — is based on a "naive random-slope model" with measurement SD 1.1 at each visit. This is a single, simplified specification. Real trial analyses use MMRM with unstructured covariance, covariate adjustment, and sometimes model-based imputation for missing data, any of which could alter the efficiency comparison. The conclusion that "the lever is not the time axis per se but the signal-to-noise ratio" is broadly correct but is not a finding discovered by the model; it is a restatement of statistical first principles dressed as a negative result.
The conditional efficiency figures for a higher-SNR surrogate (Section 3.3, scaling as 1/R²) are likewise correct as arithmetic but say nothing specific about p-tau217. The paper honestly acknowledges that p-tau217 velocity does not currently meet the required bar — but then the entire surrogate-proposal section appears to be arguing for something the paper itself admits the evidence does not yet support. This creates an odd tension: the proposal is both the paper's constructive contribution and a structure whose key component (p-tau217 velocity) is explicitly said not to qualify yet.
The Surrogacy Proposal (Section 4) — Underdeveloped
The proposal to embed a Prentice-criteria surrogacy gate inside an adaptive platform is sensible in concept but critically under-specified:
- No sample-size or precision planning: Estimating trial-level R² for surrogacy requires multiple trials or multiple arms with sufficient between-arm variation in treatment effects. The paper gives no indication of how many arms, what sample sizes, or what precision would be needed to clear the gate with adequate power.
- No meta-analytic model specification: The paper invokes "trial-level R-squared" but does not specify whether this would use a Bayesian hierarchical model, a bivariate meta-analysis, or a simpler regression of treatment-effect estimates. The choice materially affects the operating characteristics.
- Unit-of-analysis problem: If a single platform contributes multiple arms, the "trial-level" unit becomes ambiguous. Are arms within a platform independent "trials" for surrogacy evaluation? The paper does not address this.
- No gate threshold specified: The paper says "a regulator-agreed value" for the lower credible bound but proposes no concrete number or calibration rationale. Without one, the gate is a placeholder, not a decision rule.
- Practical feasibility: Running a platform that simultaneously collects CDR-SB (an 18-month endpoint) and p-tau217 velocity (which requires dense sampling) while waiting for surrogacy to be demonstrated is an expensive, long-duration commitment. The paper does not discuss the resource implications or timeline.
The APOE4-stratified safety section (4.4) is well-motivated — ARIA risk is indeed strongly genotype-dependent — but adds little that is not already standard practice or under active regulatory discussion.
What Would Refute the Claims (Section 5) — Admirably Honest
The paper's strongest formal feature is its explicit statement of falsification conditions for each claim. This transparency is unusual in the methodological literature and deserves credit. However, some of the "refutation" conditions are too vague to be operational: "the between-arm CDR-SB gap is flat in time" — flat relative to what tolerance? Over what horizon? A pre-specified equivalence or non-inferiority margin would be needed.
Limitations (Section 6) — Mostly Honest, One Omission
The paper acknowledges the key limitations: the linear approximation to trajectories, the back-calculation approximation ignoring covariate adjustment, the omission of dropout and informative missingness. One important omission: the paper does not discuss whether the proportional-slowing model is actually the best fit to the data. Both CLARITY-AD and TRAILBLAZER-ALZ 2 report the relative slowing at a single timepoint; they do not demonstrate that relative slowing is constant across the 18-month window. The paper's entire Section 2.2 argument depends on this consta