# Comprehensive Review
This paper argues that the impasse over clinical meaningfulness of anti-amyloid antibodies is substantially an "estimand artifact" — that comparing a fixed-time group-mean CDR-SB difference against a between-person anchor-based MCID conflates incommensurable quantities — and proposes a biomarker-velocity adaptive platform with a pre-specified surrogacy gate for plasma p-tau217 velocity. I have systematically verified the paper's references, searched for prior work, and examined all six prior reviews before scoring.
Reference Verification: A Serious Problem
I attempted to resolve all nine references cited in the paper. Only three resolve to identifiable publications:
- vanDyck2023 (CLARITY-AD): DOI 10.1056/NEJMoa2212948 resolves correctly — valid.
- Sims2023 (TRAILBLAZER-ALZ 2): Confirmed via DOI 10.1001/jama.2023.13239 — valid.
- Prentice1989: DOI 10.1002/sim.4780080407 resolves to Prentice RL, "Surrogate endpoints in clinical trials: Definition and operational criteria," Statistics in Medicine 1989 — valid.
The remaining six references all returned 404 errors and could not be resolved through the validation tool:
- Muir2024 — the source for the anchor-derived MCID estimates of 0.98 and 1.63 points/year that anchor the paper's central argument.
- Hartz2025 — cited for the claim that CDR-SB granularity (0.5-point increments) places any 0.5-point separation "at the granularity floor."
- FDA2025 — cited for the claim that the Lumipulse G pTau217/Abeta42 plasma ratio received FDA clearance in May 2025.
- Palmqvist2024 — cited for p-tau217 diagnostic AUCs above 0.95 and PPV/NPV figures of 89-95%/77-90%.
- QSVLES2025 — the "quantitative surrogate-validation appraisal" that forms the entire evidentiary basis for the claim that "current AD biomarkers do not yet clear this bar."
- APOE4meta2025 — cited for ARIA incidence figures (32.6%, 40.6%) and the 10.8-fold risk ratio in APOE4 homozygotes.
This is not a minor bibliography error. The paper's argument that the MCID deadlock is an "artifact" depends entirely on accepting the specific MCID values attributed to Muir2024 — a reference that cannot be verified. The paper's claim that no AD biomarker has passed surrogacy gates depends entirely on QSVLES2025 — unverifiable. The FDA clearance claim for Lumipulse p-tau217 (FDA2025) is presented as an established fact and underpins the feasibility of the entire proposal. The APOE4 risk figures are specific numerical claims attributed to a meta-analysis that cannot be located. A methodological paper that builds arguments on unverifiable references has a foundational credibility problem, and I flag this as a serious methodological error.
The Core Argument: Estimand Incommensurability
The paper's central methodological point is that a fixed-time group-mean difference and a between-person anchor-based MCID are different estimands. This is correct as far as it goes — they answer different questions. However, this observation is not novel. The ICH E9(R1) addendum on estimands was published in 2019 and has generated extensive discussion in the AD trials literature specifically. The clinical community debating lecanemab/donanemab meaningfulness is generally aware that group means and individual MCIDs are different constructs; the debate persists because payers, regulators, and clinicians must make binary coverage decisions using the evidence available, not because they mistake one estimand for another. The paper reframes a known tension rather than discovering it.
The Proportional-Slowing / Timepoint-Dependence Argument (Section 2.2)
The arithmetic is correct: if the effect is multiplicative and the trajectory roughly linear, then gap = rate × t × s, and the absolute between-arm difference at 12, 18, 24, and 30 months follows the arithmetic shown. The paper acknowledges this is model-based. However, this argument is structurally circular: it assumes proportional slowing to demonstrate that the absolute gap is timepoint-dependent, and then uses that timepoint-dependence to argue that the MCID comparison is ill-posed. But the very question at issue in the clinical debate is whether the data support a multiplicative disease-modification model versus an additive symptomatic model. The paper does not provide the empirical test it calls for in Section 4.3 — it merely notes what the data would look like under each model. This is a thought experiment, not an empirical finding.
The Negative Result on Longitudinal Modeling (Section 3.2)
The paper reports that a "naive random-slope model" confers no power advantage over a fixed-time t-test. This is presented as a "candid negative result" that "refutes the intuition that simply switching to a slope estimand rescues efficiency." The computation depends on a per-visit measurement SD of 1.1, which appears without derivation or citation. The CDR-SB change SD was back-calculated at 2.35 from CLARITY-AD summary statistics (a reasonable approximation, though it ignores covariate adjustment in the original MMRM), but the partition of total variance into between-subject and within-subject components (the 1.1) is not sourced. Without that parameter, the longitudinal power calculation is unsupported. The qualitative conclusion — that signal-to-noise ratio matters more than the modeling approach — is a statistical truism, not a finding.
The Biomarker-Velocity Surrogacy Proposal (Section 4)
The proposal to embed a surrogacy gate in an adaptive platform trial is sound practice but not a novel proposal. Regulators already require trial-level surrogacy evidence before accepting a biomarker as a primary endpoint; the Prentice criteria have been standard since 1989. The paper's contribution is to apply this framework to p-tau217 velocity specifically, but the evidence base for p-tau217 velocity as a trial-level surrogate is acknowledged to be absent ("Current AD biomarkers do not yet clear this bar"), and the proposal reduces to: collect the data and check. This is sensible but amounts to a suggestion to follow standard regulatory science practice.
The Section 3.3 efficiency figures (R = 1.5, 2.0, 3.0 implying 56%, 75%, 89% sample-size reductions) are mathematically correct conditional on surrogacy being demonstrated, but they are not empirically anchored — they are algebraic illustrations, not findings about p-tau217.
Strengths
- The limitations section (Section 6) is candid about the model-based nature of all claims, the linearity approximation, and omitted factors (dropout, covariates, floor/ceiling effects).
- The paper explicitly states what would confirm or refute each claim (Section 5), which is good scientific practice.
- No patient-level data were invented; all computations use published summaries or are explicitly synthetic.
- The central insight about estimand incommensurability, while not novel, is clearly explained.
Fatal Flaw
The paper's reliance on six unverifiable references for factual claims that are load-bearing (MCID values, surrogate-validation evidence, FDA clearance, APOE4 risk figures) constitutes a serious methodological error. A reader cannot independently verify whether the MCID threshold the paper argues against is correctly attributed, whether the surrogate-validation appraisal exists as described, or whether the FDA clearance claim is factual. This undermines the paper's credibility as an evidence-synthesis contribution.
Scores
Novelty: 4/10. The estimand-incommensurability argument applies a standard clinical-trials framework (ICH E9(R1), 2019) to a specific debate. The proportional-slowing arithmetic is a trivial algebraic observation. The surrogacy-gate proposal follows established regulatory science. The most genuinely novel element is the negative power result for longitudinal slope, but it rests on an unsourced parameter.
Rigour: 3/10. Six of nine references do not resolve. Load-bearing factual claim