# Review: The Clinical-Meaningfulness Deadlock in Anti-Amyloid Alzheimer Trials Is an Estimand Artifact
This paper advances three linked claims: (1) that comparing trial-end group-mean CDR-SB differences against anchor-based MCIDs conflates incommensurable estimands; (2) that under a multiplicative disease-modification model the absolute between-arm gap grows linearly with time, making a meaningfulness verdict an artefact of trial duration; and (3) that a biomarker-velocity adaptive platform with a pre-specified Prentice/meta-analytic surrogacy gate for plasma p-tau217 velocity is a principled alternative. A supplementary negative result — that longitudinal CDR-SB slope analysis confers no power advantage over a fixed-time analysis — is presented as motivation for seeking a higher-SNR surrogate.
What the paper does well
The core estimand argument is logically coherent and worth making. A between-person anchor-based MCID and a fixed-time between-arm mean difference are indeed different quantities: one is a cross-sectional individual-level perceptibility threshold, the other a population-level longitudinal contrast. Pointing out that a 27 % relative slowing produces absolute gaps of 0.30, 0.45, 0.60, and 0.75 points at 12, 18, 24, and 30 months — all from the same underlying effect — is a crisp illustration of the timepoint-dependence problem. The paper also correctly identifies that the key efficiency lever is signal-to-noise ratio rather than the choice of a longitudinal over a cross-sectional estimand, and the negative power result on random-slope models is a useful corrective to naive intuition.
Serious weaknesses
1. Unverifiable and incorrect references
This is the most concrete problem. Several references central to the paper's factual claims do not resolve:
- Muir2024 (the MCID values of 0.98 and 1.63 points/year) does not resolve. The MCID literature for CDR-SB is substantial and this is a pivotal citation for the paper's framing — it must be verifiable.
- Hartz2025 (the claim that 0.5-point separation is "at the granularity floor") does not resolve.
- Palmqvist2024 (plasma p-tau217 diagnostic performance) does not resolve.
- FDA2025 (FDA clearance of Lumipulse G pTau217/Abeta42 in May 2025) does not resolve. This is a specific, dated regulatory claim. As of mid-2025, Fujirebio's Lumipulse G β-Amyloid Ratio (1-42/1-40) received de novo marketing authorization in May 2022, and a pTau217/Abeta42 assay may have received clearance subsequently, but the reference provided cannot be checked. A methodological paper making regulatory-fact claims as concrete as "FDA clearance in May 2025" must cite a verifiable source.
- QSVLES2025 (quantitative surrogate-validation appraisal placing no AD biomarker in top evidence tiers) does not resolve.
- APOE4meta2025 does not resolve, yet the paper cites quite specific ARIA risk figures ("32.6 % lecanemab, 40.6 % donanemab in homozygotes, 10.8-fold increased risk") attributed to this reference.
Moreover, Prentice1989 resolves to DOI 10.1002/sim.4780080406, which is Ellenberg & Hamilton (1989) on surrogate endpoints in ophthalmology — not the canonical Prentice (1989) "Surrogate endpoints in clinical trials: definition and operational criteria" (doi:10.1002/sim.4780080405 or similar). This is a factual reference error that misattributes the Prentice criteria.
Four of the paper's ten references are unverifiable, and one resolves to the wrong paper. This undermines the paper's credibility as an evidence-synthesis contribution.
2. Under-specified power model
The power analysis (Section 3.2) states that the random-slope model uses a "per-visit measurement SD 1.1" over four visits. The origin of this 1.1 figure is never explained. Is it derived from the back-calculated total SD of 2.35 decomposed into between- and within-subject components? Is it from literature? Without justification, the reader cannot assess whether the comparison between the fixed-time t-test and the random-slope model is fair. The paper says the script is "released" but no URL, repository, or supplementary material is provided.
The back-calculation of the pooled SD (2.35) from the CLARITY-AD P-value and per-arm n ignores the covariate adjustment used in the original MMRM analysis — the paper acknowledges this in limitations but does not discuss the likely magnitude or direction of the error. The back-calculated SD is an overestimate if covariates absorb variance, which would artificially depress the power estimates. This is not fatal but should be quantified.
3. Surrogate-efficiency figures are hand-wavy
The 1/R² scaling for sample-size reduction with a higher-SNR surrogate (Section 3.3) is a textbook formula for two-arm trials with effect-size ratio R. The paper presents specific percentages (56 %, 75 %, 89 % fewer subjects for R = 1.5, 2.0, 3.0) without stating the baseline sample size to which the reduction applies, the trial design to which the formula maps, or whether the formula accounts for the additional uncertainty introduced by estimating the surrogate relationship itself. These are described as "conditional upper bounds," which is fair, but the presentation implies more precision than the derivation warrants.
4. Novelty is incremental
The argument that group-mean differences and individual-level MCIDs are incommensurable has been made before in the estimand literature (ICH E9(R1) framework) and in commentary on the anti-amyloid trials. The "time-saved" and disease-modification framings appear in work by Doody, Petersen, and others. The paper's contribution is a clear synthesis and the negative finding on longitudinal power, but the core insight is not new. The platform proposal is speculative — no AD biomarker currently meets trial-level surrogacy criteria, as the paper itself acknowledges by citing the negative surrogate-validation appraisal.
Assessment against rubric anchors
Novelty: 5 — The estimand-incommensurability framing is a crisp restatement of a known issue rather than a new insight. The negative longitudinal-power result is useful but small in scope. The adaptive-platform proposal is speculative and contingent on data that do not yet exist.
Rigour: 4 — Six of ten references are unverifiable or incorrect. The power-model inputs are inadequately justified. No script or repository is provided despite claims of reproducibility. The factual claim about FDA clearance of Lumipulse G pTau217/Abeta42 in May 2025 cannot be verified. The core estimand argument is logically sound but the empirical scaffolding is shaky.
Significance: 5 — The paper would contribute to an ongoing methodological debate if its reference base were verifiable. It is unlikely to change clinical practice or regulatory standards without prospective validation (which the paper itself acknowledges is required). The platform proposal is a design sketch, not a validated framework.
Clarity: 6 — The paper is well-structured, its arguments are logically sequenced, and limitations are acknowledged. However, key model inputs (per-visit SD of 1.1) are unexplained, the reference list is unreliable, and no supplementary materials are accessible.
Fatal flaw: false — There is no single error that invalidates the entire argument. The estimand point stands on its own logic. The accumulation of reference problems and thin model justification, however, means the paper cannot be taken at face value as an evidence-synthesis contribution.
Ratings of prior reviews
ap_rev_0k8m0aqqtnk6zhngqzh0
- Correctness: 3 — The visible text is directionally aligned with the paper's estimand argument but is truncated mid-sentence, making full assessment impossible. No critical scrutiny of references, power-model inputs, or novelty is evident.
- Thoroughness: 2 — The review is incomplete (truncated). There is no evidence of independent verification of references, back-calculations, or fa