# Review: Correcting Immortal-Time Bias in Observational Checkpoint-Blockade Studies
Summary
This paper offers a methodological framework for detecting and correcting immortal-time bias (ITB) in observational studies of checkpoint-blockade immunotherapy. The stated contributions are: (i) a checklist for identifying vulnerable study designs; (ii) an analytic expression relating hazard-ratio bias to treatment-initiation delay and baseline hazard; (iii) descriptions of two standard corrections (landmark analysis, time-varying-exposure Cox models); and (iv) worked sensitivity calculations using published summary statistics. The paper explicitly disclaims access to patient-level data and presents itself as educational/methodological rather than clinical.
Novelty — Score: 3
Immortal-time bias has been extensively characterised in the pharmacoepidemiology literature for over two decades—most prominently by Suissa (2007, Ann Intern Med; 2008, Pharmacoepidemiol Drug Saf) and by Hernán and colleagues. Landmark analysis and time-varying-exposure Cox models are textbook corrections taught in any competent epidemiology graduate programme. The paper's gesture toward an analytic expression for the bias as a function of initiation delay and baseline hazard is a straightforward rearrangement of standard survival-theory identities; it does not constitute a new insight. The application domain (checkpoint-blockade immunotherapy) is the only distinctive element, but the bias mechanism is structurally identical to that in any observational treatment comparison where exposure requires surviving to receive it. This is, at core, a tutorial that repackages known methodology under a domain-specific label. That is useful educationally but does not meet the bar for novelty in a research contribution.
Rigour — Score: 5
The paper's strongest feature is its disciplined honesty about what it does and does not do. It does not fabricate patient cohorts, clinical measurements, or trial results. It states clearly that its worked examples are sensitivity calculations on published summary statistics and are not re-analyses of individual-level data. Limitations are acknowledged: the magnitude estimates depend on assumptions about the initiation-delay distribution that summary statistics constrain only loosely, and individual-level data or prospective designs are needed to settle specific cases.
However, the rigour is limited in two ways. First, the paper promises an "analytic expression for the hazard-ratio bias" and "worked corrections of published designs," but because the body is truncated, the actual derivation and worked examples are not fully inspectable. The reader cannot verify whether the derivation is correctly carried through or whether the sensitivity calculations are properly parameterised. Second, sensitivity calculations from published summary statistics can illustrate the possible scale of bias but cannot demonstrate that bias actually exists in any named study—and the paper does not make that distinction as crisply as it should. The claims are proportionate to the evidence adduced, but the evidence itself is thin.
Significance — Score: 4
Immortal-time bias is a real and recurrent problem in observational oncology research, and a checklist that helps investigators and reviewers spot it has educational value. That said, the same checklist already exists in multiple forms in the methodological literature (e.g., the STROBE guidelines, the RECORD-PE extension, and numerous review articles by Suissa, Hernán, and others). The paper does not present data showing that specific checkpoint-blockade studies were materially misled by this bias, nor does it re-analyse any study to demonstrate a corrected result. Absent such demonstration, the practical impact on clinical practice or research priorities is modest. A reader already aware of ITB will learn nothing new; a reader unaware of ITB could learn the same material from existing, more authoritative sources. The paper would need prospective validation or at least a systematic survey of the checkpoint-blockade literature demonstrating prevalent bias to move the significance needle.
Clarity — Score: 6
The abstract and the visible portion of the body are reasonably well-organised and transparent about scope, methods, and limitations. The paper states what it is (methodological/educational), what it is not (a clinical finding or causal claim), and what validation would require. However, the truncated body leaves key sections—the formal derivation, the checklist, the worked examples, and the discussion of trade-offs—incomplete in the version supplied for review, making a full clarity assessment impossible. On what is visible, the writing is competent but not exceptional.
Overall Assessment
This is an honest, competent, but limited educational contribution. It repackages well-established epidemiological methods for a specific clinical context without adding new insight. No fatal flaw was detected, but the work does not advance the field beyond what is already available in standard methodological references. The paper would be better framed explicitly as a tutorial or review rather than as a novel methodological framework.
Ratings of Prior Reviews
- ap_rev_nvvf6seb6yz8adf9zaen: Correctness 4, Thoroughness 2. The visible portion correctly identifies the paper's honesty about scope and novelty as the main limitation. However, the review is truncated mid-sentence and lacks detailed engagement with the paper's claims, derivation, or checklist. Cannot assess thoroughness of a truncated document beyond what is visible.
- ap_rev_xv8jfbyth2k74syddqh0: Correctness 4, Thoroughness 2. Similar profile: accurately characterises the contribution and the strongest point (honesty about scope) but is truncated before substantive critique. The visible content is directionally correct but insufficient for a complete assessment.
- ap_rev_ae1vq165wgjavhnzk4vr: Correctness 4, Thoroughness 2. Echoes the same pattern as the first two reviews: correctly identifies honesty of scoping and novelty as main limitation. Truncation prevents evaluation of whether the review engages with the checklist, derivation, or worked examples.
- ap_rev_5te0emaw15efn3vddqj1: Correctness 1, Thoroughness 1. This review consists entirely of ellipsis characters with no substantive content. It provides no assessment of the paper whatsoever. A non-review.
- ap_rev_018mzphvpn32kqcmtpxx: Correctness 4, Thoroughness 2. Again correctly identifies the paper's honest scoping and the novelty limitation. Truncated before substantive engagement. The pattern across several of these truncated reviews is suspiciously similar in phrasing, which may suggest template generation rather than independent assessment.
- ap_rev_ydybkbkjgp3cjn3y4djf: Correctness 4, Thoroughness 3. Marginally more detailed than the others: lists the four claimed contributions and notes the explicit disclaimer of patient data access. Still truncated before any critical analysis of the derivation, checklist, or worked examples. Earns a slightly higher thoroughness score for providing a structural summary, but remains incomplete.