This paper addresses immortal-time bias in observational studies of checkpoint blockade immunotherapies, a well-known but still common pitfall. The authors propose a framework comprising a checklist, an analytical derivation of the bias, and worked corrections using summary statistics from published studies. Overall, the paper is well-structured and clearly written, making the concepts accessible to a broad audience.
Strengths: The topic is timely and clinically relevant. The checklist is practical and can help researchers avoid misinterpretation of observational data. The derivation linking bias to treatment-initiation delay and baseline hazard is a useful contribution, and the worked examples effectively show how even modest delays can inflate the apparent benefit. The approach of using only published summary data, without accessing individual records, is novel and broadens the paper's applicability.
Weaknesses: The main limitation is that the methodological corrections presented—landmark analysis and time-varying-exposure Cox models—are not new. They have been standard tools for decades. The paper’s contribution is therefore incremental, applying these methods to a specific therapeutic area and providing a formula for bias magnitude. However, the formula’s derivation assumes a constant baseline hazard and a simplified delay distribution; the sensitivity of results to these assumptions is not thoroughly examined. The lack of empirical validation with real patient data (or even simulated data) weakens the claims about the magnitude of bias correction in practice. Additionally, the paper does not engage with the target trial emulation framework, which provides a unified, causal approach to avoiding immortal-time bias and other biases, and would have strengthened the methodological discussion.
Detailed Comments:
- Introduction: Clear motivation, but could better situate the work within the existing literature on immortal-time bias beyond cancer.
- The Bias, Formally: The derivation is elegant, but the notation could be more clearly defined for non-specialist readers.
- Detecting Vulnerable Designs: The checklist is useful, but it is essentially a restatement of known vulnerability patterns. Adding a decision tree or flowchart could improve usability.
- Corrections: The discussion of landmark analysis and time-varying models is accurate but lacks a critical comparison; for example, when one might be preferred over the other.
- Worked Examples: These are compelling but rely on assumptions about treatment-initiation delays that are not well-justified. The authors should discuss how robust the conclusions are to alternative plausible delay distributions.
- Limits and Prospective Validation: The limitations are acknowledged, but the paper would benefit from a more extensive discussion of what exactly a prospective validation would involve and the challenges.
Recommendation: I recommend acceptance conditional on minor revisions. The authors should expand the discussion of limitations, particularly regarding the sensitivity of bias estimates to assumptions, and clarify the novel aspects relative to existing methodological literature. Incorporating a brief simulation or additional sensitivity analyses would greatly strengthen the paper, though it is not strictly necessary for a methodological note. Minor clarifications in notation and a more structured comparison of correction methods would also improve the manuscript.