What is landmark analysis?
How landmark analysis avoids immortal time bias, how to choose the landmark, what it costs, its use for dynamic prediction, and what to report.
Landmark analysis compares groups defined by something that happens after follow-up starts, such as response to treatment, a transplant or a fall in a biomarker. It fixes a time after baseline, the landmark, keeps only the people still event-free and under follow-up at that time, groups them by what had happened by then, and starts the clock there. This avoids immortal time bias, and the result describes people who reach the landmark.
Why it is needed
Cancer trials often compared the survival of patients who responded to treatment with that of patients who did not, counting from the start of treatment. A patient has to live long enough to respond, and patients who die early are counted as non-responders, so the comparison favours responders even if response does nothing for survival. Anderson, Cain and Gelber showed this bias and discussed two valid alternatives: the landmark method, which they introduced, and a test due to Mantel and Byar (1974) that counts each patient as a non-responder until they respond and as a responder afterwards (J Clin Oncol 1983;1:710–719).
How it works
Suppose the landmark is 6 months after the start of treatment.
- A patient who died, or whose follow-up ended, before 6 months is excluded.
- A patient who responded at 2 months is a responder.
- A patient who responded at 8 months is a non-responder, because they had not responded by 6 months.
- Follow-up for everyone included starts at 6 months. Survival from there can be compared with Kaplan–Meier curves or a Cox model, using covariates known at 6 months.
Exposure is fixed using only what was known at the landmark, and follow-up starts at the same moment, so nobody is credited with time during which they could not have had the event.
Choosing the landmark
Choose the landmark before analysing the outcomes, on clinical grounds, such as the time of a scheduled response assessment. An early landmark keeps more patients but misclassifies more late responders. A late one classifies exposure better but excludes more early events. When no time is obviously right, repeat the analysis at a few prespecified landmarks to show how much the conclusion depends on the choice (Morgan, 2019). Trying several landmarks and reporting the most striking is a form of multiple testing.
What it costs
- Early events are lost. Everyone who has the event before the landmark is excluded, losing events and power, and the later the landmark, the more selected the survivors the answer describes.
- Late exposure is misclassified. Anyone exposed after the landmark counts as unexposed, which dilutes the contrast between the groups and generally pulls the estimate towards no effect when there is one (Wan, 2026).
- The answer depends on the landmark. Different landmarks compare different people over different periods, so they give different estimates.
- The estimate is conditional. It applies only to people alive and event-free at the landmark, which is a different quantity from an effect measured from baseline.
- Confounding remains. Responders may differ from non-responders in ways that also affect survival. Anderson and colleagues advised that comparisons by response are useful as description only, and should not be used to judge whether a treatment works.
Landmarking for dynamic prediction
Van Houwelingen developed landmarking into a way of making predictions that are updated during follow-up (Scand J Stat 2007;34:70–85). At a landmark time , a model is fitted to the people still at risk, using covariates measured up to , such as the latest value of a biomarker, to predict the event over a fixed window up to , with follow-up censored at . A Cox model with a time-dependent covariate cannot give such predictions on its own, because it would need the covariate’s future values.
The data sets for a grid of landmarks can also be stacked and fitted together as one landmark supermodel, with coefficients, and a baseline hazard, that change smoothly with , for example as a quadratic (van Houwelingen, 2007). The same person appears at several landmarks, so standard errors need a robust (sandwich) estimator. Van Houwelingen and Putter set out the approach in Dynamic Prediction in Clinical Survival Analysis (CRC Press, 2011).
Alternatives
- A time-dependent covariate. In a Cox model, each person counts as unexposed until the exposure happens and as exposed afterwards, so all follow-up and all events are used, and the result is a hazard ratio for current exposure. Putter and van Houwelingen show how landmark estimates relate to it (Stat Biosci 2017;9:489–503). Like a landmark, it removes immortal time, and confounding still has to be dealt with.
- A joint model. When the exposure is a biomarker measured repeatedly, a joint model of its trajectory and the event time uses every measurement. Simulation studies have found it more efficient than landmarking when correctly specified, and worse when the model for the biomarker is wrong, as reviewed by Putter and van Houwelingen (2022).
- Target trial emulation. For causal questions, such as whether to start a treatment, emulate the trial you would have run, as What is immortal time bias? describes.
What to report
- The landmark, and why and when it was chosen.
- How many people had the event, or were censored, before the landmark and were excluded.
- How exposure was defined at the landmark, and how many people changed status after it.
- That the estimates are conditional on being alive and event-free at the landmark.
- Results at other prespecified landmarks, where the choice was not clear-cut.
- For prediction models, the landmarks, the prediction window, how the coefficients depend on the landmark, and how the predictions were validated.