What is non-collapsibility?
Why adjusting for a prognostic factor changes an odds ratio or hazard ratio without confounding, which measures are collapsible, and what to report instead.
A measure of effect is non-collapsible when its value in the whole population is not an average of its values in subgroups, even when there is no confounding. The odds ratio and the hazard ratio are non-collapsible, so adjusting for a covariate that predicts the outcome changes them, even in a randomised trial.
Take a randomised trial, where by design there is no confounding. Fit an unadjusted Cox model, then add a covariate that predicts the outcome but, because of randomisation, is unrelated to treatment. The hazard ratio moves, usually further from one (Gail, Wieand and Piantadosi, 1984). Nothing has gone wrong. The unadjusted estimate is a marginal effect, comparing the whole treated arm with the whole control arm. The adjusted one is a conditional effect, comparing patients with the same value of the covariate. Neither is biased for what it estimates.
Which measures are affected
- Non-collapsible: the odds ratio and the hazard ratio. If the odds ratio is the same in every subgroup of a covariate that predicts the outcome and is balanced between the groups, as in a trial, the overall odds ratio is closer to one, and the hazard ratio behaves similarly.
- Collapsible: the risk difference, the risk ratio, the difference in survival at a fixed time and the difference in restricted mean survival time. Without confounding, the overall value is a weighted average of the subgroup values.
A worked example
A trial randomises 1,000 patients 1:1. Half have a poor prognosis and half a good one, so each arm has 250 of each. The outcome is response to treatment.
- Poor prognosis. 50 of 250 respond on control (20%) and 125 of 250 on treatment (50%). The odds of response are 0.25 and 1, so the odds ratio is 4; the risk difference is 30 percentage points and the risk ratio 2.5.
- Good prognosis. 125 of 250 respond on control (50%) and 200 of 250 on treatment (80%). The odds are 1 and 4, so the odds ratio is again 4; the risk difference is 30 points and the risk ratio 1.6.
- Whole trial. 175 of 500 respond on control (35%) and 325 of 500 on treatment (65%). The odds are 0.54 and 1.86, so the odds ratio is 3.45; the risk difference is 30 points and the risk ratio 1.86 (the same number as the treated odds by coincidence, because 65% = 1 − 35%).
The odds ratio is 4 in both groups and 3.45 overall, with prognosis perfectly balanced between the arms. The overall risk ratio, 1.86, is the average of 2.5 and 1.6 weighted by the number of control responders in each group (50 and 125).
Why hazard ratios drift
With survival data, non-collapsibility comes with a selection effect. Suppose half the patients have a death rate of 0.1 per year and half 1.0 per year, and treatment halves both, so the hazard ratio is 0.5 in each group throughout. At the start, the ratio of the two arms’ overall hazards is also 0.5. High-risk patients die sooner, and more of them die on control, so the arms drift apart in their mix of patients. After two years, 14% of surviving control patients are high-risk against 29% of surviving treated patients, and the ratio of the overall hazards has risen to 0.79. An unadjusted Cox model averages over this drift.
Hernán called this the built-in selection bias of the hazard ratio (Hernán, 2010). Once events start, if treatment has an effect, the people still at risk in the two arms are no longer comparable, so a hazard ratio for a later period is not a randomised comparison. The same happens with any prognostic factor left out of the model, measured or not, so even in a randomised trial the hazard ratio lacks a simple causal interpretation, though the Cox model still gives a valid test of no treatment effect when robust standard errors are used (Aalen, Cook and Røysland, 2015).
What it changes in practice
- An adjusted and an unadjusted hazard ratio from the same trial estimate different things, so say which covariates a hazard ratio is conditional on. A difference between the two is expected, and is not evidence of confounding.
- Odds ratios or hazard ratios adjusted for different covariates estimate different quantities, so pooling them in a meta-analysis, or comparing them across studies, mixes different quantities (Daniel, Zhang and Farewell, 2021).
How it differs from confounding
Confounding happens when the groups compared differ in something else that affects the outcome. Non-collapsibility needs no such difference, because it comes from the measure itself. For the risk difference and the risk ratio, if the covariates are enough to control confounding, a change on adjustment signals confounding and nothing else, which may be why the two ideas are often confused. For odds ratios and hazard ratios they do not, so a change in the estimate when a covariate is added can come from either or both, and on its own does not show confounding (Greenland, Robins and Pearl, 1999).
What to do
- Say which effect you want. Decide in advance whether the target is a marginal effect, of treating everyone versus no one, or a conditional one, for patients with the same covariates, as the FDA’s guidance on covariate adjustment asks trial sponsors to do (FDA, 2023).
- Standardise. Fit an adjusted model, predict each patient’s outcome or survival curve under each treatment, and average the predictions over everyone. Contrasts of these averages are marginal effects, and in a trial, adjusting for prognostic covariates in this way generally improves precision, as the same guidance notes.
- Report absolute effects. Differences in survival at chosen times and in restricted mean survival time are collapsible, so marginal estimates of them can be compared across analyses adjusted for different covariates.
Want to learn more?
- RMST or a hazard ratio?
- What a hazard ratio means when hazards are not proportional
- Our hazard ratio tool
- Our survival analysis course
MethodSurvival analysis