Survival Analysis & Kaplan–Meier
Sometimes the outcome is a when, not a number: the time until a patient relapses, a customer churns, a part fails, or a participant drops out of a habit study. This is time-to-event ("survival") data, and ordinary methods fail on it for one reason. When the study ends, some events haven't happened yet. Those unfinished cases are called censored observations, and survival analysis exists to use them without bias.
Why you can't just average the times
Suppose half your patients were still relapse-free when the study closed at 36 months. What's their "time to relapse"? You only know it's more than what you observed. Averaging the observed times treats them as if they relapsed when observation stopped, and that pulls the estimate down. Throwing censored cases away is even worse: you'd be deleting the people who lasted longest. Both shortcuts bias the result, so censoring has to be built into the method.
The Kaplan–Meier idea: survive one step at a time
The Kaplan–Meier estimator builds the survival curve S(t), the probability of lasting beyond time t, one step at a time. At each observed event time it asks: of the people still at risk at that moment, what fraction got through it? Multiply those step-by-step survival fractions together and you get the familiar staircase curve. Censored people count while they're observed and leave the "at risk" pool afterward, so nothing has to be guessed about their futures.
🎮 Kaplan–Meier Playground
Two groups (say, treatment vs control), exponential event times over 36 months. Tick marks are censored subjects. Raise the hazard ratio to separate the curves, add censoring to see how the estimator handles it, and read the log-rank test below the chart.
Push the censoring slider up and a median may show "> 36 mo" in place of a number. That is correct. If the curve never falls to 0.5 inside the follow-up window, the median survival time has not been reached yet, and papers write it as not reached. Do not invent a number there, or extend the curve past your last observation to find one: that turns a short study into a claim about long survival times. When the median is out of reach, report survival at a fixed landmark, such as the 24-month rate with its confidence interval.
Comparing groups: the log-rank test
The log-rank test checks whether two survival curves differ. At every event time it asks how many of that time's events group A "should" have had, given how many people each group had at risk. Summing observed-minus-expected events across all times gives a χ² statistic with 1 degree of freedom, a survival version of the chi-square test. It weights every event equally and is most powerful when one group's hazard is a constant multiple of the other's (proportional hazards).
Beyond the curve
Kaplan–Meier describes and the log-rank test compares, but neither adjusts for covariates. For that you need Cox proportional-hazards regression, which models the hazard ratio as a function of predictors and is a close relative of the GLMs you've already met. The hazard-ratio slider in the playground is the quantity a Cox model estimates.
When the hazards aren't proportional
The log-rank test and the Cox model rely on the same assumption: whatever the hazard ratio is, it stays constant across follow-up. A treatment that helps early and wears off later violates it, and so does anything whose curves cross. The log-rank test suffers most, because an early advantage and a late disadvantage cancel in its running sum.
There are three ways to check, in rising order of formality. The first is to look: crossing or converging curves are the visible symptom. The second is to plot log(−log S(t)) against log t for each group, where proportional hazards show up as parallel curves separated by a constant vertical gap of exactly log HR. The third is to test the Schoenfeld residuals, which record at every event time how far the covariate value of the person who had the event was from the average of everyone still at risk. Under proportional hazards those residuals have no trend against time. R's cox.zph() turns that into a χ² per covariate plus a global test, and a small p-value is evidence against the assumption, not for it.
The fix depends on which variable violates it. Stratify the Cox model on that variable, and each of its levels gets its own baseline hazard while the other covariates keep a single ratio. Let its coefficient change with time. Split the follow-up into periods and report a hazard ratio for each. Or drop hazard ratios and compare restricted mean survival time, the average event-free time up to a fixed horizon, which means the same thing whether or not the curves stay proportional.
Why it matters: whenever the outcome is "time until something happens," some events will not have happened by the time the study stops. Kaplan–Meier and the log-rank test use the information in finished and unfinished observations, without guessing and without discarding anyone. For a study of your own, Plan My Analysis collects the assumptions to check and explains why survival power is counted in events rather than participants.
Problem 79 in the practice problems has ten patients, four of them censored, and a curve you can build with a calculator. It works through the whole Kaplan–Meier table and shows what the two obvious averages get wrong.
Common questions
What is censoring in survival analysis?
A censored observation is one whose event hadn't happened when you stopped observing (the study ended, the participant moved away). You know their survival time exceeded some value, but not by how much. That is still information: censored subjects count in the at-risk pool while they are observed. Standard methods assume censoring is uninformative, meaning that dropping out is not related to how soon the event would have happened.
What does a hazard ratio mean?
It is the ratio of the two groups' instantaneous event rates: HR = 2 means that at any moment, the exposed group's event rate is double the reference group's. It does not mean "twice as likely to die overall" or "half the survival time." HR < 1 means the exposure is protective. A proportional-hazards model assumes this ratio is constant over follow-up, so check that assumption before relying on it.
How many participants does a survival study need?
Count events, not people. Power depends on how many participants actually have the event, so a big sample watched briefly can have less power than a small one followed for longer. The standard planning formula asks for about 4(zα/2 + zβ)² ÷ (ln HR)² events in total. Detecting a hazard ratio of 2 at α = .05 with 80% power takes roughly 65 to 70 events, and a hazard ratio of 1.5 takes nearly 200. You work out the sample size afterwards, from the event rate you expect and the follow-up you can afford. A Cox model adds its own requirement: around ten events for every covariate you want to adjust for.