Kaplan-Meier Survival Analysis: A Practical Guide With Examples, Interpretation, R & SPSS

Say you follow 100 patients for five years to see how long they stay cancer-free. One recurs at 8 months. A dozen are still recurrence-free the day the study closes. Two move abroad at 18 months and stop answering your calls entirely.

Now try to summarize that with an average. You can’t, not honestly. Averaging follow-up time punishes the patients who did well. Calculating a simple recurrence percentage forces you to decide what to do with the two who vanished, and neither option is good; count them as recurrence-free, and you’re guessing; drop them and you’ve thrown away 18 months of real data. Kaplan-Meier analysis exists because this situation is the normal one in clinical research, not the exception. It’s built for time-to-event data where some of the follow-up gets cut short.

This guide walks through what the method is, how the estimator actually works (with a dataset small enough to do by hand), how to read the curve, how to compare groups, how to run it in R and SPSS, and the mistakes that show up most often in submitted manuscripts.

What Is Kaplan-Meier Survival Analysis?

Kaplan-Meier is a non-parametric method for estimating the probability that someone stays free of a defined event past a given point in time. Non-parametric just means it doesn’t assume your data follow a bell curve or any other fixed shape. You feed it the observed event times, and it works from those.

A handful of terms come up constantly, so here they are up front. Survival analysis, or time-to-event analysis, is the same thing; it studies how long something takes to happen. The event is whatever outcome you’re tracking. Survival time is how long each person was followed before that outcome occurred. The survival function is the curve describing the probability of staying event-free as time passes. And censoring is what happens when you never get to see someone’s event at all.

Now the part beginners consistently get wrong. Survival doesn’t mean staying alive. It means not having had the event yet, whatever you’ve defined the event to be:

  • Death
  • Cancer recurrence
  • Disease progression
  • Kidney graft failure
  • Hospital readmission
  • Treatment failure

The arithmetic is identical across all of these. Only the clinical meaning shifts. A nephrology paper using Kaplan-Meier for time to graft failure is doing exactly the same calculation as an oncology paper on mortality.

When Should You Use Kaplan-Meier Survival Analysis?

Four conditions, and you need all four:

  1. A clear starting point; diagnosis, surgery, randomization, first dose, whatever anchors time zero
  2. A measurable follow-up time for each person
  3. An event you’ve defined precisely
  4. At least some participants who never had the event during observation

That last one is the whole reason survival analysis exists. If everyone in your study had the event and you know exactly when, you don’t need Kaplan-Meier.

Picture the cancer recurrence study again. One patient recurs at 12 months, another at 24. Several reach the 60-month endpoint still recurrence-free. Two leave the country at 18 months. Those last two groups are censored. You learned something real from them, just not the whole story. Kaplan-Meier keeps what they gave you rather than discarding them.

Those four conditions, a clear starting point, measurable follow-up, a precisely defined event, and some participants who never reach it, are exactly what a well-designed cohort produces. We break down how to set one up properly in our guide to cohort study design, including how to define your starting point and follow-up window.

Key Terms in Kaplan-Meier Analysis

TermMeaning
Survival timeTime from study entry to either the event or censoring
EventThe outcome being studied
CensoringThe event was not observed during available follow-up
Right censoringThe person is observed up to a certain point, then follow-up stops
Risk setEveryone still being followed and still able to have the event at a given moment
Survival probabilityEstimated probability of staying event-free past a given time
Median survivalThe time at which estimated survival hits 50%
HazardThe instantaneous event rate among those still at risk
Follow-upThe window during which participants are observed

If you only memorize one of these, make it the risk set. Nearly everything confusing about Kaplan-Meier, such as, why the curve steps down when it does, why censoring behaves strangely, why the tail gets unstable resolves once you’re tracking who’s still in the risk set.

American Academy of Research & Academics

Censored data shouldn’t mean guesswork.

If survival curves, hazard ratios, and censoring assumptions still feel shaky, AARA’s Biostatistics Online Course builds that foundation properly, from probability basics to the methods you’ll actually use in your research.

Explore Biostatistics →

How Does the Kaplan-Meier Estimator Work?

The method handles one event time at a time and asks the same question at each: of the people still at risk right now, what fraction made it past this moment without the event?

Then it multiplies those fractions together. Hence the other name, the product-limit estimator.

$$\hat{S}(t) = \prod_{t_i \leq t}\left(1 – \frac{d_i}{n_i}\right)$$

Reading the notation: $t_i$ is a time when at least one event happened, $d_i$ is how many events happened then, and $n_i$ is how many people were at risk immediately before.

Try it with numbers. Ten patients at risk, two have the event. Eight got through, so that step contributes 8/10 = 0.80. At the next event time, eight are at risk and one has the event: 7/8 = 0.875. Multiply them and your cumulative survival estimate is 0.70.

Because the estimate only moves when an event occurs, you get a staircase rather than a smooth line. Flat stretch, sudden drop, flat stretch. That shape isn’t a rendering quirk, it’s the estimator honestly showing you that nothing was observed in between.

Kaplan-Meier Survival Analysis Example With Data

Six patients is small enough to do by hand, which is the fastest way to see what’s happening.

PatientFollow-up timeEvent
12 monthsYes
23 monthsNo (censored)
34 monthsYes
45 monthsYes
57 monthsNo (censored)
68 monthsYes

Notice what censoring did and didn’t do. It never dropped the curve. What it did was shrink the denominator for every step afterwards, which is why the later drops are so much steeper. Going from 3 patients to 2 is a bigger proportional hit than going from 6 to 5. That’s the tail instability everyone warns about, visible in a six-person dataset.

How to Read a Kaplan-Meier Survival Curve

The x-axis is follow-up time from your defined start point, usually months or years. The y-axis is the estimated probability of staying event-free, running from 1.0 down toward 0.

Each vertical drop is an event. Flat stretches mean nothing was observed during that interval. Small tick marks sitting on the curve are censored observations; they’re informative, not a problem.

The shaded confidence bands show uncertainty around the estimate, and they fan out toward the right as your sample thins.

Then there’s the risk table, the row of numbers printed underneath the plot showing how many people are still under observation at each time point. Include it. A survival estimate calculated from four remaining patients is drawn with exactly the same confident-looking line as one calculated from four hundred, and the risk table is the only thing on the page that tells you which you’re looking at. Dramatic late-curve separations that vanish once you check the risk table are a recurring theme in peer review.

One wording point. If the estimate at 12 months is 0.75, say: the estimated probability of remaining event-free through 12 months is roughly 75%. Don’t write “75% of patients survived exactly 12 months.” That sounds like a headcount taken on a particular day, and it isn’t what the estimator produced.

Kaplan-Meier Censoring and Right Censoring

Clinical studies mostly deal with right censoring, which arrives in three flavors: the study ends before the person has the event, the person is lost to follow-up, or the person withdraws.

The mechanics are simple. Before censoring, they count in the risk set. At the censoring moment, the curve doesn’t drop. After that, they’re out and have no effect on anything downstream.

There’s an assumption tucked underneath all of this, though, and it gets skipped in most beginner guides. Kaplan-Meier assumes censoring is non-informative, meaning censored participants would have had roughly the same future risk as the people still being followed. Think about whether that holds in your study. If patients dropped out because they were deteriorating, the assumption fails, and your curve will look better than the truth. Why people left matters as much as how many.

Kaplan-Meier Median Survival Time

Median survival is the time at which estimated survival first reaches 0.50.

Finding it on a plot takes three seconds: locate 0.50 on the y-axis, move right until you hit the curve, drop down to the x-axis. In the six-patient example above, survival falls to 0.417 at 5 months, so the median is 5 months.

Here’s what people don’t expect. If the curve never dips to 0.50 during your follow-up, there is no median to report. R will print NA. SPSS leaves the cell empty. This is not a bug, and there’s no setting that will produce the number, the data genuinely don’t contain it. Usually it means your follow-up was too short or your event was rare, both of which are worth saying out loud. Write “median not reached” and report survival probabilities at fixed time points instead.

Kaplan-Meier Confidence Interval

Your survival estimate is an estimate. Pair it with a 95% confidence interval so readers can judge how much weight it carries.

Something like: estimated 12-month survival was 78% (95% CI, 70%–85%).

Software gets that interval from Greenwood’s formula. You won’t need to compute it by hand, but it’s worth knowing what drives the width: as the number at risk falls, the interval widens. That’s the whole reason confidence bands flare out on the right side of nearly every published Kaplan-Meier plot. A survival estimate is only useful alongside a sense of how much weight it can carry, which is exactly what the confidence interval is for. We unpack that distinction, and how it differs from a p-value, in our guide to p-value vs. confidence interval.

Kaplan-Meier Curve Comparison and Log-Rank Test

Most questions involve two groups or more. Treatment A against Treatment B, exposed against unexposed. You plot a curve for each on shared axes.

A visible gap between them proves nothing on its own. Small samples produce gaps by chance all the time. The log-rank test is what checks whether the difference holds up.

Keep the two roles separate in your head. The Kaplan-Meier curve describes estimated survival over time. The log-rank test asks whether the survival distributions differ between groups. Different jobs.

Watch your phrasing when you write it up. “p < 0.05, so the treatment works” overstates what happened. Better: a statistically significant log-rank test indicates evidence that the survival distributions differ between the groups.

The test’s real advantage is that it uses the entire follow-up period rather than one time point you picked. Comparing survival only at 12 months throws away everything before and after, and invites the suspicion that you chose 12 months because it looked good.

The log-rank test is the right tool here, but it isn’t the only one available, and picking correctly depends on what your data actually look like. We map out that decision in our guide on which statistical test to use in medical research.

Kaplan-Meier Survival Analysis in R

Breaking that down: time is each patient’s follow-up duration, status is the event indicator (1 for event, 0 for censored, in the usual coding), and group is whatever splits your comparison arms. Surv() packages time and status into an object R recognizes as survival data. survfit() computes the estimates. survdiff() runs the log-rank test.

For plots you’d actually put in a paper, risk table included, confidence bands shaded, add the survminer package and use ggsurvplot().

And if your question involves adjusting for age, stage, comorbidity, or anything else, Kaplan-Meier isn’t the tool. You want Cox regression, which is coxph() in the same package.

Kaplan-Meier Survival Analysis in SPSS

SPSS has a dedicated menu path: Analyze → Survival → Kaplan-Meier.

From there:

  1. Drop your follow-up time variable into Time
  2. Drop your event indicator into Status
  3. Click Define Event and state which value marks the event, usually 1
  4. Put your treatment or exposure variable under Factor
  5. In Options, tick Survival plot
  6. In Options, tick Mean and median survival
  7. Click Compare Factor and choose Log-rank

The output gives you a survival table with the estimate at each event time, mean and median survival with confidence intervals, the plot itself, and hazard plots if you ask for them.

Step 3 is the one that gets skipped. Define the event value or SPSS treats every single case as censored, and you’ll get a flat line at 1.0, wondering what went wrong. That one missed step, Define Event, is the most common reason a Kaplan-Meier plot comes out flat in SPSS. We walk through this and the rest of the menu path in our full SPSS tutorial for medical research.

How to Report Kaplan-Meier Results

A structure you can adapt:

“Kaplan-Meier analysis showed an estimated median event-free survival of X months (95% CI, X-X). The estimated event-free survival at 12 months was X% (95% CI, X-X). Survival distributions were compared using the log-rank test, which showed [result].”

Your write-up should include the sample size, the number of events observed, how many were censored and ideally why, median survival or an explicit statement that it wasn’t reached, 95% confidence intervals, survival probabilities at clinically meaningful time points, the log-rank p-value if you’re comparing groups, and the number at risk.

That last item is the one reviewers ask for more than anything else on this list. Put it in the first draft and skip a revision round.

Common Kaplan-Meier Survival Analysis Mistakes

Coding censoring as an event: This is the classic data-entry error. It drags your survival estimates well below the truth, and the resulting curve looks plausible enough that nobody catches it.

Assuming survival means death: Recurrence, relapse, graft failure, readmission, all valid events. State which one you used in the methods section.

Ignoring the number at risk: The right end of a curve often rests on a handful of patients. A separation that looks enormous can be three people and a coin flip.

Reading p < 0.05 as clinical importance: A significant log-rank test says the distributions differ. It says nothing about whether the difference is large enough to change what you’d do for a patient.

Treating a gap as proof of causation: In observational data, two curves can separate because of confounding just as easily as because of the exposure.

Running Kaplan-Meier over competing risks without thinking it through: If patients can die of something unrelated before your event of interest, treating those deaths as ordinary censoring will overestimate the cumulative incidence of the event you care about. Cumulative incidence functions and Fine-Gray models exist for this. Work out what you’re actually trying to estimate before you run anything, this one is a design decision, not a software setting.

Getting the event definition and risk set right matters just as much upstream, when you’re still calculating how many patients you need to detect a real effect. We cover that groundwork in our guide to sample size calculation for clinical research.

American Academy of Research & Academics

The right method starts before you touch the data.

Choosing Kaplan-Meier, defining your event, deciding what counts as censoring, these are study design decisions, not statistical ones. AARA’s Basic Research Methodology course covers exactly this groundwork.

Explore Research Methodology →

Kaplan-Meier Survival Analysis vs Cox Regression

Kaplan-MeierCox regression
Estimates survival probability over timeModels the association between predictors and time to event
Mainly descriptive and for group comparisonHandles multiple covariates at once
Produces survival curvesProduces hazard ratios
Log-rank test for group comparisonWald, likelihood ratio, and score tests with HR estimates
No covariate adjustment in the basic analysisAdjusts for confounders

In practice they’re not rivals. Most published papers use both: Kaplan-Meier curves to show readers what the data look like, Cox regression to account for everything that muddies the comparison.

That same logic, one curve to show the data, a second method to account for what complicates it, scales up when you’re combining survival data across multiple studies instead of one. We walk through that process in our guide on how to write a systematic review and meta-analysis.

Still unsure whether Kaplan-Meier is the right fit for your data, or stuck on how to handle competing risks? The American Academy of Research and Academics works with researchers on study design, survival analysis, and reporting. Reach out and tell us where you’re stuck. 

American Academy of Research & Academics

One study’s survival curve. What about the field’s?

Once you can read a single Kaplan-Meier curve, the next step is pooling survival data across studies. AARA’s 12-week Meta-Analysis Module walks you through combining hazard ratios, forest plots, and heterogeneity testing.

Explore the Meta-Analysis Module →

Frequently Asked Questions

Q1. What is Kaplan-Meier survival analysis? 

A non-parametric method for estimating the probability of staying event-free over time, designed to handle censored follow-up.

Q2. What does a Kaplan-Meier curve show? 

The estimated probability of remaining event-free at each point in follow-up, drawn as a step function that drops whenever an event occurs.

Q3. What does censoring mean in Kaplan-Meier analysis? 

That the event was never observed for that person within your follow-up window. The study ended, they withdrew, or they were lost to follow-up.

Q4. How do you interpret a Kaplan-Meier curve? 

Read the y-axis value at any time point as the estimated probability of remaining event-free through that time. Check the risk table before you trust the tail.

Q5. What is median survival time? 

The time at which estimated survival first reaches 50%.

Q6. What does the log-rank test tell you?

 Whether there’s statistical evidence that survival distributions differ between groups, assessed across the whole follow-up period rather than at one time point.

Q7. Can Kaplan-Meier analysis be performed in SPSS? 

Yes. Analyze → Survival → Kaplan-Meier, which produces survival plots, median survival, and the log-rank test.

Q8. Can Kaplan-Meier analysis be performed in R? 

Yes, using survfit() and survdiff() from the survival package, with survminer for publication-quality plots.

Q9. What does a p-value mean in Kaplan-Meier analysis?

 It reflects how much evidence you have against the idea that the groups share identical survival distributions. It doesn’t measure the size of the difference or whether it matters clinically.

Q10. What if median survival is not reached?

 That’s a finding, not a failure. The curve simply never fell to 50% during your follow-up. Report it as “not reached” and give survival probabilities at specific time points instead.

Disclaimer:

Articles published by American Academy of Research & Academics are prepared by our team using information from direct experience, publicly available resources, and educational references. AI tools may be used to assist with drafting, proofreading, and formatting; however, all content undergoes review and approval before publication.
The information provided is intended for educational purposes only. Requirements, policies, and processes may change over time. Readers should consult official sources for the most current information.

Facebook
Twitter
LinkedIn
Email
0

No products in the basket.