Introduction
Say you’ve noticed something in your clinic and you want to know if it’s real. Do the patients on a certain drug actually run into more heart problems later? Does the exposure come first, or are you just seeing a pattern that isn’t there?
That question- exposure now and outcome later is what a cohort study is built for. You take a defined group of people, sort them by whether they’ve been exposed to something, and then watch what happens to them. This guide covers what a cohort study actually is, how prospective and retrospective cohorts differ, how to design one from scratch, what to do about bias and confounding, how incidence and relative risk fit in, and when a case-control study would serve you better.
What is a Cohort Study Design?
Cohort study definition
A cohort study is an observational design ( a study design where you just observe and don’t perform any intervention). You classify a defined group of participants by some exposure or characteristic, then follow them over time and record whether the outcome you care about happens.
That word observational is doing a lot of work. You don’t hand anyone the exposure. You find it, document it, and then wait.
What makes the design valuable is the order of events. Exposure gets measured before the outcome shows up. That sequence is why a cohort study can say something about timing, which most other observational designs struggle with.
The basic cohort study framework
Strip away the details and every cohort study looks the same:
Study population → Exposure assessment → Exposed vs. unexposed groups → Follow-up → Outcome → Comparison
Hold onto that chain. Everything below is just an expansion of one link or another.

How Does a Cohort Study Work?
Let’s walk it through with one example and keep that example for the rest of the article: does smoking increase the risk of COPD?
Define who’s in the study: Adults who don’t already have COPD. That last part matters, as you can’t watch a disease develop in someone who already has it.
Identify the exposure: Smoking.
Split by exposure status: Smokers in one group, nonsmokers in the other.
Follow them: Say ten years, with scheduled check-ins.
Record the outcome: Who developed COPD, and when.
Compare: How often did it happen in each group? That comparison is where incidence and relative risk come from, and we’ll get to both later.
Most introductory articles stop after “what is a cohort study.” The useful part is treating it as a workflow you can actually follow when you sit down to write a protocol.
Types of Cohort Studies
Prospective cohort study
Here you record the exposure first, then follow people forward and wait for outcomes to appear.
Example: You enrol smokers and nonsmokers who are free of COPD at baseline, take a detailed smoking history at enrollment, and follow everyone for ten years to see how many in each group develop the disease.
The advantage is control. You decide how smoking is measured, which baseline variables get recorded, and how often people come back. Your data are as good as your protocol.
The cost is, well, cost. Ten years is ten years. If the outcome is uncommon you’ll need a big sample, and people will drift away from the study no matter how well you plan.
Retrospective cohort study
Same logic, different starting point. You reach into records that already exist, rebuild the cohort on paper, and look at exposures and outcomes that have already played out.
Example: You pull hospital records from 2015 to 2025, flag every patient who was prescribed a particular medication, and compare their cardiovascular outcomes with patients who never got it.
Electronic health records, disease registries, insurance claims, occupational databases, these are the usual sources. When a well-defined population and decent historical exposure data already exist, a retrospective cohort is far faster and cheaper than waiting a decade.
The catch is that you’re stuck with whatever someone else bothered to write down. Missing fields, definitions that shifted over the years, and variables nobody collected are all common, and you can’t go back and fix them.
For IMGs in a research year, a retrospective chart review is usually the first project a US mentor will hand you, which is exactly this design.
Longitudinal cohort study
Longitudinal isn’t really a third type. It describes repeated measurement of the same people over time. A study can be longitudinal and prospective, or longitudinal and retrospective, depending on how the data were put together.
A note on terminology: Prospective and retrospective describe when data were collected relative to the outcome and not just whether you used a database. Analyzing data that were gathered before the outcome occurred, but analyzing them years afterwards, sits awkwardly between the two labels. Rather than arguing about which word applies, state plainly in your methods when the exposure was recorded, when the outcome occurred, and when you did the analysis.

How to Conduct a Cohort Study: Step-by-Step
Step 1: Define the research question
There’s a formula worth memorizing:
Among [population], does [exposure], compared with [comparison], affect [outcome] over [follow-up period]?
Filled in: Among adults without COPD, does smoking, compared with not smoking, increase the risk of developing COPD over ten years?
This is the PECO structure: Population, Exposure, Comparison, Outcome. And it forces you to make your design decisions now, on paper, rather than discovering them halfway through data collection.
Step 2: Define the study population
Inclusion criteria, exclusion criteria, setting, recruitment. Write all four down before you enroll anyone.
For the COPD example: adults aged 40 and above, no prior COPD diagnosis, excluding anyone with other chronic lung disease. You’ll also need a sample size estimate, and that depends on how often you expect the outcome to occur.
A fuzzy population definition is one of the main reasons nobody can replicate a study later. Be specific enough that a stranger could rebuild your cohort from your methods section.
Step 3: Define the exposure
An exposure might be a behavior, a drug, an occupational or environmental factor, something dietary, a biomarker, or just a clinical characteristic.
Whatever it is, decide exactly how you’ll measure it and how you’ll group it. Smoking alone gives you several options: ever versus never, current versus former versus never, or pack-years as a continuous variable. Each of those will give you a different answer, so commit to one and say so.
Step 4: Define the outcome
Pick one primary outcome. List secondary outcomes separately and treat them as secondary.
For each, spell out the diagnostic criteria and when it gets assessed. Incident COPD, in our example, might be defined by spirometry using a fixed threshold, confirmed at scheduled study visits.
Loose outcome definitions are quietly fatal. If a reader can’t tell what counted as a case, they can’t judge your results.
Step 5: Establish exposed and unexposed groups
Your comparison group should look like your exposed group in every way that matters except the exposure.
And here’s a distinction worth getting right: a cohort study doesn’t have a control group, strictly speaking. You never assigned anything. The unexposed group is a comparison group. Control groups come from randomization, and there’s no randomization here. Plenty of student papers get this wrong.
Step 6: Plan follow-up
Set your total duration, your assessment intervals, and how you’ll check for outcomes at each visit.
Then plan for the people who leave. Dropouts are rarely a random slice of your cohort; if sicker participants stop showing up, your results shift. Track how many you lose and at what point, and report it.
Step 7: Identify confounders
A confounder is something tied to both the exposure and the outcome, and it can bend the association you end up seeing.
In a smoking and COPD study, you’d worry about age, occupational dust exposure, air pollution, and socioeconomic status, at minimum.
Sort this out at the design stage, because you can’t adjust for a variable you never measured. That’s the whole point. Confounding is a design problem that happens to have statistical solutions, not the other way around.
Cohort Study Data Collection and Variables
Think of your variables in four buckets.
Exposure variables capture what participants were exposed to, and how much, and for how long.
Outcome variables capture what happened during follow-up and when it happened.
Covariates and confounders are everything else that might influence the relationship, age, sex, comorbidities, and so on.
Baseline characteristics describe the cohort at entry. These let a reader judge whether your two groups were comparable to start with.
Data can come from questionnaires, interviews, clinical exams, labs, imaging, EHRs, or registries.
If you’re working retrospectively, do yourself a favor and audit the database before you commit to it. Check how completely the exposure, the outcome, and your key confounders were actually recorded. Then report the gaps honestly instead of hoping nobody asks.
Cohort Study Analysis: Incidence and Relative Risk
Incidence in cohort studies
Because everyone starts without the outcome, a cohort study can count new cases as they appear. Not many designs can do that, and it’s one of the main reasons to choose this one.
Two measures show up constantly:
Cumulative incidence is new cases divided by the number of people at risk when follow-up began.
Incidence rate is new cases divided by total person-time. Use this when people were followed for different lengths of time, which in practice is most of the time.
Relative risk in cohort studies
Relative risk compares the risk of the outcome in the exposed group against the risk in the unexposed group.
RR = Risk in exposed ÷ Risk in unexposed
Reading it is simple enough. An RR of 1 means no difference between the groups. Above 1 means higher risk with exposure. Below 1 means lower risk.
Always report the confidence interval next to it, since the interval tells readers how much to trust the number. And you’ll usually report adjusted estimates from a regression model too, because that’s where all those confounders you measured in Step 7 finally earn their keep.
Biostatistics
Relative risk is the easy line. The adjusted estimate underneath it is where reviewers look.
Cumulative incidence, incidence rates with person time, confidence intervals, and the regression models that turn measured confounders into adjusted estimates. Our Biostatistics Online Course works through each one on real medical datasets, so the numbers in your results section are ones you can defend.
Advantages and Disadvantages of Cohort Studies
| Advantages | Disadvantages |
| Establishes the temporal sequence between exposure and outcome | Can be expensive to run |
| Can measure incidence directly | Prospective studies may take years |
| Can study several outcomes from one exposure | Loss to follow-up can bias results |
| Useful when the exposure is rare | Large samples needed for uncommon outcomes |
| Allows calculation of relative risk | Confounding must be measured and adjusted for |
| Can describe the natural history of a condition | Selection and information bias remain possible |
Notice that the strengths and weaknesses have the same source. Everything good about a cohort study comes from following people forward and watching outcomes appear. Everything difficult about it comes from the same thing; following people over time is slow, expensive, and you will lose some of them.
Retrospective cohorts trade away some of that cost and time, and pay for it with reduced control over data quality.
Cohort Study Bias and Confounding
Common sources of bias
Selection bias: Shows up when your exposed and unexposed participants differ in some systematic way that affects the outcome. Volunteers who are healthier than average are a classic version.
Information bias: Comes from measuring exposure or outcome inaccurately. Asking people to recall their smoking history from twenty years ago is a good example.
Loss-to-follow-up bias: Appears when the people who leave the study differ from the ones who stay especially when leaving is related to the outcome itself.
Confounding
Confounding can make an association look stronger than it is, weaker than it is, or point in the wrong direction entirely.
You have a few tools: restrict the study to a narrower population, match exposed and unexposed participants on key variables, stratify the analysis, or run multivariable regression. All of them depend on choices you made while designing the study. None of them can rescue a variable you never collected.
Cohort Study Design vs. Case-Control Study
The cleanest way to tell these apart is to ask where the study begins. A cohort study starts with exposure and looks forward. A case-control study starts with the outcome and looks backward.
| Feature | Cohort study | Case-control study |
| Starting point | Exposure | Outcome |
| Participants | Exposed vs. unexposed | Cases vs. controls |
| Direction | Exposure → outcome | Outcome → previous exposure |
| Incidence | Can be measured | Cannot be measured directly |
| Common measure | Relative risk | Odds ratio |
| Rare outcome | Often inefficient | Well suited |
| Rare exposure | Well suited | Often inefficient |
| Multiple outcomes | Possible | More limited |
| Follow-up | Central to the design | Usually not required |
A cohort design makes sense when you can clearly identify and measure the exposure, you can sort participants into exposed and unexposed, outcomes can be observed over time, and you want incidence, relative risk, or more than one outcome from the same cohort.
Look at case-control instead when the outcome is rare, the disease takes years to develop, or a cohort study would simply take too long or cost too much to be realistic.
How to Report a Cohort Study
Observational studies, cohort studies included, are reported using the STROBE statement, a checklist of what a published report should contain so readers can judge how the work was done.
For a cohort study, make sure you’ve covered:
- study design and setting
- eligibility criteria and how participants were selected
- how exposure was defined and measured
- how the outcome was defined and measured
- which confounders you considered
- follow-up duration and how complete it was
- statistical methods, including your approach to confounding
- missing data
- results, limitations, and generalizability
One clarification, because people mix this up: STROBE tells you how to report a study. It says nothing about how to design or conduct one. A badly designed study reported perfectly is still a badly designed study.
Conclusion
Reach for a cohort study when your question depends on watching an exposure and an outcome unfold over time. It works when you’ve defined your population clearly, picked an exposure you can actually measure, specified your outcome precisely, chosen a sensible comparison group, and followed everyone systematically.
Deal with confounding, bias, missing data, and loss to follow-up before you collect a single data point. Trying to fix any of them afterward rarely goes well.
If this is your first cohort study, work through the seven steps above as a checklist before you write your protocol. And if the finished study is headed for your residency application, we cover how to list research in ERAS separately.
Frequently Asked Questions
Q1. Is a cohort study the same as a longitudinal study?
Not quite, though the two overlap a lot. Longitudinal simply means you measure the same participants more than once over time. Cohort refers to how the group was assembled and compared by exposure status. Most cohort studies are longitudinal, but a study can be longitudinal without being a cohort study, for example, if it tracks one group over time without comparing exposed and unexposed participants.
Q2. Can a cohort study prove that an exposure causes an outcome?
No, and this trips up many first-time researchers. A cohort study can show that the exposure came first and that the outcome occurred more often in the exposed group, which is stronger evidence than a cross-sectional or case-control study can offer. But you didn’t assign the exposure, so some unmeasured factor may still explain the association. Cohort studies support causal reasoning. They don’t settle it on their own. That is also why systematic reviews and meta analyses sit above cohort studies on the evidence pyramid.
Q3. How many participants do I need for a cohort study?
There’s no fixed number. Sample size depends on how common the outcome is, how large a difference between groups you want to be able to detect, your planned follow-up length, and how much loss to follow-up you expect. Rare outcomes need much larger cohorts than common ones. Run a formal sample size calculation during the design stage and build in a margin for dropouts, usually somewhere between ten and twenty per cent depending on your setting.
Q4. Should I choose a prospective or a retrospective cohort study?
It usually comes down to what data already exist and how much time you have. If good records with clearly documented exposure and outcome information are already available, a retrospective cohort will get you an answer far faster and at a fraction of the cost. If the exposure or outcome isn’t recorded well in existing sources, or you need to measure something in a specific way, you’ll have to collect it prospectively and accept the longer timeline. If a multi year follow up is out of reach for your timeline, a narrative review is the usual first publication for a new researcher.
Q5. Why do cohort studies report relative risk instead of odds ratio?
Because a cohort study can measure incidence directly. You know how many people started without the outcome and how many developed it, so you can calculate actual risk in each group and compare the two. A case-control study starts with people who already have the outcome, so it can’t measure incidence, and it reports an odds ratio instead. That’s why the measure follows the design rather than the other way around.
Q6. How much loss to follow-up is acceptable?
There’s no universal cutoff, though many reviewers get uncomfortable above twenty percent. What matters more than the raw number is whether the people who left differ from the people who stayed. If dropouts are related to the exposure or the outcome, even a modest loss can distort your findings. Report your retention rate, describe who was lost and at which time points, and consider a sensitivity analysis to show how much your conclusions would change under different assumptions.
If you’re planning your first cohort study, or you have a protocol sitting half-finished, get in touch with the American Academy of Research and Academics and work through it with people who’ve done it before.
Research Methodology
Every one of the seven steps is decided before enrollment. No regression model fixes them afterward.
Population, exposure, outcome, comparison group, follow up plan, confounders. Our Basic Research Methodology course takes you through each decision with a mentor, so the protocol you submit to the review committee is one you can actually run.
Disclaimer:
Articles published by American Academy of Research & Academics are prepared by our team using information from direct experience, publicly available resources, and educational references. AI tools may be used to assist with drafting, proofreading, and formatting; however, all content undergoes review and approval before publication.
The information provided is intended for educational purposes only. Requirements, policies, and processes may change over time. Readers should consult official sources for the most current information.