Risk of Bias Assessment in Systematic Reviews: A Beginner’s Guide

A systematic review might include forty studies, but those forty studies do not carry equal weight as evidence. Some were designed in ways that could push their results in a particular direction. Risk-of-bias assessment is how you check for that. It asks whether the methods used in a study could have systematically distorted the result you are about to pool.

One thing trips up almost every beginner: risk of bias is not a verdict on whether a study is good or bad. It is a judgment about whether specific methodological problems could have affected a specific result. A well-written paper from a respected journal can still carry a high risk of bias.

What is Risk of Bias in a Systematic Review?

Bias is a systematic error. It is not random noise that averages out with a bigger sample. It is a consistent pull in one direction, caused by how the study was designed, run, analyzed, or reported.

This matters in evidence synthesis because a meta-analysis inherits the problems of its inputs. If four of six trials measured the outcome in a way that favored the treatment, pooling them produces a tidy summary estimate that is wrong. The confidence interval will narrow, and the result will look more certain than it is.

You can see this on the plot itself. How to read a forest plot covers what the diamond and its width are telling you, and why a narrow one is not automatically a reassuring one.

Assessment happens at the level of the study result, not the paper. A single trial may report a low-risk result for mortality and a high-risk result for pain, because the problems that threaten each outcome are different.

A narrow p-value does not fix any of this. Statistical significance tells you about precision, not about direction of error.

That distinction is worth getting straight before you assess anything. P value versus confidence interval covers what each statistic measures and, more usefully, what neither of them can tell you.

What does a risk-of-bias assessment actually tell you?

The question you are answering is narrow and specific:

Could the way this study was designed, conducted, analyzed, or reported have systematically influenced this result?

That is more useful than asking whether the study was “high quality,” because quality is vague and the answer changes depending on who you ask. The bias question has a target. You are looking for a mechanism that could have shifted the number in a predictable direction.

RoB 2 is built around this idea. You assess one result for one outcome at a time, rather than stamping a single label across an entire trial.

Risk of Bias vs Study Quality: Are They the Same?

No, though the terms get used interchangeably all the time.

Risk of bias asks a narrow question: could systematic errors have distorted this estimated result?

Study quality is a broader idea. It can cover reporting completeness, sample size, statistical precision, ethical conduct, and how well the study matches your review question. A small trial with wide confidence intervals is imprecise, but imprecision is not bias. A poorly written paper may be hard to appraise without being biased.

Keeping these separate saves you from a common error: assuming that a study which reports everything clearly must also be free of bias.

Why “quality scoring” can be misleading

Checklist scores flatten important differences. Adding up ten items and reporting a total treats every item as equally serious, which they are not.

A study scoring 8/10 is not automatically more trustworthy than one scoring 7/10. The study with 8 points may have lost its points on inadequate allocation concealment, which can seriously distort an effect estimate. The study with 7 points may have lost its points on incomplete reporting of baseline characteristics, which is untidy but rarely changes the result.

The number hides which problems were present. This becomes relevant when we get to the Newcastle-Ottawa Scale later on.

When Should You Perform Risk-of-Bias Assessment?

Assessment usually happens after you have identified your eligible studies and before you interpret the synthesis. It sits between screening and conclusions.

The order generally looks like this:

  1. Define eligibility criteria
  2. Select the appropriate risk-of-bias tool
  3. Extract the study information you need to make judgments
  4. Assess each study or result
  5. Resolve disagreements between reviewers
  6. Carry the judgments into your synthesis and interpretation

That last step is the one people skip.

Should risk of bias be assessed before or after data extraction?

Teams do it both ways. Some assess bias in the same pass as data extraction, others in a separate round.

What matters more than the order is that you pre-specify your approach in the protocol, including the tool, the version, and how you will handle unclear reporting. You also want to avoid a situation where knowledge of a study’s results shapes your judgment of its methods. It is very easy to be harsher on a trial whose findings you disagree with.

Pre-specifying means writing it into the registered protocol, not just agreeing it with your co-reviewer. Registering a systematic review on PROSPERO covers where the tool and version go on the form, and what happens if you change them later.

How Do You Choose the Right Risk-of-Bias Tool?

The tool follows the study design. You do not pick one tool and apply it to everything.

Study designCommon tool
Randomized controlled trialRoB 2
Non-randomized intervention studyROBINS-I
Diagnostic accuracy studyQUADAS-2
Cohort or case-control studyNewcastle-Ottawa Scale, depending on review methodology
Non-randomized study of exposuresROBINS-E

RoB 2 is intended for randomized trials. ROBINS-I is designed for non-randomized studies of interventions, where the absence of randomization creates threats that RoB 2 was never built to catch. QUADAS-2 addresses diagnostic accuracy studies, which have their own structure of index test and reference standard.

If your review includes more than one design, you use more than one tool and report them separately.

If any of those designs are unfamiliar, start with the designs rather than the tools. Cross sectional study design covers what separates one observational design from another and why the distinction changes everything downstream.

Can you use the same risk-of-bias tool for every study?

No. Study design determines which sources of bias are possible in the first place.

A randomized trial cannot suffer from baseline confounding in the way a cohort study can, because randomization is what distributes unmeasured differences between groups. A cohort study has no randomization process to assess, so applying RoB 2 to it would leave you rating a domain that does not exist while missing confounding entirely.

Using one tool everywhere produces judgments that look consistent and mean very little.

Research Methodology

Study design determines which sources of bias are even possible.

Which means you cannot choose a tool until you can identify a design on sight, and tell a cohort from a case control from a controlled before after study. Learn the designs properly and the tool selects itself.

Start with Methodology →

RoB 2 Domains for Randomized Controlled Trials

RoB 2 works through five domains. Each is reached using signalling questions, which are factual questions about what the study did. Your answers to those questions lead to a domain-level judgment of low risk, some concerns, or high risk.

1. Bias arising from the randomization process

You are looking at how the allocation sequence was generated and whether it was concealed until assignment. A sequence produced by computer random number generator is fine. Alternate allocation by day of admission is not, because it is predictable.

Baseline imbalance between groups is a signal here. Large differences in prognostic factors can suggest that the randomization process failed or was subverted.

2. Bias due to deviations from intended interventions

Participants do not always receive what they were assigned. Some cross over to the other arm, some stop treatment, some receive extra co-interventions.

Your judgment depends on which effect the trial was trying to estimate. If the question is about assignment to treatment, an intention-to-treat analysis handles this reasonably well. If the question is about adherence to treatment, you need a more careful analysis than simply excluding the people who dropped out.

3. Bias due to missing outcome data

Outcome data goes missing through loss to follow-up, withdrawal, or measurements that were never taken.

The concern is not the amount of missing data on its own. It is whether the missingness is related to the true outcome, and whether it differs between arms. Ten percent dropout distributed evenly is less worrying than five percent dropout that is concentrated in one group because of side effects.

4. Bias in measurement of the outcome

This is where blinding of outcome assessment sits.

Lack of blinding does not automatically mean high risk. It depends on the outcome. All-cause mortality is objective, so an unblinded assessor is unlikely to change the count. Pain scores, symptom severity, and clinician-judged improvement are subjective, and knowledge of the assigned arm can shift them.

Ask what the outcome is, who measured it, and whether that person’s expectations could have influenced the measurement.

5. Bias in selection of the reported result

This is selective reporting bias at the level of the result.

Researchers often have several ways to analyze the same data: multiple time points, multiple scales, adjusted and unadjusted models. If they ran several and reported the most striking one, the published estimate is not a fair representation.

Compare the paper against its protocol or trial registration where you can. Look for outcomes that were registered but never reported, or reported in a form nobody planned.

ROBINS-I for Non-Randomized Studies

ROBINS-I is the tool for non-randomized studies of interventions: cohort studies, controlled before-after studies, and similar designs where the researcher did not assign the treatment.

Its logic is worth understanding. ROBINS-I compares each study against a hypothetical target trial, meaning the randomized trial that would have answered the same question. Every domain asks how far the actual study falls short of that ideal.

Why is confounding so important in ROBINS-I?

Because without randomization, the groups being compared are usually different from the start.

Suppose patients receiving a new treatment are older and have more severe disease than patients receiving standard care. If age and disease severity are not adequately accounted for, the apparent treatment effect partly reflects those differences rather than the drug. The new treatment might look worse than it is.

The ROBINS-I confounding domain asks whether the authors identified the important confounders, measured them well, and adjusted for them appropriately.

Two forms matter:

Adjustment happens in the analysis, usually through a regression model. Which statistical test to use in medical research covers when a model with multiple predictors is the right approach and what it can realistically account for.

  • Baseline confounding: the groups differed before the intervention started
  • Time-varying confounding: something that changed during follow-up influenced both the treatment received and the outcome

The framework uses seven domains in total:

  1. Confounding
  2. Selection of participants into the study
  3. Classification of interventions
  4. Deviations from intended interventions, including co-interventions
  5. Missing data
  6. Measurement of outcomes
  7. Selection of the reported result

A note for 2026: a version 2 of ROBINS-I has been circulating as a draft. Check the current official documentation before you start a new review, and state clearly in your methods which version you used.

Newcastle-Ottawa Scale: How Does NOS Scoring Work?

The Newcastle-Ottawa Scale is a star-based tool for cohort and case-control studies. It is quicker to apply than ROBINS-I, which is why it appears in so many published reviews.

It covers three broad areas:

  • Selection: how participants were chosen, and whether the comparison group was drawn from a suitable population
  • Comparability: whether the study controlled for important confounders in its design or analysis
  • Exposure or outcome: how exposure was ascertained in case-control studies, or how outcome was assessed and followed up in cohort studies

You award stars against the items in each area and report a total, usually out of nine.

Is a higher Newcastle-Ottawa Scale score always better?

Not reliably. Converting a methodological assessment into a single number loses the information that made the assessment useful.

Two studies can both score 7. One lost its stars on confounding control, which can reverse an effect estimate. The other lost its stars on follow-up duration being slightly short. Those are not comparable problems, but the score treats them as if they were.

There is also no universally agreed cutoff that separates good studies from poor ones. Thresholds vary between reviews and are often chosen after the fact. Follow the methodology you specified in your protocol, and report the domain-level stars alongside the total so readers can see where the study actually struggled.

QUADAS-2 for Diagnostic Studies

QUADAS-2 is used for primary diagnostic accuracy studies, where an index test is compared against a reference standard.

It has four domains:

  • Patient selection: how patients were enrolled, and whether the sample avoided inappropriate exclusions
  • Index test: how the test under evaluation was conducted and interpreted, including whether its threshold was set in advance
  • Reference standard: whether the comparator test correctly classifies the target condition, and whether it was interpreted without knowledge of the index test result
  • Flow and timing: the interval between tests, whether all patients received the same reference standard, and whether all patients were included in the analysis

QUADAS-2 also introduces something the other tools do not emphasize: applicability concerns. These are assessed for the first three domains.

Risk of bias vs applicability in QUADAS-2

These two judgments answer different questions, and beginners often merge them.

Risk of bias: could the study methods have distorted the accuracy estimate? A study where the reference standard was interpreted with knowledge of the index test result may report inflated sensitivity.

Applicability: does the study match your review question? A well-conducted study in a tertiary referral center may be at low risk of bias and still tell you very little about test performance in primary care, where disease prevalence is far lower.

A study can be at low risk of bias with high applicability concerns. Report them separately.

Key Risk-of-Bias Domains Every Beginner Should Understand

Some concepts appear across several tools under different names. Learning these once makes every tool easier.

Allocation concealment

Allocation concealment means the person enrolling a participant cannot know which arm that participant will be assigned to before they are enrolled.

This is not the same as blinding, and the distinction matters. Allocation concealment happens before and during assignment. Blinding concerns whether people know the assigned intervention after allocation has happened.

Why it matters: if a clinician can see that the next envelope says “placebo,” they may steer a sicker patient toward the treatment arm instead. That single decision undermines randomization.

Concealment is always possible in a randomized trial. Blinding sometimes is not, such as in surgical trials.

Blinding of outcome assessment

Ask whether the outcome is objective or subjective. Death, laboratory values, and imaging read by an independent radiologist are hard to influence. Pain, quality of life, and clinician-rated improvement are not.

Missing outcome data

Dropout introduces bias when the people who left differ systematically from those who stayed, in ways connected to the outcome. Check whether the reasons for missingness were reported, and whether they differed between arms.

Selective reporting bias

Three patterns to look for: outcomes that were registered but never published, outcomes reported in a different form than planned, and analyses where only one of several possible models appears in the paper.

Confounding

A third variable that influences both the exposure and the outcome, creating an association that is not causal. This is the central concern in observational studies and the reason ROBINS-I puts it first.

Step-by-Step Risk-of-Bias Assessment Workflow

Step 1: Identify the study design. Randomized trial, cohort, case-control, or diagnostic accuracy study. Everything else follows from this.

Step 2: Choose the appropriate tool. Do this before you have seen the results, not after. Choosing a tool once you know which studies favor your hypothesis is a route to biased judgments of bias.

Step 3: Pre-specify your approach in the protocol. State the tool, the version, how many reviewers, and your decision rules for handling unclear reporting.

Step 4: Assess the relevant domains. Work through the signalling questions rather than jumping straight to an overall impression.

Step 5: Record the supporting evidence. Do not write “high risk” and move on. Write what the study said, or did not say, that led you there. Quote the sentence and note the page. Keeping that evidence in a structured form from the start saves rebuilding it later. REDCap or a shared extraction sheet both work, provided every reviewer uses the same fields.

Step 6: Have two reviewers assess independently. Independent means without discussing each study first.

Step 7: Resolve disagreements. Discuss, then bring in a third reviewer if you cannot reach consensus.

Step 8: Use the judgments when interpreting the synthesis. This is the point of all the preceding steps.

How Should You Handle Disagreements Between Reviewers?

The standard process runs like this: reviewer 1 and reviewer 2 assess independently, then compare. Where they differ, they discuss and try to reach consensus. If they still disagree, a third reviewer arbitrates.

Disagreement is not a sign that something went wrong. Primary studies are often vague about their methods, and two careful readers can interpret the same ambiguous sentence differently. Record how many disagreements occurred and how you resolved them, since this is part of your methods.

How to Create a Traffic Light Plot for Risk of Bias

A traffic light plot is the standard way to display risk-of-bias results.

The structure is simple:

  • Rows: individual studies or results
  • Columns: the risk-of-bias domains from your tool
  • Colored circles: the judgment for each study in each domain, typically green for low risk, yellow for some concerns, and red for high risk

The plot lets a reader scan a column and see immediately that, for example, six of your nine trials had problems with outcome measurement. A table of the same data takes far longer to read.

The usual way to produce one is robvis.

What is Robvis?

Robvis is a tool built specifically for visualizing risk-of-bias assessments. It exists as an R package and as a free web application, so you do not need to code to use it.

The workflow is short. Enter your domain-level judgments into a spreadsheet, with one row per study and one column per domain. Upload or import that file. Select the template matching your tool. Download the figure.

Templates are available for RoB 2, ROBINS-I, and QUADAS-2, along with a generic option for other tools.

Traffic light plot vs summary plot

A traffic-light plot shows every study individually. You can trace one study across all domains.

A summary plot, sometimes called a weighted bar plot, shows the distribution of judgments within each domain across all studies. You lose the study-level detail and gain a quick overview.

Most reviews include both. The summary plot goes in the main text, the traffic-light plot often goes in a supplement.

7 Common Risk-of-Bias Assessment Mistakes Beginners Make

  1. Using one tool for every study design. RoB 2 on a cohort study will miss confounding entirely.
  2. Treating risk of bias as a quality score. They answer different questions.
  3. Assessing bias after seeing the meta-analysis. Knowing which studies favor your hypothesis will color your judgments, even when you try not to let it.
  4. Recording a judgment with no supporting evidence. A reader cannot check “high risk” on its own.
  5. Assuming lack of blinding means high risk. For objective outcomes, it often does not.
  6. Ignoring confounding in observational studies. It is usually the largest threat present.
  7. Reporting risk of bias and then never mentioning it again. A table in the results that has no effect on the discussion is decoration.

How Does Risk of Bias Affect a Meta-Analysis?

Risk-of-bias judgments are meant to change what you conclude, not just to fill a table before the discussion section.

There are several ways to use them:

  • Interpret certainty of evidence. Risk of bias is one of the domains that lowers certainty in a GRADE assessment.
  • Run sensitivity analyses. Repeat the meta-analysis excluding high-risk studies and see whether the pooled estimate holds. If it moves substantially, your readers need to know.
  • Explore heterogeneity. Differences in methodological quality can explain why studies disagree.
  • Conduct subgroup analysis where appropriate. Stratifying by overall risk of bias is a recognized approach.
  • Temper your conclusions. If every included trial is at high risk in the same domain, the summary estimate should be presented with that caveat attached.

RoB 2 guidance is explicit that overall judgments should influence interpretation, and it discusses stratifying meta-analysis by risk of bias as a way to do that.

Meta-Analysis

A meta-analysis inherits the problems of its inputs.

Sensitivity analysis, stratifying by risk of bias, exploring heterogeneity, feeding judgments into GRADE. These are the steps that turn a bias table into a conclusion, and they all sit on the analysis side. Twelve weeks working through real datasets with a mentor.

See the 12-Week Module →

How to Report Risk of Bias in Your Systematic Review

In your methods, state:

  • The tool used and its version
  • The number of reviewers and whether they assessed independently
  • How disagreements were resolved
  • Any predefined criteria or decision rules you applied

In your results, report:

  • Domain-level judgments for each study
  • The rationale supporting each judgment, usually in a supplementary table
  • Overall judgments where the tool provides them
  • A traffic-light plot, a summary plot, or both

In your discussion, explain:

  • How risk of bias affects your confidence in the findings
  • Whether sensitivity analyses changed the result

Risk-of-Bias Assessment Checklist for Students

Frequently Asked Questions

Q1. What is risk of bias assessment in a systematic review?

It is a structured evaluation of whether the way a study was designed, conducted, analyzed, or reported could have systematically distorted its results. It is done for each included study, and often for each specific result within a study, using a tool matched to the study design.

Q2. What are the RoB 2 domains?

RoB 2 has five: bias arising from the randomization process, bias due to deviations from intended interventions, bias due to missing outcome data, bias in measurement of the outcome, and bias in selection of the reported result. Each domain is judged as low risk, some concerns, or high risk.

Q3. What is the ROBINS-I confounding domain?

It assesses whether differences between the compared groups, present before or arising during the study, could explain the observed effect. Reviewers check whether the authors identified the important confounders, measured them adequately, and adjusted for them properly. It covers both baseline and time-varying confounding.

Q4. How is Newcastle-Ottawa Scale scoring done?

Stars are awarded across three areas: selection, comparability, and exposure or outcome, depending on whether the study is case-control or cohort. Most versions allow a maximum of nine stars. Report the stars by area rather than the total alone, since the total hides which items were lost.

Q5. What is a traffic light plot in systematic reviews?

A figure with studies as rows and risk-of-bias domains as columns, where each cell is a colored symbol showing the judgment. It lets readers see patterns across studies and domains at a glance.

Q6. What is robvis used for?

robvis generates risk-of-bias figures from your domain-level judgments. It produces traffic-light plots and weighted summary bar plots, with built-in templates for RoB 2, ROBINS-I, and QUADAS-2. It is available as an R package and as a free web app.

Q7. Which tool is used for diagnostic accuracy studies?

QUADAS-2. It assesses four domains, patient selection, index test, reference standard, and flow and timing, and separately rates applicability concerns for the first three.

Final Thoughts

Risk-of-bias assessment feels intimidating at first, mostly because of the number of tools and the amount of terminology attached to them. The underlying question is simpler than the paperwork suggests: could the way this study was run have pushed its result in a particular direction? Once you can answer that for a single study, applying RoB 2 or ROBINS-I becomes a matter of working through the domains in order. Start with one study, take notes on your reasoning, and compare with a second reviewer. That first round of disagreements will teach you more than any checklist. 

Systematic Review

A table in the results that has no effect on the discussion is decoration.

That is the mistake nobody catches until peer review, and it happens because risk of bias gets treated as a box to tick rather than a decision that shapes the conclusion. Same with the protocol, the search strategy, and the screening log. Each one only works if it was planned to carry weight.

The American Academy of Research & Academics supports students and researchers through every stage of a systematic review, from the first question through protocol, screening, appraisal and submission. Get in touch to find out how we can help with yours.

See the Full Course →

Disclaimer:

Articles published by American Academy of Research & Academics are prepared by our team using information from direct experience, publicly available resources, and educational references. AI tools may be used to assist with drafting, proofreading, and formatting; however, all content undergoes review and approval before publication.
The information provided is intended for educational purposes only. Requirements, policies, and processes may change over time. Readers should consult official sources for the most current information.

Facebook
Twitter
LinkedIn
Email
0

No products in the basket.