Insights / Topic · Bias & quality appraisal

Risk of Bias and Quality Appraisal in Systematic Reviews

How to choose and apply risk of bias tools, assess small-study effects and publication bias, and report the judgements so reviewers can follow them.

Notes
5
Total reading
23m
Overview

A meta-analysis is only as credible as the studies it pools. Risk of bias assessment asks whether the design and conduct of each included study could have pushed its result away from the truth, and it is the step statistical reviewers check first when they decide whether a pooled estimate can be believed.

The tool follows the design. RoB 2 is the current Cochrane tool for randomised trials and works outcome by outcome, across domains for the randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome and selection of the reported result. ROBINS-I covers non-randomised studies of interventions and asks the reviewer to compare each study with a hypothetical target trial. Diagnostic accuracy studies use QUADAS-2. Generic quality scores that add up points are discouraged, because a total hides which domain carries the problem.

Bias also operates across studies. When small studies report larger effects than large ones, a funnel plot looks asymmetric and Egger's test may flag it. Asymmetry is evidence of small-study effects, not proof of publication bias: genuine heterogeneity, differences in study quality and chance produce the same picture, and the test has little power below about ten studies.

The objections reviewers raise here are predictable. Assessments done by a single reviewer, a tool that does not match the study design, an overall judgement with no domain-level detail, or a high-risk study left in the main analysis with no sensitivity analysis excluding it. Each is avoidable at the planning stage, and each is far harder to fix once the manuscript is under review, because the assessments then have to be redone across every included study.

The notes in this topic cover choosing a tool, running assessments in duplicate, presenting traffic-light and summary plots, testing for small-study effects, and carrying risk of bias into sensitivity analyses and GRADE so that the judgements change the conclusions rather than sitting in a supplementary table.

Notes in this topic