Insights / Topic · Bias & quality appraisal
How to choose and apply risk of bias tools, assess small-study effects and publication bias, and report the judgements so reviewers can follow them.
A meta-analysis is only as credible as the studies it pools. Risk of bias assessment asks whether the design and conduct of each included study could have pushed its result away from the truth, and it is the step statistical reviewers check first when they decide whether a pooled estimate can be believed.
The tool follows the design. RoB 2 is the current Cochrane tool for randomised trials and works outcome by outcome, across domains for the randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome and selection of the reported result. ROBINS-I covers non-randomised studies of interventions and asks the reviewer to compare each study with a hypothetical target trial. Diagnostic accuracy studies use QUADAS-2. Generic quality scores that add up points are discouraged, because a total hides which domain carries the problem.
Bias also operates across studies. When small studies report larger effects than large ones, a funnel plot looks asymmetric and Egger's test may flag it. Asymmetry is evidence of small-study effects, not proof of publication bias: genuine heterogeneity, differences in study quality and chance produce the same picture, and the test has little power below about ten studies.
The objections reviewers raise here are predictable. Assessments done by a single reviewer, a tool that does not match the study design, an overall judgement with no domain-level detail, or a high-risk study left in the main analysis with no sensitivity analysis excluding it. Each is avoidable at the planning stage, and each is far harder to fix once the manuscript is under review, because the assessments then have to be redone across every included study.
The notes in this topic cover choosing a tool, running assessments in duplicate, presenting traffic-light and summary plots, testing for small-study effects, and carrying risk of bias into sensitivity analyses and GRADE so that the judgements change the conclusions rather than sitting in a supplementary table.
Explainer
The step most teams skip is specifying the target trial. Without it, confounding has nothing to judge against and every rating becomes impressionistic.
Explainer
Counting domains or averaging judgements misuses the tool. How the overall RoB 2 rating is actually derived, and what to do with high-risk studies.
Explainer
AMSTAR 2 is not scored, and a single critical flaw drops a review to critically low. Which domains are critical, and what that means for an umbrella review.
Explainer
overlap, umbrella review, PICO, specification, search methods
Explainer
Egger's test detects small-study effects, which have several causes. How to run it, when it is uninformative, and what to claim from the result.