RoB 2 Is Not a Score

Counting domains or averaging judgements misuses the tool. How the overall RoB 2 rating is actually derived, and what to do with high-risk studies.

5 min

The most common error in risk-of-bias assessment is arithmetic. A team assesses five domains, finds three at low risk and two with some concerns, and reports the trial as moderate quality, or assigns points and produces a total.

RoB 2 does not work that way, and the algorithm it does use is not a compromise between the domains. It is closer to a weakest-link rule.

How the overall judgement is derived

RoB 2 assesses five domains for a randomised trial: the randomisation process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. Each domain is judged low risk, some concerns, or high risk, and each judgement is reached through signalling questions rather than global impression.

The overall judgement then follows three rules. Low risk requires every domain to be low risk. Some concerns applies when at least one domain raises some concerns and none is high risk. High risk applies when any single domain is high risk, or when multiple domains at some concerns substantially lower confidence in the result.

One high-risk domain makes the trial high risk. Four excellent domains do not offset it, because bias in a single domain is sufficient to distort the effect estimate, and a well-conducted randomisation sequence does not repair an outcome measured by unblinded assessors on a subjective scale.

The assessment is per outcome, not per trial

RoB 2 is applied to a specific result, not to a paper. The same trial can be at low risk for all-cause mortality and high risk for patient-reported pain, because blinding of outcome assessors matters enormously for one and not at all for the other.

Reviews that report a single risk-of-bias rating per study have skipped this. It matters most in exercise, behavioural and rehabilitation trials, where blinding participants is usually impossible and the outcomes range from objective performance measures to self-reported questionnaires within a single study.

If your review has three outcomes in its Summary of Findings table, you have three sets of risk-of-bias judgements for each trial.

Two domains that get rushed

Domain 2, deviations from intended interventions, requires you to decide first which effect you are estimating: assignment to intervention, or adherence to it. Nearly every review wants the effect of assignment, which is the intention-to-treat question. Signalling questions differ depending on which you choose, and answering them without making the choice explicit produces judgements that cannot be interpreted.

Domain 5, selection of the reported result, is the one most often marked "some concerns" by default because nobody looked. Answering it properly means comparing the reported outcomes against a prespecified plan: a registry entry, a published protocol, or a statistical analysis plan. If no protocol is available, that is itself the finding, and it should be stated rather than converted silently into a neutral rating.

Tooling helps here. TrialScout, published in August 2026, uses a language model to match published trial reports to their registry entries, which removes most of the tedium from locating the registration in the first place.

What to do with high-risk studies

Excluding them is one option and it needs prespecifying, because a decision to exclude taken after seeing which way the high-risk trials point is not a methodological decision.

The more common approach is to include them and handle the risk in two places. A prespecified sensitivity analysis restricted to low-risk studies shows whether the pooled result depends on the weaker evidence. And the GRADE assessment downgrades for risk of bias where the studies contributing most weight are at high risk.

Those two steps together answer the question a reader has, which is not "were the trials perfect" but "does the conclusion survive if the weakest trials are removed".

A worked example

Twelve trials of a supervised exercise intervention for knee osteoarthritis, outcome self-reported pain at 12 weeks.

Participants cannot be blinded to whether they are exercising. Outcome assessors are the participants themselves. Domain 4 is therefore at high risk for most of these trials, and no amount of allocation concealment changes that.

Reporting all twelve as high risk is correct but uninformative on its own. Useful reporting adds that the high-risk judgement stems from a single unavoidable domain common to the field, that the same trials are at low risk for the objectively measured six-minute walk distance, and that a sensitivity analysis restricted to trials with blinded assessors of objective outcomes gives a pooled estimate in the same direction. GRADE downgrades once for risk of bias, and the discussion says why.

References

Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., & Welch, V. A. (Eds.). (2024). Cochrane handbook for systematic reviews of interventions (Version 6.5). Cochrane. https://training.cochrane.org/handbook

Sterne, J. A. C., Savović, J., Page, M. J., Elbers, R. G., Blencowe, N. S., Boutron, I., Cates, C. J., Cheng, H.-Y., Corbett, M. S., Eldridge, S. M., Emberson, J. R., Hernán, M. A., Hopewell, S., Hróbjartsson, A., Junqueira, D. R., Jüni, P., Kirkham, J. J., Lasserson, T., Li, T., … Higgins, J. P. T. (2019). RoB 2: A revised tool for assessing risk of bias in randomised trials. BMJ, 366, l4898. https://doi.org/10.1136/bmj.l4898

von Schreeb, S., Bruckner, T., DeVito, N. J., Axfors, C., Ioannidis, J. P. A., & Nilsonne, G. (2026). TrialScout: A large language model tool for matching trial reports to registry entries. Journal of Clinical Epidemiology. https://doi.org/10.1016/j.jclinepi.2026.112484

Common questions

Can I convert RoB 2 judgements into a numerical score for meta-regression?
No. The judgements are ordinal categories reached through a structured algorithm, and treating them as an interval scale assumes distances between categories that the tool does not define. If you want to examine whether bias relates to effect size, use the domain-level judgements as subgroups, or restrict to a prespecified low-risk subset.
Do I need two independent assessors?
It is the expected standard and reviewers will ask. Where a review is conducted by a single assessor, that should be stated as a limitation in the methods rather than left for the reader to infer. Some teams use a second assessor on a random sample and report agreement, which is weaker than full duplicate assessment but considerably better than nothing.
Should I use RoB 2 or the original RoB tool?
RoB 2 for randomised trials, unless you are updating a review that used the original tool and consistency across included studies matters more than currency. Mixing the two within one review makes the results incomparable across studies.