ROBINS-I Starts with a Target Trial

The step most teams skip is specifying the target trial. Without it, confounding has nothing to judge against and every rating becomes impressionistic.

4 min

ROBINS-I assesses risk of bias in non-randomised studies of interventions by comparing each study against a hypothetical randomised trial that would have answered the same question. That comparison is the mechanism of the whole tool.

Which is why the step teams skip most often, specifying the target trial, is the step that makes the rest of the assessment meaningful. Without it, the confounding domain has nothing concrete to judge against, and the assessor falls back on a general impression of study quality, which is what the tool was built to replace.

Specifying the target trial

Before assessing any study, write down the trial you would run if ethics and funding were no obstacle: the eligible population, the intervention and its comparator, how assignment would happen, when follow-up would start, and the outcome and its timing.

Fifteen minutes of work, and it changes the assessment in three ways.

It makes the confounding domain answerable. Confounding is defined relative to the assignment mechanism in the target trial, so you can now list the prognostic factors that randomisation would have balanced. Each included study is then judged on whether it measured and adjusted for them.

It exposes time-zero problems. In the target trial, follow-up begins at assignment. Observational studies frequently start the clock somewhere else, and the resulting immortal time bias is invisible unless you have the comparison in front of you.

It clarifies the comparator. "No intervention" in an observational dataset is often a mixture of people who declined, people never offered, and people receiving something else. The target trial forces that to be named.

The seven domains, and where the weight sits

ROBINS-I covers confounding, selection of participants, classification of interventions, deviations from intended interventions, missing data, measurement of outcomes, and selection of the reported result.

The first three concern issues that arise before or at the start of the intervention; the rest concern what happens afterwards. In practice, confounding does most of the work. A study that has not adjusted for the factors your target trial would have balanced is at serious risk regardless of how well the remaining domains score.

The rating scale differs from RoB 2 and the difference is deliberate: low, moderate, serious, critical, or no information. Low risk means the study is comparable to a well-conducted randomised trial. That is a demanding standard, and most observational studies do not meet it. Moderate is the realistic ceiling for a sound observational study, and there is no reason to treat that as a failure.

Critical means the study is too problematic to provide useful evidence. Cochrane guidance is that critical-risk studies should not be included in the synthesis. This is worth deciding at protocol stage, because dropping studies after seeing their results is not a defensible sequence.

Confounding and adjustment are not the same thing

A frequent assessment error is to treat any multivariable model as adequate adjustment. The question is not whether the authors adjusted, but whether they adjusted for the right things, measured them adequately, and used a method that can plausibly handle them.

Three checks. Were the confounders on your prespecified list actually measured, or were proxies used? Were they measured before the intervention, or could they have been affected by it, in which case adjustment introduces bias rather than removing it? Is there residual confounding the authors acknowledge, such as unmeasured socioeconomic factors or baseline fitness?

A study adjusting for twenty variables, none of which is the main confounder your target trial identified, is at serious risk of bias, and its long covariate list should not be mistaken for rigour.

References

Hernán, M. A., & Robins, J. M. (2016). Using big data to emulate a target trial when a randomized trial is not available. American Journal of Epidemiology, 183(8), 758–764. https://doi.org/10.1093/aje/kwv254

Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., & Welch, V. A. (Eds.). (2024). Cochrane handbook for systematic reviews of interventions (Version 6.5). Cochrane. https://training.cochrane.org/handbook

Sterne, J. A. C., Hernán, M. A., Reeves, B. C., Savović, J., Berkman, N. D., Viswanathan, M., Henry, D., Altman, D. G., Ansari, M. T., Boutron, I., Carpenter, J. R., Chan, A.-W., Churchill, R., Deeks, J. J., Hróbjartsson, A., Kirkham, J., Jüni, P., Loke, Y. K., Pigott, T. D., … Higgins, J. P. T. (2016). ROBINS-I: A tool for assessing risk of bias in non-randomised studies of interventions. BMJ, 355, i4919. https://doi.org/10.1136/bmj.i4919

Common questions

Can I pool randomised and non-randomised studies in one meta-analysis?
You can, and usually you should not, at least not as the primary analysis. The two designs answer subtly different questions and carry different bias structures. The usual approach is to synthesise them separately and compare, which is more informative than a single pooled estimate and avoids a randomised effect being pulled by observational studies with residual confounding. If you do combine them, prespecify it and present the separate estimates as well.
What if a study reports no information for a domain?
Use the "no information" rating rather than defaulting to serious. It is a distinct judgement and it tells the reader something real: the study did not report enough to assess. Do not treat it as equivalent to low risk when deriving the overall rating.
Is ROBINS-I appropriate for cross-sectional studies?
Generally not. ROBINS-I is built for studies of interventions with a clear time sequence between exposure and outcome, which cross-sectional designs lack. ROBINS-E covers exposure studies, and for prevalence or descriptive cross-sectional work, JBI checklists or the Newcastle-Ottawa Scale are more commonly used. Say in the protocol which tool applies to which design.