AMSTAR 2 Has Seven Critical Domains

4 min

If your review includes other systematic reviews as its unit of analysis, you need a tool for appraising them. AMSTAR 2 is the usual choice, and it is used incorrectly about as often as it is used.

The most common error is scoring. AMSTAR 2 has 16 items, and adding them up to produce a total out of 16 misuses the instrument. Shea and colleagues (2017) were explicit that an overall score is not appropriate, because the items are not equally important.

The seven critical domains

Weakness in a critical domain undermines the validity of the review's conclusions regardless of how the other items score. The developers identified seven:

Protocol registered before the review commenced (item 2). Adequacy of the literature search (item 4). Justification for excluding individual studies (item 7). Risk of bias assessment of the included studies (item 9). Appropriateness of the meta-analytical methods (item 11). Consideration of risk of bias when interpreting results (item 13). Assessment of publication bias and its likely impact (item 15).

How the overall rating works

Four categories, derived by rule rather than by arithmetic.

High: no or one non-critical weakness. Moderate: more than one non-critical weakness, no critical flaws. Low: one critical flaw, with or without non-critical weaknesses. Critically low: more than one critical flaw.

Notice the consequence. A review with a superb search, careful risk-of-bias assessment and appropriate methods, but no registered protocol and no assessment of publication bias, has two critical flaws and rates critically low. That feels harsh and it is deliberate: those two omissions leave open the possibility that the review's analytic choices were shaped by its results.

What to do with critically low reviews in an umbrella review

Three options, and one of them should be prespecified.

Exclude them, which requires the decision to be made in the protocol rather than after seeing which reviews point which way.

Include them and stratify the synthesis, presenting findings by AMSTAR 2 rating and examining whether conclusions differ by quality.

Include them and downweight them narratively, stating that conclusions rest primarily on the higher-rated reviews.

The second is usually the most informative, because whether low-quality reviews reach different conclusions from high-quality ones is a finding in itself, and one that umbrella reviews are well placed to report.

The item people fudge

Item 7 asks whether the authors provided a list of excluded studies with justification for each exclusion at the full-text stage. Almost no review does this, and it is a critical domain.

The result is that a large proportion of published systematic reviews rate low or critically low on AMSTAR 2. That is a real finding about the literature rather than a defect of the instrument, and it is worth stating plainly in an umbrella review rather than adjusting the criteria until the ratings look more comfortable.

It also has an obvious implication for your own reviews: publish the list of full-text exclusions with reasons as a supplementary file. It takes one table and it is the cheapest available improvement to how your review will be appraised.

References

Shea, B. J., Reeves, B. C., Wells, G., Thuku, M., Hamel, C., Moran, J., Moher, D., Tugwell, P., Welch, V., Kristjansson, E., & Henry, D. A. (2017). AMSTAR 2: A critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ, 358, j4008. https://doi.org/10.1136/bmj.j4008

Whiting, P., Savović, J., Higgins, J. P. T., Caldwell, D. M., Reeves, B. C., Shea, B., Davies, P., Kleijnen, J., & Churchill, R. (2016). ROBIS: A new tool to assess risk of bias in systematic reviews was developed. Journal of Clinical Epidemiology, 69, 225–234. https://doi.org/10.1016/j.jclinepi.2015.06.005

Pieper, D., Antoine, S.-L., Mathes, T., Neugebauer, E. A. M., & Eikermann, M. (2014). Systematic review finds overlapping reviews were not mentioned in every other overview. Journal of Clinical Epidemiology, 67(4), 368–375. https://doi.org/10.1016/j.jclinepi.2013.11.007

Common questions

Can I use AMSTAR 2 on non-randomised reviews?
Yes. AMSTAR 2 was designed for reviews including randomised trials, non-randomised studies, or both, and several items have parallel wording for each. Note in the methods which version of each item you applied.
Is ROBIS an alternative?
ROBIS assesses risk of bias in a review rather than methodological quality, which is a related but distinct construct, and it is domain-based rather than item-based. Either is acceptable; using both is redundant. State which you chose and why, since a reviewer may have a preference.
Should two people apply AMSTAR 2 independently?
Yes, as with any appraisal instrument. Several items involve judgement, particularly items 4, 11 and 15, and agreement between assessors is worth reporting. Where one assessor is working alone, say so in the limitations.