What SUCRA Can and Cannot Tell You

Ranking statistics are the most misread output of a network meta-analysis. A treatment can rank first on evidence that cannot distinguish it from fourth.

4 min

Network meta-analysis produces rankings, and rankings are what readers remember. The surface under the cumulative ranking curve, SUCRA, condenses a treatment's ranking distribution into a single number between 0 and 1, with 1 meaning certainly best.

The number is easy to read and easy to over-read.

What the statistic is

SUCRA is derived from the probability that a treatment occupies each possible rank, accumulated across ranks. P-score is the frequentist analogue developed by Rücker and Schwarzer (2015) and is interpreted the same way.

Two properties matter for interpretation. It is relative: it depends entirely on which other treatments are in the network, and adding or removing a comparator changes every SUCRA value. And it takes no account of effect magnitude: a treatment can achieve a high SUCRA on a trivial advantage over the alternatives.

Why rankings mislead

A treatment supported by one small imprecise trial can rank first, because its point estimate is favourable and its uncertainty is wide enough to give it substantial probability of being best. A treatment supported by four large trials with a slightly smaller effect ranks below it.

Trinquart and colleagues (2016) examined uncertainty in treatment rankings and found it commonly large enough that the ranking cannot be relied on. Salanti, Ades and Ioannidis (2011), who introduced SUCRA, were clear that it is a summary of the ranking distribution and not a substitute for the effect estimates.

The result is that the sentence readers take away from a network meta-analysis, that treatment X ranked highest, is frequently not supported by the analysis it came from.

How to report rankings responsibly

Present effect estimates first, rankings second. The league table of pairwise estimates with confidence intervals is the primary result. SUCRA is a summary of it.

Report SUCRA alongside the certainty rating for each comparison. A high SUCRA with very low CINeMA confidence should not be reported without that qualifier attached.

Say how many treatments could plausibly be best. Where the top four SUCRA values are 0.81, 0.78, 0.74 and 0.71, the analysis does not distinguish them, and the honest statement is that four treatments were compatible with being the most effective.

Include the rankogram, or the cumulative ranking curves, rather than only the SUCRA values. It shows the spread of the ranking distribution, which the single number hides.

Avoid superlatives in the abstract. "Treatment X had the highest SUCRA value (0.82), though estimates for the four highest-ranked treatments overlapped substantially" is accurate. "Treatment X was the most effective intervention" is not.

Mbuagbaw and colleagues' framework

Mbuagbaw and colleagues (2017) set out approaches to interpreting and choosing among treatments in network meta-analyses, and their central point is that ranking alone is an inadequate basis for a recommendation.

A defensible interpretation combines the effect estimate against a clinically meaningful threshold, the certainty of the evidence for that comparison, and the ranking, in that order. Where a decision must be made, factors outside the network usually matter too: cost, availability, acceptability and harms, most of which are not in the model.

References

Mbuagbaw, L., Rochwerg, B., Jaeschke, R., Heels-Ansdell, D., Alhazzani, W., Thabane, L., & Guyatt, G. H. (2017). Approaches to interpreting and choosing the best treatments in network meta-analyses. Systematic Reviews, 6, 79. https://doi.org/10.1186/s13643-017-0473-z

Rücker, G., & Schwarzer, G. (2015). Ranking treatments in frequentist network meta-analysis works without resampling methods. BMC Medical Research Methodology, 15, 58. https://doi.org/10.1186/s12874-015-0060-8

Salanti, G., Ades, A. E., & Ioannidis, J. P. A. (2011). Graphical methods and numerical summaries for presenting results from multiple-treatment meta-analysis: An overview and tutorial. Journal of Clinical Epidemiology, 64(2), 163–171. https://doi.org/10.1016/j.jclinepi.2010.03.016

Trinquart, L., Attiche, N., Bafeta, A., Porcher, R., & Ravaud, P. (2016). Uncertainty in treatment rankings: Reanalysis of network meta-analyses of randomized trials. Annals of Internal Medicine, 164(10), 666–673. https://doi.org/10.7326/M15-2521

Common questions

Is a SUCRA of 0.9 meaningfully better than 0.7?
Not necessarily. SUCRA depends on the treatments included and on the precision of each estimate, so a difference in values does not translate into a difference in effect. Compare the effect estimates and their intervals, not the SUCRA values.
Should I report P-score or SUCRA?
Whichever your framework produces: SUCRA in Bayesian analyses, P-score in frequentist ones. They are near-equivalent in interpretation. Report one, name it, and do not present both as though they were independent evidence.
Can rankings change if I add a treatment to the network?
Yes, and substantially, because ranking is relative to the set of comparators. This is a reason to define the network's scope in the protocol and to avoid adding treatments once the results are visible.