Transitivity Is an Argument, Not a Test
Statistical tests check consistency. Transitivity is the clinical assumption underneath, and it has to be defended in prose before any network is fitted.
Network meta-analysis works by combining direct and indirect evidence. If trials compare A with B, and other trials compare B with C, the network estimates A against C even though no trial made that comparison.
That inference rests on transitivity: the assumption that participants in the A-versus-B trials could plausibly have been randomised into the B-versus-C trials instead. Put differently, that the trials differ in which treatments they compare but not in ways that modify the treatment effect.
Transitivity cannot be tested statistically. What can be tested is consistency, the agreement between direct and indirect estimates where both exist, and consistency is the observable footprint of transitivity rather than transitivity itself. A network can pass every consistency check and still violate the assumption, particularly in sparse networks where there is little direct evidence to disagree with.
So the defence has to be made in prose, and it has to be made before the model is fitted.
What the argument has to cover
Three groups of characteristics, each compared across the comparisons in the network rather than across studies generally.
Population. Are the participants in trials of one comparison systematically different in severity, age, comorbidity or setting? Sports and rehabilitation networks often mix trained athletes with sedentary clinical populations, and if one treatment has only been tested in one of those groups, transitivity is doubtful.
Intervention definition. Is the common comparator genuinely the same thing throughout? "Usual care" is the standard offender: usual care in a 2004 trial in one health system is not usual care in a 2023 trial in another, and if the shared node is heterogeneous the indirect comparisons running through it are unreliable. The same applies to dose and duration, which are often collapsed into a single node when they should be separate.
Trial characteristics. Year of publication, length of follow-up, outcome definition and timing, and risk of bias. A treatment evaluated only in older, smaller trials is compared indirectly against one evaluated in recent large trials, and the comparison inherits every difference between those eras.
How to present it
Build a table with one row per treatment comparison and one column per potential effect modifier, populated with the distribution of that modifier across trials making the comparison. Mean age, baseline severity, proportion female, follow-up duration, publication year.
The table shows the reader what you looked at and lets them judge for themselves. If a modifier is clearly imbalanced across comparisons, say so and describe what you did about it: restricting the network, splitting a node, or carrying it into a network meta-regression.
This table belongs in the manuscript rather than in supplementary material. It is the evidence for the assumption that the entire analysis depends on.
Then test consistency
With the argument made, the statistical checks follow.
Node-splitting compares direct and indirect estimates for each comparison where both exist, and it is the most interpretable check because it localises any problem.
Design-by-treatment interaction provides a global test of inconsistency across the network.
The loop-specific approach examines inconsistency within closed loops and is useful for identifying which part of the network is behaving oddly.
A caution about interpretation. These tests are underpowered in sparse networks. A non-significant result is weak evidence of consistency, not a clean bill of health, and should be reported as such rather than as confirmation that the assumption holds.
Rating certainty across the network
CINeMA provides a structured framework for assessing confidence in network meta-analysis results, covering within-study bias, reporting bias, indirectness, imprecision, heterogeneity and incoherence. It is the current standard and reviewers in methodologically strong journals will expect either CINeMA or an equivalent structured approach rather than an informal narrative.
One practical point: certainty is rated per comparison, not for the network as a whole. A network can support a high-certainty conclusion about A against B and a very low-certainty one about A against D, and collapsing that into a single statement about "the network" loses the information a reader needs.
The ranking problem
Treatment rankings, SUCRA values and rankograms are the output most likely to be misread. A treatment can rank first while its estimate is imprecise and statistically indistinguishable from the second, third and fourth.
Rankings should be reported alongside the effect estimates and their intervals, never on their own, and the discussion should resist language implying that the top-ranked treatment is established as best. The honest statement is usually that several treatments are compatible with being the most effective and the evidence does not separate them.
References
Chaimani, A., Higgins, J. P. T., Mavridis, D., Spyridonos, P., & Salanti, G. (2013). Graphical tools for network meta-analysis in STATA. PLoS ONE, 8(10), e76654. https://doi.org/10.1371/journal.pone.0076654
Hutton, B., Salanti, G., Caldwell, D. M., Chaimani, A., Schmid, C. H., Cameron, C., Ioannidis, J. P. A., Straus, S., Thorlund, K., Jansen, J. P., Mulrow, C., Catalá-López, F., Gøtzsche, P. C., Dickersin, K., Boutron, I., Altman, D. G., & Moher, D. (2015). The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions: Checklist and explanations. Annals of Internal Medicine, 162(11), 777–784. https://doi.org/10.7326/M14-2385
Nikolakopoulou, A., Higgins, J. P. T., Papakonstantinou, T., Chaimani, A., Del Giovane, C., Egger, M., & Salanti, G. (2020). CINeMA: An approach for assessing confidence in the results of a network meta-analysis. PLoS Medicine, 17(4), e1003082. https://doi.org/10.1371/journal.pmed.1003082
Salanti, G. (2012). Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: Many names, many benefits, many concerns for the next generation evidence synthesis tool. Research Synthesis Methods, 3(2), 80–97. https://doi.org/10.1002/jrsm.1037
Common questions
- How many studies do I need for a network meta-analysis?
- There is no minimum, but a network with several comparisons supported by a single small trial each will produce estimates too imprecise to be useful, and consistency will be untestable in most loops. Assess whether the network is connected and whether the key comparisons have any direct evidence before committing to the design.
- Can I do a network meta-analysis if the network is disconnected?
- Not as a single network. Disconnected components cannot be compared. Options are to analyse the components separately, to find studies that bridge them, or in some cases to use matching-adjusted or unanchored methods, which carry much stronger assumptions and should be presented as exploratory.
- Does high heterogeneity rule out network meta-analysis?
- Not automatically, but it complicates the transitivity argument, since the sources of heterogeneity are frequently the same characteristics that would act as effect modifiers. Investigate them, and if a modifier is distributed unevenly across comparisons, consider network meta-regression or a restricted network.
