Subgroup Analysis Will Always Find Something

Run enough subgroups and one will reach significance. What makes a subgroup finding credible, how many is too many, and how to report the ones that fail.

5 min

Split a meta-analysis by sex, age band, intervention duration, intensity, setting, risk of bias and geography, and you have run seven tests. At a 5% significance level, the probability that at least one comes back significant by chance alone is around 30% even when no true effect modification exists.

This is why subgroup findings are treated with suspicion, and the suspicion is warranted. Oxman and Guyatt made the argument more than thirty years ago and the situation has improved only modestly since.

The practical question is what separates a subgroup finding worth reporting from a coincidence.

Prespecification is necessary and not sufficient

Prespecifying subgroups in the protocol removes the most obvious problem: choosing which comparison to highlight after seeing which one worked. It does not remove the multiplicity problem. Seven prespecified subgroups carry the same false-positive arithmetic as seven post hoc ones.

So prespecify, and prespecify few. A protocol listing three subgroup analyses, each with a stated rationale and a stated expected direction, is a stronger design than one listing twelve. The direction matters more than people expect. If you can say in advance that you expect the effect to be larger in supervised than unsupervised programmes, and it is, that is a genuine prediction confirmed. If you have no prior expectation, a difference in either direction can be rationalised after the fact, and the analysis has told you nothing.

The test is between subgroups, not within them

The most common reporting error: presenting a significant pooled effect in one subgroup and a non-significant effect in another, then concluding that the intervention works only in the first.

That comparison is invalid. A subgroup with fewer or smaller studies has wider confidence intervals and may fail to reach significance while having an identical point estimate. What you need is a formal test for interaction, or in a random-effects framework a test of the difference between subgroup estimates, reported with its own confidence interval.

If the interaction test is non-significant, the correct statement is that the review found no evidence of effect modification, not that the intervention worked in one group and not the other.

Seven questions that decide credibility

The ICEMAN instrument (Schandelmaier et al., 2020) formalises this and is worth using in full. Its core questions:

  1. Was the subgroup variable measured at baseline, before the intervention? Post-baseline variables can be affected by the intervention, which makes any subgroup difference uninterpretable.
  1. Was the comparison within studies or between them? Within-study comparisons, where each trial contributes participants to both subgroups, are far more credible. Between-study comparisons, where whole trials are allocated to one subgroup, are confounded by every other way those trials differ.
  1. Was the direction predicted in advance?
  1. Was there a plausible mechanism, ideally supported by external evidence?
  1. How many subgroup analyses were conducted in total?
  1. Is the interaction significant, and how large is it?
  1. Is the effect consistent across related outcomes?

A subgroup effect that passes most of these is worth a sentence in the abstract. One that fails several belongs in the discussion, labelled as hypothesis-generating.

A worked example

Twenty trials of a school-based physical activity programme, prespecified subgroup by programme duration, under 12 weeks against 12 weeks or more.

Shorter programmes: SMD 0.18, 95% CI −0.02 to 0.38, 8 trials. Longer programmes: SMD 0.41, 95% CI 0.22 to 0.60, 12 trials.

The tempting reading is that only longer programmes work. The interaction test gives p = 0.11.

The defensible reading: point estimates favour longer programmes, the difference is not statistically significant, the comparison is between studies rather than within them and so is confounded with everything else that differs between short and long programmes, and the finding is consistent with a dose-response mechanism but does not establish one. Duration warrants investigation in a trial designed to test it.

That paragraph will not make the abstract. It is also the only version that survives a statistical review.

Reporting the subgroups that found nothing

Report all of them. Every prespecified subgroup analysis, including those with null results, and a count of any post hoc ones with a clear label.

Selective reporting of subgroups is the mechanism by which the multiplicity problem becomes invisible. A reader who sees one significant subgroup out of one reported has no way to know it was one of nine. A reader who sees one out of nine can weigh it appropriately, which is what you want if the finding is real.

References

Oxman, A. D., & Guyatt, G. H. (1992). A consumer's guide to subgroup analyses. Annals of Internal Medicine, 116(1), 78–84. https://doi.org/10.7326/0003-4819-116-1-78

Higgins, J. P. T., & Thompson, S. G. (2004). Controlling the risk of spurious findings from meta-regression. Statistics in Medicine, 23(11), 1663–1682. https://doi.org/10.1002/sim.1752

Schandelmaier, S., Briel, M., Varadhan, R., Schmid, C. H., Devasenapathy, N., Hayward, R. A., Gagnier, J., Borenstein, M., van der Heijden, G. J. M. G., Dahabreh, I. J., Sun, X., Sauerbrei, W., Walsh, M., Ioannidis, J. P. A., Thabane, L., & Guyatt, G. H. (2020). Development of the Instrument to assess the Credibility of Effect Modification Analyses (ICEMAN) in randomized controlled trials and meta-analyses. CMAJ, 192(32), E901–E906. https://doi.org/10.1503/cmaj.200077

Common questions

How many subgroup analyses is too many?
There is no threshold, but three to five prespecified analyses in a review of moderate size is a reasonable working range. What matters more than the count is that the total is reported, so readers can adjust for multiplicity themselves.
Is meta-regression better than subgroup analysis?
For continuous modifiers such as mean age, duration or baseline severity, yes, because dichotomising loses information and the cut point is usually arbitrary. Meta-regression needs studies to be reasonably numerous; a common rule of thumb is at least ten studies per covariate, and even then the analysis is observational at the study level and subject to aggregation bias.
Can I use subgroup analysis to explain high heterogeneity?
That is the appropriate use, provided the subgroups were prespecified. Investigating heterogeneity is a better reason for subgrouping than hunting for effects. If the analyses do not explain it, say so and reflect the unexplained variation in the GRADE assessment.