What I² Actually Tells You about Heterogeneity

4 min

Almost every meta-analysis reports I². Far fewer interpret it correctly, and the mistake is consistent enough that statistical reviewers have a stock objection ready. This note covers what the statistic measures, the three errors that appear most often, and what to write when a reviewer challenges a high value.

What I² measures

I² is the percentage of total variation across studies that is due to real heterogeneity rather than sampling error. It is derived from Cochran's Q:

I² = 100% × (Q − df) / Q

where Q is the chi-squared heterogeneity statistic, and df is the degrees of freedom. When Q is smaller than its degrees of freedom, I² is set to zero rather than reported as negative. The critical word is percentage. I² is a ratio, not a quantity. It tells you what share of the variability is real, and says nothing about how large that variability is in absolute terms.

The three common errors

Treating 75% as a threshold for abandoning the analysis

The familiar 25/50/75 tiers for low, moderate and high heterogeneity come from the original Higgins and Thompson paper, where they were offered as tentative descriptive labels. They were never intended as decision rules, and the authors said so explicitly.

A high I² is a signal to investigate, not a verdict. The appropriate response is subgroup analysis or meta-regression to explain the variation, not withdrawal of the pooled estimate.

Reading I² as a measure of absolute inconsistency

Because I² is relative, it rises as the studies within a meta-analysis become more precise, even when the true spread of effects is unchanged. Pool ten large trials and a modest real difference between them produces a high I². Pool ten small trials with the same real difference and sampling error dominates, so I² falls.

This is why tau² and the prediction interval matter more for interpretation. Tau² is expressed in the units of your effect measure and answers the question a clinician actually has: how much do true effects differ?

Reporting I² without its uncertainty

I² is an estimate, and with few studies it is a very imprecise one. A meta-analysis of five trials reporting "I² = 0%" is frequently reporting a point estimate whose confidence interval runs from 0% to 80%. Report the interval.

What to write when a reviewer objects

The objection usually arrives in one of two forms. The first is "heterogeneity is substantial (I² = 78%); pooling is inappropriate." The second is "the authors should explain the source of heterogeneity."

Both are answerable, and neither requires abandoning the meta-analysis. A response that works has three parts:

  1. Restate the basis for pooling. Pooling was justified on clinical and methodological grounds, decided at protocol stage, before any effect estimates were seen. Say so and cite your registration.
  2. Give the absolute measure. Report tau² and the 95% prediction interval. The prediction interval is the honest summary of what a new study might find, and it communicates heterogeneity far better than I².
  3. Show the investigation. Present the prespecified subgroup analyses or meta-regression. If heterogeneity remains unexplained, say that plainly and reflect it in the GRADE assessment by downgrading for inconsistency.

The last point is where many responses fail. A reviewer is rarely asking you to make the heterogeneity disappear. They are asking you to demonstrate that you noticed it, investigated it, and let it affect your confidence in the conclusion.

A worked example

Suppose twelve trials of an exercise intervention give a pooled standardised mean difference of −0.42 (95% CI −0.61 to −0.23), with I² = 76% and tau² = 0.09.

The confidence interval describes uncertainty around the average effect. The prediction interval, which incorporates tau², runs roughly −1.04 to 0.20. That interval crosses zero.

These two facts are not in conflict, and stating both is the correct reporting. The average effect is clearly beneficial. A new trial in a new population could nonetheless find no benefit. A reader planning to apply this evidence needs the second fact as much as the first, and it is invisible if you report I² alone.

Common questions

Is an I² of 80% too high to pool studies?
Not automatically. I² describes the proportion of observed variation attributable to real differences between studies rather than chance. A high value tells you the effect varies across studies; it does not tell you the pooled estimate is wrong or that pooling was inappropriate. The decision to pool rests on clinical and methodological similarity, which you judge before seeing the data.
What is the difference between I² and tau²?
Tau² estimates the absolute variance of true effects between studies, in the units of your effect measure. I² is a relative percentage and depends heavily on the precision of the included studies. Two meta-analyses with identical tau² can report very different I² values if one contains larger trials. Report both.
Does a low I² mean my studies agree?
Not necessarily. I² is calculated from the studies you included, and with few small studies the confidence interval around it is very wide. An I² of 0% from four small trials is weak evidence of consistency. Report the confidence interval around I² rather than the point estimate alone.