Insights / Topic · Heterogeneity
What Q, I², tau² and prediction intervals each measure, how to choose between fixed and random effects, and how to answer a reviewer about heterogeneity.
Studies asking the same question rarely return the same answer. Some of the spread is sampling error; the rest is heterogeneity, real differences in populations, interventions, comparators, outcomes or methods. How a review measures that variation, and what it does with it, shapes both the pooled estimate and how far a reader can apply it.
Each statistic answers a different question. Cochran's Q tests whether the variation exceeds what chance would produce, but it has low power when there are few studies and excessive power when there are many. I² describes the proportion of observed variability that reflects true differences rather than chance; it is not a measure of how large those differences are, and it rises as studies become more precise. Tau² estimates the between-study variance on the scale of the effect, and a prediction interval translates it into the range of effects a new study might show, which is often the most useful number for a clinical reader.
Model choice is a separate decision. A common-effect model assumes one true effect; a random-effects model assumes a distribution of true effects. The choice should follow from what the studies are expected to estimate and be stated in the protocol, not be switched after seeing I². Estimator choices matter as well: REML is generally preferred over DerSimonian-Laird for tau², and the Hartung-Knapp adjustment gives more honest confidence intervals when studies are few.
Reviewers return to heterogeneity more than any other topic. The usual requests are to justify the model choice, to report tau² and a prediction interval alongside I², to explain a high I² rather than only describe it, and to stop treating an I² threshold as permission to pool or a reason to refuse. A clear, pre-specified plan for each of these answers most of them before they are asked, and turns a long response letter into a short one.
The notes in this topic explain how to interpret these quantities, how to report them, when heterogeneity undermines pooling and when it does not, and how to move from describing heterogeneity to explaining it through pre-specified subgroup analysis and meta-regression.
Explainer
Adding 0.5 to every cell is the default in most software, and it isn't always right. What continuity corrections do to your estimate, and the alternatives.
Explainer
DerSimonian-Laird is the default in most software and performs poorly with few studies. What the alternatives do, and why the Hartung-Knapp adjustment matters.
Explainer
Run enough subgroups and one will reach significance. What makes a subgroup finding credible, how many is too many, and how to report the ones that fail.
Explainer
A confidence interval describes uncertainty around the average effect. A prediction interval describes what a new study might find.
Explainer
Ten studies per covariate is the working rule, and most reviews break it. What meta-regression can support and how to report a model that found nothing.
Explainer
I² above 75% does not invalidate your meta-analysis. What the statistic measures, what it does not, and how to answer a reviewer who objects.
Explainer
The model choice is an assumption about what your studies estimate, not a response to heterogeneity. Why reviewers check, and how to justify it.