Random or Fixed Effects: Choose Before You See the Data

4 min

Reviewers ask about model choice more often than almost any other statistical point, and the reason is that the answer is frequently given backwards. Researchers run a fixed effect model, observe high heterogeneity, switch to random effects, and report the second result. That sequence is the problem.

What the two models assume

The models differ in what they suppose your studies are estimating.

A fixed effect model assumes there is one true effect, identical across all included studies, and that observed differences arise entirely from sampling error. The pooled estimate is the best guess at that single value.

A random effects model assumes true effects vary between studies, drawn from a distribution. The pooled estimate is the mean of that distribution, and tau² describes its spread.

This is a claim about the world, not about your data. It follows from how similar your studies are in population, intervention, comparator, and outcome measurement, all of which you know at protocol stage.

Why the order matters to reviewers

If you choose the model after seeing heterogeneity, the choice is data dependent. You have made an analytical decision conditional on the result, which is the same class of problem as choosing an outcome after seeing which one reached significance.

In practice, switching to random effects almost always widens the confidence interval, so the switch usually makes the finding less impressive rather than more. The objection is not that you gained an advantage. It is that the analysis was not prespecified, which weakens everything else you prespecified.

State the model in the registration. One sentence is enough.

Which to choose

Random effects is the appropriate default for most clinical and public health reviews, because studies differ in ways that plausibly change the true effect: different populations, different doses, different follow-up, different settings.

A fixed effect model is defensible when the included studies are genuinely near-identical. Multiple sites of one multicentre trial, or a small set of

replications run to a common protocol, can justify it.

There is one further case worth naming. With very few studies, typically fewer than five, tau² is estimated so poorly that random effects intervals become unreliable. The answer is not to fall back to fixed effect, which assumes away the problem. Options include the Hartung-Knapp adjustment, which widens the interval appropriately, or reporting the studies narratively without pooling.

Choosing an estimator

Having chosen random effects, you also choose how tau² is estimated, and this detail is increasingly queried.

DerSimonian-Laird remains the most widely used and is the default in much software. It is known to underestimate uncertainty when studies are few or heterogeneity is substantial.

Restricted maximum likelihood (REML) is the better general default and is what most current methodological guidance recommends. Paule-Mandel is a reasonable alternative for binary outcomes.

Pairing REML with the Hartung-Knapp adjustment is a defensible standard choice for a review of moderate size, and naming both in the protocol closes off a whole line of reviewer questioning.

What to report

Report the model, the estimator, and the adjustment. Report tau² alongside I², because tau² is in the units of your effect measure and is interpretable. Report the prediction interval for any random effects analysis with more than a handful of studies.

A sensitivity analysis using the alternative model is genuinely useful and is not the same thing as switching. Present the prespecified model as the primary result and the alternative as a sensitivity check. If the two agree, that is reassuring and worth a sentence. If they disagree substantially, that itself is a finding: it usually means a small number of studies carry disproportionate weight, and identifying which ones is more informative than either pooled estimate.

Common questions

Should I switch to random effects because my I² is high?
No, and doing so is a recognised error. The model should follow from whether you believe the included studies estimate one common effect or a distribution of effects, and that judgement is made at protocol stage before any results are seen. Switching models after observing heterogeneity is a data-dependent decision that reviewers treat as a form of selective analysis.
Which model gives wider confidence intervals?
Random effects, whenever heterogeneity is present, because the interval incorporates between-study variance in addition to within-study error. When tau² is estimated as zero the two models give identical results. Random effects is therefore the more conservative default in most applied settings.
How do the models weight small studies differently?
Fixed effect weights are proportional to precision, so large studies dominate. Random effects weighting is more even, which gives small studies relatively more influence. This matters when small studies are systematically different from large ones, as they are when publication bias or small-study effects are present.