Funnel Plot Asymmetry Is Not Proof of Publication Bias

4 min

The funnel plot is one of the most reproduced figures in evidence synthesis and one of the most over-interpreted. An asymmetric funnel is routinely reported as demonstrating publication bias. It does not demonstrate that, and reviewers who know the literature will say so.

What the test actually measures

Egger's test regresses the standardised effect estimate on its precision. A significant result indicates that effect size is associated with study size: smaller studies report systematically different effects from larger ones.

The technical name for this is a small-study effect. Publication bias is one cause. It is not the only one.

The four explanations

When Egger's test is significant, at least four mechanisms could be responsible, and distinguishing them requires judgement rather than another test.

Publication bias. Small studies with null results are less likely to be published. This is real and well documented, and it is the explanation most readers assume.

Genuine clinical heterogeneity. Small trials often differ systematically from large ones. They tend to recruit more selected populations, deliver interventions more intensively, and follow participants more closely. Each of these can produce a genuinely larger effect. Nothing is hidden; the small trials are measuring something different.

Methodological quality. Smaller trials are, on average, less likely to conceal allocation adequately or to blind outcome assessment. Both inflate effect estimates. Here the asymmetry reflects bias within studies rather than bias in which studies exist.

Mathematical artefact. For odds ratios and standardised mean differences, the effect estimate and its standard error are not independent. This induces funnel asymmetry even with no bias of any kind. It is most pronounced when events are rare or baseline risk varies. For binary outcomes, the Peters test or the arcsine test are better behaved than Egger's for this reason.

When not to run it

With fewer than ten studies, the test has too little power to be informative, and a non-significant result says almost nothing.

The honest reporting is to state that the assessment was not performed because the number of studies was insufficient, and to note that publication bias therefore could not be excluded. This is a limitation, and writing it plainly is better than running an underpowered test and implying reassurance.

Substantial heterogeneity also undermines the test, because the regression assumes a common underlying effect. A significant Egger's test in the presence of I² of 85% is difficult to interpret in any direction.

What to write

The wording that survives review sets out the finding and the range of explanations without overclaiming:

Funnel plot asymmetry was assessed using Egger's regression test across the fourteen included trials. The test indicated evidence of small-study effects (p = 0.03). Publication bias is one possible explanation, although the smaller trials in this analysis also delivered a more intensive version of the intervention and were less likely to report allocation concealment, either of which could produce the same pattern. The certainty of evidence was downgraded one level for publication bias.

Three things make this work. It names the test and the number of studies. It offers competing explanations grounded in the actual characteristics of the included trials, not as boilerplate. And it carries the finding through to the GRADE assessment, which is where a reviewer looks to see whether the limitation changed anything.

On trim-and-fill and related methods

Trim-and-fill imputes studies to make the funnel symmetric, then recomputes the pooled estimate. It assumes the asymmetry is caused by missing studies, which is precisely the question at issue.

Report it as a sensitivity analysis. If the adjusted estimate is close to the original, that is worth a sentence and mildly reassuring. If it differs substantially, that is worth reporting too, as evidence that the conclusion is sensitive to assumptions about missing data.

What it is not is a corrected result. The unadjusted estimate remains your primary finding, and presenting a trim-and-fill estimate as the headline number is a reliable way to attract a methodological objection.

Common questions

How many studies do I need before running Egger's test?
At least ten as a working minimum, which is the threshold Cochrane recommends. Below that the test has very low power, and a non-significant result carries almost no information. With fewer than ten studies the appropriate action is to say the analysis was not conducted and explain why, rather than to run it and report a null finding as evidence of no bias.
Is a significant Egger's test proof of publication bias?
No. The test detects an association between effect size and study precision, which publication bias can produce but so can genuine clinical differences between small and large trials, poorer methodological quality in smaller studies, or the mathematical properties of certain effect measures. Publication bias is one explanation among several and should be presented as such.
Should I use trim-and-fill to correct my estimate?
Use it as a sensitivity analysis, never as a correction that replaces your primary result. Trim-and-fill imputes hypothetical missing studies under a specific and often wrong assumption about the mechanism. Report it as one exploration of robustness alongside the unadjusted estimate.