Forty-two reviews of one question, sharing almost no trials

6 min

A preprint posted on 21 August takes a question that has been reviewed to exhaustion and asks why the reviews disagree.

Weibel and colleagues examined 42 systematic reviews of corticosteroids in sepsis published between 2015 and 2025, covering 121 unique randomised trials. More than half of the pairwise comparisons between those reviews shared no trials at all. Only three pairs showed high overlap.

Of the 38 reviews that included a meta-analysis of short-term mortality, 39% found a benefit and 61% found no effect.

The finding that makes this more than a curiosity is the last one. Discordance occurred exclusively among reviews of broad corticosteroid strategies. Reviews of narrowly specified regimens agreed with each other.

What that pattern rules out

The obvious explanations for reviews disagreeing are statistical: different effect measures, different heterogeneity models, different handling of zero-event trials, different inclusion of unpublished data. Those explanations predict disagreement scattered across the whole set. They do not predict disagreement that lands entirely on one side of a line drawn by how tightly the intervention was defined.

What the pattern does fit is an eligibility explanation. "Corticosteroids in sepsis" is not one question. It is a family of questions covering different drugs, doses, timings relative to shock onset, durations, tapering schedules, co-interventions and severity thresholds. Each review author drew that boundary somewhere. Different boundaries produced different trial sets, and different trial sets produced different answers, all of them arguably correct about the question actually asked and none of them answering the same question.

Which means the reviews are not really in conflict. They are talking past each other, and the meta-analytic machinery downstream is working faithfully on inputs that were never comparable.

The practical consequence for anyone planning a review

Three things follow, and all of them happen before a single record is screened.

Check the overlap first. If reviews of your question already exist, calculate how much they share before deciding whether another one is needed. The corrected covered area is the standard measure and it takes an afternoon once the citation matrices are assembled. A high result tells you the ground is covered. A low result tells you something more interesting: that existing reviews have been asking different questions, and that the useful contribution may be an overview that explains the divergence rather than a forty-third review that adds to it.

Specify the intervention at the level at which it is actually delivered. Broad definitions are attractive because they produce more included studies and, in the abstract, a more generalisable claim. This preprint is a direct argument that they also produce answers that do not replicate. If dose and timing plausibly change the effect, they belong in the eligibility criteria rather than in a subgroup analysis added later.

Decide before screening, not during. Eligibility boundaries that shift mid-screening in response to what the search returned are the mechanism by which two competent teams, working honestly on the same question, end up with non-overlapping evidence bases. Registering the protocol is the discipline that makes the boundary auditable.

If you proceed anyway, report the overlap

Sometimes a further review is justified even where several exist: the older ones have aged out, a large trial has since reported, or the existing set genuinely does answer a different question. In those cases the overlap analysis you did at the planning stage is worth publishing rather than filing.

A short paragraph in the introduction stating how many prior reviews exist, how much they share, and where your eligibility boundary sits relative to theirs does three things at once. It pre-empts the reviewer question about why another review was needed. It tells readers which of the earlier reviews yours supersedes and which it sits alongside. And it makes the discrepancy interpretable if your conclusion differs from a well-known predecessor, because the reader can see whether you were working from the same trials.

Most reviews of crowded questions assert novelty in a sentence. Showing the arithmetic takes a paragraph and is considerably more convincing.

Where the evidence on searching currently stands

A systematic review of search methods published in Research Synthesis Methods on 27 August is a useful companion piece. Clark, Beller, Forbes, Furuya-Kanamori and Sanders screened 3,520 records down to 14 studies covering 12 alternative approaches to searching.

Four approaches may improve searches: using seed studies to design the search strategy, applying study-design filters, backwards citation searching, and automation tools for translating a strategy across databases. Six may make searches worse. Two were inconclusive.

The caveat carries as much weight as the findings. Most included studies were at high or unclear risk of bias on a modified QUADAS-2, so these are provisional signals rather than recommendations. It is a striking state of affairs: the search is the foundation of every systematic review, and the evidence base for how to conduct one well consists of 14 studies, most of them methodologically fragile.

Abundant evidence is not the same as current evidence

A third item from late August illustrates the gap from another direction. Cochrane published an evidence and gap map on dengue prevention on 24 August, built with Cochrane Crowd, EPPI-Reviewer and EPPI Mapper.

The headline finding is not about dengue. Hundreds of primary studies exist. Community education alone accounts for 47 randomised trials and 140 non-randomised studies. Only a fraction of that evidence sits inside a current, high-quality synthesis. Reviews of adult vector control are outdated apart from work on Wolbachia, and behavioural and implementation outcomes for vaccines are close to absent.

That combination shows up repeatedly in mature fields: a large primary literature, several syntheses of varying age and quality, and no clear picture of which questions have actually been answered. Gap maps are unglamorous and slow, and they are the tool that tells you whether a proposed review is filling a gap or adding to a pile.

What to take from all this

If reviews of your question disagree, work out whether the disagreement is empirical or definitional before designing anything. An empirical disagreement is a research finding worth investigating. A definitional one means the reviews answered different questions, and the fix is a tighter question rather than a bigger search.

And if you are about to argue that your review is needed because the existing evidence is conflicting, check that the conflict is real. In the sepsis case, it substantially was not. Forty-two reviews, 121 trials, and most pairs sharing nothing at all.

References

Clark, J., Beller, E., Forbes, C., Furuya-Kanamori, L., & Sanders, S. (2026). Methods for conducting searches in healthcare reviews: A systematic review. Research Synthesis Methods, 1-18. https://doi.org/10.1017/rsm.2026.10112

Cochrane. (2026, August 24). New Cochrane gap map highlights big gaps in dengue prevention research. https://www.cochrane.org/about-us/news/new-cochrane-gap-map-highlights-big-gaps-dengue-prevention-research

Pieper, D., Antoine, S.-L., Mathes, T., Neugebauer, E. A. M., & Eikermann, M. (2014). Systematic review finds overlapping reviews were not mentioned in every other overview. Journal of Clinical Epidemiology, 67(4), 368-375. https://doi.org/10.1016/j.jclinepi.2013.11.007

Weibel, S., Dungfelder, H., Pscheidl, T., Krone, M., & Meybohm, P. (2026). Discordant evidence on corticosteroids in sepsis: A meta-research study [Preprint]. medRxiv. https://doi.org/10.64898/2026.08.20.26360343