SMD or MD: Decide Before You See the Scales
The standardised mean difference solves one problem and creates two more. When to use it, when to convert back, and why the choice belongs in the protocol.
Twelve trials measure the same construct with six different instruments. The mean difference cannot pool them, so the standardised mean difference does, by dividing each study's difference by its standard deviation.
That solves the problem of incompatible scales. It also introduces two that are less obvious.
Problem one: the denominator carries the population
An SMD is a ratio, and the denominator is the standard deviation within the study population. A trial in a narrow, homogeneous sample has a small SD, so a modest raw difference becomes a large SMD. A trial in a heterogeneous sample produces the reverse.
Which means SMDs are not straightforwardly comparable across trials whose populations differ in spread, even when both used the same instrument. Some apparent heterogeneity in SMD meta-analyses is variation in the denominator rather than variation in treatment effect.
Problem two: nobody can interpret it
An SMD of 0.42 means nothing to a clinician deciding whether to offer an intervention. Cohen's conventional thresholds of 0.2, 0.5 and 0.8 are widely quoted and were offered as rough descriptive labels for a different purpose, not as clinical decision rules.
The fix is to convert back. Multiply the pooled SMD and its interval by the standard deviation of a representative, familiar instrument, and report the result in the units of that scale alongside the SMD. Murad and colleagues (2019) set out this and several other approaches, including converting to a minimal important difference unit or reporting the proportion of patients achieving a threshold response.
The conversion belongs in the Summary of Findings table. That is where the reader deciding something will look.
When to use which
Use the mean difference when every study reports the outcome on the same scale, with the same direction and comparable timing. Same scale means the same instrument, not the same construct. Two depression inventories are not the same scale.
Use the SMD when studies measure the same construct on different instruments and you have a defensible argument that the instruments are measuring the same thing. That argument matters. Pooling a symptom checklist with a clinician-rated severity scale produces a number, and whether it means anything depends on whether the instruments correlate well.
Use the ratio of means as an alternative where the outcome is positive and ratio-scaled, such as distance walked or time to exhaustion. It sidesteps the denominator problem and is interpretable as a percentage change, though it is less familiar to reviewers and will need a sentence of justification.
The direction problem
When scales point in opposite directions, higher being better on one and worse on another, multiply one set of means by −1 before pooling. This is routine, easy to forget, and produces a pooled estimate near zero when it is forgotten, because the effects cancel.
Build the direction check into the extraction form as an explicit field rather than trusting it to the analysis stage.
Prespecify it
The choice between MD and SMD should sit in the protocol, along with the rule for what happens if the anticipated situation does not materialise: "The mean difference will be used where at least three studies report outcomes on [instrument]; otherwise the standardised mean difference will be used and back-converted to [instrument] units for the Summary of Findings table."
That sentence takes thirty seconds to write and removes an entire class of reviewer question.
References
Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., & Welch, V. A. (Eds.). (2024). Cochrane handbook for systematic reviews of interventions (Version 6.5). Cochrane. https://training.cochrane.org/handbook
Hedges, L. V. (1981). Distribution theory for Glass's estimator of effect size and related estimators. Journal of Educational Statistics, 6(2), 107–128. https://doi.org/10.3102/10769986006002107
Murad, M. H., Wang, Z., Chu, H., & Lin, L. (2019). When continuous outcomes are measured using different scales: Guide for meta-analysis and interpretation. BMJ, 364, k4817. https://doi.org/10.1136/bmj.k4817
Common questions
- Which SMD should I use, Cohen's d or Hedges' g?
- Hedges' g, essentially always. It applies a small-sample correction to Cohen's d, the correction is negligible in large studies, and it matters in the small ones that populate most meta-analyses. Most software defaults to it; check rather than assume, and state which you used.
- Can I pool change scores and final values in the same analysis?
- For the mean difference, no, unless the baseline means are comparable. For the SMD it is more defensible, since standardisation partially accommodates the difference, but it is a source of heterogeneity worth a sensitivity analysis. Extract both where reported and prespecify which takes precedence.
- The trial reports standard errors, not standard deviations. What now?
- Convert: SD equals SE multiplied by the square root of the sample size. The same applies to confidence intervals, from which an SD can be recovered given the sample size. Record in the extraction form which values were reported and which were derived, since reviewers ask.
