Meta-Analysis & Forest Plots
No single study settles a question. Samples are small, contexts differ, and chance plays a large part. Meta-analysis treats the studies themselves as data points. It pools every estimate of an effect, weighting each by its precision, to get the field's best combined answer. It also measures how much the studies disagree beyond what chance alone would produce.
The one rule: weight by precision
A 400-person trial should count for more than a 30-person pilot. Meta-analysis formalizes this with inverse-variance weighting: each study's weight is 1/SE². Studies with small standard errors, which are usually the large ones because the standard error shrinks as n grows, dominate, and noisy studies contribute what little they know. Before any weighting, the studies have to be on one scale. That usually means converting each result to a common effect size. If a literature reports a mix of d, r and odds ratios, the effect-size converter does the arithmetic. The pooled effect is the weighted average, and its standard error shrinks as studies accumulate, as if the whole literature were one large study.
Reading a forest plot
The standard picture is the forest plot: one row per study, a square at its effect estimate (sized by weight), whiskers for its 95% CI, and a diamond at the bottom for the pooled effect. The diamond's width is its confidence interval. At a glance you can see the effect sizes, which studies were precise, and whether the studies agree.
🎮 Build Your Own Literature
Each row is a simulated study of the same true effect μ (standardized mean difference). Raise between-study heterogeneity τ and the studies spread out, I² climbs, and the random-effects diamond widens while the fixed-effect one stays too narrow.
Fixed vs random effects (and I²)
A fixed-effect model assumes every study estimates the same true effect, differing only by sampling noise. A random-effects model allows the true effect to vary across studies (different populations, doses, protocols) with between-study standard deviation τ. The Q statistic tests whether studies scatter more than chance allows, and I² converts it to a percentage: roughly, how much of the visible spread reflects real differences between studies and not sampling noise. When I² is large, the random-effects model, with its wider diamond, is the defensible summary. The interesting question then becomes why the effect varies. Moderator analysis, the meta-analysis version of interactions, addresses that.
The diamond is not a forecast
The diamond answers one question: where is the average true effect? It says nothing about the true effect in the next study, though readers often take it to mean both. When the true effects vary across studies, those two answers are far apart.
The second question has its own interval. A prediction interval covers the plausible true effect of a new study drawn from the same population of studies, so it includes the between-study spread:
PI = d̂ ± t(k − 2) × √(τ̂² + SE(d̂)²)
Take a literature shaped like the playground's defaults: k = 8 studies, a pooled d of 0.40 with a standard error of 0.10, and τ̂ = 0.20, which is an I² of about 50%. The diamond runs from 0.20 to 0.60, a clearly positive average effect. The 95% prediction interval runs from −0.15 to 0.95. A new trial finding nothing whatsoever would be entirely consistent with this literature, and the forest plot as usually drawn gives no hint of that.
The interval is wider than the distribution of true effects it describes (which spans 0.01 to 0.79 here) because μ and τ are both estimated. With 6 degrees of freedom the multiplier is 2.45 rather than 1.96. A small number of studies also makes τ̂ itself imprecise, so a prediction interval built on five studies can be wide enough to look almost useless. That width is still informative: it shows how little a literature that small can tell you about the next study. In R, metafor's predict() prints the interval beside the pooled estimate.
The elephant: publication bias
Meta-analysis can only pool what got published, and significant results are more likely to be published. If null results stay in file drawers, your pooled estimate comes from a sample of the evidence that is biased toward larger effects. Funnel plots, Egger's test, and trim-and-fill are the standard diagnostics. The deeper fixes are changes in research practice that the replication crisis pushed the field toward: preregistration and registered reports, which put a study on record before its result is known. A meta-analysis is only as unbiased as the literature it draws on. That is the same "how did this sample come to be?" question you've asked since lesson 1.1.
Why it matters: a single study rarely settles a question, but a body of studies often can. Meta-analysis combines noisy, conflicting results into one weighted estimate with its uncertainty, and I² tells you when the disagreement between studies is itself the main finding.
Common questions
What does I² tell you in a meta-analysis?
It is the share of the visible between-study variation that reflects real differences in effects and not sampling noise. Rough bands: 25% low, 50% moderate, 75% high heterogeneity. High I² is a finding in its own right. It means the effect varies across populations or protocols, the fixed-effect summary is too confident, and the interesting question becomes what moderates the effect.
How many studies do I need for a meta-analysis?
Two will produce a number, so the arithmetic is never the constraint. The problem with few studies is the between-study variance τ². With fewer than about five studies it is estimated so poorly that I² and the width of a random-effects interval are close to guesswork, and the pooled result carries that uncertainty without showing it. From around ten studies the estimates become more stable, and moderator analyses need considerably more, since each moderator is in effect a regression on a sample of studies. The Hartung-Knapp adjustment is the usual protection for a small set, because it widens the interval to reflect how poorly τ² is estimated. Comparability matters more than the count, though. Five studies asking the same question can be pooled meaningfully, while twenty that measure subtly different things produce a precise average that describes none of them.
What is a funnel plot and what does asymmetry mean?
It plots each study's effect size against its precision. Large, precise studies cluster at the top and small, noisy ones spread out below, symmetrically if all results were published. A missing lower corner (typically small null studies) suggests publication bias, which you can test with Egger's regression and probe with trim-and-fill. Asymmetry can have innocent causes too, such as small studies using different populations.