Effect Size & Power
A p-value answers one narrow question: is there an effect at all? It says nothing about how big the effect is, or whether your study was even capable of finding it. Those two questions, answered by effect size and power, separate a meaningful result from a misleading one.
Effect size: how big, not just whether
With a huge sample, a trivially small difference can be "statistically significant." With a tiny sample, a huge difference can miss significance. Significance is tangled up with sample size, so we need a measure of the effect that isn't. That's effect size. For a difference between two means, the standard one is Cohen's d: the gap between the means measured in standard deviations.
d = (mean₁ − mean₂) / standard deviation
Rough conventions: d ≈ 0.2 is small, 0.5 medium, 0.8 large. It's the same currency as a z-score (a standardized distance), so it's comparable across studies and scales.
Which standard deviation goes on the bottom? For two groups it is the pooled one, both groups' spreads combined into a single yardstick, which takes for granted that the two spreads are similar to begin with. There is also a small-sample wrinkle worth knowing about: d calculated this way comes out slightly too large when samples are small, by about 4% with ten people per group and 2% with twenty. Multiplying by a correction factor removes that bias and gives you Hedges' g, which is the version meta-analyses pool. Past roughly fifty per group the two agree to two decimal places, so the distinction earns its keep only in small studies.
🎮 The Power Playground
Top panel: the two populations: this is what an effect of size d looks like in the world. Bottom panel: what your test sees, the distribution of results you could observe if there were no effect (H₀) vs. the effect being real (H₁). Power is the green area: the share of possible study outcomes that lands past the significance cutoff. Turn all four dials and watch the areas fight.
Power: could you even detect it?
Power is the probability that your study correctly rejects the null when a real effect exists — your chance of not missing it. The convention is to aim for at least 80%. Power analysis is really a four-way negotiation; the playground gives you all four dials:
- Effect size (d). Slide it up and the two populations pull apart and the H₁ curve below slides away from the cutoff: the green area swells. Tiny effects are genuinely hard to detect.
- Sample size (n). The populations don't move an inch, but the sampling curves below sharpen dramatically. Same effect, narrower curves, less overlap across the cutoff → more power. This is the dial you usually control.
- Significance level (α). A stricter α (say .001) drags the cutoff away from zero, shrinking the false-alarm area but eating your power with it. There's no free lunch: α and β trade off.
- One- vs. two-tailed. Committing to a direction moves the cutoff closer to zero and buys power, at the price of being blind to an effect in the other direction.
Try this: set d = 0.30 (a realistic effect in psychology), α = .05, two-tailed. At n = 30 the power readout is a coin flip at best. Now drag n until the green area hits 80% — that's a power analysis, and it's exactly the calculation you should run before collecting a single data point. Then set d = 0 and notice power equals α: with no real effect, "detections" are pure false alarms.
Read the bottom panel like a story. Your one study will land somewhere under the orange H₁ curve. If it lands past the dashed cutoff, you celebrate a significant result; if it lands short, you miss a real effect (β). Power is simply how much of your future self's luck is green.
How precisely do you know d?
The playground treats d as a fact about the world. Your study only ever hands you an estimate of it, and that estimate carries a margin which APA 7 asks you to report. A bare "d = 0.50" reads like a measurement when it is closer to a reading with a range attached.
The range is not d plus or minus something. Because d is a ratio of two quantities you estimated, its sampling distribution is a noncentral t rather than a normal curve, so the interval has to be found by asking which true effects would leave your observed t where it landed. The APA Results Formatter runs that search and writes the answer into the sentence for you; the APA cheat sheet has the finished pattern if you just need the shape of it.
Try it on this playground's own dials. At the defaults, d = 0.50 with 30 per group, the test gives t(58) = 1.94 and p = .058, so the study misses, and the interval on d runs [−0.02, 1.01], which still holds open the possibility that the effect points the other way. Now take the 64 per group the readout above asks for at 80% power, observe the same d = 0.50, and the test gives t(126) = 2.83, p = .005. Its interval is [0.15, 0.85]. The study worked, and it still cannot separate a small effect from a large one.
A result sitting exactly on the significance boundary has a 95% interval whose lower limit is exactly zero, so "significant at .05" and "the interval excludes zero" are two ways of saying one thing. Significance tells you that zero is out, and nothing at all about size. Pinning d down to ±0.20 around 0.50 takes 199 per group, three times what 80% power costs. Detecting an effect and measuring it are different jobs, and the second is the expensive one.
Why underpowered studies are dangerous
If power is only 40%, you'll miss a real effect more often than you find it — and a "non-significant" result tells you almost nothing. Worse, the few significant results that do squeak through tend to overestimate the effect. This is why researchers run a power analysis before collecting data: pick the effect size you care about, the power you want (say 80%), and solve for the sample size you need. Want the exact number for your own study? The power & sample-size calculator does it for t-tests, ANOVA, correlation, and more, and the complete worked project shows the calculation being made for a real design before a single participant is recruited.
Before is doing the work in that sentence. Running the same calculation afterwards on the effect size you happened to observe produces a number that contains no information you didn't already have from the p-value, and power analysis for complex designs sets out why that is arithmetic rather than bad luck, along with the after-the-fact questions that are worth asking.
Why it matters: reporting an effect size alongside the p-value, and planning for adequate power, is the difference between research that replicates and research that doesn't. Significance is the start of the story, never the whole of it.
Problem 9 of the practice problems works through a newsletter's misreading of p = .030, then asks the question this lesson is for: a repeat trial on eight plots comes back p = .21, so has anything been shown?
Common questions
What is a good effect size?
Cohen's benchmarks (d ≈ 0.2 small, 0.5 medium, 0.8 large) are rough field-wide defaults, not laws. What counts as meaningful depends on context: d = 0.2 on mortality is enormous; d = 0.5 on a novel lab task may be routine. Compare against typical effects in your literature, and translate d into overlap or percentile terms with the effect-size converter to build intuition.
What does 80% power mean?
If the true effect is exactly the size you assumed, a study with 80% power has an 80% chance of returning a significant result — and a 20% chance of missing it (β = 0.20). It's a property of the design, chosen before data collection: the conventional compromise between missing real effects and the cost of ever-larger samples.
I found a significant effect with a small sample. Doesn't that make it more impressive?
It usually makes the estimate less trustworthy, and the reason is worth seeing in numbers. Take a real effect of d = 0.30 and run it with twenty people per group: that study has about a 15% chance of reaching significance at all, and across the runs that do reach it, the average observed d is 0.82, nearly three times the truth. A small sample clears the significance bar only when the noise happens to push in the helpful direction, so the estimates that survive are the exaggerated ones. That is why one significant small study is a reason to look again rather than a reason to believe a large effect, and why meta-analyses treat published effects as inflated until shown otherwise.