Section 3.1

Effect Size & Power

A p-value answers one narrow question: is there an effect at all? It says nothing about how big the effect is, or whether your study was even capable of finding it. Effect size answers the first of those questions and power the second.

Effect size: how big, not just whether

With a huge sample, a trivially small difference can be "statistically significant." With a tiny sample, a huge difference can miss significance. Because significance depends on sample size, you also need a measure of the effect that does not: an effect size. For a difference between two means, the standard one is Cohen's d: the gap between the means measured in standard deviations.

d = (mean₁ − mean₂) / standard deviation

Rough conventions: d ≈ 0.2 is small, 0.5 medium, 0.8 large. Like a z-score, d is a standardized distance, so it can be compared across studies and scales.

Which standard deviation goes on the bottom? For two groups it is the pooled one, which combines both groups' spreads into a single SD. Pooling assumes the two spreads are similar to begin with. In small samples, d calculated this way also comes out slightly too large: by about 4% with ten people per group and 2% with twenty. Multiplying by a correction factor removes that bias and gives you Hedges' g, the version meta-analyses pool. Past roughly fifty per group the two agree to two decimal places, so the correction matters only in small studies.

🎮 The Power Playground

Top panel: the two populations, which show what an effect of size d looks like in the world. Bottom panel: the distribution of test results you could observe if there were no effect (H₀) and if the effect were real (H₁). Power is the green area: the share of possible study outcomes beyond the significance cutoff. Change the four controls and watch how the areas trade off.

Power (chance to detect)—
β: Type II (miss) risk—
Smallest significant d̂—
n/group for 80% power—
Population overlap—

Power: could you even detect it?

Power is the probability that your study correctly rejects the null when a real effect exists — your chance of not missing it. The convention is to aim for at least 80%. Four things set a study's power, and the playground has a control for each:

  • Effect size (d). Slide it up and the two populations pull apart and the H₁ curve below slides away from the cutoff: the green area swells. Tiny effects are hard to detect.
  • Sample size (n). The populations do not move, but the sampling curves below become much narrower. With the same effect, narrower curves overlap less across the cutoff, so power rises. Sample size is usually the one you control.
  • Significance level (α). A stricter α (say .001) moves the cutoff away from zero. That shrinks the false-alarm area but also lowers power, so α and β trade off.
  • One- vs. two-tailed. Committing to a direction moves the cutoff closer to zero and raises power, but the test can then never detect an effect in the other direction.

Try this: set d = 0.30 (a realistic effect in psychology), α = .05, two-tailed. At n = 30 the power readout is well below 50%. Now drag n until the green area reaches 80%. That is a power analysis, the calculation you should run before collecting any data. Then set d = 0: power equals α, because with no real effect every "detection" is a false alarm.

Reading the bottom panel. If the effect is real, your one study's result is a random draw from under the orange H₁ curve. A result beyond the dashed cutoff is significant. A result short of it misses a real effect (β). Power is the share of the H₁ curve that is green.

How precisely do you know d?

The playground treats d as a fact about the world. A study only gives you an estimate of it, and APA 7 asks you to report a confidence interval around that estimate. A bare "d = 0.50" hides how wide that interval can be.

The interval is not simply d plus or minus a fixed amount. Because d is a ratio of two quantities you estimated, its sampling distribution is a noncentral t rather than a normal curve. To find the interval, you search for the true effects under which your observed t would not fall in either 2.5% tail. The APA Results Formatter runs that search and writes the interval into your results sentence. The APA cheat sheet shows the finished sentence pattern.

Try it with the playground's settings. At the defaults, d = 0.50 with 30 per group, the test gives t(58) = 1.94 and p = .058, so the study misses. The interval on d runs [−0.02, 1.01], so an effect in the other direction is still possible. Now take the 64 per group that the readout above gives for 80% power and observe the same d = 0.50. The test gives t(126) = 2.83, p = .005, and the interval is [0.15, 0.85]. The study detected the effect, but it still cannot tell a small effect from a large one.

A result exactly on the significance boundary has a 95% interval whose lower limit is exactly zero. So "significant at .05" and "the interval excludes zero" say the same thing. Significance tells you that zero is outside the interval, and nothing about size. Estimating d to within ±0.20 of 0.50 takes 199 per group, about three times the 64 that 80% power needs.

Why underpowered studies are dangerous

If power is only 40%, you'll miss a real effect more often than you find it, and a "non-significant" result tells you almost nothing. Worse, the few significant results that do squeak through tend to overestimate the effect. This is why researchers run a power analysis before collecting data: pick the effect size you care about, the power you want (say 80%), and solve for the sample size you need. Want the exact number for your own study? The power & sample-size calculator does it for t-tests, ANOVA, correlation, and more. The complete worked project shows the calculation for a full study design, made before any participant is recruited.

The calculation has to come before the data. Run it afterwards on the effect size you happened to observe, and the power it reports tells you nothing the p-value did not. Power analysis for complex designs explains why, and which questions you can usefully ask after the study.

Why it matters: significance alone says that the data rule out zero. It does not say how large the effect is, and in an underpowered study a non-significant result means little. So report an effect size alongside every p-value, and plan for adequate power before you collect data.

Problem 9 of the practice problems works through a newsletter's misreading of p = .030. It then asks the question this lesson answers: a repeat trial on eight plots gives p = .21, so has anything been shown?

Common questions

What is a good effect size?

Cohen's benchmarks (d ≈ 0.2 small, 0.5 medium, 0.8 large) are rough field-wide defaults, not laws. What counts as meaningful depends on context: d = 0.2 on mortality is enormous; d = 0.5 on a novel lab task may be routine. Compare against typical effects in your literature, and translate d into overlap or percentile terms with the effect-size converter to build intuition.

What does 80% power mean?

If the true effect is exactly the size you assumed, a study with 80% power has an 80% chance of returning a significant result — and a 20% chance of missing it (β = 0.20). It's a property of the design, chosen before data collection: the conventional compromise between missing real effects and the cost of ever-larger samples.

I found a significant effect with a small sample. Doesn't that make it more impressive?

It usually makes the estimate less trustworthy. Take a real effect of d = 0.30 and run it with twenty people per group. That study has about a 15% chance of reaching significance at all, and among the runs that do reach it, the average observed d is 0.82, nearly three times the truth. A small sample reaches significance only when sampling error has pushed the estimate well above the true value, so the significant estimates are the exaggerated ones. One significant small study is therefore a reason to look again, not a reason to believe in a large effect. It is also why meta-analyses treat published effects as inflated until shown otherwise.