Section 1.12

Hypothesis Testing Logic

A hypothesis test is a structured way to answer "could this result just be a fluke?" The logic mirrors a courtroom: the null hypothesis is presumed innocent, and we only convict (reject it) when the evidence would be genuinely surprising if it were true.

The four moves

  1. State two hypotheses. The null (H₀) is the boring "nothing's going on" claim (no effect, no difference). The alternative (H₁) is what you suspect instead. (Turning a vague research question into these two crisp statements is a skill of its own. From Question to Hypothesis walks the whole path.)
  2. Assume H₀ is true. This gives you a null distribution, the range of results you'd expect from chance alone if there really were no effect.
  3. Measure surprise. Find where your actual result falls on that distribution. The p-value is the probability of getting a result this extreme or more, just by chance, if H₀ were true.
  4. Decide. If the p-value is below a pre-set threshold α (usually 0.05), the result is "too surprising to be chance," so you reject H₀. Otherwise, fail to reject it.

Where the statistic comes from

The explorer below starts with a z already in hand, which is the right place for a lesson about the logic. An exam hands you a claim and a sample instead, and the first job is turning those into that single number. When the population standard deviation σ (sigma) is known, the arithmetic is short. Take the gap between what you saw and what the claim predicts, then measure that gap in standard errors:

z = (x̄ − μ₀) / (σ/√n)

μ₀ (mu-nought) is the value the null hypothesis names, so it is a number you are given rather than one you work out. The denominator is the ordinary sample-to-sample wobble of a mean. Dividing by it gives z a scale that travels: however the measurement was made, z answers "how many standard errors of ordinary wobble away from the claim did this sample land?"

🔢 From numbers to z

The four numbers an exam hands you, turned into the statistic the explorer draws. Edit any of them and the explorer below jumps to the z they produce; the button sends the numbers as they stand.

Standard error of the mean—
Gap from the claim—
Test statistic z—

🎮 P-Value Explorer

The bell is the null distribution: what chance alone would produce. Slide your observed result and watch how much of the curve is "this extreme or worse" (the shaded p-value).

p-value0.134
Threshold α0.05
DecisionFail to reject H₀

What a p-value is — and isn't

Slide the statistic out toward the tails and the p-value shrinks: extreme results are unlikely under chance, so they count as evidence against H₀. But the p-value is slippery, and two misreadings are everywhere:

A p-value is NOT the probability that H₀ is true. It's the probability of data this extreme assuming H₀ is true. Those are different questions. p = 0.03 does not mean "3% chance there's no effect."

"Fail to reject" is NOT "proven true." A courtroom's "not guilty" doesn't mean "definitely innocent," only "not enough evidence." Likewise, a non-significant result means we couldn't rule out chance, not that H₀ is correct. (Writing that up honestly, without "trend toward significance," is covered in Writing About Non-Significant Results.)

The five steps, in exam order

The four moves above are the logic. On paper, an examiner wants them written in a fixed order with the arithmetic in the middle, and marks are given for the steps rather than the answer:

  1. Write H₀ and H₁ as statements about the population parameter, in symbols, before looking at the data. H₀ names one value; H₁ says what "not that value" means for this question.
  2. Fix α. It is a decision about how much risk of a false alarm you will accept, and it belongs here, not after the p-value has appeared.
  3. Compute the test statistic, showing the substitution as well as the result.
  4. Find the p-value from a table or a calculator, matching the tail or tails to H₁.
  5. Compare p with α and write the conclusion in the words of the original question, not as "reject H₀" alone.

The five steps on a cereal line. A filling machine is set to 227 g a box, and the factory's records give σ = 5 g. Four boxes off this morning's line average 222 g. (1) H₀: μ = 227 against H₁: μ ≠ 227, since "has the machine drifted" does not say which way. (2) α = .05. (3) z = (222 − 227) / (5/√4) = −5 / 2.5 = −2.00. (4) Two-sided, so p = 2 × P(Z < −2.00) = .0455. (5) Because .0455 is below .05, reject H₀: this morning's boxes are lighter than the machine is meant to make them. At α = .01 the same evidence falls short, and an honest report says which threshold it was judged against. The 95% interval built from these very numbers came out as [217.1, 226.9], which excludes 227 for exactly the reason the test rejects: at a matching α, a two-sided test and a confidence interval always agree.

One tail or two, and where the choice comes from

The explorer's alternative setting is not a matter of taste. The question sets it, and it is set before any data arrive.

Wording that names a direction points at one tail: "is the new fertilizer better", "does the drug lower blood pressure", H₁: μ > μ₀ or H₁: μ < μ₀. Wording that only asks whether something has moved points at both: "is the machine off target", "is the printed label incorrect", "do the two groups differ", H₁: μ ≠ μ₀. Two-sided is the cautious default, and it is what software prints unless you ask otherwise.

Which one you picked also changes how the p-value is read off a printed table. A left-tailed test reads the area below your z straight from Table A. A right-tailed test reads that area and subtracts it from 1, or equivalently looks up −z and reads the area below that. A two-sided test finds one tail and doubles it.

Picking the tail after seeing which way the result went is cheating. It halves the p-value for free and lifts the true false-alarm rate from 5% to 10%, which is why it appears on every list of questionable research practices. If the direction was not predicted in advance, the test is two-sided. A one-tailed test is a promise made in advance that a result in the other direction, however large, would have counted as nothing.

Significant, and possibly trivial

A small p-value says the data would be surprising if nothing were going on. It says nothing whatever about how much is going on, and sample size drives a wedge between the two questions: with enough data, a difference far too small to care about will clear any threshold you like.

Picture a weight-loss app tested on forty thousand users, reporting a mean loss 0.2 kg greater than the control group's at p < .001. The p-value is doing its job and the difference is almost certainly real. It is also a fifth of a kilogram. Statistical significance is a claim about evidence; practical significance is a claim about size, and only the second one answers "should anyone change what they do?"

The habit that keeps them apart is reporting an effect size and a confidence interval beside every p-value, so a reader sees a magnitude with its uncertainty instead of a verdict. When a paper reports a metric you do not think in, the effect-size converter will restate it as one you do.

Absence of evidence is not evidence of absence

"Study finds no link between X and Y" is a headline written off a p-value above .05, and it almost always claims more than the study earned. Failing to reject H₀ means the data were compatible with no effect. They were also compatible with a small effect, and frequently with a large one, because a test that fails to reject has not measured anything precisely.

The confidence interval is what separates those cases. Around a difference, an interval of [−0.1, +0.2] genuinely does rule out anything big. An interval of [−4.0, +4.5] rules out nothing at all and reports only that the study was too small to tell. Both results are non-significant; only the first one supports the headline.

Writing this up without overclaiming in either direction has a lesson of its own, Writing About Non-Significant Results, which also covers equivalence testing: the procedure to reach for when "no meaningful difference" is the claim you actually want to make, rather than the one you fall back on.

Two ways to be wrong, and how often

A test can fail in two directions. A Type I error rejects a true H₀, announcing an effect that was never there, and its rate is exactly α. That is why α is the one error rate you choose outright: set it at .05 and you will cry wolf on 5% of the occasions when there is nothing to see. A Type II error is the opposite failure, missing an effect that is really present, and its rate is written β (beta). Nobody picks β. It falls out of three things at once: how big the real effect is, how variable the measurement is, and how much data you gathered.

What gets quoted is β's complement. The power of a test, 1 − β, is the probability of detecting an effect of a stated size when that effect genuinely exists. Both quantities become concrete the moment you draw two sampling distributions on one axis: the one H₀ predicts, and the one the truth would actually produce.

🎮 Two Curves

The left curve is what tray means do if the null is right and the true mean is 22.0 cm. The right curve is what they do if the truth is μ₁ instead. Red slivers are α, split across both tails of the null. The gray block is β, the part of the true distribution that still lands inside the keep-H₀ zone. Everything green is power.

Power (1 − beta)—
Beta, the miss rate—
Reject outside—
Effect, in standard errors—

Walk the seedlings through by hand and the number stops being abstract. A nursery's seedlings have averaged 22.0 cm at four weeks for years, with σ = 0.5 cm. A new light schedule might lift the mean to 22.6. Testing five trays at α = .05, two-sided, the cutoffs sit at 22.0 ± 1.960 × 0.5/√5, which is 22.0 ± 0.438, so the test rejects for a tray mean below 21.562 or above 22.438. If 22.6 really is the new mean, the chance of landing beyond that upper cutoff is .765. A power of .765 means a study built this way walks past a real effect roughly one time in four.

Every lever in the picture behaves the way intuition says it should once the two curves are visible. More data narrows both of them and pulls the overlap apart. A larger true effect slides them further from each other. A smaller α pushes the cutoffs outward and costs power, which is the trade the two error rates make against one another; there is no setting that shrinks both at once. Choosing n so that power reaches .80 before any data exist is the whole business of a power analysis, run here by the power calculator and taught properly in Effect Size & Power.

Why it matters: every test you'll meet (t-tests, ANOVA, chi-square, regression) runs this exact playbook. Master the logic once here, and the rest is just swapping in different null distributions.

Two practice problems run on this lesson: Problem 9 hands you a newsletter sentence that misreads p = .030 as "a 3% chance the feed makes no difference", and Problem 18 runs a σ-known test end to end, with the tail read off the wording and the verdict flipping between α = .05 and α = .01.

Common questions

What does p < 0.05 actually mean?

It means: if there were truly no effect (H₀ true), data this extreme would occur less than 5% of the time by chance alone. Since that's rare, we treat the result as evidence against H₀. It does not mean there's a 95% chance the effect is real, and it says nothing about how large or important the effect is. That's the job of effect sizes.

Why learn the z-test when σ is almost never known?

It usually isn't, and §1.13 swaps σ for the sample's own s about a week later. The z version earns its place anyway. It isolates one idea at a time: with σ known, the mean is the only thing being estimated, the sampling distribution is exactly normal, and every step of the logic stays visible without the extra layer that estimating a spread adds. It is also less artificial than it sounds, since a long-running measurement process (a calibrated instrument, an assay a lab has run for years, a standardized test whose SD is published) really does have a spread that is known while its mean is the thing in question. The two procedures then converge as n grows, because a t distribution with many degrees of freedom is the normal distribution. That is why the bottom row of a printed Table D reads 1.645, 1.960 and 2.576, the same three z* values the σ-known interval uses.

Why is the significance level set at 0.05?

Convention, not law of nature. R. A. Fisher suggested in the 1920s that one-in-twenty was a convenient benchmark for "surprising" (roughly the chance of landing beyond ±2 SDs), and it stuck. Nothing magical happens between p = .049 and p = .051, which is why fields with different stakes choose differently: particle physics demands "5 sigma" (about 1 in 3.5 million), genome-wide studies use 5 × 10⁻⁸, and some journals now suggest .005 for new discoveries. What matters is fixing α before you look at the data, and remembering that crossing it says nothing about how large or important the effect is.