Section 1.15

Inference for Proportions

A proportion is a kind of mean. Score every success as 1 and every failure as 0, average them, and you have p̂. Every tool from the last few lessons applies to it, with one adjustment: the spread of a proportion depends only on the proportion itself, so you never have to estimate it separately.

The sampling distribution of p̂

Take a random sample of size n from a population in which a fraction p would say yes. The count of yes answers is binomial, and the sample proportion is that count divided by n. Dividing by a constant rescales the mean and the standard deviation, which gives:

mean of p̂ = p SD of p̂ = √(p(1 − p) / n)

So p̂ is unbiased, and its spread shrinks with √n, just as the sample mean's does. The shape needs more care. A binomial count becomes close to normal as n grows. But a proportion near 0 or 1 is pressed against a boundary, so it needs a much larger n than one near the middle before its shape is close to normal.

Conditions for inference

Three conditions must hold before you can trust the methods below, and the first is not a formula. The data must come from a random sample of the population you want to describe. Nothing on this page can correct a sample that was not random.

The sample must also be small next to the population, so that drawing people without replacing them barely changes the spread. The usual rule is a population at least ten times the sample, and Moore, McCabe and Craig set it at twenty. Survey 400 people in a town of 900 and you have broken it. The real standard error is then smaller than the formula reports, so the interval you print is wider than it needs to be.

Third, the counts of successes and of failures both have to be large enough for the normal approximation to work. Here the interval and the test use different rules, and so do the textbooks.

For a confidence interval, count what you actually observed. Introduction to the Practice of Statistics asks for at least 15 successes and 15 failures. OpenIntro Statistics and the AP-style texts ask for 10 of each. Both numbers are in wide use, so answer with the one your own course teaches. The disagreement is itself a warning: an interval built on twelve successes rests on a borderline approximation.

For a test, count what H₀ predicts, because a test assumes H₀ is true: np₀ ≥ 10 and n(1 − p₀) ≥ 10. The two texts agree on that one.

The condition light in the interactive below uses the 10 rule for both counts. Read it as a minimum: amber means you have failed even the loosest rule in use, and green means you have passed that rule, which may be looser than your course's.

The condition is about both counts, never about n alone. A sample of 500 with 4 successes fails it. So a rare outcome needs a large study before you can use a z method.

One more point about the interval. Even when both counts are well above these minimums, p̂ ± z*√(p̂(1 − p̂)/n) covers the true proportion slightly less often than its stated level. As n grows, the shortfall oscillates and does not steadily disappear. The simplest fix is the plus-four interval: add two successes and two failures to the data, then run the same formula on the new counts. The questions at the foot of this page say more about the Wilson interval, the version most software now reports.

🎮 One Proportion, or Two

Set the counts and read off the interval and the test. The condition light turns amber whenever a group holds fewer than 10 successes or fewer than 10 failures.

p̂—
95% CI for p—
z—
two-sided p—
Conditions—

A confidence interval for one proportion

The interval has the shape every interval in this course has, an estimate plus or minus a margin:

p̂ ± z* √(p̂(1 − p̂) / n)

Suppose 232 of 400 sampled students say they would use a late-night bus. Then p̂ = 0.58, the standard error is √(0.58 × 0.42 / 400) = 0.0247, and with z* = 1.96 the margin of error is 0.0484. The interval runs from 0.532 to 0.628, so somewhere between 53% and 63% of the student body would use the bus.

A test for one proportion

Testing works the same way, with one change that students lose marks on every year. A hypothesis test assumes H₀ is true, and for a proportion H₀ names an exact value p₀, which fixes the standard deviation as well as the center. So the test uses p₀ in the standard error where the interval used p̂:

z = (p̂ − p₀) / √(p₀(1 − p₀) / n)

For the bus survey against H₀: p = 0.5, the denominator is √(0.25/400) = 0.025, giving z = (0.58 − 0.50)/0.025 = 3.20 and a two-sided p-value of 0.0014. The majority is unlikely to be a sampling accident. The two standard errors differ here only in the third decimal. When p̂ is far from p₀ they can differ enough to matter, and using the wrong one is marked wrong even when the conclusion stays the same.

Comparing two proportions

Two independent samples, two proportions, and the quantity of interest is their difference p̂₁ − p̂₂. Variances of independent quantities add, so the standard error of the difference is the square root of the sum:

SE = √(p̂₁(1 − p̂₁)/n₁ + p̂₂(1 − p̂₂)/n₂)

Take 45 successes out of 120 in one group and 72 out of 150 in the other, so p̂₁ = 0.375 and p̂₂ = 0.480. The difference is −0.105 with a standard error of 0.0601, and the 95% interval runs from −0.223 to 0.013. It contains zero, so the data are consistent with no difference at all.

The test again uses a different standard error. H₀ says the two population proportions are equal, so under H₀ there is only one proportion, and the best estimate of it uses everyone: the pooled proportion p̂ = (x₁ + x₂)/(n₁ + n₂) = 117/270 = 0.433. That single value goes into both halves of the standard error:

z = (p̂₁ − p̂₂) / √(p̂(1 − p̂)(1/n₁ + 1/n₂))

which gives 0.0607 for the denominator and z = −1.73, two-sided p = 0.084. The interval and the test agree here: the interval covers zero and the test fails to reject.

Why χ² on a 2×2 table is this same test

The same 270 people can be written as a two-by-two table of successes and failures by group. A chi-square test of independence on that table asks the same question as the pooled z test. In fact the two are algebraically the same test, and

χ² = z² with 1 degree of freedom

Here z = −1.7301 and χ² = 2.9932, which is 1.7301² to every decimal either statistic prints, and both report p = 0.084. The identity holds for any 2×2 table, and it is the same one that links the two distributions in the statistical tables. The one practical difference is that the z version gives you a signed difference and a confidence interval, and χ² gives you neither. For a 2×2 table, prefer the two-proportion z.

Why it matters: surveys, clinical trials and A/B tests mostly produce proportions. The interval says how precisely you know the percentage, and the test says whether a difference is larger than chance alone would explain. Both need the conditions above, and neither can correct a sample that was not random.

Problem 22 in the practice problems runs the pooled test and the unpooled interval side by side on 450 clinic reminders. Its fourth part asks why the two standard errors are not the same number.

Common questions

Why does the confidence interval use p-hat but the test uses p-nought?

Because a proportion's spread depends on the proportion itself. The standard deviation of p̂ is the square root of p(1−p)/n, so naming a value for p also names the standard error. A hypothesis test assumes H₀ is true, and H₀ supplies p₀, so the test has a value to use and should use it. A confidence interval assumes no value for p, so it can only use the data's own p̂. When p̂ is close to p₀ the two standard errors barely differ, and when it is far away they can differ enough to change the conclusion.

My sample has only 4 successes out of 300. Can I still build an interval?

Not with the normal-based formula, because the success-failure condition needs at least 10 of each, however large the sample. With 4 successes the sampling distribution is pressed against zero and cannot be symmetric. There are two good alternatives. Report an exact binomial (Clopper-Pearson) interval, which most software offers and which stays valid at any count, or use the Wilson score interval, which behaves far better than the textbook formula near the boundaries. Both will be asymmetric around p̂, and that asymmetry is correct, not a defect.

If chi-square and the two-proportion z test are the same test, which should I report?

For a 2×2 table, prefer the z test. The two are algebraically identical, χ² equals z squared on one degree of freedom, and they return the same p-value to every decimal. They differ in what else they report. The z test gives a signed difference and a confidence interval for it, so the reader of your paper can judge whether the gap matters. Chi-square gives a positive number with no direction and no interval. Once the table is larger than 2×2, the z test no longer applies and you use chi-square.