Section 1.16

Chi-Square Tests

Everything so far has compared means of numeric data. But what about counts: how many people picked option A vs. B, recovered vs. didn't, voted left vs. right? The chi-square (χ²) test works with frequencies in categories, and its most common use is to ask: are two categorical variables associated, or independent?

Observed vs. expected

The whole test rests on one comparison. You have the observed counts (what actually happened). Then you compute the expected counts: what you'd see if the two variables were completely unrelated (each cell's expected count is just its row total × column total ÷ grand total). The bigger the gap between observed and expected, the more evidence of a real association:

χ² = Σ (observed − expected)² / expected

The Σ (sigma) is an instruction to add: work the fraction out for one cell, then do it for every cell and total them.

That formula, the expected-count rule, the df and Cramér's V are four consecutive rows on the printable formula sheet. Keep it beside you the first few times.

Work one cell, then the next, then add

To get the sum right on paper, leave the adding until last. Take the cells in any order. For each one, write down its expected count, then subtract, square and divide. That gives you a short list of numbers called the components, and χ² is their total.

Start with the trial in the builder below. Its first cell holds 45 recoveries where independence expects 65 × 70 ÷ 130 = 35.0, so that cell contributes (45 − 35.0)² ÷ 35.0 = 2.86. The other three come to 3.33, 2.86 and 3.33, and χ² = 12.38.

Writing out the components has two benefits. An arithmetic slip stays visible in one cell and does not disappear into the total. And the components are a result in themselves: the largest ones mark the cells furthest from independence, and those are the cells your results paragraph should name.

🎮 Build a Table

A treatment trial to start with. Type over any count, rename any row or column, and grow the table to four rows by four columns. Every cell prints the count you observed, the count independence expects, and that cell's component of χ². The bars underneath show each row's split beside the split of everyone together. Independence predicts that every row has that overall split.

Each cell reads: the observed count you can edit, then E the expected count, then the component (O − E)² ÷ E. The orange tint grows with a cell's share of χ².

χ² statistic—
Degrees of freedom—
p-value—
Table F reads—
Conditions—

Reading the result

Degrees of freedom for a table are (rows − 1) × (columns − 1), which is 1 for the 2×2 trial and 6 for a three-by-four. A χ² far from zero, with p < 0.05, means counts like these would be surprising if the two variables were independent, so you conclude they are associated. In the trial, treatment and recovery are related.

The bars under the table show the same thing without arithmetic. Each row's split across the columns is that row's conditional distribution, and the faint bar at the foot is the split of everyone together. Independence means that every row's conditional distribution matches that bottom bar, so χ² measures how far the rows are from it. Rows with identical splits give a χ² of zero, however many cases the table holds.

On paper the p-value comes as a bracket, not a single number. Take χ² and its df to Table F, find the df row, and move along it until your statistic falls between two entries. The p-value is between those two column headings. The builder prints that reading beside the exact one, so you can check a hand answer.

For a bigger example, take three neighborhoods by four ways of getting to work, 300 people in all. Load it into the builder. A loaded table arrives with generic names, so type in the real ones: rows center, inner ring, outer ring, and columns walk, bike, transit, car. χ² comes to 18.77 on 6 degrees of freedom. Table F's df 6 row has 18.55 under .005 and 20.25 under .0025, so the paper answer is "between .005 and .0025" and the exact p is .0046. The tint shows where it comes from: walking in the center and driving in the outer ring together account for more than half of χ², while the middle row contributes almost nothing.

Association, not amount. Like correlation, a significant χ² tells you a relationship exists, not how strong it is. For strength, report an effect size such as Cramér's V or the odds ratio alongside the test. To compare an odds ratio with the rest of your results, the effect-size converter will turn it into the more familiar d or r.

Where the curve comes from

The χ² distribution was not invented for tables. Square a standard normal variable and you have a χ² distribution on one degree of freedom; add k independent squares and you have χ² on k. A contingency table's statistic is a sum of squared, standardized differences, so the same family describes it. That origin explains the rest of its behavior. χ² can never be negative, a table that matches its expected counts exactly gives zero, and every p-value here is a right-tail area, because only large values are evidence against independence. The distribution playground includes the family among its nine curves. Move its df slider and the shape shifts right and becomes more symmetric as df grows.

Two flavors of chi-square

  • Test of independence is the one above: are two categorical variables related? (Treatment × outcome, gender × preference.)
  • Goodness-of-fit: does one categorical variable match an expected distribution? (Is this die fair? Do customers pick the four flavors equally?)

What you're assuming

Counts have to be independent: each observation falls in exactly one cell, and one person's category does not affect another's. Sixty people should produce sixty entries in the table.

The expected counts also have to be big enough for the χ² curve to describe the statistic well. Textbooks state that in three parts, and the builder checks all three live as you type:

  • every expected count is at least 1;
  • no more than 20% of the expected counts fall below 5;
  • on a 2×2 table, where 20% of four cells rounds down to none of them, all four expected counts reach 5.

The rule is about expected counts, not observed ones, and that catches many students out. An observed zero is not a problem by itself. A cell where independence expects only two people is. When a table fails, you can merge categories that belong together or collect more data. On a 2×2 you can also switch to Fisher's exact test, which computes the probability of the table directly and needs no approximation. If the study is still ahead of you, the power calculator will tell you how many observations a given association needs.

The 2×2 table you have already met

Cross a two-group variable with a yes/no outcome and you have exactly the data a two-proportion test uses. The two procedures are algebraically the same test, with χ² = z² on one degree of freedom and the same p-value to every decimal. Inference for proportions works through that identity on 270 people. It also explains why the z version is usually the one to report: it gives a signed difference and a confidence interval for it, while χ² gives only a positive number. For tables larger than 2×2 the z test no longer applies, and you use χ².

When the same person appears twice

Two different meanings of "independent" meet in this test, and only one of them is the null hypothesis. The test asks whether the two variables are independent. The assumption above is about the observations: sixty people should produce sixty entries, not a hundred and twenty.

Sixty students say whether they feel confident reading a results section, once before a workshop and once after. Stack the two occasions and you have a 2 × 2 table of 120 answers, which looks like an ordinary chi-square problem. It is not one. Each student appears in that table twice, so the table has only 60 independent units, and a test that assumes 120 independent answers gives the wrong p-value.

The fix is to count each person once. Cross each student's before-answer with their own after-answer, and every student is in exactly one of four cells: yes both times, no both times, or one of the two kinds of switcher. Students who answered the same way twice tell you nothing about whether the workshop changed anything. So McNemar's test leaves them out and uses only the switchers, which textbooks call the discordant pairs:

χ² = (b − c)² ⁄ (b + c), with 1 degree of freedom

where b and c are the two off-diagonal counts. Suppose 18 students went from no to yes and 6 went the other way. Then χ² = (18 − 6)² / 24 = 6.00 and p = .014. SPSS and R apply a continuity correction by default, giving 5.04 and p = .025. With only 24 switchers you can skip the approximation and use the exact binomial test, which gives p = .023. All three lead to the same conclusion, and when b + c is small, quote the exact one.

What does stacking the occasions do? In a simulation of 60 students whose two answers correlate at .6, with no real change, the ordinary chi-square on the 120 rows rejects about 1% of the time, not 5%. So it is too cautious, not too eager, which surprises most people. Give the same students a real shift, from 60% confident to 78%, and McNemar detects it 78% of the time, against 60% for the stacked test. If the two answers are unrelated, so that the pairing carries no information, the two tests have the same power, 58% each. Pairing helps when the two answers are linked, and ignoring it here wastes participants but does not create false findings.

With three or more occasions on the same people, the same logic goes by a different name: Cochran's Q, which generalizes McNemar to k repeated yes/no measurements. The printable assumptions poster names McNemar as the fix for this case. Mixed & multilevel models explains why it works, and what to do when the repetition is people inside schools rather than occasions inside people.

Why it matters: categorical data is everywhere (survey responses, A/B test conversions, medical outcomes, demographic breakdowns), and chi-square is the standard test of whether two categorical variables are related. The other association tests, and when each one applies, are in the third panel of the printable test-chooser poster.

Problem 21 in the practice problems works a 2 × 3 table the paper-exam way: the margins, six components, a bracket from Table F, and the two row percentages for the results sentence. Problem 40 takes a different 2 × 3 table on to Cramér's V.

Common questions

My chi-square is significant. Which cell is responsible?

The test itself will not tell you, because it adds every cell into one number. Look at the adjusted standardized residuals, which SPSS offers under Crosstabs → Cells and R returns as chisq.test(tab)$stdres. Each one is roughly a z-score for its cell under independence. A value beyond about ±2 marks a cell with noticeably more or fewer cases than independence predicts, and the sign tells you which. Read the pattern before you write the sentence: a significant χ² on a 3 × 4 table usually comes from one or two cells, not from the whole table at once. With many cells, treat the residuals as exploratory, not as a set of formal tests.

Should I use Yates' continuity correction on a 2 × 2 table?

Yates' correction subtracts 0.5 from each |O − E| before squaring, which shrinks χ² and raises the p-value. The idea is to compensate for describing discrete counts with a continuous curve, and SPSS prints it on every 2 × 2 table as Continuity Correction, one row under Pearson. On this lesson's trial table the uncorrected statistic is 12.38 with p = .00043 and the corrected one is 11.17 with p = .00083. Fisher's exact test, which uses no approximation at all, gives .00076, so the correction has overshot the exact answer. That is the usual pattern. So most current advice is to leave the correction off when the expected counts are large enough, and to use Fisher's exact test when they are not. If your course asks for the corrected value, report it and say which one you quoted.

Can a chi-square test tell me how strong the association is?

χ² itself can't; it grows with sample size, so a huge study can produce an enormous χ² from a trivial association. Pair the test with an effect size: Cramér's V (0 = none, 1 = perfect) for general tables, or the odds ratio for 2×2 tables, which our effect-size converter can translate into other metrics.