Categorical Predictors & Dummy Coding
Regression works with numbers: it multiplies each predictor by a coefficient. What do you do with a predictor like treatment group, country, or education level, where the values are categories, not amounts? You can't just label them 1, 2, 3 (that would falsely claim group 3 is "three times" group 1). You use dummy coding.
How dummy coding works
For a categorical variable with k groups, you create k − 1 new 0/1 indicator variables. One group is left out as the reference; each dummy variable flags membership in one of the other groups. Three groups (Control, Drug A, Drug B) with Control as reference becomes two dummies:
🎮 Coefficients Are Just Group Differences
Set each group's average outcome. The intercept equals the reference group's mean; every other coefficient is simply that group's distance from the reference (the dashed line).
Reading the coefficients
Each coefficient has a simple meaning:
- The intercept is the predicted outcome for the reference group (all dummies = 0).
- Each dummy coefficient is the difference between that group's mean and the reference group's mean.
- Each coefficient's p-value tests whether that group differs significantly from the reference.
The key fact: a regression with a single categorical predictor is a t-test (for two groups) or a one-way ANOVA (for more). So t-tests, ANOVA, and regression are versions of the same linear model.
Choosing the reference
Switch the reference group above and watch the coefficients change, but notice that the predicted group means stay identical. The reference is just the baseline you measure everything against, so pick the one that makes your comparisons most meaningful: the control condition, the placebo, the "standard" category. The model is the same either way, and only the meaning of the individual coefficients changes.
The same analysis twice, on the study-methods data
Here is the same analysis done both ways, on one file. study-methods.csv holds 105 students, 35 in each of three revision conditions, with an exam score each. Section 2.2 runs it as a one-way ANOVA. Here the same file is analyzed as a regression.
Three groups means two dummies. With rereading as the reference, d₁ is 1 for a flashcards student and 0 for everyone else, and d₂ is 1 for a practice-testing student and 0 for everyone else. A rereading student is 0 on both, so the reference group needs no column of its own. Fit ŷ = b₀ + b₁d₁ + b₂d₂ and the Coefficients box shows:
| Model | B | Std. Error | t | Sig. | 95% CI |
|---|---|---|---|---|---|
| (Constant) | 66.143 | 1.649 | 40.12 | <.001 | [62.87, 69.41] |
| d₁ flashcards | 8.286 | 2.331 | 3.55 | <.001 | [3.66, 12.91] |
| d₂ practice testing | 10.686 | 2.331 | 4.58 | <.001 | [6.06, 15.31] |
Every number in that table comes from the group means. The group means are 66.143, 74.429 and 76.829. The intercept is the rereading mean exactly. 8.286 is 74.429 − 66.143 and 10.686 is 76.829 − 66.143. So each coefficient is a distance from the reference, and a group's predicted mean is the intercept plus its own coefficient. The two standard errors are identical because the groups are the same size. Both are √(MSE(1/35 + 1/35)) = 2.331, the standard error of the pooled pairwise comparison in Section 2.3.
The model's ANOVA table, printed above the coefficients, is:
| Source | SS | df | MS | F | Sig. |
|---|---|---|---|---|---|
| Regression | 2200.305 | 2 | 1100.152 | 11.57 | <.001 |
| Residual | 9701.829 | 102 | 95.116 | ||
| Total | 11902.133 | 104 |
That is the one-way ANOVA table, cell for cell. The regression sum of squares is SSG and the residual sum of squares is SSE. The two df are I − 1 and N − I, and F(2, 102) = 11.57, p < .001 is the number Section 2.2 reports on this same file. R² = .185 is η². The residual standard error √95.116 = 9.753 is sp, the pooled SD that every follow-up comparison uses. Nothing has been approximated: it is one analysis printed in two layouts.
Compared with the ANOVA table, the regression layout gives every difference a confidence interval. You can say that flashcards beat rereading by between 3.66 and 12.91 points, not just that the difference is significant. It also gives each dummy a t, which is the pairwise comparison of that group against the reference on N − I degrees of freedom. What it leaves out is the third comparison. No row in the table compares flashcards with practice testing, because neither of them is the reference. That gap is 2.40 points, t(102) = 1.03, p = .31, and you get it by refitting with a different reference or through a contrast. The F test does not cover it either, because the F test never tests any single comparison.
Why the third dummy has to go
Add d₃, flagging rereading, and the model no longer has a unique fit. For every student in the file d₁ + d₂ + d₃ = 1, and the intercept is itself a column of 1s, so the four columns are exactly linearly dependent. This is perfect collinearity. The estimates are not just unstable: infinitely many sets of coefficients fit these data equally well, and the arithmetic cannot choose between them. Software responds by dropping a column and fitting the model without it, sometimes with only a blank row or a footnote to show it. This is the dummy variable trap, and using k − 1 dummies avoids it.
The other coding scheme you'll meet
Dummy coding (software often calls it treatment coding, and it is R's default) is not the only way to turn groups into numbers. Effect coding uses the values −1, 0 and 1, and it changes the baseline. The intercept becomes the grand mean of all groups, not one group's mean, and each coefficient becomes that group's distance from the grand mean. Both schemes fit the same model and give the same predictions and R². They differ only in what the individual coefficients mean. The difference matters once you add an interaction. Under dummy coding the lower-order terms describe the reference group, not an average, and reading them as averages is the most common misreading of a factorial regression.
Usually you code dummies so that you can add covariates beside them in the next model. The guide One Question, an ANOVA and a Regression takes a three-group study through that sequence: ANOVA first, then the dummies alone to show that the two agree, then the controls. Assignments usually ask for that order, and it keeps the coefficients easy to read.
Why it matters: almost every real dataset mixes numeric and categorical predictors. With dummy coding, a single regression can handle "dose (mg) and treatment group and sex and region" all at once, and the rest of applied modeling builds on that.
Common questions
Why do I create k − 1 dummy variables instead of k?
Because the k-th dummy is redundant. If a case is 0 on all the others, it must be in the reference group, so a full set of k dummies would be perfectly collinear with the intercept (the "dummy variable trap"), and software would drop one of them anyway. The reference group's mean is the intercept, and each dummy coefficient measures a difference from it.
What if my categorical variable has 20 or 50 levels?
Dummy coding still works, but it needs many parameters. 50 countries means 49 dummies and 49 degrees of freedom, and the coefficients for the small categories rest on a handful of cases each. There are three alternatives, roughly in order of preference. Collapse levels into meaningful groups you can defend in advance (region rather than country), which is a coding decision and belongs in your codebook. Fit the variable as a random effect in a multilevel model, which pools information across levels to estimate each one. Or, if the levels have a real order and roughly even spacing, enter the variable as a single numeric predictor, which uses one degree of freedom, not 49.
Is ANOVA just a special case of regression?
Yes. Run a regression with one dummy-coded categorical predictor and you get exactly the same F, p, and group means as the one-way ANOVA. With two groups it reduces further to the t-test. ANOVA, t-tests, ANCOVA, and regression are one linear model in different notation, so once you understand regression you understand all of them.