Section 2.9

Multiple Regression

Income depends on education and experience and region. Health depends on diet and exercise and age. Multiple regression extends the simple line to several predictors at once. Its most important idea is holding everything else constant.

What a coefficient means now

The model is just a line with more terms:

ŷ = b₀ + b₁x₁ + b₂x₂ + … + bₖxₖ

Each coefficient is now read differently. b₁ is the effect of x₁ on y while holding every other predictor fixed. It's the "all else equal" effect: the part of x₁'s influence that the other variables don't already explain. Because of this, regression can control for confounders, and controlling for one can completely change the result.

🎮 The Power of "Controlling For"

Outcome vs. x₁, with a lurking second variable shown by color (two groups). The gray line ignores it, and the colored lines are the model that holds it constant. The two sliders set the two parts of the bias separately: how much x₂ shifts the outcome, and how far it pulls the groups apart on x₁.

Naive slope (ignoring x₂)—
Controlled slope (holding x₂ fixed)—
Naive − controlled—
b₂ (x₂’s effect on y)—
δ (x₂ regressed on x₁)—
Bias b₂ × δ—
What omitting x₂ does—

When the confounder is strong, the naive slope (gray) can point the opposite way from the true within-group effect (colored). The gray line says "higher x₁ goes with higher y", but within each group the relationship is negative. This is Simpson's paradox. A multiple regression that holds the group constant avoids it.

The headline idea: a multiple-regression coefficient answers "what's the effect of this variable, among people who are otherwise alike?" Leaving out an important predictor (an omitted variable) can bias the others, sometimes enough to reverse their sign.

Omitted-variable bias, and the sign you can work out in advance

"Can bias the others" can be made precise, because the bias has a formula. Suppose the true model is

y = β₀ + β₁x₁ + β₂x₂ + ε

and you fit y on x₁ alone, because x₂ was never measured. Write δ (delta) for the slope you would get by regressing the missing x₂ on the included x₁. Then the coefficient you do get is not estimating β₁:

E[b₁] = β₁ + β₂ · δ

The bias is a product, so it is zero when either factor is zero. Leaving out a variable that has no effect on the outcome does no harm, and neither does leaving out one that is unrelated to the predictor you kept. Only a variable related to both can bias the estimate. So "control for everything you have" is the wrong rule, and choosing the controls is a question about causal structure.

The sign of a product follows from the signs of its two factors, and you can often argue both signs from what you know about the subject, without any data at all:

β₂: does x₂ raise or lower y?δ: does x₂ rise or fall with x₁?Bias β₂δSo your b₁ is
raises it (+)rises with it (+)positivetoo high
raises it (+)falls with it (−)negativetoo low
lowers it (−)rises with it (+)negativetoo low
lowers it (−)falls with it (−)positivetoo high

"Too high" means too far in the positive direction, which for a negative β₁ means closer to zero, or past it.

For example, a researcher regresses monthly wage on years of education and reports b₁ = €156 per year of schooling. Motivation was never measured, but there is little doubt that it belongs in the model. More motivated people earn more, so β₂ is positive, and more motivated people stay in education longer, so δ is positive too. Positive times positive is positive, so €156 is too high. The discussion should say that a year of education is worth at most about €156 a month, and probably less.

Suppose a later study measures motivation on a 1-to-10 scale and can estimate both parts. Holding motivation constant, education is worth €120 a year. Motivation itself is worth β₂ = €90 a point. Regressing motivation on education gives δ = 0.4 points per year of schooling. The numbers add up: 120 + 90 × 0.4 = 156. The first study computed its own number correctly but misread what it meant, and it had enough information to say in which direction it was wrong.

The widget above shows this identity. Its two sliders set the two parts separately, and the readouts print the estimated β₂, the δ between the two predictors, their product, and the gap between the naive and controlled slopes. The last two always agree to the last decimal, because in a sample the relation is an algebraic identity, not a statement about long-run averages. Drag either slider and watch the controlled slope: it does not move at all. Omitting a variable does not add noise to the estimate. It shifts the estimate by β₂δ, and collecting more data does not shrink that shift.

Reading the whole model

With p predictors and an intercept the model estimates p + 1 parameters. Everything that used n − 2 degrees of freedom in a one-predictor fit now uses n − p − 1. The residual standard error divides by it, every coefficient's t = b/SE(b) is tested on it, and every confidence interval takes its t* from it. Section 2.7's n − 2 is the case p = 1. Write the df out on any table you read: a model fitted on 200 people with six predictors is tested on 193, not on 199.

Each coefficient is a partial effect, holding the other predictors constant, and its p-value asks whether that predictor adds anything beyond what the others already explain. R² now reports how much of the outcome's variation all the predictors explain together.

Above the coefficients, software prints one test of the model as a whole. Its null hypothesis is that every slope is zero at once:

F = (R²/p) / ((1 − R²)/(n − p − 1)),    on F(p, n − p − 1)

Rejecting it says that at least one slope is not zero, but not which one. So a significant F above a table where no single coefficient reaches .05 is not a contradiction, and it is not rare. It usually means the predictors overlap so much that none of them is distinguishable on its own. That overlap is called multicollinearity. The opposite case, a nonsignificant F with a respectable R², usually means the model uses up degrees of freedom on predictors that explain little.

Software prints one more number beside R² and labels it R. The multiple correlation is √R², and it is the ordinary correlation between the observed y and the model's fitted values ŷ. This is a useful way to picture a many-predictor model. It combines every predictor into a single prediction, ŷ, and comparing ŷ with y is a one-variable problem again. Least squares makes ŷ correlate positively with y, so R is never negative, however many of the coefficients are negative.

🎮 More Predictors Always Look Better

The first two predictors are real, and the rest are pure noise. As you add them, plain R² keeps rising, but adjusted R² rises only for the real predictors, then flattens out. The widening gap between the two curves is the part of R² that comes from fitting noise.

R²—
Adjusted R²—
Last predictor added—

Plain R² can only go up as you add predictors, because adding a column never makes the best in-sample fit worse, even if that column is random noise. Adjusted R² subtracts a penalty for each predictor, so adding a predictor raises it only if that predictor's t is larger than 1 in absolute value:

adjusted R² = 1 − (1 − R²) · (n − 1)/(n − p − 1)

The penalty is in that fraction. It is always greater than 1 and it grows as p grows, so it enlarges the unexplained share before the subtraction. A predictor that explains nothing leaves R² where it was while the penalty grows. So adjusted R² can fall, and can even go negative, which plain R² never does. That makes adjusted R² the better measure for comparing models of different sizes, and it is why adding predictors just to raise R² is a mistake. When the comparison matters, model comparison has better tools than either number.

One continuous predictor and one dummy

The most common model in a real report mixes a measured predictor with a categorical one. Predict grip strength from height in centimeters and from sex, coded 0 for one group and 1 for the other, and the fit is

ŷ = b₀ + b₁·height + b₂·sex

Both coefficients are read the same "all else equal" way. The two sentences sound different only because the predictors are different kinds of variable. The height coefficient b₁ says that, holding sex constant, one more centimeter of height goes with b₁ more units of grip strength. That is a comparison between two people of the same sex. The sex coefficient b₂ says that, holding height constant, the group coded 1 averages b₂ units above the group coded 0. That is a comparison between two people of the same height. Neither sentence is about any particular person. Both describe the shape of one fitted surface.

If you draw it, the surface is two parallel lines, one per group, b₂ apart. The lines are parallel because the model assumes it, not because the data show it: this model cannot let height matter more in one group than in the other, even if it does. To let the slopes differ, add an interaction, which adds one more coefficient. The coding itself, including what to do with three groups or twenty, is covered in Section 2.10.

Putting the predictors on one scale

Each b is measured in the units of its own predictor, so their sizes say nothing about which predictor matters more. Weekly study hours and nightly sleep hours are both "hours" and still aren't comparable: study time varies enormously between students, sleep barely at all. The usual solution is the standardized coefficient, written β:

β = b × (SD of the predictor) ÷ (SD of the outcome)

You get the same numbers by z-scoring every variable before fitting. β then reads as "a one-standard-deviation rise in this predictor moves the outcome by β standard deviations". With a single predictor it is exactly the correlation r. Software usually prints it next to b (SPSS labels the column Beta), and APA style writes it with no leading zero: β = .42.

β depends on how much your particular sample happened to vary, so the same variable gets a larger β in a diverse sample than in a narrow one. That is why βs from different studies aren't interchangeable. β also means little for a dummy-coded group, whose "standard deviation" depends only on how many people happened to be in each group. Use β to rank predictors within one model, and quote the raw b when you want a sentence a reader can picture. If two predictors are strongly correlated with each other, neither coefficient is stable enough to rank at all. That problem is multicollinearity.

A caution

"Holding constant" is a statistical operation, not a real experiment. Regression can only adjust for variables you measured and put in the model. It can't control for confounders you forgot. And it never proves causation by itself: which variables belong in the model is a question about causal structure, and you settle it before you fit anything.

Sometimes you know the confounder but cannot measure it. The omitted-variable formula above still describes the bias, and instrumental variables are one remedy. You find a source of variation in the predictor whose causes are unconnected to the outcome, and you estimate the effect from that variation alone.

Why it matters: multiple regression is used throughout the social and health sciences, and almost everything left in this course builds on it: interactions, categorical predictors, logistic regression, and more. The fourth panel of the printable test-chooser poster matches each of these models to the kind of outcome it needs. None of the results can be trusted unless the assumptions hold. If you're planning a model of your own, Plan My Analysis turns the design into a test, a sample size, and a write-up skeleton.

Problem 63 of the practice problems takes a two-predictor model whose raw coefficients point one way and whose standardized coefficients point the other, and asks which one a journalist should have quoted.

The guide One Question, an ANOVA and a Regression applies the omitted-variable formula to a three-condition study. Adding one control variable cuts a group difference from 9.12 points to 5.09, and the guide shows the arithmetic that accounts for the missing 4.03 exactly.

Problem 72 shows two nested models and asks for the df, the multiple correlation, the overall F and both adjusted R² values, in a case where R² rises and adjusted R² falls. Problem 73 applies the sign table to three published coefficients whose authors could not measure the variable they left out.

Common questions

What does 'controlling for' a variable actually mean?

It means comparing like with like, statistically: the coefficient for x₁ describes how y differs between cases that differ in x₁ but are identical on every other predictor in the model. It's an adjustment computed from the data, not a real experiment, and it only works for confounders you actually measured and included.

How many predictors can I put in a regression?

A common rule of thumb is at least 10–15 observations per predictor. With fewer, the model starts fitting noise, so with n = 50 you should stay around 3–5 predictors. It is a guideline, not a law (effect sizes and collinearity matter more), but models that break it produce coefficients and R² values that do not hold up on new data. To check a model, use adjusted R² and cross-validation.

My R² is low but my predictors are significant. Is the model useless?

Those two numbers answer different questions. Significance asks whether a coefficient is different from zero. R² measures how closely the data follow the fitted surface. With enough people, a coefficient can be clearly significant while R² stays tiny: at n = 1,000, an R² of only .01 still gives F(1, 998) = 10.08, p = .002. For human behavior, an R² between .10 and .30 is ordinary in psychology or epidemiology, while an engineering calibration might reach .99. Such a result says two things: your predictor is reliably related to the outcome, and the outcome also varies for many reasons you did not measure. A model like that can teach you something real and still be poor at predicting individuals.