Section 2.11

Interactions in Regression

Sometimes the effect of one predictor depends on another. A training program might help beginners a lot and experts barely at all. A drug's benefit might grow with the dose only for younger patients. The factorial-ANOVA interaction has a regression version: the interaction term.

Adding a product term

A model with two predictors and their interaction looks like this:

ŷ = b₀ + b₁·x + b₂·group + b₃·(x × group)

The last term is the product of the two predictors. Its coefficient b₃ measures how much the slope of x changes between groups. If b₃ = 0, both groups share one slope (the lines are parallel, just a level difference). If b₃ ≠ 0, the groups have different slopes: x matters more for one than the other.

🎮 The Interaction Term Tilts the Lines

Two groups, one continuous predictor x. When the interaction is zero the lines are parallel. Move the slider away from zero and the orange group's slope differs more and more from the blue group's.

Group 1 slope—
Group 2 slope—
Interaction b₃—
Effect of x…—

Why the main effects get tricky

When an interaction is present, the "main effect" of x is no longer a single number: it's different in each group, so quoting one average slope can be misleading. With a product term in the model, b₁ becomes the slope of x specifically when group = 0 (the reference), not an overall effect.

Center your predictors. Interpreting interactions is far easier when continuous predictors are mean-centered first. Then b₁ and b₂ are effects "at the average" of the other variable. Without centering they are effects at zero, which is often a meaningless value. It also reduces multicollinearity between the predictors and their product.

Continuous × continuous interactions

Interactions aren't just for groups. Two continuous predictors can interact too: the effect of advertising spend on sales might depend on the product's price. These interactions are harder to picture, because the "lines" become a twisted surface. So you usually probe them with simple slopes: plot the effect of x at a few representative values of the moderator (say, low, medium and high), as the widget does with two groups.

Why it matters: many effects differ between people or conditions. An interaction term lets you test whether an effect depends on another variable, and describe exactly how it changes.

If you plan a study around an interaction, remember that it is estimated less precisely than the main effects. A study powered to find a main effect is usually not powered to find the interaction. Power analysis for complex designs works out how much larger the sample needs to be.

Problem 64 of the practice problems puts a group difference and a pretest covariate in one model and asks what the raw 6.4-point advantage was measuring. Problem 66 shows an abstract whose headline is a main effect and asks what the moderator does to that claim.

Common questions

Why should I center variables before creating an interaction term?

Centering helps in two ways. With a product term in the model, b₁ is the effect of x₁ when the other variable equals zero, and zero is often a meaningless value (age 0, income 0). Centering moves "zero" to the mean, so b₁ becomes the effect at a typical value. It also reduces the artificial correlation between x and x×z, so the estimates suffer less from multicollinearity.

Do I have to keep the main effects if only the interaction is significant?

Yes, keep them. This is the principle of marginality: a model containing x·z should also contain x and z. If you drop a main effect, you force that variable's effect to be exactly zero wherever the other variable equals zero. The remaining coefficients then no longer mean what their names suggest, and the interaction picks up whatever the dropped term was explaining. The fit also starts to depend on how you coded the variables, so rescaling a predictor can change your conclusion. A non-significant main effect alongside a significant interaction is not a problem to fix. It is the usual result when an effect is real in one condition and absent in the other, and that pattern is your finding.

What are simple slopes?

Simple slopes are the effect of x computed at chosen values of the moderator, usually its mean and ±1 SD. They turn a product coefficient into direct statements: "among low-experience users the feature gains 4 points; among high-experience users, 0.5." Simple-slope tests then tell you at which moderator values the effect is significantly different from zero.