Section 2.12

Multicollinearity & Variable Selection

Multiple regression gives each predictor its own "holding the others constant" coefficient. But that works only if the predictors can be told apart. When two predictors are highly correlated with each other (multicollinearity), the data cannot separate their effects, and the coefficients become very unreliable.

The problem: overlapping predictors

Imagine predicting salary from "years of experience" and "age." They rise together almost perfectly. If you ask the model "what's the effect of experience holding age fixed?" you're asking about a situation that barely exists in your data: people with lots of experience but low age are rare. With so little independent variation in the data, the estimate is very imprecise.

🎮 Watch the Coefficients Go Haywire

The true effect of x₁ is exactly +1.00 (dashed line). Each dot is x₁'s estimated coefficient from one fresh sample. Raise the correlation between the predictors and the estimates spread out more and more.

Variance Inflation Factor—
Average estimate of b₁—
Spread of estimates (instability)—
Severity—

Perfect, and merely bad

Multicollinearity ranges from perfect to merely high, and only the second needs judgment.

Perfect collinearity is when a predictor is an exact linear function of the others: weight in kilograms entered alongside weight in pounds, three subscale scores entered alongside their own total, or a full set of k dummies entered alongside the intercept. Then the coefficients have no unique solution, because infinitely many sets of coefficients fit the data equally well. Software does not warn you clearly. It drops one predictor and prints the model without it, sometimes with only a footnote. If a variable is missing from your Coefficients table, check for perfect collinearity.

For high but imperfect collinearity, start with the correlation matrix and look for any pair beyond roughly ±.8. Beyond ±.9, almost everyone agrees it is a problem. Then ask a question the matrix cannot answer, because it only looks at pairs: is any one predictor well predicted by all the others together? The VIF answers that question.

The tell-tale signs

  • Huge standard errors: coefficients with enormous uncertainty, so individual predictors look "not significant" even though the model as a whole fits well.
  • Unstable, surprising coefficients: estimates that flip sign or become very large, and change a lot when you add or drop a variable.
  • A high VIF. The Variance Inflation Factor measures how well the other predictors explain this one. The recipe and the thresholds are in the next section.
  • A low tolerance. Software usually prints a second column beside the VIF called tolerance. It is the reciprocal of the VIF: tolerance = 1 − R², the share of a predictor's variance that is its own, not explained by the other predictors. So VIF > 5 is tolerance < .20, and VIF > 10 is tolerance < .10. A predictor with tolerance .08 has only 8% of its variance left once the other predictors are accounted for.

Working out a VIF

Computing a VIF once by hand shows where the number comes from. Take predictor j. Make it the outcome. Regress it on every other predictor in the model, ignoring y entirely. Keep the R² from that side regression and call it R²j. Then

VIFj = 1 / (1 − R²j)

Repeat for each predictor, so a four-predictor model means four side regressions and four VIFs. R²j = 0 gives VIF 1, the minimum, and means this predictor shares nothing with the rest. R²j = .50 gives 2, .80 gives 5, .90 gives 10 and .99 gives 100. The curve is steep at the top. Raising R²j from .90 to .99 takes the VIF from 10 to 100, while raising it from 0 to .50 takes it only from 1 to 2.

The name describes what the number does. The variance of bj is multiplied by exactly this factor, compared with what it would have been if predictor j were unrelated to the others, and its standard error by √VIF. At VIF 10 the standard error is about 3.2 times what it could have been, and the confidence interval is 3.2 times wider.

The textbook threshold is VIF > 10, where R²j passes .90, and a more cautious convention uses 5. Both are conventions. It is more useful to read the VIF beside the coefficient it belongs to. A VIF of 12 on a control variable you will not interpret is not a problem, while a VIF of 6 on the predictor your argument rests on deserves a paragraph in your write-up.

Key reassurance: multicollinearity inflates the uncertainty of individual coefficients, but it does not hurt the model's overall predictions or its R². If you only care about predicting y, you can often ignore it. It matters only when you need to interpret individual coefficients.

What to do about it

  • Drop or combine redundant predictors (keep one of experience/age, or average several items into a single index).
  • Use domain knowledge to choose which variable to keep, not a blind algorithm.
  • Regularize: ridge regression shrinks coefficients toward zero to reduce the instability, and the related lasso can drop a redundant predictor automatically.

Two ways to choose the predictors

Courses teach two approaches to choosing predictors, and it helps to know which one your course uses.

Theory-first specification means writing the model down before anyone looks at the data. A predictor is included because there is a reason it belongs: the treatment itself, the confounders the causal story names, the controls a reviewer is certain to ask about. Nothing is added or dropped because of a p-value. As a result, every test in the output has its stated error rate, and you can make the omitted-variable argument of Section 2.9 in advance, not improvise it afterwards.

General-to-specific works the other way, and many courses teach it as the standard procedure. Start from a general model with every candidate predictor, remove the weakest one, refit, and repeat. Judge each step by the coefficient's t, by adjusted R² and by AIC or BIC. You stop when removing any remaining predictor would make the model worse by those criteria. There are real arguments for it. A general starting model has less omitted-variable bias than any of the smaller models nested inside it, and the sequence of decisions is written down where a reader can see it.

Its problems are those of automated stepwise selection. Choosing predictors by their p-values runs many hidden tests and keeps whatever happened to fall below .05. So the surviving p-values are too small, R² is too high, and a small change in the data produces a different final model. The procedure ignores theory, so it can drop a confounder that was in the model to remove bias. And because it is driven by significance, not by causal structure, it can leave you with a tidy model whose coefficients have the omitted-variable bias that Section 2.9 taught you to anticipate.

In practice, let theory choose the predictors, and use general-to-specific to report a sequence of models, not to discover one. If you need automatic selection for prediction, not explanation, use a penalized method such as the lasso with cross-validation.

Why it matters: real datasets are full of related measurements, so some multicollinearity is almost always present. Recognizing it stops you from over-interpreting a coefficient that is mostly noise.

Problem 69 of the practice problems shows the textbook symptom in one output block: R² = .31 with F(3, 146) = 21.87, p < .001, no significant predictor, and a negative sign on years of experience. You work out the tolerance and VIF from each predictor’s R² on the other two, then weigh three possible fixes.

Common questions

What VIF value indicates a problem?

Common alarm thresholds are VIF > 5 (cautious) or VIF > 10 (lenient). A VIF of 5 means that predictor's coefficient variance is inflated 5-fold because other predictors largely explain it. Treat these as warnings, not verdicts. A high VIF between a predictor and its own squared term is normal, while a VIF of 4 between two conceptually distinct predictors may still deserve thought.

Every pairwise correlation is below .6. Am I safe from multicollinearity?

No. The correlation matrix shows only overlap between pairs, and collinearity can involve many variables. A predictor can be almost perfectly predicted by a combination of the others while correlating only modestly with each one. Take four predictors where the fourth is roughly the sum of the first three. Every pairwise correlation is about .55, which looks harmless, yet the fourth predictor's VIF is 13 and its tolerance is .08. That is why the VIF is a separate diagnostic: it regresses each predictor on all the others at once.

Why is stepwise regression frowned upon?

Because selecting predictors by significance amounts to many tests you never report. Whatever chanced below .05 stays in, so the reported p-values are biased low and R² is inflated, and a slightly different sample would often give an entirely different "final" model. It also ignores theory. Choose predictors from subject knowledge, and if you need automatic selection for prediction, use a penalized method such as the lasso with cross-validation.