Assumptions of Regression
A regression coefficient and its p-value can be trusted only if the model's assumptions hold. Ordinary least squares has four, and their initials spell LINE. There is also a separate problem: a single data point with enough leverage to move the whole line.
The LINE assumptions
- L · Linearity: the relationship between predictors and the mean of
yis actually a straight line (or plane). Check with residual plots. - I · Independence: the errors are independent of each other. Violated by time series and clustered/repeated data.
- N · Normality of residuals: the errors are roughly normal. Matters most for small samples; check with a Q-Q plot of residuals.
- E · Equal variance (homoscedasticity): the residuals have constant spread across all fitted values. A funnel shape means this assumption is violated.
Leverage vs. influence vs. outlier
These three are often confused, but they are different. An outlier has an unusual y (far from the line). A high-leverage point has an unusual x (far from the other predictors). A point is influential only when it's both: extreme in x and off the trend, so that removing it noticeably moves the line. Drag the point below to see the difference.
🎮 Leverage & Influence Playground
Drag the ringed point, or focus the chart and nudge it with the arrow keys (hold Shift for bigger steps). The solid line is fit with it; the dashed line is fit without it. At the edges (high leverage), small moves shift the whole line.
Put the point in the middle (average x) and even a big vertical move barely changes the line: low leverage. Drag it far left or right and the line follows it: high leverage. A point changes the fit a lot only when it has both high leverage and a large residual.
Cook's distance bundles leverage and residual size into a single number measuring how much the whole fit shifts when a point is removed. Points with a large Cook's distance deserve a hard look. But never delete data just because it's influential. Investigate first: is it an error, or a real and important case?
How high is "high"?
The readout above gives the dragged point's leverage as a number and as a multiple of the average leverage. The second figure matters because the total leverage is fixed. Across all n points the leverages always sum to p, the number of parameters the model estimates (an intercept plus one slope here, so p = 2). So the average leverage is p/n, whatever the data look like. This chart has 13 fixed points plus the one you drag, so its average is 2/14 = 0.14. A point can rise above that only by taking leverage from the other points.
The conventional cutoffs are 2p/n, and sometimes a stricter 3p/n, which come to 0.29 and 0.43 for this chart. Put the point at the center and its leverage falls to 0.07, half the average. Drag it to either edge and it reaches about 0.27, 1.9 times the average and still under the cutoff. With thirteen well-spread points, one point cannot gain much leverage. Leverage is relative: the same x value is unremarkable in a wide sample and extreme in a narrow one.
Residuals need adjusting too, because of leverage. A high-leverage point pulls the line toward itself, so its raw residual is too small to compare with the others. The fix is a studentized residual, which divides each residual by its own standard error, s√(1 − h). SPSS's Casewise diagnostics box flags points using the plain standardized residual. The studentized version is a separate tick in the same dialog's Save… panel, and it is the better one to use. Because of its leverage, a point can look unremarkable by the first measure and extreme by the second.
Put the two together and you get Cook's distance in closed form: D = (r²/p) × h/(1 − h), where r is the studentized residual. The first factor comes from the residual and the second from the leverage. Because D is a product, it is near zero whenever either factor is, so influence requires both. The usual cutoffs are D > 1 and the stricter 4/n, which for these 14 points is also 0.29.
These numbers tell you which rows to look at, not what to do with them. Whether a row is an error or a real, interesting case is a question about the study, not about the arithmetic.
When assumptions fail
- Nonlinearity → add polynomial or transformed terms.
- Non-independence → use models built for it (time-series, mixed/multilevel models).
- Non-normal or heteroscedastic errors → transform
y, or use robust and bootstrapped standard errors. - Influential points → report the fit with and without them, and explain the difference.
"Robust" in the third bullet has a specific meaning, and your software will list it by an abbreviation. A heteroscedasticity-consistent standard error (the HC family, usually HC3 by default) computes each coefficient's standard error from the individual squared residuals, not from one pooled variance. The slope estimates do not change at all. Only their standard errors do. That is the right fix, because unequal variance distorts the standard errors but leaves the coefficients unbiased.
Why it matters: the software will report a slope and a p-value however badly the assumptions are violated. Only by checking them, and looking for influential points, can you tell whether those numbers mean what they claim. The printable assumption-checks sheet puts the whole routine, including this lesson's four assumptions, on a single page.
Problem 43 of the practice problems applies both measures to one case. In a 40-property regression, case 31 has leverage h = 0.31 and Cook’s D = 0.68, and you have to say whether either number is unusual enough to act on.
Common questions
Do my predictors need to be normally distributed?
The normality assumption concerns the residuals, not the predictors or even the raw outcome. Believing otherwise is one of the most persistent regression myths. Skewed predictors, binary dummies and lumpy x-distributions are all fine. Fit the model, then check a Q-Q plot of the residuals. That is the only normality that matters, and it matters mostly in small samples.
What is leverage in regression?
Leverage is a point's potential to move the line, and it depends only on how unusual its predictor values are. A point far from the center of the x's has high leverage, like a weight at the end of a seesaw. Leverage alone is not a problem: a high-leverage point right on the trend makes the fit more stable. Trouble comes from leverage plus a large residual, a combination called influence.
Is there a real test for equal variances, or do I just eyeball the residual plot?
There are two standard tests, and your software already has them. The Breusch–Pagan test (car::ncvTest() in R, het_breuschpagan() in statsmodels) regresses the squared residuals on the fitted values and asks whether the spread changes with them. White's test does the same but also allows for curvature. Run them, but don't rely on them alone, because their power grows with n like any other test's. At n = 2,000 a violation too small to matter comes back significant, and at n = 25 a serious funnel can go undetected. Look at the plot, run the test, and if they disagree, trust the plot.