Section 2.7

Inference for Regression

Fitting a line to a scatterplot is description. Section 1.18 did that: least squares, the slope, the intercept, the residuals. This lesson asks how much the slope would change from sample to sample, and whether it is different from zero. It answers with the same t statistic you already know.

The population regression model

Inference needs a model of where the data came from. The standard model says that each value of x has a whole population of y values. Their means fall exactly on a straight line, and they scatter around it by the same amount everywhere:

y = β₀ + β₁x + ε ε ~ N(0, σ)

β₀ (beta-nought) and β₁ (beta-one) are the parameters, fixed and unknown, and σ (sigma) is the size of the scatter. The line you fit gives b₀ and b₁, which estimate them.

σ is estimated by the residual standard error, sometimes just called s, which is the typical size of a residual:

s = √(Σ(y − ŷ)² / (n − 2))

The divisor is n − 2 because fitting the line used up two degrees of freedom, one for the slope and one for the intercept. Every test below uses those n − 2 degrees of freedom.

The conditions, and what they guarantee

That equation comes with five conditions. Linearity: the population means of y fall on a straight line. Independence: one observation's deviation says nothing about the next one's. This depends on how the data were collected, and a scatterplot cannot settle it. Equal variance, usually called homoscedasticity: the scatter around the line is the same size at every x. Normality: the deviations ε are roughly normal. Section 2.8 shows you how to look for problems with these four in a residual plot.

The fifth is exogeneity, and no plot will show it to you. It says the predictor is unrelated to the error term: whatever else moves y, it does not also move x. The residuals cannot test it. They are measured from a line fitted to these same data, so they are uncorrelated with x by construction, whether the condition holds or not. You can only argue for exogeneity from how the data were collected.

If the first three conditions and exogeneity hold, a theorem says that least squares is BLUE, the best linear unbiased estimate of β₁. Read the acronym backwards. Unbiased means that across repeated samples b₁ averages out to β₁ exactly. Linear means b₁ is a weighted sum of the y values, which makes its mean and standard error easy to derive. Best means that among all estimators that are both linear and unbiased, least squares has the smallest standard error. Normality is not needed for any of the three. It is needed only to make the t distribution below exact when n is small.

Exogeneity is the condition that fails most often, and its failure is the most serious, because it makes b₁ biased. The other failures make the estimate less precise. It usually fails because a variable that belongs in the model was left out, and Section 2.9 gives a formula for the resulting bias, including its sign.

🎮 One Line, Two Bands

The sliders change the true slope and the scatter but reuse the same random draws, so the points keep their pattern. The inner band is the confidence band for the mean response, and the outer band is the prediction band for one new case.

b₁—
SE(b₁)—
t—
p—
95% CI for β₁—
R²—

Two ways to write the same slope

Before the standard error, one more look at the slope itself. The covariance of two variables measures how they move together. Its formula is the variance formula with one of the two copies of x replaced by y:

cov(x, y) = Σ(x − x̄)(y − ȳ) / (n − 1)

A point above both means contributes a positive product. So does a point below both, since the product of two negatives is positive. A point on one side of x̄ and the other side of ȳ contributes a negative product, and the sign of the sum says which kind of point dominates. The total does not tell you how strong the relationship is, because its units are the product of the two variables' units. On the screen-time data it is −0.681 sleep-quality points times hours, a unit that is hard to interpret.

Two familiar numbers rescale the covariance so that it can be interpreted. Divide it by both standard deviations and every unit cancels, which gives the correlation r. Divide it by x's variance alone and only the x units cancel. What is left is a rise in y per unit of x, which is the slope:

b₁ = cov(x, y) / var(x)

Section 1.18 wrote the same number as b₁ = r · (sy/sx), and one line of algebra connects them: r = cov(x, y)/(sxsy), so r · sy/sx = cov(x, y)/sx², and sx² is var(x). Check it on the 60 screen-time people. The covariance is −0.681 and var(x) = 2.904, so the ratio is −0.2346. The correlation is −.205 and sy/sx = 1.9467/1.7041 = 1.1424, so the product is −0.2346. On an exam, use whichever form matches the numbers you are given. The covariance form is also the one that extends to several predictors, where a matrix of covariances replaces the single number.

The standard errors

How precisely you know the slope depends on two things, and the formula puts them side by side:

SE(b₁) = s / √(Σ(x − x̄)²)

Tight scatter around the line makes s small and the slope precise. Widely spread x values make the denominator large and do the same. The second point matters when you design a study. If you can choose where to measure, spreading your x values out makes the slope more precise at no extra cost. A study that samples only a narrow range of the predictor has little power to detect a slope. Raise the scatter slider in the widget and the bands widen while the points keep the same left-to-right pattern.

The intercept has a standard error too, built from the same two ingredients and one extra term:

SE(b₀) = s · √(1/n + x̄² / Σ(x − x̄)²)

Tight scatter and spread-out x values help here as well. The x̄² term is new, and it comes from extrapolation. b₀ is the line's height at x = 0. The further your data are from zero, the further the line has to be extended to reach it, and a small error in the slope moves the far end a long way. Center the predictor by subtracting x̄ from every value and that term drops out, leaving SE(b₀) = s/√n, the standard error of a mean. That makes sense, because a centered intercept is just the mean of y.

On the screen-time data SE(b₀) = 0.597, so the intercept 6.845 has t(58) = 11.46 and a 95% interval of [5.65, 8.04]. That t is huge but tells you almost nothing, because the hypothesis it rejects is that sleep quality is zero among people with zero screen hours. Most intercept tests are like that. Software prints the row, but researchers rarely report it.

Testing and estimating the slope

The usual null hypothesis is that the predictor tells you nothing about y, H₀: β₁ = 0. As always, the statistic is estimate over standard error, and it is a t on n − 2 degrees of freedom:

t = b₁ / SE(b₁) CI: b₁ ± t* SE(b₁)

Take screen-time.csv from the practice datasets, 60 people with daily screen hours and a sleep-quality score. The fitted line is ŷ = 6.845 − 0.235x, with s = 1.922 and SE(b₁) = 0.147. That gives t(58) = −1.598, p = .115, and a 95% interval for β₁ of [−0.529, 0.059].

The interval tells you more than the test. It contains zero, so it agrees with the test, but it also shows what the data do and do not rule out. A drop of half a point of sleep quality per screen hour is compatible with these 60 people, and so is a small increase. So the study is too small to settle the question. That is a different statement from "there is no effect", and Section 3.5 explains how to write about the difference.

The ANOVA table for a regression

Software prints an analysis-of-variance table above the coefficients, and it uses the same decomposition as one-way ANOVA. Total variation in y splits into the part the line explains and the part it leaves behind:

SST = SSR + SSE

SourceSSdfMSF
Regression9.4319.4322.554
Residual214.16583.692
Total223.5959

R² is SSR/SST, here 9.43/223.59 = .042, so the line accounts for about 4% of the variation in sleep quality. The F statistic is exactly the square of the t from the slope test: 2.554 = (−1.598)², and both report p = .115. In simple regression the F test adds nothing, because there is only one predictor to test. It becomes useful in multiple regression, where it tests whether the predictors as a set explain any of the variation in y.

Reading the output

SPSS prints a regression as three stacked boxes, and every number in them has already appeared above. Here they are on the screen-time fit, in SPSS's own column names.

Model Summary gives R, R Square, Adjusted R Square and Std. Error of the Estimate: .205, .042, .026 and 1.922. The first column is the multiple correlation, the correlation between y and the model's own predictions, which with one predictor is the absolute value of r and so never comes out negative. The last column is s, the residual standard error, under a different name.

ANOVA is the table in the previous section, with Sig. = .115 for F(1, 58) = 2.554.

Coefficients is the box people actually read:

ModelBStd. ErrorBetatSig.95% CI LowerUpper
(Constant)6.8450.59711.463<.0015.6508.040
screen_hours−0.2350.147−.205−1.598.115−0.5290.059

B is the coefficient in the outcome's own units, Std. Error is its standard error, and t is B divided by its standard error: −0.235/0.147 = −1.598. Doing that division once on a printed table is a good habit, because it shows that the t and Sig. columns follow from the first two. Sig. is the two-tailed p, and the two CI columns are b ± t*·SE with t*(58) = 2.002. The Beta column is the standardized coefficient, which in a one-predictor model equals r exactly. The (Constant) row leaves it blank because the intercept of a standardized regression is always 0.

Three cross-checks connect the boxes, and exam questions often ask for one of them. The Coefficients t squared is the ANOVA F: (−1.598)² = 2.554. R Square is SSR/SST from the ANOVA box. And s in Model Summary is √MSresidual = √3.692 = 1.922. If two boxes disagree, you have misread a row.

Confidence band or prediction band?

Both bands are centered on the same fitted line, but they answer different questions.

A confidence interval for the mean response at a given x asks where the average y for everyone at that x lies. A prediction interval gives a range for the y of one new individual at that x. The second must also allow for that individual's own deviation from the line, so its standard error has an extra 1:

mean: s√(1/n + (x₀ − x̄)²/Σ(x − x̄)²) new case: s√(1 + 1/n + …)

At 6 screen hours the line predicts 5.44. The confidence interval for the mean runs [4.60, 6.28], a width of 1.68. The prediction interval for one person runs [1.50, 9.37], a width of 7.87, nearly five times wider. It barely narrows as n grows, because that leading 1 does not shrink with sample size. More data pins down the line, but it does not make one person's value more predictable.

Both bands are narrowest at x̄ and widen toward the edges, because a small error in the slope matters more the further you are from the center. That widening is also a reason not to extrapolate. Far from x̄ the line is poorly determined, and outside the range of the data the linear model may not hold at all.

Why it matters: a slope on its own does not tell you how much it would change in another sample. Its standard error does. This lesson adds inference to the regression you met in Stats 1. Every extension that follows, from several predictors to a yes/no outcome, keeps the same structure: estimate, standard error, t, interval.

Problem 71 of the practice problems shows a four-line printout and asks you to rebuild the rest: the standard error from s and Σ(x − x̄)², the t and its interval, the F that is its square, and the intercept row.

Common questions

What is the difference between a confidence interval and a prediction interval in regression?

Both are centered on the same fitted value, but they answer different questions. The confidence interval is about the average outcome of everyone with that x. The prediction interval is about one new person, so it has to cover that person's scatter around the line as well. That scatter is the extra 1 under the square root in its standard error. In the screen-time data at 6 screen hours, the first interval is 1.68 points wide and the second is 7.87, nearly five times wider. A bigger sample narrows the confidence interval a lot and the prediction interval very little, because one person's scatter around the line stays the same however many people you measure.

Why is the slope test on n − 2 degrees of freedom rather than n − 1?

Because the line has two estimated parameters. Residuals are measured from a line whose slope and intercept were both estimated from the same data. That imposes two constraints, so only n − 2 of the residuals are free to vary. The residual standard error divides SSE by n − 2 for that reason. The slope's standard error is built from s, so its t has the same n − 2 degrees of freedom. In multiple regression this becomes n − p − 1 for p predictors plus an intercept.

My regression has a tiny R-squared but a significant slope. Which one should I believe?

Both, because they measure different things. The t test asks whether the slope is different from zero, and with a large enough sample even a very shallow slope is significant. R-squared measures how much of the variation in y the line accounts for, and a real but small effect accounts for little. A significant slope with an R-squared of .04 means there is probably a relationship, but it is nearly useless for predicting any one person. Report the slope with its confidence interval so a reader can see how big the effect is, not only that it was detected.