ANCOVA: Controlling for Covariates
A training program can look very effective until you notice that its students were already stronger on the pretest. ANCOVA (Analysis of Covariance) compares group means after statistically holding a covariate constant, so you compare like with like. It is the method behind many of the "controlling for age / baseline / IQ" statements you read in papers.
Why compare like with like
An ordinary ANOVA asks: do the group means differ? But raw means mix two things together: the effect of the treatment and any differences the groups already had at the start. If the treatment group started ahead on a variable that predicts the outcome (a covariate), the raw comparison makes the treatment look better than it is. If they started behind, it can hide a real effect.
ANCOVA fixes this by fitting a regression line from the covariate to the outcome within each group, then comparing the groups at the same covariate value (conventionally the overall mean). Those are the adjusted means, and the vertical gap between the parallel lines is the adjusted effect.
🎮 The Fair-Comparison Machine
Each dot is a person: pretest score across, outcome up. The lines are the within-group regression fits. The dashed vertical line is the overall pretest mean, and the big markers on it are the adjusted means. Give the orange group a head start: the naive comparison then gets the effect wrong, while the ANCOVA estimate stays close to the true effect. (The same people stay on screen as you drag, and "New sample" redraws them.)
What ANCOVA actually does
It does two things:
- It corrects for baseline imbalance. The adjusted difference is the raw difference minus the part explained by the covariate gap: adjusted = raw − slope × (covariate gap between groups). Set the head start to +10 with a real effect of 0, and the naive test finds a significant effect while ANCOVA correctly reports an effect near 0.
- It reduces noise. Even with zero imbalance (perfect randomization), the covariate explains outcome variation that would otherwise be part of the error term. Removing it shrinks the residual variance, so the same effect gets a smaller p-value at no cost. Try it: head start 0, effect 5, slope high, and compare the two p-values.
Try two settings: set the true effect to 0 and the head start to +10. The naive t-test "finds" an effect that is entirely baseline difference, while ANCOVA reports ≈ 0. Now set the head start to −10 with a true effect of 6: the naive test misses a real effect that ANCOVA recovers. In both cases the naive comparison is wrong, in opposite directions.
The assumptions behind ANCOVA
- Linearity: the covariate–outcome relationship should be roughly a straight line within each group.
- Homogeneity of regression slopes: the group lines must be parallel: the covariate's effect is the same in every group. If treatment changes the slope itself, a single "adjusted difference" doesn't exist. Fit the interaction model instead and report how the effect varies.
- The covariate is measured before (and unaffected by) treatment. Adjusting for a variable the treatment itself moved throws away part of the effect you're trying to measure.
- Plus the usual assumptions from ANOVA: independent observations, roughly normal residuals, similar residual spread.
"Controlling for" has limits. In a randomized experiment, ANCOVA improves precision. In observational data it removes only the confounding due to the covariates you measured, never the ones you didn't. And adjusting for a post-treatment variable or a collider can create bias. When the question is causal, draw the DAG first, then decide what belongs in the model.
How big is the adjusted difference?
An F and a p say the adjusted difference is unlikely to be zero. Neither says how large it is. The usual effect size is partial eta squared, written ηp²: the share of outcome variation the grouping factor explains, after the covariate's share is removed from the denominator.
ηp² = SSgroup ÷ (SSgroup + SSerror)
You can also read it off the F table without any sums of squares, since ηp² = F·df₁ ÷ (F·df₁ + df₂). An ANCOVA reported as F(1, 57) = 6.84 therefore has ηp² = .11. Both SPSS and JASP print it under "Estimates of effect size", and APA writes it upright with no leading zero.
The word partial matters. The covariate's variance has already been removed from the denominator. So when randomization has balanced the covariate, ηp² comes out larger than the plain η² an unadjusted ANOVA would give on the same outcome. When the groups differed on the covariate at the start, the adjustment shrinks the numerator too, and ηp² can come out smaller. Either way the two numbers answer different questions, so don't pool them or compare them as if they were the same. State which one you are reporting and which covariates were in the model. Also give the adjusted difference in the outcome's own units, because that is the number a reader will remember.
One family: ANCOVA, ANOVA, regression
ANCOVA is multiple regression with a dummy-coded group and a continuous covariate: outcome = b₀ + b₁·group + b₂·covariate. The b₁ coefficient is the adjusted difference. ANOVA, ANCOVA, and regression are one linear model written three ways. When you have several outcome variables to adjust at once, the same idea extends to MANCOVA.
Why it matters: covariate adjustment is standard in group comparisons in psychology, medicine, and education: pre/post designs, clinical baselines, demographic controls. If you know what the adjustment does and what it can't do, you can tell when an adjusted result is misleading. For a design of your own, Plan My Analysis lays out the assumptions to check and the write-up before you collect anything. Choosing Statistics for Your Dissertation answers the form of this question that supervisors hear most: "but I need to control for age".
The adjustment is four lines of arithmetic once you have the slope. Problem 64 in the practice problems works it through on a reading program whose 6.4-point advantage turns out to be mostly a head start.
Common questions
What is the difference between ANOVA and ANCOVA?
ANOVA compares raw group means. ANCOVA first removes the part of the outcome explained by a continuous covariate (pretest score, age, baseline severity), then compares the adjusted means. This has two benefits: it corrects for baseline imbalance between groups, and it reduces noise, which often shrinks p-values even in perfectly randomized experiments.
Can ANCOVA fix pre-existing group differences in observational studies?
Only partially. ANCOVA can adjust only for variables you measured and put in the model, so any unmeasured confounder is still there. And when groups differ a lot on the covariate, the adjusted comparison extrapolates to covariate values where one group has no data (this is the core of Lord's paradox). ANCOVA works best in randomized experiments, where it makes the estimate more precise. In observational data its estimate depends on strong assumptions.
Should I analyze change scores (post minus pre) instead of using ANCOVA?
In a randomized experiment ANCOVA is the better default, and the reason is arithmetic. A change-score analysis is ANCOVA with the pretest slope fixed at 1, while ANCOVA estimates that slope from the data. Whenever the true pre-post correlation r is below 1, the estimated slope removes more noise. The standard error of the treatment effect is then √((1 + r) / 2) times the change-score version's: roughly 13% smaller when pre and post correlate .5, and 30% smaller when they hardly correlate at all. The two agree exactly only at r = 1. Outside a randomized design they can disagree about the direction of the effect as well as its precision. That disagreement is known as Lord's paradox.