Model Comparison
You can almost always make a model fit your data better by adding more predictors or more flexibility. But fitting the data you have is not the same as capturing the real pattern. If you aim only for the first, you get overfitting: a model that fits the noise in your sample and predicts new data badly.
Why R² can't pick the winner
Adding any predictor (even pure random noise) never decreases R². So if you choose models by R² alone, you'll always pick the most complicated one. You need a criterion that rewards good fit but penalizes unnecessary complexity. The widget below shows a model overfitting as you add complexity.
🎮 Overfitting in Action
Fit a polynomial of rising degree to the blue training points. As the degree rises, the curve bends to pass near every training point, while its error on the orange test points (new data it was not fitted to) gets worse.
Training R² rises steadily toward 1.0 as the curve passes closer to every training point. But test error first falls, reaches its minimum near the true complexity, and then rises again as the model starts fitting noise. That U-shape is the bias–variance trade-off: a model that is too simple underfits, one that is too complex overfits, and the best model is in between.
Tools for comparing models
- Adjusted R²: like R², but with a penalty for each predictor added, so it can go down when a new predictor adds little.
- AIC & BIC: information criteria that score fit minus a complexity penalty (BIC penalizes harder). Lower is better, and they let you compare non-nested models.
- Nested F-test: for models where one is a special case of the other, tests whether the extra terms significantly improve the fit.
- Cross-validation: the most direct check: measure error on held-out data, like the orange test points here.
Prefer the simpler model. When two models explain the data about equally well, choose the simpler one. It's easier to interpret, less likely to be overfit, and more likely to hold up on new data. In prediction, Occam's razor has a measurable advantage, not only a philosophical one.
Why it matters: you rarely fit a model just to describe the sample you collected. You want it to work for cases you haven't seen, and model comparison chooses models with that goal. The same idea connects classical statistics to machine learning, a field built largely around it.
Problem 80 of the practice problems gives only the log-likelihoods and parameter counts of four nested models. You compute AIC and BIC yourself, weigh the best-fitting model against the simplest one, and say which you would report.
Common questions
What is the difference between AIC and BIC?
Both score a model as its fit minus a complexity penalty (lower is better), but BIC's penalty grows with sample size (k·ln n vs AIC's 2k). So BIC favors simpler models, especially in large samples. AIC aims to minimize prediction error, while BIC aims to identify the true model among the candidates. In practice, report both, and look more closely when they disagree.
Can I compare AIC across models fitted to different data?
No. AIC comparisons are valid only for models fitted to exactly the same rows. AIC is a relative score with no meaning on its own, so the likelihoods being compared must refer to the same observations. People often break this rule without noticing. One model includes a predictor with missing values, the software silently drops those cases, and the two models end up fitted to different samples. Their AICs are then not comparable, however carefully you computed them. Check that n is identical in every model you rank. If it is not, refit them all on the complete cases or handle the missing values first. The same rule applies to BIC. For the same reason, AIC cannot compare two models with different outcome variables or different transformations of the outcome.
What is a nested model?
One model is nested in another when it is a special case of it, obtained by deleting predictors (setting their coefficients to zero). For example, y ~ x₁ + x₂ is nested in y ~ x₁ + x₂ + x₃. Nesting matters because it allows an exact significance test (the nested F-test, or likelihood-ratio test) of whether the extra terms improve the fit. Models that are not nested must be compared with AIC/BIC or cross-validation.