Cross-Validation & Overfitting
A model's job is to predict data you don't have yet, not merely to fit the data you have. Give a model enough flexibility and it will "explain" every random wiggle in your sample. It then scores very well on the training data and badly on the next dataset. That failure is overfitting, and cross-validation is the standard way to detect it.
The trap: fitting the noise
Every dataset is signal plus noise. A too-simple model misses part of the signal (underfitting); a too-flexible one chases the noise (overfitting). Fit statistics computed on the training data, such as R² or mean squared error, always improve as the model gets more flexible, so on its own training data the wiggliest model always looks best. The training data alone cannot show you that a model is overfitting.
Held-out data tells the truth
The fix is to evaluate the model on data that was not used to fit it. k-fold cross-validation does this without permanently setting any data aside. Split the data into k chunks ("folds"), fit the model on k−1 of them, and measure prediction error on the fold left out. Repeat until every fold has been the test set once, then average the k error estimates. That average estimates how well the model will predict new data. With k = 5 this is the 5-fold CV used below.
🎮 The Overfitting Playground
The dots are a noisy sample from the dashed teal curve (the truth). Raise the polynomial degree and watch the orange fit: training error (teal line, bottom panel) only ever falls, but 5-fold CV error (orange line) turns back up the moment the model starts memorizing noise.
As you raise the degree, training error keeps falling. CV error traces a U: it falls while the extra flexibility captures real signal, then rises once the model starts fitting noise. The degree at the bottom of the U is the complexity the data can actually support. Press "New sample" a few times and the best degree changes a little from sample to sample, because the choice of model is subject to sampling error too.
The mistake that puts the optimism back
A cross-validation error is an unbiased estimate only if you did not use it to choose anything. Suppose you use the CV error to pick the polynomial degree above. The winner was picked partly because it was lucky on those particular folds, so its CV error is the best of several and overstates how well the chosen model will perform. The same applies to anything else tuned by looking at the folds: a regularization penalty, a variable-selection step, even the decision of which outcome to model. There are two standard fixes. One is to hold out a final test set and use it exactly once, at the very end, after every choice is fixed. The other is nested cross-validation, where an inner loop does the tuning and an outer loop measures a model whose tuning it never saw. The naive estimate is usually only slightly too optimistic when one parameter was tuned, and far too optimistic when several were.
The bias–variance trade-off, in one sentence
Simple models are wrong in a stable way (high bias, low variance). Flexible models are right on average but change a lot from sample to sample (low bias, high variance). Prediction error is the sum of both. As flexibility grows, bias falls and variance rises, so the CV curve falls and then rises, and "more flexible" does not mean "better."
Where you'll meet this idea again
Cross-validation is used to choose predictors in variable selection, to compare models alongside AIC/BIC from model comparison, and to tune almost every machine-learning method. The ML course builds on it in train/test splits and generalization. The bootstrap uses a related idea to measure uncertainty: it judges a procedure by how it behaves across repeated resamples of the data.
Why it matters: fit on the training data makes every model look better than it is, so it cannot decide between models. Evaluating on held-out data, with a test set or k-fold cross-validation, is the most reliable way to build models that still predict well on new data.
Problem 80 of the practice problems puts this side by side with the information criteria: four nested models, and the question of whether the one that fits best is the one to report.
Common questions
What value of k should I use for k-fold cross-validation?
k = 5 or k = 10 is the standard, well-studied choice. Each fold's training set is nearly the full data, which keeps bias low, and you avoid the variance and computing cost of leave-one-out (k = n). With a small dataset, use k = 10 or repeated CV (several random fold splits, averaged) to make the estimate more stable. There's rarely a reason to deviate.
What is the difference between a validation set and a test set?
The validation set is used during modeling (comparing candidate models, tuning complexity), so decisions get optimized against it, and its error estimate becomes optimistic. The test set is opened exactly once, after all decisions are final, to measure performance on unseen cases. Cross-validation usually takes the place of the validation set, but an untouched test set is still best practice.
Can I use cross-validation on time-series data?
Not the ordinary kind. Random folds put future observations in the training set and past ones in the test set. The model then learns from the period it is supposed to be predicting, and the score comes out far too good. Use a scheme that keeps time in order. Rolling-origin (also called forward-chaining) validation trains on everything up to a cut-off and tests on the stretch that follows. It then moves the cut-off forward and repeats, so the model is always tested on later data than it was trained on. If neighboring observations are correlated, also leave a gap between the training window and the test window: an observation from an hour before the test period leaks nearly as much as one inside it. The same care applies to any data whose structure random splitting would break, including repeated measurements on the same person: split by person, not by row.