Missing Data & Imputation
Real datasets have holes: participants skip questions, sensors fail, people drop out. The quick fixes (analyze whoever's left, or fill gaps with the average) seem harmless. Whether they are harmless depends on why the data are missing. Statisticians sort that "why" into three mechanisms, and which one applies decides whether your results are fine, slightly noisy, or systematically wrong.
Three ways data go missing
- MCAR: Missing Completely At Random. The holes are pure chance, unrelated to anything (a dropped test tube). Analyses of the remaining data are unbiased; you just lose precision.
- MAR: Missing At Random (a confusing name). Missingness depends on other observed variables, not on the missing value itself: for example, students with heavy course loads (observed) skip your exam more often. Complete-case analysis can be biased, but the observed data contain enough information to correct it.
- MNAR: Missing Not At Random. Missingness depends on the missing value itself: the students who'd have scored worst are the ones who didn't show. This is the dangerous one: the observed data alone can't fully fix it.
🎮 The Missingness Machine
200 students: study hours (X, always observed) vs exam score (Y, sometimes missing). Hollow orange points are the values you never saw. Switch mechanisms and watch the estimated mean (orange line) drift from the truth (teal dashed). Then turn on mean imputation and check what happens to the SD and the correlation.
Why mean imputation is a trap
Replacing every hole with the observed average looks tidy, but you've invented values with zero variability, all exactly at the mean. The imputed sample's standard deviation shrinks, and every correlation weakens, because all those identical values pull r toward 0. Worst of all, your software now believes it has a full n, so standard errors and p-values come out overconfident. Switch the widget to mean imputation and the SD and r readouts fall while the mean stays the same.
What your software calls all this
In software menus, complete-case analysis is called listwise deletion, or "exclude cases listwise". The other option in the same dialog, pairwise deletion, uses a case in every calculation it has the values for, so each cell of a correlation matrix can come from a different subset of people. It keeps more data than listwise deletion, but it has two problems. You no longer have a single n to report. And the matrix is no longer guaranteed to be internally consistent. With 35% of values missing at random, a pairwise correlation matrix comes out mathematically impossible in roughly one simulated dataset in forty: it has a negative eigenvalue, which no real correlation matrix can have. Factor analysis or regression then either refuses to run or gives meaningless results.
Neither is a modern default, but you need to recognize both, because they are what the software offers first.
What to do instead
The modern default is multiple imputation. Fill each hole several times with plausible values drawn from a model of the data, including its random spread. Analyze each completed dataset, then pool the results, so your standard errors include the extra uncertainty that comes from the missing values. Maximum-likelihood approaches (like FIML, standard in mixed models software) do the same job. Both are valid under MCAR and MAR. Under MNAR no method fixes the problem automatically, and you need a sensitivity analysis.
Why it matters: deleting incomplete cases or plugging in averages makes the strongest assumption possible, that the holes were harmless, without stating it. Name the mechanism first (MCAR, MAR, MNAR), then choose a method built for it. Even a few missing values can bias a conclusion when they are not missing completely at random.
Two habits during data collection and cleaning make all of this easier. First, record missingness so that the reason reaches your analysis software: one documented code per reason in the codebook, not a blank cell that could mean anything. Second, make the decision about who gets dropped in a cleaning script, where it can be read and re-run, and not by hand in a spreadsheet.
Problem 70 of the practice problems applies this lesson to a trial: 24 of 180 participants never came back for the final weigh-in, 18 of them from the intervention arm. The analyst simply dropped them.
Common questions
How much missing data is too much?
The mechanism matters more than the amount. 5% missing not-at-random (MNAR) can bias conclusions more than 30% missing at random handled with multiple imputation. Still, past ~10% you should report sensitivity analyses, and past ~40% the results depend heavily on the imputation model. Always report how much was missing, why you believe it went missing, and how you handled it.
How do I report missing data in a paper?
Report four things, before the results and not in a footnote after them. First, how much is missing, per variable and per group, in counts as well as percentages: "8% missing" can hide the fact that it was 14 in one arm and 2 in the other. Second, why you think it went missing, named as a mechanism and defended in a sentence, because your choice of method depends on that assumption. Third, what you did about it, including the imputation model and the number of imputations if you imputed. Fourth, whether the conclusion holds under a different reasonable assumption. A sensitivity analysis answers that. Trials also report the flow of participants through the study, and most of this belongs in the results section. Reviewers rarely object to missing data that is clearly described. They do object when they find it in a degrees-of-freedom count that does not match the stated sample size.
How many imputations should I use in multiple imputation?
The old advice of m = 5 dates from when computing was expensive. The modern rule of thumb is at least as many imputations as the percentage of incomplete cases (30% incomplete → m ≥ 30). m = 20–50 covers most studies and keeps standard errors and p-values stable across reruns. Computing is cheap now, so when in doubt, use more.