Causal DAGs & Confounding
"Correlation isn't causation" is where most stats courses stop. Causal inference asks the next question: under what assumptions would this association reflect a cause? It starts with a picture, a DAG (directed acyclic graph) of what you believe causes what, and the picture tells you which variables to adjust for. Adjusting for the wrong variable can do worse than fail to help. It can create bias where there was none.
Confounders: a common cause
A confounder Z is a common cause of both your predictor X and outcome Y (Z → X and Z → Y). It opens a "backdoor path" X ← Z → Y that mixes a spurious association into the X–Y relationship. The naive regression of Y on X then reports the causal effect plus that spurious association. The fix is the one you know from multiple regression: include Z in the model. That closes the backdoor, and the X coefficient recovers the true effect.
Colliders: a common effect
Now reverse the arrows. A collider C is a common effect: X → C ← Y. A path through a collider is blocked to begin with, so a collider biases nothing if you leave it alone. But condition on C (adjust for it, select on it, stratify by it) and you open the path, creating a spurious association between X and Y that wasn't there. So "control for everything you measured" is bad advice. Whether adjusting for a variable helps or hurts depends on the causal structure, and the DAG writes that structure down.
🎮 Adjust, or Don't: Confounder vs Collider
The graph shows the true causal structure generating the dots below (colored by the third variable: teal = low, indigo = mid, orange = high). Compare the naive slope to the adjusted one against the true effect you set. Then switch structures and watch the right move become the wrong one.
The backdoor rule of thumb
The general rule (Pearl's "backdoor criterion") comes down to this: adjust for common causes, never for common effects, and leave mediators alone unless you specifically want the direct effect (see mediation). Colliders also explain several well-known paradoxes: selection bias, "why are attractive people jerks in my dating pool," and Berkson's hospital paradox. Each of them comes from conditioning on a collider.
The alternative to adjusting at all
Everything above applies when the exposure was not under your control. Random assignment solves the same problem at the design stage, and solves it better. If a coin decides who gets X, nothing that came earlier can be a cause of X, so every backdoor path into X is closed before you collect any data. That includes paths through variables you never thought to measure. Adjustment cannot do that, because the backdoor criterion only closes paths through variables you have. When randomization is impossible, quasi-experimental designs use a naturally occurring assignment rule to get some of the same protection, and you use a DAG to argue that the rule did its job.
The hardest case is a backdoor through something no study could measure: motivation, ability, health before anyone was recruited. Adjustment cannot close that path, and the next three lessons each offer a way around it. Instrumental variables stops trying to block the confounded comparison. It uses only the part of the variation in X that comes from a source unrelated to the confounders. Regression discontinuity uses a rule. When a treatment goes to everyone above a strict cutoff on some score, the people just either side of the cutoff are alike in everything that changes smoothly with the score, measured or not. Difference-in-differences uses time. It follows a group that got the treatment and a group that did not, before and after, so anything about either group that stayed constant drops out of the comparison.
What a DAG is for
A DAG will not tell you whether your arrows are right: that's subject-matter knowledge. What it does is make your assumptions public and checkable, and turn "which covariates should I include?" from a fishing expedition into a derivation. Two researchers who agree on the graph must agree on the adjustment set; two who disagree can now argue about the actual disagreement.
A DAG also helps you read published causal claims. When a paper relies on an instrument, a cutoff or a policy date, it is telling you which backdoor paths it could not close by adjusting for covariates. Reading a Causal-Inference Paper goes through the questions a seminar asks of those three designs, with a printable checklist at the end.
Why it matters: adjust for the variables your causal model says to adjust for, and no others. Controlling for a confounder removes bias, and controlling for a collider creates it. The data alone can't tell you which is which. Only a causal model of how the data came to be can.
Problem 95 of the practice problems starts from a press release claiming that an optional online module "raises exam scores" by nine marks. It asks you to name the confounding that self-selection introduced and say what design would have answered the question properly.
Common questions
What is a collider in simple terms?
A variable caused by two others: X → C ← Y. Left alone, it transmits nothing. But select or adjust on it and you create a spurious X–Y association. The classic example: among hospitalized patients (being hospitalized is the collider), two diseases look negatively correlated even if they are independent in the population, because having either one is enough to get you admitted.
Should I control for every variable I measured?
No. Putting every measured variable into the model ("kitchen-sink regression") can create bias as easily as remove it. Adjusting for confounders (common causes) removes bias, adjusting for colliders (common effects) creates it, and adjusting for mediators removes the part of the effect that runs through them. The data alone can't tell these apart, so the covariate list must come from a causal diagram of how the data was generated.
What if I can't tell whether a variable is a confounder or a collider?
The data will not tell you. Adjusting for either one changes your estimate, and the output looks equally plausible both ways. The decision has to come from what you know about how the variables arose. The most useful practical guide is time: a variable measured before X was in place cannot be a common effect of X and Y, so it can be a confounder but not a collider. A variable measured after both X and Y may be a collider. When the order in time does not settle it, do not split the difference. Draw both graphs and report the adjustment set implied by the one you think is right. Then show the other as a sensitivity analysis, so a reader who disagrees with your arrow can see what their assumption would have produced.