Reading a Causal-Inference Paper
Advanced courses tend to finish with a journal club. Somebody hands out a paper that claims a cause, everyone reads it on the train, and an hour later you are expected to say something useful about it. The skill being examined there is not running the analysis; it is refereeing a design, which has its own order of operations. This guide is that order. Seven questions, asked in sequence, then the whole route walked once on a study built for the purpose so every number can be checked, then a checklist short enough to take into the room.
1 · What is the causal question?
Start by writing the paper's question down in your own words, as a comparison between two states of the world. Not a topic, not a pair of variables, a comparison: these units, under this treatment, against the same units under its absence, over this stretch of time. A paper that estimates the effect of a tutoring program on test scores is asking what a given child's score would have been had they not attended, and the whole apparatus exists because that second number does not exist anywhere. Causal DAGs & Confounding is where the site sets out that vocabulary.
Abstracts hedge, so read the hedge. "We examine the relationship between" is a different promise from "we estimate the effect of", and papers sometimes make the first promise in the abstract and the second in the conclusion. Note which one the paper is making, because everything you do next is a test of whether it earned the stronger one.
Four things belong in your sentence, and a paper that leaves one out has left you work to do. What the treatment is, precisely enough that you could administer it. Who the units are. What they are being compared with. And over what horizon, since a policy measured after six months and a policy measured after six years are different questions with the same title.
2 · Which of the three designs is it?
Almost every causal claim from observational data in this literature rests on one of three designs, and each announces itself in a single paragraph of the methods section. An instrumental variable is a variable the authors claim shifts the treatment while having no business in the outcome equation. A regression discontinuity is a strict rule applied to a continuous score, with a cutoff. A difference-in-differences is a policy that arrived somewhere at a known moment and did not arrive somewhere else.
Find that paragraph before you read anything else, because it tells you which assumption you are going to spend the hour on. If you cannot find it, the paper may be doing plain covariate adjustment and calling the result an effect, which is a fourth and much weaker position, and the thing to ask about then is the list of controls and what is missing from it.
| Design | Where the variation comes from | The assumption no test reaches | What a paper offers instead |
|---|---|---|---|
| Instrumental variables (§3.13) | A variable that moves the treatment for reasons unconnected to the outcome | Exclusion: the instrument touches the outcome only through the treatment | The story of where the instrument came from, balance across its levels, falsification samples, an overidentification test if there are several instruments |
| Regression discontinuity (§3.14) | A strict cutoff on a continuous running variable | Continuity: everything except the treatment varies smoothly through the cutoff | Covariate jumps at the cutoff, a McCrary density test, placebo cutoffs, the estimate across a range of bandwidths |
| Difference-in-differences (§3.15) | A treatment arriving at a known date in some units and not others | Parallel trends: the comparison group's path is what the treated group would have followed | An event study over the pre-periods, placebo dates, a second comparison group, results on both scales |
3 · Where the variation comes from
Every one of these designs works by finding a slice of the treatment that arrived for a reason the outcome had nothing to do with. So locate the slice. Which units, at which moment, were treated differently, and why? The answer should be a short account of something that happened in the world: a lottery was drawn, a rule was written, a budget ran out, a border moved, a date was chosen by a legislature for reasons of its own.
Then ask the awkward follow-up. Who could have influenced that reason, and did they have a motive to? A rule written by an agency whose staff also deliver the treatment is not the same as a rule written into a statute years earlier. An eligibility score a caseworker can round up is not the same as an exam mark typed in by a machine. A policy date chosen because outcomes were already deteriorating is the worst case of all, because the choice was made using the very thing the design is trying to measure.
Also ask who is not in the comparison. Designs like these usually discard most of the sample: everyone far from the cutoff, everyone the instrument did not move, everyone in a region with no policy change. That is a legitimate price, and it is a price a reader should be told about rather than left to infer from a sample-size row.
4 · The assumption, and its defense
This is the part of the hour worth protecting. Write the identifying assumption out in one plain sentence, in the paper's own setting, with no symbols. For the tutoring study further down, it reads: living near a club site changes a child's test score only by getting them into the club, and by nothing else at all. Once it is in words rather than notation, a room full of people can argue about whether it is true, which is the entire purpose of the exercise.
The assumption is not testable. It is a statement about an outcome nobody observed, so no dataset contains the column that would settle it, and no diagnostic, goodness-of-fit measure or residual plot carries information about it. All three lessons say this in their own vocabulary, and it is the reason a causal paper is refereed rather than checked.
What an honest paper does instead is put the assumption in positions where it could visibly fail. There are four such offers, and it is worth naming which ones the paper made.
- An account of the variation good enough that a skeptic can attack the story rather than the statistics.
- Balance on things fixed before the treatment existed, which the treatment cannot have caused.
- A falsification: a sample, a period or a cutoff where the assumption implies nothing should happen.
- A sensitivity calculation saying how large a violation would have to be before the conclusion changes.
Now judge the offer, because this is the step that separates refereeing from nodding along. A falsification test that comes back insignificant has told you something only if it could have caught the violation that matters. Work out what size of effect it had the precision to detect, compare that with the size of violation that would overturn the result, and you will often find the test was never in a position to fail. The worked study below is exactly that case, and the arithmetic takes one line.
Balance tables invite the same misreading in reverse. A row with a large standard error and a small t is not a row that says zero; it is a row that says nobody knows. Read the point estimates, ask whether the imbalances line up into one story, and ask which direction that story would push the headline. An imbalance that works against the paper's finding is very different news from one that works in its favor, and papers rarely spell out which they have.
5 · Reading the tables
Results sections are laid out to be read forward and are better read backward. The headline table is the one the abstract is about; the tables that decide whether to believe it usually sit in front of it or in an appendix. Go to those first, and arrive at the coefficient last.
For an instrumental-variables paper, three things should be visible. The first stage with its F statistic, which tests relevance and nothing else. The reduced form, the plain regression of the outcome on the instrument, which is the whole finding before any dividing happens. And the ordinary least squares estimate beside the two-stage one, so a reader can see how far apart they are and in which direction. A paper that reports only the two-stage coefficient has hidden the arithmetic: a ratio can be large because its numerator was large or because its denominator was small, and those are different papers.
Balance and placebo tables are the assumption's public exposure, and they are the tables to read with a pencil. Robustness columns are a different animal: work out what is being varied across them, and then ask whether the specification closest to the identifying concern is also the weakest one. That pattern is common and it is rarely flagged, since the column the authors trust least is the column they are least inclined to discuss.
Two habits cover most of the rest. Check that the sample size is the same across columns that claim to be comparable, because a specification that silently drops a third of the data has changed more than its controls. And check whether the standard errors were clustered at the level the treatment was assigned at, which for a policy that arrived by region means the region rather than the person. Difference-in-Differences works through why that choice can move a standard error by a factor of two or three.
6 · The magnitude, and whose it is
Read the coefficient in the units of the outcome, out loud, as a sentence. Then convert it into something a person can hold: a share of the outcome's standard deviation, a fraction of a year of progress, a percentage of the baseline mean. Stars tell you about the standard error and nothing about whether the effect is worth having.
An implausibly large estimate is a diagnostic rather than a triumph. Designs like these often produce big numbers, partly because they are estimating the effect on a particular subgroup and partly because a noisy estimate that clears significance has to be large to have done so. If an after-school club appears to be worth two thirds of a standard deviation, the first thing to suspect is the design, not the club.
Then say whose effect it is, because none of these designs estimates the average effect in the population. An instrument recovers the average among the compliers, the units it actually moved. A cutoff recovers the effect for units sitting at the cutoff. A difference-in-differences recovers the effect on the treated units in the periods they were treated. The sentence you want the paper to have written is of the form "this is the effect of X on Y, among Z", and if the paper did not write it, write it yourself before the seminar and see whether anybody disagrees.
7 · What would change your mind
Before you speak, name the one finding that would flip your verdict. If you are inclined to believe the paper, name the result that would make you stop: a covariate that jumps at the cutoff, a pre-trend, a comparison group that gives a different answer. If you are inclined to doubt it, name what would win you over.
Doing this converts a summary into a discussion, and it has a useful side effect: it exposes the cases where nothing would change your mind. When that happens you are not refereeing the paper, you are reporting a prior, and the room deserves to know which one is on offer. It also gives the authors, if they are ever in the room, something they can act on, which a list of grievances does not.
The standing objections, by design
These are questions rather than accusations, and the difference in a seminar is not cosmetic. Each one has a good answer available; asking it gives the presenter the chance to give it.
Instrumental variables
- Where did the instrument come from, and who decided it?
- What is the first stage, and what does the reduced form look like on its own?
- Which direct routes from the instrument to the outcome have you ruled out, and on what evidence?
- Who are the compliers here, and roughly what share of the sample are they?
- If the instrument is a policy, is anything else in that policy also touching the outcome?
Regression discontinuity
- What does the density of the running variable do at the cutoff?
- Which pre-treatment covariates jump there, and how precisely were they measured?
- How does the estimate move across bandwidths, and does the specification force one slope across the cutoff?
- Is this threshold used by any other rule or program?
- How many distinct values does the running variable take near the cutoff?
Difference-in-differences
- How many pre-treatment periods are there, and how precisely was each one estimated?
- Why this comparison group, and what happens with a different one?
- Did units adopt at different dates, and if so which estimator was used?
- Levels or logs, and does the answer survive the other choice?
- What else happened at that date in the treated units alone?
The route walked once
The study below is invented, along with every number in it. Nothing here describes or criticizes a real paper or a real author, which is deliberate: a worked critique is only honest if the thing being critiqued cannot be misrepresented. Courses differ on how much help a student may take with an assigned reading, AI tools included, and the rules that bind you are the ones in your own handbook. Read them before you bring anything from this page to a paper you were set. Using AI Tools Ethically covers where that line usually falls and the reasoning behind it.
The paper. A city has run free after-school homework clubs for decades, sited in the 1970s under a population rule and unchanged since. Using administrative records on 6,000 children aged 8 to 11, the authors ask whether attending a club raises the end-of-year test score. Attendance is voluntary, so they instrument it with living within 800 meters of a club site. The outcome is scored out of 100 with a standard deviation of 15 in this sample.
Question 1. The comparison is a child who attended a club against the same child in a year where they did not, measured at the end of that school year. The paper says "effect of attendance", so it has made the strong promise.
Question 2. Instrumental variables. The cutoff language about 800 meters is tempting to read as a discontinuity, but no rule allocates anything at 800 meters; it is the authors' own dividing line on a continuous distance, used to define the instrument.
Question 3. The variation is the accident of where a family's address sits relative to buildings chosen half a century ago. Nobody currently involved chose it, which is the good part. Families do choose where to live, which is the part to worry about.
Question 4. The assumption in words: being near a site changes a child's score only by getting them into the club. Here is the paper's balance table, its main defense.
Table 1 · Characteristics by distance to the nearest club site
| Measured before the year began | Near (n = 2,400) | Far (n = 3,600) | Difference | SE | p |
|---|---|---|---|---|---|
| Age at the start of the year | 9.41 | 9.39 | +0.02 | 0.030 | .50 |
| Girls | .501 | .489 | +.012 | .013 | .36 |
| Previous year's test score | 58.92 | 58.58 | +0.34 | 0.40 | .39 |
| Free school meals | .287 | .312 | −.025 | .012 | .038 |
| Household income (thousands) | 33.5 | 32.4 | +1.10 | 0.37 | .003 |
Three rows pass and two fail, and the two that fail tell one story: households near a site are better off. The highlighted row is the one the paper treats as reassuring.
The paper reports that the previous year's score is balanced, t = 0.86, and moves on. Look at the number rather than the verdict. Near children score 0.34 points higher before the club year begins, and the income gap is exactly the size that would produce it: 1.10 thousand at a gradient of roughly 0.30 points per thousand gives 0.33. So the balance table's most reassuring row is quantitatively consistent with a direct path from the instrument to the outcome worth about a third of a point. Hold that number.
The paper's falsification is a good one. Older siblings in the same households, aged 13 to 15 and never eligible for the clubs, show a near-versus-far score difference of +0.21 with a standard error of 0.58, t = 0.36. If distance were a proxy for neighborhood quality, it ought to move their scores too, and it does not. The offer is real. It is also weaker than it looks: to reach significance that test needed a gap of about 1.14 points, and the violation that would matter is 0.34, which would have shown up as t = 0.59. The falsification passes and could never have failed.
Question 5. Now the tables, read in the order that decides things.
Table 2 · First stage and reduced form
| (1) Attended a club | (2) End-of-year score | |
|---|---|---|
| Near a site | 0.280 | 1.12 |
| (0.012) | (0.395) | |
| Mean of the far group | 0.170 | 62.65 |
| First-stage F | 557.6 | not applicable |
| Observations | 6,000 | 6,000 |
Column 1 is the first stage and column 2 the reduced form. Between them they contain the whole estimate, before any dividing.
Table 3 · The effect of attending a club
| Ordinary least squares | Two-stage least squares | |
|---|---|---|
| Attended a club | 9.00 | 4.00 |
| (0.430) | (1.41) | |
| 95% confidence interval | [8.16, 9.84] | [1.24, 6.76] |
| Observations | 6,000 | 6,000 |
4.00 is 1.12 divided by 0.280, which is the Wald ratio worked by hand. The ordinary estimate is more than twice as large and far more precise, and it is the one nobody should believe.
The first stage is enormous and answers the easiest of the three conditions. Nobody should be impressed by an F of 558, and a reader who treats it as a validation of the design has confused relevance with exclusion. The reduced form is the number that carries the finding, and at 1.12 points with a standard error of 0.395 it is a real but modest result, t(5998) = 2.83.
Now use the 0.34 from the balance table. If a third of a point of that 1.12 arrived by a route other than attendance, the honest numerator is 0.78, and dividing by the same first stage gives 2.79 rather than 4.00. The bias is 0.34 divided by 0.280, which is 1.21 points, and it is worth noticing that a weaker instrument would have magnified it: against a first stage of .08 the same leak would have been worth 4.25 points, more than the entire reported effect. That is the calculation to bring to the seminar.
Table 4 · Robustness
| Specification | Estimate | SE | 95% CI |
|---|---|---|---|
| (1) No controls | 4.00 | 1.41 | [1.24, 6.76] |
| (2) Household controls added | 3.12 | 1.44 | [0.30, 5.94] |
| (3) Neighborhood fixed effects added | 2.85 | 1.79 | [−0.66, 6.36] |
| (4) Near defined as 600 meters | 4.31 | 1.86 | [0.66, 7.96] |
| (5) Excluding households that moved | 3.94 | 1.49 | [1.02, 6.86] |
Every interval is the estimate plus or minus 1.96 standard errors. Column 3 is the one that speaks to the identifying worry, and it is the one whose interval covers zero.
Read the columns in the order of how much each one addresses the concern, rather than left to right. Adding household controls pulls the estimate down toward the 2.79 the sensitivity calculation predicted, which is a point in favor of the diagnosis. Neighborhood fixed effects pull it further and cost enough precision to cross zero. The paper's abstract quotes column 1.
Question 6. Four points on a test with a standard deviation of 15 is 0.27 SD, and the paper describes a year of ordinary progress on this test as about 11 points, so the club buys something like a third of a school year. That is a substantial claim but not an absurd one, which is a reasonable place for an estimate to land. The naive comparison of 9.00 points, at 0.60 SD, would have been absurd, and its absurdity is a useful signal about what selection into an optional program does.
Whose effect it is follows from the first stage. Seventeen percent of children attend wherever they live, 55% attend nowhere, and the 28% in between are the compliers: children who went because a site was close and would not have gone otherwise. The 4.00 belongs to those 1,680 children and to nobody else. A city deciding whether to build more sites is asking about precisely that group, so the number is the policy-relevant one here. A claim about what clubs do for children in general it is not.
Living within 800 meters of a club site raised attendance by 28.0 percentage points, first-stage F = 557.6, and was associated with a 1.12-point difference in end-of-year scores, t(5998) = 2.83, p = .005. The implied effect of attendance is 4.00 points, SE = 1.41, 95% CI [1.24, 6.76], or 0.27 standard deviations, among the 28% of children whose attendance the proximity of a site changed.
Question 7. Two things would change the verdict. A falsification with the power to detect a third of a point, on a sample where the assumption implies a clean zero, would make the imbalance a curiosity rather than a live threat. And a neighborhood-fixed-effects estimate that held near 4.00 rather than falling to 2.85 would say the income gradient is not doing the work. In the other direction, an estimate that moved toward zero as the distance threshold tightened would suggest neighborhood rather than club, since a tighter ring should be more like its surroundings rather than less. The verdict to offer the room is that the design is plausible, the finding is probably real, and the reported magnitude is likely too large by somewhere between a quarter and a half.
A checklist for the room
One page, whichever of the three designs the paper uses. Fill in the middle column while you read and the right-hand column is what you say.
| Q | Write this down | The tell that it is missing |
|---|---|---|
| 1 | The question as a comparison: effect of what, on whom, against what, over how long | The paper says "relationship" and never names the comparison |
| 2 | The design in one word: instrument, cutoff or policy date | The methods section names a software command instead of a design |
| 3 | Where the variation came from, and who could have influenced it | The variation is described but never argued to be as good as random |
| 4 | The identifying assumption in one plain sentence, plus which of the four defenses the paper offers | The assumption appears in symbols and never in words |
| 5 | First stage and reduced form, or balance and density, or the pre-period coefficients, with their numbers | Only the headline specification is shown |
| 6 | The estimate in outcome units, as a share of an SD, and the group it belongs to | The result is described by its stars |
| 7 | The one finding that would change your verdict | You cannot name one |
The size of the violation
For the defense that is doing the most work, write down two numbers: how big a violation would have to be to overturn the conclusion, and how big a violation that test could have detected. If the second is larger than the first, the test passed without ever being in danger.
| If it is | Ask first | Then ask |
|---|---|---|
| Instrumental variables | What is the reduced form on its own? | Which direct route to the outcome is ruled out, and how? |
| Regression discontinuity | What does the density do at the cutoff? | How does the estimate move across bandwidths? |
| Difference-in-differences | How many pre-periods, and how precise? | What happens with a different comparison group? |
The two robustness checks that would most change your verdict, written before the discussion starts, are the most useful sentence you can bring.
Keep going
- Instrumental Variables & 2SLS: relevance, independence and exclusion, the Wald ratio, weak instruments and compliers.
- Regression Discontinuity Design: the running variable, the bandwidth decision, sharp and fuzzy cutoffs, and the three checks.
- Difference-in-Differences: four cells, the interaction coefficient, clustered standard errors, event studies and staggered adoption.
- Causal DAGs & Confounding: the vocabulary of backdoor paths and colliders that every argument above is written in.
- Quasi-Experiments & Natural Experiments: the same trade made with other accidents, from the Research Toolkit.
- Practice Problems: the Stats 3 set works an encouragement design, a cutoff and a two-by-two policy table by hand.
- One Study, Start to Finish: the other side of the desk, where you are the one making the claim.