The Replication Crisis
In 2015, a team of 270 researchers tried to repeat 100 published psychology experiments. Fewer than four in ten reproduced the original result. This wasn't a story about fraud. Almost every original study was run by honest scientists using accepted methods. It was a story about how ordinary, well-meaning flexibility in data analysis, combined with a publishing system that rewards surprising positive results, manufactures findings that aren't real. Understanding how is the best inoculation against doing it yourself.
Why findings evaporate
Several forces push in the same direction, toward a literature fuller of "discoveries" than the truth can support:
- p-hacking, trying analyses until one crosses p < .05: dropping "outliers," adding a covariate, testing a subgroup, collecting a few more participants and re-checking. Each step feels reasonable in isolation. These are the questionable research practices the Ethics course dissects in detail.
- The garden of forking paths. Even without conscious fishing, the many defensible choices a researcher would have made for whatever the data happened to show inflate the false-positive rate. You don't have to run every path; it's enough that the path you chose depended on the data.
- Publication bias. Journals prefer significant, novel results, so the file drawer fills with null findings that never see daylight. The published record is a biased sample of what was actually found, the same asymmetry a meta-analysis funnel plot is designed to expose.
- Low power. Underpowered studies not only miss real effects, they make the significant results they do produce more likely to be flukes and to badly overestimate the true effect size, an inflation that follows directly from the effect size and power arithmetic.
- HARKing — Hypothesizing After the Results are Known: presenting a pattern you found by exploration as though you'd predicted it all along, which turns a chance blip into a "confirmed" theory.
Play the garden of forking paths
Below is a study with no true effect — two groups drawn from the same population, so honestly there is nothing to find. But you have four ordinary analytic choices to make (which outcome measure, whether to trim outliers, whether to adjust for age, whether to focus on the "engaged" subgroup). That's 2⁴ = 16 analysis paths. Go hunting for p < .05, then read the headline showing how often pure noise hands you a "significant" result once you're free to pick the path.
🌱 The Garden of Forking Paths
Same null dataset, sixteen defensible analyses. Flip the choices, hunt for significance, and watch the false-positive rate climb.
All 16 paths for this dataset. Green marks a path that reached p < .05; click one to jump there.
The lesson lives in the gap between two numbers. Commit to one analysis before seeing the data and your false-positive rate is exactly what statistics promises — 5%. Reserve the right to choose among these four small forks after the fact, and it climbs to roughly one in four. Add more forks (a fifth outcome, a second cut-off, an interaction to probe) and it keeps rising. No single choice was dishonest; the flexibility itself was the problem.
The crisis is science working
It's tempting to read all this as "science is broken." The opposite is truer. The replication crisis was discovered, quantified, and acted on by scientists policing their own field, the self-correction that distinguishes science from dogma. The response has been concrete and constructive: preregistration to nail down the analysis in advance, bigger and better-powered samples, open data and materials, registered reports that accept a study on its design before the results exist, and a new norm of valuing a solid null over a flashy fluke. The fixes are the subject of the next lesson.
Researcher degrees of freedom. The umbrella term for every small, defensible choice you make while analyzing data: exclusions, transformations, which covariates, which of several measures, when to stop collecting. Each is individually reasonable; collectively they are the raw material of the garden. The antidote isn't to make no choices, it's to make them (and declare them) before the data can whisper which one gives you the answer you want.
Why it matters: the person your future methods must protect against is not usually a fraudster. It's ordinary, hopeful, pattern-seeking you, six months into a project you badly want to succeed. That's exactly why the safeguards are structural: preregister, power your study, share your data, and treat an exploratory finding as a hypothesis to test next time, not a result to announce now.
Problem 40 of the practice problems puts four of these practices into one short manuscript excerpt in which nothing is fabricated and every number is real, and asks you to find them.
Common questions
Is p-hacking always intentional fraud?
Usually not. Most p-hacking is motivated flexibility rather than deliberate deceit: an honest researcher who wants a project to work makes a string of individually defensible choices (dropping an 'outlier', adding a covariate, checking a subgroup) that happen to nudge the result across p < .05. This is the 'garden of forking paths': the analysis you'd have run depended on the data, so the true false-positive rate is far above 5%. It's a systemic problem to design against with preregistration, not a character flaw to accuse people of.
Does a failed replication mean the original finding was wrong?
Not on its own, and treating one failed replication as a verdict repeats the same mistake that produced the problem: reading a single study as definitive. Several things can produce a non-replication. The original may have been a false positive, which is the possibility everyone jumps to. The replication may itself be underpowered, in which case it has simply failed to detect something. The effect may be real but smaller than first reported, so a replication powered for the inflated original estimate misses it. Or the effect may depend on something that differed between the two, a population, a procedure, a moment in time, which is a finding rather than a failure and one worth chasing. What settles it is neither study alone but the accumulation: several well-powered replications, ideally preregistered, and a meta-analysis that pools them honestly rather than an argument about whose study was better.
How many psychology studies actually replicate?
In the landmark 2015 Reproducibility Project, the Open Science Collaboration repeated 100 published psychology studies and found that only about 36–39% produced a significant result in the same direction, with replication effect sizes roughly half the originals'. Rates vary by field and by how 'replication' is scored, and later large-scale projects found similar or somewhat better figures. The precise number matters less than the lesson: a sizeable fraction of published findings don't hold up, and better methods, not more accusations, are the fix.