Fraud & How Science Self-Corrects
At the far end of the line that began with honest flexibility and ran through questionable research practices sits the thing everyone fears and few commit: outright fraud, or making the data up. It's rare, but it does real damage, and the reassuring part of the story is what happens next. Science has built a growing toolkit for catching it. This lesson covers what fraud is, how one of its most famous cases unraveled, how the record gets corrected, and, hands-on, one of the arithmetic checks that can flag a fabricated result from the published numbers alone.
The two Fs: fabrication and falsification
Research misconduct has a narrow, serious definition, usually three items: fabrication (inventing data or results that were never collected), falsification (altering real data: changing values, deleting inconvenient cases, doctoring an image), and plagiarism (covered in its own lesson). The first two are about the numbers themselves: reporting findings that don't correspond to what actually happened. This is categorically different from the QRPs of the last lessons, which bend the analysis of real data; fraud reports data that never existed. It's the difference between arguing your case dishonestly and forging the evidence.
Anatomy of a fraud: the Stapel case
Diederik Stapel was a celebrated Dutch social psychologist with dozens of clean, elegant, widely cited papers. Too clean, it turned out: for years he simply invented his datasets, handing collaborators and students spreadsheets of data that no participant had ever produced. The tells were there: effects that were suspiciously large and tidy, studies whose "data" arrived without the messiness of real collection. In the end it was junior colleagues and students who raised the alarm, not an automated system. Investigations led to more than fifty retractions. Two lessons stand out: fabricated data often looks better than real data, because a fabricator produces the result they expect rather than the noise reality delivers; and the people best placed to notice are insiders, which is exactly why a culture where students can safely speak up matters so much.
Correcting the record: retractions
When a published result can't be trusted, whether through fraud, a serious honest error, or unreliable data, the formal fix is a retraction: the paper is withdrawn from the scientific record. It usually stays visible online (stamped "RETRACTED", linked to a notice explaining why) so the correction is transparent, but it should no longer be cited as evidence. Crucially, retraction isn't a synonym for fraud: plenty of retractions are honest mistakes, and a good notice says which. Services like Retraction Watch catalog them, and the annual count has climbed steadily, a sign that detection and post-publication scrutiny are improving, not necessarily that misconduct is rising.
Whistleblowing and its costs
Because insiders catch most fraud, self-correction depends on people willing to raise concerns, and doing so is rarely cost-free. Whistleblowers, often junior, risk their relationships, their references, and sometimes their careers to report someone more powerful, frequently over months of uncertainty before anything is confirmed. That the Stapel case broke at all is a credit to students who took that risk. It's why institutions increasingly build protected channels for raising integrity concerns: the health of the whole system rests on making honesty the safe choice, not the brave one.
The detection toolkit: where the hope is
The uplifting turn in this story is that fabricated numbers increasingly leave fingerprints, and anyone can learn to look. A family of consistency checks reads the published statistics against the laws of arithmetic:
- GRIM checks whether a reported mean of whole-number data is even possible for its sample size (you'll try it below).
- statcheck recomputes p-values from the reported test statistic and degrees of freedom, flagging results that don't add up. The APA results formatter runs exactly this check on anything you paste into it, and warns when the p you reported and the p implied by your statistic and degrees of freedom disagree.
- Data forensics hunts for too-even digit patterns, duplicated rows, or variances that are implausibly small for real measurement.
None of these proves fraud on its own; an inconsistency can be a typo or a rounding slip. But they narrow the search, and more importantly, they mean fabricated data is a riskier bet than it used to be.
Try the GRIM test
The arithmetic is short. A mean is a total divided by a head-count. If every response is a whole number (a count, or a single Likert item) then the total is a whole number too, so the only means that can physically occur are fractions with denominator n. Report a mean that falls in the gap between two of those reachable values, and it cannot have come from that many integer responses. Enter a reported mean and its sample size to see which values are actually reachable.
🔢 The GRIM test
For a mean of whole-number responses (a count, or one Likert item): is that mean even arithmetically reachable?
Arithmetic doesn't lie. GRIM can't tell you why a mean is impossible, since a typo and a fabrication look identical from the outside, but it can tell you a reported number could not have come from the sample it claims. That's the spirit of the whole detection toolkit: not to accuse, but to check, cheaply and publicly, whether the numbers are internally consistent. A field where anyone can run that check is a field that's harder to fool.
Why it matters: the headline is not "science is full of frauds." It's the opposite. Fabrication is rare, most researchers are honest, and the system increasingly catches what fraud there is. What keeps it that way isn't suspicion but structure: transparency, open data, protected whistleblowing, and consistency checks that make dishonest numbers a losing gamble. The most useful takeaway for your own work is simpler still — report real numbers, share your data, and you never have to worry what a GRIM test would say about your results.
Common questions
I think I have found an error in a published paper. What should I do?
Assume a mistake before you assume misconduct, because the overwhelming majority are mistakes, and proceed in a way you would be comfortable defending if you turn out to be the one who is wrong. Write it down precisely first: which number, on which page, why it cannot be right, and what you did to check. Vague suspicion helps nobody and ages badly. Then contact the corresponding author directly and neutrally, asking whether you have misread something, and give them a real chance to answer, since typos and mislabeled tables are the usual explanation and authors generally want them fixed. If there is no response or the answer does not hold up, the journal editor is next, and PubPeer exists for public post-publication comment. Keep it about the arithmetic rather than the person throughout, and if you are a student, talk to a supervisor you trust before you send anything, both for advice and because raising concerns upward is where the personal cost usually lands.
What happens when a paper is retracted?
A retraction is the formal withdrawal of a published paper from the scientific record because it can no longer be trusted — whether through fraud, a serious honest error, or unreliable data. The article usually stays online so the record is transparent, but it's stamped 'RETRACTED', linked to a notice explaining why, and should no longer be cited as valid evidence. Retraction is not the same as an accusation of fraud: many retractions are for honest mistakes, and the notice ideally says which. Databases like Retraction Watch track them, and the number has grown as detection tools and post-publication review have improved, which reflects a system catching more problems, not necessarily more misconduct happening. If you find a key source has been retracted, don't build on it.
How common is research fraud?
Outright fabrication or falsification is rare. In anonymous surveys pooled by Fanelli (2009), about 2% of scientists admitted to having fabricated, falsified, or modified data at least once, and higher fractions said they had seen colleagues do it. The far more common problem is the gray zone of questionable research practices (selective reporting, optional stopping, HARKing), which many more researchers admit to and which do more cumulative damage to the literature than the rare fraud case. The reassuring framing is that fraud is uncommon, that most researchers are honest, and that the fixes are structural: better incentives, transparency, open data, and detection tools mean fabricated results are increasingly likely to be caught.