Section 3.8

The Final Checklist

Your study is done, the sections are written, the interpretation is calibrated. One task remains, and it's the one most students skip: the consistency pass. Graders and reviewers rarely re-run your analysis. They can't. What they can do, in seconds, is check whether your own document agrees with itself. Does the sample size in the abstract match the method? Do the degrees of freedom fit the design? Is every figure referenced, every citation listed? An internal contradiction is a red flag that something deeper may be wrong, and it's the easiest kind of error to catch and to lose marks for.

Do the numbers match everywhere?

A single result appears in several places: the abstract, the results sentence, maybe a table, maybe the discussion. Every one of those copies must be identical. An abstract that says r = .28 while the Results report r = .31 tells the reader that at least one number was typed from memory, and now they doubt all of them. The same goes for sample sizes, means, and percentages. Pick the true value once, then make every mention agree.

Your degrees of freedom tell on you

Degrees of freedom are a quiet audit of your design. For an independent two-group t-test, df = N − 2; for a one-way ANOVA, the denominator df = Nk; for a correlation, df = N − 2. So if you report two groups of 60 but write t(116), the arithmetic doesn't close: 60 + 60 − 2 = 118, not 116. A grader who knows the formula reads your df straight back to your n, and a mismatch signals either a typo or a misunderstanding. Check that every reported df is consistent with the sample size and design you described.

The cross-reference checks

Three quick sweeps catch most of the rest:

  • Every table and figure is referenced in the text ("as Figure 1 shows…"), and every one you reference actually exists. An unreferenced figure is either clutter or a broken promise.
  • Every citation is in the reference list, and every reference is cited. A name in the text that's missing from the list (or vice versa) is the classic reference-manager slip.
  • Every statistic is complete: test, degrees of freedom, the statistic's value, an exact p, and an effect size, with p floored at p < .001, never p = .000 (see Reporting Statistics in APA Style).

Two of those sweeps are automated now

The p-value check and the mean-plausibility check are mechanical enough that software does them, and increasingly software does them to you before a human reads a word. statcheck reads a manuscript, pulls out every result written in APA form, recomputes p from the test statistic and its degrees of freedom, and flags the pairs that disagree. Several psychology journals, Psychological Science among them, now run it as part of peer review. The survey that prompted it looked at more than 250,000 p-values across roughly 16,700 articles: 49.6% of articles reporting null-hypothesis tests contained at least one inconsistent p, and 12.9% contained one large enough to flip the significance decision. Just under a tenth of all reported p-values did not match their own statistic.

You can run that check on yourself without installing anything. Paste a result into the APA results formatter; it recomputes p the same way and says so when the value you typed and the value your statistic implies disagree.

GRIM does the equivalent job for means. When a measure is a whole number for each person (items recalled, a sum of 1-to-7 ratings), the mean of n people has to be some whole number divided by n, which leaves most decimals unreachable. At n = 20 the possible means step by 1/20, so a reported 4.32 could not have come from that sample: 4.32 × 20 = 86.4, and nobody scored four tenths of an item. The test only bites when the sample is small relative to the decimals reported. The mock paper below gives recall means of 28.4 and 24.1 from groups of 60, and both are perfectly reachable (1,704 and 1,446 items), so GRIM has nothing to say about them. That is the ordinary outcome, worth seeing once so that a clean pass doesn't read as the check having failed.

Neither tool says anything about honesty, and both find far more typos than misconduct; fraud and self-correction covers the case where they are pointed at someone else's paper. What they do for you is finish the arithmetic half of the final pass in a couple of minutes, which leaves your attention for the half no software can take over: whether the sentence you wrote is the sentence your data will support.

Find the six inconsistencies

Here is a one-page paper in miniature — abstract, method, results, a table, a figure, and references. Six internal inconsistencies are hidden in it: numbers that disagree, degrees of freedom that don't fit, a display that's never referenced, a citation that's missing. Read like a grader and click each problem. Some things you click will turn out to be fine. That's part of the job too.

🔎 Consistency checker

Click the internal contradictions. 0 of 6 found.

Abstract

We tested whether a daytime nap improves memory. A total of 120 undergraduates were randomly assigned to a nap or no-nap group and completed a word-recall test. The nap group recalled significantly more words, and nap duration correlated with recall, r = .28. Napping may support memory consolidation.

Method

A total of 118 students (aged 18–24) took part in exchange for course credit. The study used a between-subjects design with a nap and a no-nap condition. Recall was scored as the number of word pairs correctly remembered.

Results

The nap group (n = 60) recalled more word pairs than the no-nap group (n = 60), t(116) = 2.84, p = .005, d = 0.52. Mean recall by condition is given in Table 1. Nap duration was positively correlated with recall, r = .31, p = .000. These findings echo earlier work on sleep and memory (Walker, 2017) and on retention (Smith, 2019).

Table 1. Mean recall by condition.
Nap28.4 (5.1)
No nap24.1 (5.4)
📊
Figure 1. Distribution of nap durations across the nap group.

References

Walker, M. (2017). Why we sleep. Scribner.

A document that agrees with itself earns trust. Nobody re-runs your analysis; they check whether your own numbers, degrees of freedom, and cross-references line up. Do that pass yourself, last of all — one careful read as an adversary — and you'll catch the cheap mistakes before a grader does.

Why it matters: the difference between a good grade and a great one is often not the analysis but the polish. An internal contradiction costs marks out of proportion to the effort of fixing it, because it plants doubt about everything else. Reading your own paper as a skeptical grader is the highest-return hour in the whole project. You did the hard part; don't let a mistyped df undo it. For the formatting half of the pass, keep the printable APA reporting cheat sheet beside you. And if you want to see everything this course has covered assembled into one project, the complete worked project carries a single study from its first vague question to its finished paragraph.

Common questions

A checker says my p-value is inconsistent, but my output disagrees. Which one is wrong?

Often neither, and the reason is rounding. A consistency checker only sees the statistic you printed, so it recomputes from the rounded value and compares that to the p you typed. Take t(48) = 2.01: reported to two decimals, the real statistic lies anywhere in [2.005, 2.015], which puts the true p anywhere in [.0495, .0506]. Your output might legitimately say .049 while the checker computes .050 and flags it. The flags worth acting on are the ones a rounding band cannot explain, and above all the ones that cross a decision boundary, where a reported 'significant' turns out to sit on the other side of .05. When a flag is genuine, fix the manuscript rather than the checker, and if the two disagree by more than a rounding step, recompute from your raw output before you trust either.

What do graders check first?

Internal consistency, whether your paper agrees with itself, because it is quick and revealing. Nobody re-runs your analysis; instead a grader glances across sections to see whether the numbers match (does the abstract's n equal the method's?), whether the degrees of freedom fit the design, and whether every table, figure, and citation is accounted for. A single contradiction plants doubt about everything else, so it costs marks out of proportion to the effort of fixing it. Do that same adversarial read yourself, last of all.

How do I make sure the numbers in my abstract match the rest of the paper?

Treat one place as the source of truth (usually your results output) and copy from it, never from memory. Every statistic appears in several places (abstract, results sentence, tables, discussion), and all of them must be identical: an abstract that says r = .28 while the results say r = .31 tells the reader a number was mistyped and casts doubt on the others. On your final pass, list each reported value (sample sizes, means, effect sizes, percentages) and confirm it reads the same everywhere it appears.