Writing About Non-Significant Results
Sooner or later a test comes back p = .08, and the temptation is instant: call it "a trend toward significance" and move on. Don't. A non-significant result is not a failure, not an embarrassment, and, crucially, not proof that the effect is zero. It's a specific, reportable finding, and reporting it honestly is one of the clearest signs of a careful researcher. This lesson is about what those results actually mean and how to write them without over- or under-claiming.
"Not significant" is not "no effect"
The deepest mistake is treating p > .05 as evidence that nothing is there. But absence of evidence is not evidence of absence. A high p-value can mean the effect really is near zero, or it can mean your study was too small to detect a real effect that's genuinely there (see The Logic of Hypothesis Testing). The p-value alone can't tell those two situations apart. That's why "we found no effect" is almost always the wrong sentence: what you actually found is that you couldn't rule the null out, which is a much humbler claim.
There is no "trend toward significance"
"Marginally significant," "approaching significance," "a trend toward…": these phrases try to sneak a non-significant result across a line it didn't cross. A p of .06 is not "nearly .05 significant"; significance is a yes/no threshold you chose in advance, and .06 is on the "no" side. Worse, the phrase is only ever used in one direction — nobody writes that p = .04 "only marginally" made it. Report the exact p-value and let it speak: "p = .06" is honest and informative; "a trend toward significance" is spin.
What to write instead
A good non-significant sentence carries three things the bare p-value hides:
- The effect size and its confidence interval. "d = 0.15, 95% CI [−0.29, 0.59]" tells the reader both your best guess and how much wiggle room it has. The CI is the hero here — it shows the whole range of effects your data are compatible with.
- Power or precision context. Was the study big enough to catch the effect you cared about? A wide CI that spans from "meaningfully negative" to "meaningfully positive" is your signal to say the study was underpowered, not that the effect is absent.
- Honest framing. "The difference was not significant," then describe the interval. No apology, no spin, no burying it.
When you can say "no meaningful effect"
Sometimes a null result is informative — when the confidence interval is tight and hugs zero. If d = 0.05 with a 95% CI of [−0.15, 0.25], you've effectively ruled out anything beyond a small effect, and you're entitled to say the difference is negligible. This is the intuition behind equivalence testing: instead of asking "is the effect exactly zero?", you ask "can I rule out effects big enough to matter?" A precise null answers yes; a wide null answers "I don't know yet". The same p > .05 can mask completely different stories, and only the CI reveals which one you have.
And remember: a non-significant result is not unpublishable and not a wasted study. Well-designed nulls correct the record, feed meta-analyses, and stop others chasing effects that aren't there.
Testing for equivalence: TOST
The paragraph above stops one step short. "Rule out effects big enough to matter" is a test you can actually run, and it has a name: two one-sided tests, or TOST. You supply the bound, the procedure supplies the verdict.
The bound is yours to choose and has to be chosen before you look: the smallest effect size of interest, the point below which you would not care even if the effect were real. That decision is substantive rather than statistical, which is why nobody can hand you a number for it. Say d = 0.3 for the tight null above. TOST then runs two one-sided tests at once, one against "the effect is at least −0.3" and one against "the effect is at most +0.3", and reports the larger of the two p-values. Reject both and you have evidence that the effect, whatever its sign, is too small to matter.
Run it on the two studies this lesson has been comparing, both of them non-significant:
- The tight one (d = 0.05, 200 per group, 95% CI [−0.15, 0.25]). Against ±0.3 the TOST p is .006. Equivalence declared: the data are incompatible with anything you would call an effect.
- The wide one (d = 0.15, 40 per group, 95% CI [−0.29, 0.59]). Same bound, TOST p = .25. Nothing is declared. The study cannot rule out a difference worth having, and it never could.
Two identical verdicts of "not significant," two opposite answers to "so is there nothing there?" There is a shortcut worth knowing, because it also explains a number that surprises people: at α = .05, TOST succeeds exactly when the 90% confidence interval falls entirely inside your bounds. Not the 95% one. The two one-sided tests each spend 5% in a single tail, which is the 90% interval's own definition. For the tight study that interval is [−0.11, 0.21], comfortably inside ±0.3; for the wide study it runs [−0.22, 0.52] and spills straight past the upper bound.
Neither test can be run after the fact on a bound picked to make the numbers work. Set the bound when you set the hypothesis, and say in the Methods that you did (Preregistration & Open Science is where that promise gets recorded).
Pick the honest sentence
Five studies, all non-significant, but they don't all mean the same thing. Read each result card (the p-value, the effect size, and especially the confidence interval), then choose the one sentence that reports it honestly. The others overclaim, dismiss, or reach for "a trend."
⚖️ Which sentence is honest?
Same "not significant," five different stories. Let the interval decide.
The confidence interval is the whole message. A wide interval that straddles zero means "underpowered, can't tell," not "no effect." A tight interval hugging zero means "we can rule out anything that matters." Two identical p > .05 results can demand opposite sentences; only the interval tells you which.
Why it matters: how you write a null result is a character test that graders and reviewers watch for. Dress a p = .08 up as "a trend" and you've told the reader you'll bend language to get the result you wanted. Report it straight — effect size, interval, honest framing — and you've shown you can be trusted with the significant results too. Nulls aren't the absence of a finding; they're a finding that asks for precise, humble language.
Problem 43 of the practice problems hides this lesson inside an APA formatting exercise: three of the four things wrong with its sentence are formatting, and the fourth is the phrase “were more accurate” attached to p = .062.
Common questions
How do I report a non-significant result in APA style?
Report it exactly like a significant one, and never hide or soften it. Give the test statistic, degrees of freedom, and the exact p-value, then the effect size and its confidence interval: "The groups did not differ significantly, t(58) = 1.30, p = .20, d = 0.34, 95% CI [−0.18, 0.84]." The confidence interval is the important part — it shows the range of effects your data are compatible with. Frame it honestly ("not significant") and, if the interval is wide, note that the study was underpowered rather than claiming there is no effect.
Can I say a result was 'marginally significant' or 'a trend toward significance'?
No. Significance is a yes/no threshold you set in advance, so a p of .06 is on the 'no' side. It is not 'nearly significant', and there is no such thing as 'marginally significant'. Tellingly, the phrase is only ever used to nudge a non-significant result upward; nobody writes that p = .04 'only marginally' made it. Report the exact p-value, the effect size, and the confidence interval, and interpret them honestly. If the point estimate looks promising but the interval includes zero, say the result is inconclusive and needs a larger, ideally preregistered, study.
I want to claim there is no effect. Equivalence test or Bayes factor?
Either will do what a p-value cannot, which is give you evidence for a null rather than a failure to reject it. They ask different questions, so pick the one whose question is yours. An equivalence test (TOST) asks whether the effect is smaller than a bound you name in advance, and it answers with an ordinary p-value; the hard part is the bound, because you have to decide what counts as too small to care about before you look. A Bayes factor asks how much more probable your data are under one hypothesis than the other, and its hard part is the prior, because "the effect exists" has to be pinned down as a distribution before the ratio means anything. Both are in JASP: the Equivalence T-Tests module offers a frequentist TOST and a Bayesian version side by side. Reviewers accept either, provided you say which bound or which prior you chose, and when you chose it.