Discussion & Limitations
The Discussion is where you finally get to say what your numbers mean — and where the most papers go off the rails. After paragraphs of careful, hedged results, the temptation is to reward yourself with one big confident claim. Resist it. A great Discussion interprets exactly as far as the evidence reaches and not one verb further; its limitations are specific, not boilerplate; and its future directions actually follow from what you found. The whole skill is calibration — matching the size of your claim to the size of your evidence.
The Discussion recipe
A clean discussion moves through five steps, roughly in this order:
- Restate the main finding in plain words: no statistics this time, just what happened: "students who slept more tended to report higher wellbeing."
- Relate it to prior work. Does it agree with, extend, or contradict what others found? This is where your Introduction's literature pays off.
- Interpret. Now you may say what it means, why it might have happened, and what mechanism could explain it. This is the one section where interpretation belongs.
- Limitations that matter: not a ritual "the sample was small," but the specific weaknesses that could actually change the conclusion, and how.
- Future directions that follow: the next study this one implies, not a generic "more research is needed."
Limitations with teeth
Every thesis lists "a limitation was the small sample size" and every grader's eyes glaze over. A real limitation says what it could have changed. Compare:
- ❌ "A limitation is that the sample was small."
- ✅ "With 38 participants the study was underpowered to detect small effects, so the non-significant interaction should not be read as evidence of no interaction."
The second names the limitation, its direction, and its consequence for a specific claim. Good limitations often point at the exact sentence in your Discussion they undercut — that's honesty, and graders reward it.
One questionnaire, two variables
There is a limitation with teeth that almost no student thesis names, and it applies to the most common student design there is: both variables measured by self-report, on the same form, at the same sitting. Anything shared by the two measures other than the constructs themselves becomes part of their correlation. Mood on the day, how the items were worded, a habit of ticking the middle box, wanting to look consistent to the researcher: these push both answers in the same direction at once. That shared push is called common-method variance.
The arithmetic is unforgiving and worth seeing once. Suppose a method factor accounts for a share m of each measure's variance, and the two traits themselves correlate ρ. The correlation you observe is (1 − m)ρ + m. Set ρ to zero and it collapses: the observed correlation is just m. A shared method accounting for a tenth of each measure's variance manufactures r = .10 between two variables with no relationship whatsoever; a quarter manufactures r = .25.
Turn that on the study in the panel below. The observed r = .28 came from one questionnaire. If a shared method carries 10% of the variance in each measure, the true correlation implied is .20. If it carries 25%, the implied truth is .04, and the finding all but evaporates. Nothing in the data can tell you which. A limitation written this way names an alternative account, gives its direction (upward, always), and puts a number on how much of the conclusion is at stake.
It is also a limitation with a fix, which makes it worth raising in the future-directions move as well: measure the two variables by different methods or at different times, and the shared push mostly goes away. A sleep diary kept nightly and a wellbeing scale completed weeks later cannot share a mood or a response style the way two blocks of one survey can. Survey & Questionnaire Design covers the item-level half of this, and Bias & Blinding the participant-level half.
Calibrate your verbs to your design
The fastest way to overclaim is a single verb. Correlational data cannot license "causes," "leads to," "boosts," or "improves" — those are causal verbs, and only a design that supports causal inference earns them (see Causal DAGs & Confounding). From an observational study you may write "was associated with," "predicted," "tended to," "was linked to." Two more calibration traps:
- Strength words. A correlation of r = .28 is not "a strong effect." Match the adjective to the number.
- Universal quantifiers. A group-level result doesn't mean "every" participant, and "all students" or "will feel better" turns an average into a guarantee.
Underclaiming is a mistake too — dismissing a genuine finding as "no meaningful relationship" wastes real information (see Writing About Non-Significant Results). Calibration cuts both ways: say exactly as much as your data support.
Catch the overclaims
Below is the result card from one study. Read it, then judge each discussion sentence: is it calibrated to this evidence, does it overclaim (says more than the data support), or does it underclaim (says less)? When you're wrong (or right), the exact word that breaks calibration lights up.
🔍 Overclaim detector
One modest, correlational finding. Rate each sentence against it.
Calibration is a verb-level discipline. "Was associated with" and "causes" describe the same data but promise wildly different things. Before every claim in your Discussion, ask: does my design earn this verb, does my effect size earn this adjective, and does my sample earn this "all"? If not, dial the word back to what the evidence supports.
Why it matters: the Discussion is the part everyone reads and remembers — and the part where one over-reaching sentence can undo a whole careful study. Graders and reviewers read it with a calibration meter running, watching whether your claims track your evidence. Interpret boldly where the design allows and hedge honestly where it doesn't, name the limitations that could actually bite, and your conclusions will outlast the ones that oversold. Say what you found — precisely that much.
Common questions
My result contradicts a published study. How do I write that?
Carefully, and without either apologizing or declaring victory. Start by checking whether the two results actually disagree, because two studies whose confidence intervals overlap substantially are telling the same story with different amounts of noise, and calling that a contradiction is a misreading. If they genuinely diverge, offer the boring explanations before the exciting one: different populations, different measures, a different manipulation strength, or one of the two studies being small enough that its estimate was never precise. Only then reach for "the effect may not be robust," and never for "their finding was wrong." The sentence that reads best is the one that says exactly what differed and what would settle it, which also gives you a real future direction instead of a generic one.
Can I use the word 'prove' in my discussion?
Almost never. A single study provides evidence that shifts our confidence; it never proves anything. 'Prove', 'proves', and 'proven' overclaim by design, and reviewers flag them instantly. Reach instead for calibrated verbs: the data 'suggest', 'support', 'are consistent with', or 'provide evidence that'. The same discipline applies to causal language: from correlational data you may write 'was associated with' or 'predicted', but not 'causes' or 'leads to' unless your design supports causal inference (see Causal DAGs & Confounding).
How do I write limitations that aren't just 'small sample size'?
Say what the limitation could have changed. Every weak limitation names a flaw and stops; every strong one names the flaw, its likely direction, and its consequence for a specific result. Instead of "the sample was small," write what that cost you: "the study was underpowered, so we can't rule out a real small effect." Instead of "the design was correlational," write "because the design was correlational, an unmeasured factor such as workload could drive both variables, so the association should not be read causally." A limitation that tells the reader exactly which sentence in your Discussion to trust less is doing its job.