Designing Surveys & Questionnaires
A survey question is a measuring instrument, and like any instrument it can be miscalibrated. The wording, the answer options, even the order of items can nudge respondents toward an answer they wouldn't otherwise give. The result looks like clean data (numbers, means, percentages), but it's measuring your phrasing, not their views. Good questionnaire design is the unglamorous work of getting out of the respondent's way.
Wording that tilts the answer
Each item should ask exactly one clear thing, in neutral language, that any respondent can answer honestly. The classic wording sins:
- Leading questions signal the "right" answer: "Don't you agree that…?" nudges people to agree.
- Loaded / emotive wording smuggles in a judgment: asking whether money is being "wasted" presumes it is.
- Double-barreled questions ask two things at once ("the food and the prices"), so anyone who feels differently about the two halves can't answer.
- Double negatives and jargon make people guess what you meant, adding noise: "don't you think it shouldn't…".
Response options matter as much as wording
How you let people answer shapes the answer. For rating scales (Likert items, whose measurement level decides what you may later do with them), the evidence favors 5 to 7 points: fewer throws away real distinctions, more offers precision people can't actually use. The scale should be balanced (as many positive as negative options around a neutral middle), because an all-positive scale (Good / Very good / Excellent) forces even a critic to compliment you. Whether to include a midpoint is a genuine choice: an odd number of points offers a neutral option (honest, but a refuge for the lazy); an even number forces a lean. And label the points with words, not just numbers: people interpret "4" very differently, but everyone reads "Agree" the same way.
Fix this survey
Below is a genuinely terrible six-item questionnaire. Each item has one dominant flaw. Diagnose each one, pick the problem and see the repaired version, and get all six.
🔧 Diagnose & Repair the Questionnaire
A student union survey, riddled with problems. For each item, click the flaw that dominates. Get it right and the fixed version appears; get all six and the clean survey assembles itself.
Flaws found: 0 / 6
✔ The repaired survey
The subtler biases
Even clean items can be undermined by the survey as a whole:
- Order effects: an earlier question can prime a later one. Ask about a specific worry, then overall happiness, and happiness drops — the worry is now top of mind. Put general questions before specific ones, and randomize order where you can.
- Acquiescence bias: some people just agree with whatever you say. The defense is reverse-coded items (a few worded in the opposite direction), so straight-line "agree" responding contradicts itself and can be caught. (You then flip those items back before scoring; this ties directly to a scale's internal consistency.)
- Social desirability bias: on sensitive topics — cheating, drinking, prejudice — people answer to look good rather than truthfully. Guaranteeing (and clearly stating) anonymity, softening the framing, and never attaching identifiers to sensitive items all reduce it.
Always pilot. Before a questionnaire goes live, have a handful of people from the target group fill it in and talk you through their thinking (a "cognitive interview"). You'll discover that a question you thought was crystal clear is read three different ways — cheap to fix now, impossible to fix after the data is collected. Every item is really an operationalization of a construct, and piloting is how you check it measures what you meant. It is also worth designing with the end in mind: the reverse-coded items, the "other" boxes and the optional questions you add here are precisely what someone has to untangle later, which the clean-survey-data guide walks through on a real export.
Why it matters: a biased question produces biased data no statistic can un-bias afterward: garbage in, garbage out. The most sophisticated analysis in the world can't rescue a number that was measuring the researcher's phrasing all along. Time spent sharpening your questions is the highest-leverage hour in the whole study.
Problem 39 of the practice problems is a council survey that collected 4,300 responses and still could not answer its own question, because everything that mattered went wrong before the first answer arrived.
Common questions
Should I make every question required?
Forcing an answer does not create information, it relocates the problem. A respondent who cannot answer honestly will either pick something arbitrary, which becomes noise you cannot distinguish from data, or abandon the survey entirely, which turns an item-level gap into a whole missing person and biases the sample toward people who did not mind the question. Both are worse than a blank. Make items required only where an answer is genuinely necessary and the question is genuinely answerable by everyone you are asking, give the others an explicit escape (a 'prefer not to say' or 'not applicable' option rather than a silent skip, so you can tell refusal from oversight), and record which is which in your codebook. What you are really deciding here is the missingness mechanism you will have to defend later, and a forced-choice survey with heavy dropout has chosen the worst one without meaning to.
What is a double-barreled question?
A double-barreled question asks about two things in a single item, so one answer can't honestly cover both — for example, 'How satisfied are you with the pay and the working hours?' Someone happy with the hours but not the pay has no valid response, and you can't tell which half their answer refers to. The fix is simple: split it into two separate questions, one per idea. The same rule applies whenever an item smuggles in an 'and' or an 'or' that respondents might answer differently.
What is social desirability bias, and how do you reduce it?
Social desirability bias is the tendency to answer sensitive questions in a way that makes one look good rather than truthfully — under-reporting cheating or drinking, over-reporting voting or exercise. It contaminates exactly the topics researchers most want honest data on. You reduce it by guaranteeing and clearly stating anonymity, never attaching identifiers to sensitive items, softening the framing so the undesirable answer feels acceptable, and using indirect techniques (like list experiments) for the most delicate questions.