Section 1.6

Producing Data & Sampling Design

Every statistical method in this course assumes the data came from a sound design. How the data were collected limits what they can show, and no amount of careful analysis afterwards can repair a sample that was gathered the wrong way.

Population, sample, parameter, statistic

The population is the whole group you want to describe: every student at the university, every patient with the diagnosis, every trial a participant could have run. A sample is the part of it you actually measure. The number you want is a parameter, a fixed property of the population, and the number you get is a statistic, computed from your sample and different every time you repeat the study.

Those four words are the basic vocabulary of inference. Greek letters conventionally mark parameters (μ, pronounced mu, for a population mean; σ, sigma, for its standard deviation) and Latin letters mark statistics (x̄ for the sample mean, s for the sample standard deviation). You will meet that convention again in every formula in Section 1.11 onward.

Observational study or experiment?

An observational study measures people as you find them. You record who exercises and who does not, then compare their blood pressure. An experiment assigns the condition yourself: you decide, by a chance mechanism, which participants get the exercise program.

The difference determines what you may claim. In an observational study the groups differ in the thing you measured, and also in whatever led people into one group or the other. So a difference in outcome has more than one possible explanation. A variable that differs between the groups alongside the one you care about is a confounder. Random assignment rules out confounders, because it makes the groups comparable on every variable at once, including the ones you did not think of. Observation supports a claim about association; a randomized experiment supports a claim about cause.

A third case falls between the two, and it is common in real research. Sometimes a policy change, a strike, a lottery or a border determines who gets the treatment. You did not do the assigning, but whatever did the assigning had nothing to do with the outcome. Studies built on such events are called natural experiments, and they are one kind of quasi-experiment: a design that compares groups you did not randomize. They can be far more convincing than an ordinary observational study, but they are not randomized experiments, so you need to know which claims each design supports.

🎮 Four Ways to Choose 100 Students

A campus of 800 students is asked whether the university should run a late-night bus. On-campus students mostly say no; commuters mostly say yes. The true proportion is fixed and known, so you can compare each design's estimates with it.

The 800 students. Filled dots want the bus; the block on the right commutes. One sample is ringed.

What 200 repeats of this design estimate. The line marks the truth.

True support—
Average estimate—
Bias—
Standard error—

What the four designs are doing

A simple random sample gives every group of n students the same chance of being the one you get. It is the reference design. The histogram shows why: the estimates center on the true 40%, so the bias readout is near zero whatever sample size you choose.

A stratified sample splits the population into groups that differ from each other, then draws a separate random sample inside each one. Here the strata are the residents and the commuters, sampled in proportion to their sizes. The estimates still center on the truth, and they cluster more tightly: at n = 100 the standard error falls from about 4.6% to about 3.9%, a 15% reduction with no extra participants. Stratifying improves precision when the groups differ on the thing you are measuring, and not otherwise. To choose the strata, you need to know something about the population first. Multistage designs extend the same idea, sampling schools and then classes and then pupils, because a national list of every pupil rarely exists.

You will meet the other two designs often in real studies. A convenience sample takes whoever is easy to reach. On this campus that means the café at lunchtime, and the café is full of residents. The histogram moves down to 25% and stays there. A voluntary response sample lets people put themselves forward. The students who want the bus are the ones motivated to click the link, so the estimate rises to about 67%.

A larger sample does not fix a biased design. Push the slider right and look at the two flawed designs. Their spread shrinks like everyone else's, so the estimate becomes more precise but stays wrong. Sample size reduces variability; only the design fixes bias.

Randomized comparative experiments

The same logic applies when you are assigning conditions rather than choosing people. A randomized comparative experiment takes the participants you have, splits them by a chance device, and gives each group a different treatment. Comparison rules out anything that would have happened anyway, such as recovery with time or practice on the task. Randomization rules out the systematic differences between the people in each group. A control group with no treatment, or a placebo, provides the baseline for the comparison.

Blocking is stratification applied to assignment: sort participants into similar blocks first, then randomize inside each block, so a known source of variation is spread evenly across the conditions. A matched pairs design takes blocking as far as it goes: each block is a pair of similar participants, or one person measured under both conditions. The paired t-test in Section 1.13 analyzes this design.

Four ways a survey goes wrong

Undercoverage leaves part of the population off the list you sample from, so those people could never have been chosen. A phone survey misses households without a phone, and the campus café misses commuters entirely. Non-response happens after selection: you picked a proper random sample and some of them declined. If the people who decline differ from the people who respond, your random sample has effectively become a voluntary one.

Response bias is about the answers rather than the people. Participants understate their drinking, overstate their voting, and adjust what they say to suit the interviewer in front of them. Question wording is the easiest to overlook, because a leading phrase can change answers without anyone noticing. Asking whether the university should "finally do something about late-night safety" pushes people toward yes, so the answers reflect the wording as much as the opinion.

None of these show up in a standard error. A confidence interval reports how much your estimate would vary across repeated samples, and it is computed as though the design were sound. The interval around that 67% is computed correctly, but the 67% itself is 27 points from the truth.

Going deeper

This lesson is the course-level treatment. The Research Toolkit covers the same ground in more depth, one design decision at a time. Sampling Methods covers probability and non-probability sampling frames and how to report a sampling strategy. Observational Designs separates cohort, case-control and cross-sectional studies. Experiments & Random Assignment goes further into allocation, blocking and blinding. Reliability & Validity separates two questions people often confuse: whether a measure agrees with itself when you repeat it, and whether it measures the thing you actually named. A bathroom scale three kilos out is perfectly reliable and not valid at all. Quasi-Experiments extends the middle ground above into interrupted time series, regression discontinuity and difference-in-differences.

Why it matters: the rest of Stats 1 assumes a random sample from the population you care about. Sampling distributions describe how a statistic varies across those repeats, and that description is only true if the sampling was random in the first place.

Common questions

What is the difference between an observational study and an experiment?

In an observational study you record the condition people are already in; in an experiment you assign it yourself, using a chance mechanism. That difference determines what you may conclude. People who chose to exercise differ from people who did not in income, health and many things that were never measured, so a blood-pressure gap has several possible explanations. Random assignment makes the groups comparable on every variable at once, including the ones you never thought of. So only an experiment supports a claim about cause.

Is a bigger sample always better?

Only for precision. Sample size controls how much an estimate varies from study to study, and it does nothing about a biased design. However large it grows, a convenience sample of 5,000 estimates only the population it can reach, more and more precisely. The famous case is the 1936 Literary Digest poll, which collected 2.4 million responses from car and telephone owners and called the presidential election for the wrong man. Fix the design first, then spend money on n.

How do I choose between simple random and stratified sampling?

Stratify when you know the population splits into groups that differ on what you are measuring, and when you can tell which group each person belongs to before sampling. Stratifying improves precision because it removes the between-group variation from the sampling error. If the groups turn out to have similar means, the extra effort gains you nothing, and if you cannot classify people in advance, you cannot stratify at all. A simple random sample is the safe default and needs no prior knowledge.