Section 3.10

Mixed & Multilevel Models (GLMMs)

Data is often grouped: students within schools, patients within hospitals, repeated measurements within people. Those observations aren't independent, and treating them as independent makes your standard errors wrong. Mixed (multilevel) models model the grouping directly, and they let groups borrow strength from each other: each group's estimate uses information from the others.

Fixed effects vs. random effects

A mixed model has two kinds of terms. Fixed effects are the usual coefficients: effects assumed constant across the whole population. Random effects let each group have its own intercept (and sometimes its own slope), drawn from a distribution the model estimates. The model does not estimate each group in isolation. It estimates how much the groups vary and uses that to inform each group's estimate.

How much of the variation is the grouping?

A fitted mixed model gives two spreads: τ (tau), the standard deviation between group means, and σ (sigma), the spread of individuals within a group. They combine into the intraclass correlation, ICC = τ² / (τ² + σ²). The ICC is the share of the total variation that is between groups rather than within them. Equivalently, it is the correlation you expect between two observations from the same group.

In the playground below, σ is fixed at 11, so moving the τ slider changes the ICC. At the default τ = 8 the ICC is .35, so a third of the variation is between groups. Move τ down to 1 and the ICC falls to .008: the groups barely differ, and shrinkage pulls every estimate to the grand mean. Move τ to 30 and it reaches .88, where each group is nearly its own population. The ICC also determines how much clustering costs a survey, and sampling methods uses it to compute the cost of cluster sampling.

What ignoring the grouping costs

The assumptions poster tells you to rule out "hidden clustering" before you trust an ordinary test. That advice is easy to agree with and hard to act on. Clustering is hidden whenever the rows of your file are not the units you sampled: pupils inside the classes you recruited, trials inside the people you tested. Analyzing those rows as if each were a separate participant is called pseudoreplication, and you can compute its cost in advance.

Survey statisticians worked this out first. The design effect from sampling methods, deff ≈ 1 + (m − 1)ρ, converts your row count into an effective sample size, with m observations per cluster and an ICC of ρ. Twenty classes of twenty pupils is 400 rows, but at an ICC of .10 the design effect is 2.9, so those rows carry about as much information as 138 independent pupils. A test that treats all 400 as independent gives a standard error too small by a factor of √2.9, which is 1.7.

A simulation shows how large the problem is. With the treatment given to whole classes and no true effect at all, an ordinary two-sample t-test across the 400 pupils is significant 24.7% of the time. At an ICC of .05 it is 16.0%, and at .20 it is 37.4%. Average each class into a single number first, so that the 20 units in the test are the 20 units that were randomized, and the false-positive rate returns to 4.8%. A five-percent test that rejects a quarter of the time is the main reason cluster-randomized trials need a cluster-level or mixed-model analysis.

The direction of the error depends on where the comparison is made. If you split the treatment within every class, so each class contributes both conditions, the same ignored clustering makes the ordinary test slightly conservative. It gives 4.1% false positives and loses about two points of power (70.9% against the mixed model's 73.2% for a difference of a quarter of a standard deviation). So ignoring clustering inflates false positives in between-cluster comparisons and costs a little power in within-cluster comparisons. The same arithmetic makes a paired test more powerful than an unpaired one on the same people counted twice.

To check, fit the model with no predictors at all, read off τ and σ, and compute the ICC. An ICC much above zero means the grouping accounts for real variance and the rows are not independent. Even an ICC of .02 matters when clusters are very large, because the cost depends on m and ρ together, not on ρ alone.

The three ways to handle groups

  • Complete pooling: ignore the groups entirely; one estimate for everyone. Simple, but throws away real group differences.
  • No pooling: estimate each group on its own. Each estimate is unbiased for its own group, but tiny or noisy groups get unreliable estimates.
  • Partial pooling is the mixed-model compromise: each group's estimate is pulled toward the overall mean, more so when the group has little data or lots of noise. This is shrinkage.

🎮 Shrinkage in Action

Each row is a group. The open circle is its own raw mean (no pooling); the filled purple circle is the partial-pooled estimate. Drag the group-variation slider down and watch the estimates pulled toward the overall mean (dashed line); small groups move most.

Overall mean—
Average shrinkage toward it—
Pooling regime—

Why shrinkage is a good thing

Moving a group's estimate away from its own data can feel like cheating, but it improves accuracy on average. A group of 3 with a very high mean was probably partly lucky, so an estimate pulled toward the overall mean is usually closer to the truth. For the same reason, a rookie's hot first week does not predict their whole season. Partial pooling formalizes regression to the mean.

In the playground, when τ is large (groups differ a lot), there is little shrinkage, and each group's own data dominates. When τ is small (groups are mostly alike), the estimates move most of the way to the overall mean. At any setting, the small groups shrink more than the big ones, because they carry less independent information. A random-effects meta-analysis uses the same partial pooling with published studies as the groups, so a small study's estimate is pulled toward the pooled effect.

The "G" in GLMM

Combine this grouping machinery with a link function and you get a generalized linear mixed model (GLMM): random effects on top of logistic or Poisson regression. That covers a large share of real research: yes/no outcomes measured repeatedly within people, event counts nested within regions, and far more.

Why it matters: mixed models are now the default for grouped and repeated-measures data across psychology, ecology, medicine, and education. They model the structure of the data and make better predictions for small groups.

Problem 67 of the practice problems starts one step earlier, with a researcher deciding whether to run 90 people through one design each or 30 people through all three. That design choice decides whether you need a mixed model at all.

Common questions

What is the difference between fixed and random effects?

Fixed effects are coefficients estimated for effects you care about specifically and would keep in a replication (treatment, age, condition). Random effects model sampled clusters (these particular schools, participants, litters) as draws from a population, estimating how much clusters vary rather than each one in isolation. Litmus test: would new data bring the same levels (fixed) or new ones (random)?

Why doesn't my mixed model print p-values?

Because for a mixed model there is no agreed answer for the denominator degrees of freedom. In a balanced classical ANOVA the error df can be counted. Once you have unequal group sizes, crossed random effects and partial pooling, the effective df fall between two defensible numbers, and the F ratio is only approximately F-distributed. So R's lme4 does not print p-values, which surprises most people the first time. There are three accepted options. Report the fixed effects with their confidence intervals and base the inference on those. Use Satterthwaite or Kenward-Roger approximate df: the lmerTest package adds them, and SPSS and JASP use Satterthwaite by default, so they show p-values and lme4 does not. Or compare nested models with a likelihood-ratio test. Say in the write-up which one you used, because the three do not always agree at the margin.

How many groups do I need to fit random effects?

The usual rules of thumb: with fewer than about 5 groups, don't, because the model cannot estimate a between-group variance from 3 numbers (use fixed dummy codes instead). With 10–20 groups it works, but the variance components are estimated only roughly. Thirty or more is comfortable, and 50 or more is preferred for random slopes or when the variance components are themselves the research question.