Generalized Linear Models
Ordinary regression assumes a continuous outcome with normal, constant-variance errors. Counts, yes/no outcomes, and rates break those rules. Generalized linear models (GLMs) extend one framework to cover all of them, and you have already met two members of the family.
Two ingredients
Every GLM is built from the same two pieces:
- A linear predictor: the familiar straight line,
η = b₀ + b₁x. - A link function that connects that line to the mean of the outcome,
g(μ) = η, or equivalentlyμ = g⁻¹(η).
The link keeps the predicted mean in a sensible range: a straight line on the "link scale" becomes whatever shape the data need on the response scale.
🎮 One Line, Three Shapes
Pick an outcome type. The model is always a straight line on the link scale, but the link bends it into the right shape (and keeps it in range) for the data.
The family you already know
- Continuous outcome → identity link → ordinary regression. The link does nothing; the line stays a line.
- Yes/no outcome → logit link → logistic regression. The line becomes an S-curve that stays between 0 and 1, and the slope is a log-odds (so eb₁ is an odds ratio).
- Count outcome → log link → Poisson regression. The line becomes an exponential curve that can't go negative, and eb₁ is a multiplicative rate ratio.
The unifying idea: coefficients always act linearly, but on the link scale, not the response scale. So interpreting them means undoing the link: exponentiate for logistic and Poisson, read directly for ordinary regression. Every GLM follows the same pattern.
Fitting and assumptions
GLMs are fit by maximum likelihood rather than least squares, and each family has its own mean–variance relationship. For Poisson, the variance equals the mean, so watch for overdispersion, when real counts vary more than that. The modeling workflow (pick predictors, fit, check, interpret) is the same one you have been practicing all along.
Counts are not rates
Poisson regression models a count, and a count means something only once you know how long you watched. Two clinics report falls: clinic A logs 40 over 500 patient-days, clinic B logs 60 over 1,500. Fit the raw counts and the model reports that B has 1.5 times A's rate, because 60/40 = 1.5. Per patient-day, A's rate is 0.08 and B's is 0.04, so B has half A's rate. The same data give the opposite finding.
The fix is an offset: a predictor entered into the model with its coefficient fixed at 1. Use log(patient-days) as the offset and the log link does the rest, since log(count) − log(days) is log(rate). The model still predicts counts, but the coefficients now describe rates. Fitted on the two clinics above, the offset version gives b₁ = log(0.5) = −0.693 and a rate ratio of exactly 0.50, while the same model without it gives 0.405 and 1.50. In R it is one argument, offset = log(days); in SPSS it is the Offset variable box in Generalized Linear Models. Without the offset, a count model compares counts when the question is about rates. Both offset and deviance are defined in the glossary.
Unequal exposure is also a common hidden cause of the overdispersion mentioned above, because counts collected over very different time windows vary more than a single Poisson rate allows.
The standard check for overdispersion uses the deviance, the quantity a GLM minimizes in place of the sum of squared residuals. It measures how far your fitted model is from a hypothetical one that reproduces every observation exactly, and for a well-specified Poisson model it should be close to its degrees of freedom. Divide one by the other: a deviance/df ratio near 1 is what equidispersion looks like, and a ratio well above 1 says the counts are more variable than Poisson allows. The 412.6 on 178 df in the practice problem gives 2.32, high enough to call for a negative binomial model. Even for counts that follow a Poisson distribution exactly, the ratio is noisy in small samples. For 32 counts drawn from a true Poisson, like those in the widget above, simulated fits average 1.07, and 90% of them fall between 0.67 and 1.55. So a small excess means little, while a ratio of 2 is a real finding.
Why it matters: many real outcomes are not continuous. GLMs let you use the whole regression toolkit (multiple predictors, interactions, categorical variables) on binary, count, and rate data, in one framework. They are among the most widely used models in applied statistics.
Problem 78 of the practice problems is a family-and-link drill on five hospital outcomes (counts with many zeros, a yes/no readmission, right-skewed costs in euros, a score out of 20, and blood pressure). It then asks you to interpret one fitted Poisson model on the rate scale.
Common questions
What is a link function in simple terms?
It connects a straight line to an outcome that can't follow one. The linear predictor b₀ + b₁x can take any value, but a probability must stay in (0, 1) and a count rate must stay positive. The link transforms the outcome's mean onto an unlimited scale where a straight line fits (logit for probabilities, log for counts). The linear part is the same in every GLM, and only the link changes.
What is overdispersion and how do I handle it?
Poisson regression assumes variance = mean, and real counts are almost always more variable, with more zeros and longer tails, because events cluster within people, days and sites. The symptoms are a deviance far above its degrees of freedom and standard errors that are too small. The standard fixes are a quasi-Poisson model, which scales up the standard errors, or, more commonly, a negative binomial model with its own dispersion parameter.
My people were followed for different lengths of time. Can I just divide and model the rate?
Dividing first throws away information the model needs. A count of 2 events in 2 months and 20 in 20 months are both a rate of 1 per month, but the second carries ten times the information. A Gaussian model on the ratio treats them as equally certain, allows negative predictions, and has a variance that grows as exposure shrinks. Keep the count as the outcome and add log(exposure) as an offset: a predictor whose coefficient is fixed at 1 rather than estimated. The coefficients then read as rate ratios, and the model still uses the information about who was followed longest. In R it is offset = log(months); in SPSS it is the Offset variable box in Generalized Linear Models.