Logistic Regression
So far the outcome y has always been a number. But often it's a yes/no: passed or failed, clicked or didn't, recovered or didn't. Fit a straight regression line to a 0/1 outcome and it predicts probabilities like −0.3 or 1.4, which are impossible. Logistic regression fixes this.
Bend the line into an S
Logistic regression does not predict y directly. It predicts the probability that y = 1, by passing the linear part through the logistic (sigmoid) function, which maps any number into the range 0–1:
P(y=1) = 1 / (1 + e−(b₀ + b₁x))
The result is a smooth S-curve: near-0 for low x, near-1 for high x, with a steep transition in between. It never goes below 0 or above 1, so the predictions are always valid probabilities.
🎮 Fit the S-Curve
Dots at the top had the outcome (y=1), and dots at the bottom didn't (y=0). Use the midpoint and steepness sliders to fit the logistic curve. The gray straight line goes below 0 and above 1, which no probability can do.
Log-odds and odds ratios
Logistic regression is a true regression model: it is linear, on the log-odds scale. The model says the log-odds of the outcome is a straight line, b₀ + b₁x. That means the coefficient b₁ is a change in log-odds per unit of x, and exponentiating it gives the odds ratio:
- Odds ratio = eb₁: how the odds of the outcome multiply for each one-unit increase in x.
- An odds ratio of 1 means x has no effect; above 1 means it raises the odds; below 1 means it lowers them.
- Both lines of that translation, the log-odds model and OR = eb, are on the formula sheet, a row apart.
Medical and epidemiological papers usually report odds ratios, which are hard to compare with the Cohen's d common in psychology. The effect-size converter translates between the two when you need to put an odds ratio next to a standardized mean difference.
Probabilities aren't linear, but log-odds are. A one-unit change in x multiplies the odds by a constant factor. Its effect on the probability depends on where you are on the curve: it is largest in the steep middle and tiny in the flat tails. So a logistic model has no single "effect on probability" number.
Is the model any good?
Ordinary regression gives you R² and an F-test. Logistic regression has neither, and software prints two substitutes under unfamiliar names. An omnibus test compares your model with an intercept-only model on the likelihood scale and reports a χ². It answers the question "does the model predict better than no predictors at all?". A pseudo-R² (Nagelkerke, Cox & Snell, McFadden) rescales the same likelihood improvement onto a 0-to-1 range so it looks like R². Treat it as a rough number for comparing models, not as variance explained, because the three versions give different values for the same model.
For individual predictors most output shows a Wald test, which is fine for modest coefficients and unreliable for large ones. When a predictor is important to your conclusions, refit the model without it and compare the two models by likelihood ratio.
From probability to decision
To classify cases, you pick a threshold (often 0.5): predict "yes" when the curve is above it and "no" below. Moving the threshold trades false positives against false negatives, and the right choice depends on which mistake is more costly (missing a disease vs. a false alarm). Logistic regression is widely used in credit scoring, medical risk models, and machine-learning classification. Choosing the threshold, and measuring the errors it produces, is the subject of classification metrics.
Why it matters: binary outcomes are everywhere, and logistic regression is the standard, interpretable tool for them. It extends naturally to multiple predictors (numeric and categorical) exactly like ordinary regression, with the same linear model and a new link function.
To fit one yourself, admissions.csv holds 160 applicants with GRE scores, GPAs and an admitted/rejected outcome, plus the coefficients to check your own against.
Problem 65 of the practice problems shows the difference between an odds ratio and a probability change. It takes the same two-unit step at two places on the curve and gets .215 at one and .017 at the other.
Common questions
How do I interpret an odds ratio?
An odds ratio is the multiplier on the odds for each one-unit increase in the predictor. OR = 1.5 means each extra unit multiplies the odds of the outcome by 1.5 (+50%). OR = 0.8 shrinks them by 20%, and OR = 1 means no effect. Keep two things in mind: "odds" are p/(1−p), not probability, and a large OR can mean a tiny absolute change when the baseline risk is low.
Why is my odds ratio in the millions, with a confidence interval to match?
Your data is almost certainly separated: some predictor (or combination of them) splits the outcome perfectly, so every case above a cut-off is a 1 and every case below is a 0. No finite coefficient fits best, because a steeper curve always fits better than the last one. The estimate keeps growing until the software stops iterating and prints wherever it stopped. In a worked fit of eight perfectly split points, the odds ratio passes 4 million by iteration 10, and by iteration 20 the algorithm breaks down entirely. Separation tells you something about the data. A predictor that good is usually a proxy for the outcome, a category with no cases in one cell, or a sample too small to have any overlap. The standard remedies are to merge the sparse categories, drop the offending predictor, or fit a penalized model (Firth logistic regression) that keeps the estimate finite.
What is the difference between odds and probability?
Probability is successes ÷ all attempts, and odds are successes ÷ failures. A 75% probability is odds of 3 (three successes per failure). The two are nearly equal for rare events (p = .01 is odds ≈ .0101) and far apart for common ones (p = .5 is odds = 1). That is why odds ratios approximate risk ratios only when the outcome is uncommon.