Non-Linear Relationships & Transformations
Fitting a straight line assumes that the relationship is straight. When the relationship bends, the fitted line is wrong in a systematic way, and you can see the pattern in a residual plot. Often a single transformation fixes it, without a bigger model.
The residual plot is where curvature shows up
Fit a line to curved data and it will run above the points at the ends and below them in the middle, or the reverse. On the scatterplot that can be hard to see, because a line through a cloud of points tends to look right. On a plot of residuals against fitted values it is obvious: you get an arch or a valley, not a formless band around zero.
That plot is the first thing Section 2.8 teaches you to draw, and curvature is the violation it catches most reliably. Set the widget's model to the plain line and its truth to any of the curved shapes, then look at the lower panel rather than the upper one.
🎮 Straighten It Out
The top panel is the data with the fitted model drawn back in the original units. The bottom panel is residuals against fitted values, where curvature is obvious. Choose the wrong model and the arch appears.
Residuals against fitted values, in the original units of y
Two transformations do most of the work
Taking the log of the predictor fits diminishing returns: y climbs quickly at first and then flattens, so equal multiples of x go with equal increases in y. Income against life satisfaction behaves this way, and so does practice against performance.
Taking the log of the outcome fits constant proportional growth: each extra unit of x multiplies y by a fixed factor rather than adding a fixed amount. Populations, compound interest and untreated epidemics do this. Logging the outcome also shortens a long right tail and often makes the variance more even, so it can fix two assumption problems at once.
Adding x², a polynomial term, does not transform either variable. It adds an extra predictor, and it fits a relationship that turns around: rising then falling, or falling then rising. Anxiety against performance is the standard example, and so is any dose that helps up to a point.
The widget's R² readout is always computed on the original scale of y. For the log-outcome model, the predictions are turned back into raw units first. This matters because an R² from a model of log y is not comparable to an R² from a model of y. The two explain variation in different quantities, and software will print both without warning you. One caveat on the back-transform: exponentiating a fitted value estimates the median of y, not its mean, so it runs slightly low. This gap is called retransformation bias, and you should mention it in a write-up whenever the predictions themselves are the point.
Reading a coefficient after a transformation
Write-ups often get this wrong, because the slope no longer means "the change in y per unit of x".
With log x, the slope is the change in y per unit of log x, a unit that is hard to picture. Convert it: doubling x changes y by b₁ × ln 2 ≈ 0.693b₁. If b₁ = 8, every doubling of x adds about 5.5 points to y. A 10% rise in x adds b₁ × ln(1.10) = 0.76. That is close to b₁/10, so the shortcut is safe for small percentages.
With log y, each one-unit rise in x multiplies y by eb₁. For b₁ = 0.08 that factor is 1.083, an increase of 8.3%. The familiar shortcut, "read 100b₁ as a percentage", gives 8% here and is off by a third of a point. It works well while b₁ is under about 0.1 and gets worse quickly after that: at b₁ = 0.5 the shortcut says 50% and the true figure is 64.9%. So exponentiate, and report the factor.
With x² in the model, neither coefficient means anything on its own, because the effect of x now depends on where you are. Report the turning point, at x = −b₁/(2b₂), and the direction on each side of it. For b₁ = −6 and b₂ = 0.35 the curve reaches its minimum at x = 8.57 and rises after that. Section 2.11 has the same issue: there, too, a coefficient describes the effect at one value of another variable. Mean-centering x before squaring it helps here for the same reason it helps there.
When to transform, and when not to
Transform when the curvature has a shape you can name and a mechanism behind it. Diminishing returns and proportional growth are real mechanisms, and a model that encodes one of them will predict sensibly outside the data you happened to collect. A transformation also keeps the model at one predictor, which keeps the interpretation and the degrees of freedom simple.
Use a larger model when the shape has several bends, when the outcome is a count or a yes/no rather than a quantity, or when the relationship differs by group. A yes/no outcome needs logistic regression, not a log. Counts need a generalized linear model, and a curve with several bends needs a flexible fit, not a different power.
Logs are undefined at zero and below, so a variable containing zeros needs a deliberate decision, not an automatic log(x + 1). Also, choosing a transformation by trying every option and keeping the best R² is a forking path: the winner is partly fitted to noise. Decide from the residual plot and the mechanism.
Going deeper
This lesson covers transformations as part of fitting a regression model. Transformations & Recoding in the Research Toolkit covers the same operations as data preparation: which transformation to apply to a skewed variable, how to recode and bin sensibly, and how to document what you did so the analysis stays reproducible.
Why it matters: fitting a straight line to a curved relationship is a serious error. The slope understates the effect where it is strong and overstates it where it is weak. Every prediction outside the middle of the data is biased in a predictable direction. A residual plot shows the problem at a glance, and usually one transformation fixes it.
Common questions
Should I log the predictor or the outcome?
It depends on the shape of the curve, and each choice matches a different mechanism. Log the predictor when equal MULTIPLES of x go with equal increases in y. That is diminishing returns: income against life satisfaction, practice against performance. Log the outcome when each extra unit of x MULTIPLIES y by a fixed factor. That is proportional growth: populations, compound interest, an untreated epidemic. The residual plot can help you decide. If the spread of the residuals fans out as the fitted values rise, logging the outcome usually fixes the curvature and the unequal variance at once. Logging the predictor cannot do that.
Can I compare R-squared between a model of y and a model of log y?
No, and software will print both without warning you. The two explain variation in different quantities, so their totals differ and the ratios are not comparable. A logged outcome usually has the higher R-squared simply because logging compresses a long tail, and that is not evidence that it predicts better. To compare the models, turn the log model's predictions back into the original units and compute R-squared there, as the widget on this page does. The simple back-transform, exponentiating the fitted value, estimates the median rather than the mean, so it runs a little low.
How do I know whether to transform or to fit a more complicated model?
A transformation is the better choice when you can name the shape of the curve and explain why the data should have it, as with diminishing returns or proportional growth. Such a model keeps making sensible predictions outside the range you sampled, and it needs only one predictor. A bigger model is needed for a curve with several bends, for a count or yes/no outcome, or for a shape that differs between groups. Do not try every transformation and keep the one with the best R-squared. That is a forking path, and the winner will be partly fitted to your noise. Decide from the residual plot and the mechanism, then report what you chose and why.