Correlation
Two things vary together: taller people tend to weigh more, more study hours tend to mean higher grades. Correlation puts a single number on that co-movement: the correlation coefficient r, which runs from −1 to +1 and captures both the direction and the strength of a straight-line relationship.
Reading the number
- Sign = direction. Positive r: as one goes up, so does the other. Negative r: as one goes up, the other goes down.
- Size = strength. Near ±1 the points hug a straight line; near 0 they're a shapeless blob with no linear trend.
What to look at in the picture
Look at the scatterplot before the number, and describe four things about every one you meet, in this order. The form: straight, curved, or broken into clumps. The direction: up to the right or down to the right. The strength: a tight band around that form, or a loose haze. Last, the deviations from the pattern, meaning the points that do not follow it. The first three describe most of the points. The fourth is where many interesting questions come from, and naming it separately keeps you from ignoring it.
Draw the plot with care too. Stretch it tall and a weak relationship starts to look convincing; flatten it and a real one turns into a smear. Keep the plotting area roughly square, let each axis cover about the range of its variable, and leave no wide empty bands at the edges. Then your impression comes from the relationship, not from the shape of the plot.
The fastest way to build intuition is to see what an r of 0.3 and an r of 0.9 look like, and to break a tidy relationship by dragging a single point.
🎮 Correlation Explorer
Set a target strength to generate a cloud, then drag any point. The coefficient r is computed live from whatever is on screen.
Computing r by hand
On a small table you can compute the coefficient in five minutes with a basic calculator. The definition says: turn every value into a z-score, multiply each pair of them together, and average the results.
r = 1 / (n − 1) · Σ [(x − x̄) / sx] · [(y − ȳ) / sy]
Standardizing puts both columns on the same scale: standard deviations from their own mean. A person above average on both contributes a positive product, and so does a person below average on both. A person high on one and low on the other contributes a negative product. Add the products up, and the sign of the total tells you which kind of person dominates. Divide by n − 1, the same divisor the standard deviation uses, and the answer always stays between −1 and +1.
Five music students below, with the hours each spent practicing in a week and the mark their teacher gave the piece. Step through the four buttons and the table fills in column by column.
📏 Correlation, Step by Step
Five students, their practice hours and their marks. Step 1 measures how far each student is from the two means, step 2 puts those distances in standard deviations, step 3 multiplies each pair, and step 4 averages the products into r.
The arithmetic runs: x̄ = 3.0 hours and ȳ = 66.0 marks, sx = 1.58 and sy = 9.62, and the five products come out as 1.447, −0.263, 0.000, −0.066 and 1.841. Their total is 2.959, and 2.959 divided by 4 gives r = 0.740. Look at where that total comes from. The two students at the ends of the practice range contribute 3.288 between them, more than the whole total, because the two in the middle go against the trend and subtract some of it. One of those two practiced an hour above average and scored a mark just below it. The student who practiced exactly the average three hours contributes zero however well or badly they did, because their standardized x is zero, and zero times anything is zero. So the further a point is from the means, the more it affects r. That is the arithmetic behind the outlier warning below.
Facts about r
Several useful properties follow from that formula. Because it is built from standardized values, r has no units and does not change when you change the units. Switch practice from hours to minutes, or marks from out of 100 to out of 20, and the coefficient is identical. Because x and y enter the product symmetrically, correlating y with x gives exactly the same number as correlating x with y. This is a first sign that r is a description, not a prediction. The mean of standardized products always stays within −1 to +1, and it reaches an endpoint only when all the points are on one straight line. Both columns have to be quantitative: there is no correlation between eye color and blood type, even if software will compute one from numbered categories.
Two more facts get their own section below. The formula measures only straight-line relationships, and a single unusual point can change the answer a lot.
Three things people get wrong
1. Correlation measures only straight-line relationships. A perfect U-shape (y = x²) can have r ≈ 0 even though the variables are perfectly related. Low r means "no linear trend," not "no relationship." Always look at the scatter.
2. Correlation is not causation. Ice-cream sales and drowning deaths rise together (both driven by hot weather). A strong r tells you two things move together, never why. If you can measure the lurking variable, a partial correlation recomputes the correlation with that variable held constant. Multiple regression generalizes the same idea of holding the others constant. Hold temperature constant and the ice-cream link disappears; the links that remain are the interesting ones.
3. A single outlier can fake a correlation, or hide one. Drag one point far away in the playground and watch r jump. A careful analysis always checks whether one point is driving the result.
All three warnings lead to one habit: plot your data. The classic demonstration is Anscombe's quartet: four datasets that share almost every summary number yet look nothing alike. (For two numeric variables the scatterplot is nearly always the right chart. When one of them is categorical, the chart chooser suggests a display that fits.)
🎮 Same r, Four Different Stories
Four datasets with identical means, correlation (r = 0.82), and best-fit line. Flip between them to see how different they look.
Every dataset above gives r = 0.82 and the same line, but only the first is the tidy linear relationship that number suggests. The second is a smooth curve, the third is a perfect line pulled off course by one outlier, and the fourth has no trend at all except for a single distant point that creates one.
How good is your eye at the reverse task? Reading a strength off a cloud is a skill that improves with feedback, and almost everybody starts out guessing too high. The ten plots below are fixed, so a whole class gets the same ten in the same order and can compare results afterwards.
🎮 Guess the r
Ten scatterplots, the same ten every visit. Type the correlation you think you are looking at, press Enter to check, then press Enter again for the next one.
The running score is a mean absolute error: for each plot you have checked, the distance between your guess and the true value, ignoring the sign, averaged over the plots so far. It answers "how far off am I, typically?", not "am I too high or too low?". Because the misses never cancel, a run of overestimates cannot hide in it. A score under about .10 after all ten plots means a good eye. Most people score nearer .15 on their first pass and improve quickly.
The exercise teaches two things. A cloud that looks like a clear relationship is often around .5, not the .8 it feels like. And the visual difference between .0 and .3 is much smaller than the difference between .6 and .9. Both are reasons to report the number, not just describe the picture in words.
Two reasons r comes out too small
A weak correlation sometimes means the relationship is weak. Check two other explanations first. Both make a real relationship look smaller than it is.
Restriction of range. Correlation depends on spread, so looking at only part of the range makes it smaller. Suppose an admissions test correlates .60 with later grades across all applicants. Compute that correlation among the students you admitted, the top quarter of the test distribution, and .60 shows up as roughly .35; in the top 10% it falls to about .30. The underlying relationship did not change. You removed the variation that made it visible. So selective samples (admitted students, hired employees, patients sick enough to be referred) give correlations that look disappointing. A validity study run inside the selected group underestimates the validity of the test used for the selection.
Unreliable measures. Every measurement contains noise, and noise is uncorrelated with everything. The shrinkage this causes is called attenuation. If your two variables have reliabilities rxx and ryy, the correlation you observe is the true one shrunk by a factor of √(rxx · ryy). Two questionnaires with an acceptable α = .70 each turn a true correlation of .60 into an observed .42. The same arithmetic sets a ceiling: with those instruments, even a perfect underlying relationship would give an observed r of at most .70.
Before writing "the variables were only weakly related," check both. Did you compute r inside a selected slice of the range, and how reliable were the measures? A .30 from a restricted sample with noisy instruments is compatible with a substantial relationship in the population.
Undoing attenuation, and why it bites
Attenuation has an obvious-looking inverse. If unreliability shrinks r by a factor of √(rxx · ryy), dividing by that factor should recover the correlation between the true scores:
rcorrected = robserved ⁄ √(rxx · ryy)
That is the correction for attenuation, and on the example above it works exactly: .42 divided by .70 gives back the .60. The formula does not show that the divisor is itself estimated from the same modest sample. Dividing by an uncertain number below 1 magnifies everything, including sampling error.
A simulation shows how much. Take two 8-item scales at α ≈ .70, 100 respondents and a true correlation of .60, with both alphas estimated from the same sample. The corrected estimate averages .60, so the correction is close to unbiased. But 90% of its values fall between .41 and .78, a spread about a third wider than the raw r's. Raise the true correlation to .90 and the same procedure gives a corrected value above 1.0 in roughly one sample in ten.
Report the correction alongside the observed value, never in place of it: the observed r, the reliabilities you used, and the corrected value labeled as corrected. Your reader can then judge whether "stronger than it looks" rests on the data or on dividing by .70. If the corrected value comes out above 1, say so. It means the reliability estimates are too low for the correlation you observed. That tells you about your instruments, not about the constructs.
The interval around r
An r of .42 from 60 people is one sample's answer, and you will want to know how far it would move if you ran the study again. The usual plus-or-minus formula cannot answer it. A correlation cannot exceed 1, so as the true value rises, the sampling distribution of r piles up against that ceiling and gets a long tail on the low side only. Symmetric limits could go past the end of the scale.
Fisher's transformation removes the ceiling. Writing z = ½ · ln[(1 + r) / (1 − r)] stretches the correlation scale so that it runs from minus infinity to plus infinity. On that scale the sampling distribution is close to normal, with a standard error of 1/√(n − 3) that depends only on the sample size. Add and subtract 1.96 standard errors there, then convert both limits back with r = tanh z.
On this lesson's own numbers: r = .42 with n = 60 gives z = .448 and a standard error of 1/√57 = .132, so the interval on the transformed scale is .448 ± .260, or [.188, .707]. Converted back, that is 95% CI [.19, .61]. The interval is not symmetric: the estimate is .23 above its lower limit and .19 below its upper one. The Correlation & Regression Calculator prints this interval beside r under the label "Fisher z", and the APA Results Formatter writes it into the sentence.
Two things follow. First, the interval is wide: .19 to .61 runs from below Cohen's "medium" to above his "large". An r reported without its interval invites a reader to take .42 more literally than 60 people can support. Narrowing the half-width to .10 around this estimate would take n = 263. Second, the transformation is used beyond reporting. Correlations are averaged on the z scale, not the r scale, for the same reason, and that is one of the conversions behind pooling results across studies.
Why it matters: correlation is the first measure of "do these two move together?" And it's the direct stepping stone to regression, which turns that relationship into a predictive line.
Got two columns of your own? Paste them into the Correlation & Regression Calculator for a live scatterplot, r with its p-value and 95% CI, Spearman's ρ, and an outlier-robustness check.
Want a real cloud to try this on? screen-time.csv holds 60 people, one influential point, and the worked answer for what r does when you drop it.
Problem 41 of the practice problems computes r from six pairs of numbers by hand, then asks what besides causation could have produced it.
Problem 23 runs the standardized-product table above on five hospital wards, ending on a negative r and the ward that contributes nothing to it.
Common questions
What counts as a strong correlation?
Common benchmarks: |r| ≈ .10 small, .30 medium, .50+ large. Psychology rarely sees field correlations above .5, while physics laughs at anything below .95. Context is everything: r = .3 between a cheap screening question and job performance is valuable; r = .8 between two versions of the same questionnaire is unremarkable. Always interpret r against what's typical for your domain.
What is the difference between Pearson and Spearman correlation?
Pearson's r measures linear association using the actual values; Spearman's ρ replaces values with ranks first, so it measures monotonic association ("consistently increasing, straight or not"). Spearman is robust to outliers and fine for ordinal data — if the two disagree sharply, that's a clue your relationship is curved or an outlier is steering Pearson.
My two groups both show a correlation, but one is bigger. Can I test whether they really differ?
You can, but two correlations usually have to be far apart before the test finds a difference. Put both correlations on the Fisher z scale, subtract, and divide by √(1/(n₁−3) + 1/(n₂−3)); the result is a z you read off the normal table. Take r = .55 in one group of 60 and r = .30 in another: z = 1.65, p = .099. A gap that looks large is not distinguishable from sampling noise, and the two intervals ([.34, .71] and [.05, .51]) overlap across most of their length. Detecting that difference with 80% power would take 168 people per group, twice what either correlation on its own needs. Correlations measured on the same people (does X predict Y better than Z does?) are a different test again, because the two estimates are not independent.