Correlation

A unit-free number between −1 and 1 that measures the strength and direction of a linear relationship between two variables.

Prerequisites: Covariance, Variance and Standard Deviation.

The correlation coefficient measures how closely two variables follow a straight-line relationship. It is a single number between −1-1 and 11: values near 11 mean the points lie close to a rising line, values near −1-1 mean they lie close to a falling line, and values near 00 mean there is no linear trend.

Correlation is one of the most widely reported statistics, and one of the most widely misread. It measures only linear association, it can be dominated by a single unusual point, and it says nothing on its own about cause and effect. This article explains what it measures and where it misleads.

The version described here is Pearson’s correlation coefficient, the one meant when people say “correlation” without qualification.

Intuition

Covariance already tells us whether two variables tend to be above and below their means together. Its problem is scale: measuring a variable in minutes instead of hours multiplies the covariance by 60, so its size says nothing about how tight the relationship is.

Correlation fixes this by measuring each variable in units of its own standard deviation before comparing them. Convert every value to a z-score (how many standard deviations it lies above or below its mean), then ask how well the two sets of z-scores agree. If they match perfectly, the points lie exactly on a rising line and the correlation is 11. If one is always the negative of the other, the points lie on a falling line and the correlation is −1-1.

Six scatter plots. Five show clouds of points ranging from a tight falling band (r = −0.91) through a shapeless cloud (r = 0.03) to a tight rising band (r = 0.90). The sixth shows points lying on a clear U-shaped curve, yet its correlation is only 0.02.
Simulated samples of 60 points each, labelled with their sample correlation r. The first five were drawn from bivariate normal distributions with population correlation −0.9, −0.5, 0, 0.5, and 0.9. The last follows y = x² plus small noise, with x spread evenly between −2 and 2: a strong relationship that correlation does not detect.

Notice in the figure that r=0.5r = 0.5 is still a fairly loose cloud; the band only looks tight near ±0.9\pm 0.9. And the last panel is a reminder that r≈0r \approx 0 does not mean “no relationship”.

Definition

Population correlation

For random variables XX and YY with standard deviations σX>0\sigma_X > 0 and σY>0\sigma_Y > 0, the correlation is the covariance divided by both standard deviations:

Corr⁡(X,Y)=ρXY=Cov⁡(X,Y)σX σY.\Corr(X, Y) = \rho_{XY} = \frac{\Cov(X, Y)}{\sigma_X \, \sigma_Y}.

The Greek letter ρ\rho (rho) is the usual symbol for a population correlation. It is undefined if either variable is constant, because then a standard deviation is zero.

Sample correlation

From nn pairs (x1,y1),…,(xn,yn)(x_1, y_1), \dots, (x_n, y_n), with sample means xˉ\bar{x}, yˉ\bar{y}, sample standard deviations sxs_x, sys_y, and sample covariance sxys_{xy}, the sample correlation is

r=sxysx sy=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2 ∑i=1n(yi−yˉ)2.r = \frac{s_{xy}}{s_x \, s_y} = \frac{\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum_{i=1}^n (x_i - \bar{x})^2}\,\sqrt{\sum_{i=1}^n (y_i - \bar{y})^2}}.

The second form follows from the first because the factors 1/(n−1)1/(n-1) in sxys_{xy}, sxs_x, and sys_y cancel. It is the easiest form to use by hand.

An equivalent way to write it makes the “agreement of z-scores” idea explicit. With zx,i=(xi−xˉ)/sxz_{x,i} = (x_i - \bar{x})/s_x and zy,i=(yi−yˉ)/syz_{y,i} = (y_i - \bar{y})/s_y,

r=1n−1∑i=1nzx,i zy,i.r = \frac{1}{n - 1} \sum_{i=1}^n z_{x,i}\, z_{y,i}.

So rr is (almost) the average product of z-scores: positive products come from points that are on the same side of both means, negative products from points that are on opposite sides.

Key properties

Worked example

We reuse the free-throw data from the covariance article: five players’ practice hours xx and free throws made out of 10, yy.

xix_i yiy_i xi−xˉx_i - \bar{x} yi−yˉy_i - \bar{y} (xi−xˉ)2(x_i - \bar{x})^2 (yi−yˉ)2(y_i - \bar{y})^2 product
1 2 −2-2 −2-2 4 4 4
2 4 −1-1 00 1 0 0
3 5 00 11 0 1 0
4 4 11 00 1 0 0
5 5 22 11 4 1 2
Sum 10 6 6

Here xˉ=3\bar{x} = 3 and yˉ=4\bar{y} = 4.

Step 1: plug the column sums into the formula.

r=610 6=660≈0.77.r = \frac{6}{\sqrt{10}\,\sqrt{6}} = \frac{6}{\sqrt{60}} \approx 0.77.

Step 2: check against the covariance route. The sample covariance is 6/4=1.56/4 = 1.5, and sx=10/4≈1.58s_x = \sqrt{10/4} \approx 1.58, sy=6/4≈1.22s_y = \sqrt{6/4} \approx 1.22. Then 1.5/(1.58×1.22)≈0.771.5 / (1.58 \times 1.22) \approx 0.77, the same answer.

Step 3: change units. If practice is recorded in minutes, the covariance becomes 90 and sxs_x becomes about 94.9, but rr is still about 0.77. The correlation describes the shape of the point cloud, not the units on the axes.

Step 4: add one unusual point. Suppose a sixth player reports 10 hours of practice but makes only 1 throw (perhaps a data-entry error). Recomputing with all six players gives r≈−0.44r \approx -0.44. A single point has turned a clear positive association into an apparently negative one. This is the most important practical lesson about rr: plot the data before trusting the number.

Correlation measures only linear association

The correlation coefficient asks one specific question: how well does a straight line describe the points? It can miss, or misdescribe, other patterns.

A classic demonstration is Anscombe’s quartet (Anscombe, 1973): four small datasets with almost identical means, variances, correlations (about 0.82), and fitted regression lines, which look completely different when plotted. One is a reasonable linear cloud, one is a smooth curve, one is a perfect line spoiled by a single outlier, and one has all but one point at the same xx value.

Spearman’s rank correlation

When the relationship is monotonic (always increasing or always decreasing) but not linear, or when outliers are a concern, a common alternative is Spearman’s rank correlation. Replace each xix_i by its rank among the xx values and each yiy_i by its rank among the yy values, then compute Pearson’s rr on the ranks. It equals 11 whenever yy strictly increases with xx, even along a curve, and it is less sensitive to a few extreme values. It still cannot detect a U-shaped pattern like the one in the figure.

Correlation and causation

A correlation between two variables does not show that one causes the other. Three common alternatives:

Establishing causation usually requires a randomized experiment, or careful study design and assumptions that rule out the alternatives. Correlation is often a good reason to investigate, not a conclusion.

Correlation and regression

Correlation is closely tied to simple linear regression. The least-squares line that predicts yy from xx has slope

β^1=r sysx,\hat{\beta}_1 = r \, \frac{s_y}{s_x},

so the slope has the same sign as rr. Moreover, in simple regression the fraction of the variance of yy explained by the line, R2R^2, equals r2r^2. For the free-throw data, r2=36/60=0.6r^2 = 36/60 = 0.6: the fitted line accounts for 60% of the variation in throws made.

This gives a more honest feel for the size of rr. A correlation of 0.50.5 means the line explains only 25%25\% of the variance; a correlation of 0.30.3, often described as “moderate”, explains only 9%9\%.

Common misunderstandings

“r = 0 means the variables are unrelated.” It means there is no linear trend. Strong curved relationships, such as the U shape in the figure, can have rr near 00.

“A correlation of 0.5 means the variables are 50% related.” There is no such interpretation. If anything, r2=0.25r^2 = 0.25 is the fraction of variance a linear fit explains.

“A larger correlation means a steeper slope.” No. Correlation measures how tightly points cluster around a line, not how steep the line is. Points exactly on a nearly flat line have r=1r = 1.

“Correlation implies causation.” See above. A confounder, reverse causation, or chance can all produce a correlation.

“The number is enough; there is no need to look at the data.” Anscombe’s quartet and the outlier in the worked example show otherwise. Always look at a scatter plot.

Further reading