Covariance

A measure of whether two variables tend to be above and below their means at the same time.

Prerequisites: Expected Value, Variance of a Random Variable, Variance and Standard Deviation.

Covariance measures whether two variables tend to move together. It is positive when large values of one variable tend to go with large values of the other, negative when large values of one tend to go with small values of the other, and near zero when there is no such linear tendency.

Covariance is rarely reported on its own, because its size depends on the units of measurement. But it is the building block of correlation and linear regression, and it appears whenever we add random quantities together, so it is worth understanding well.

Intuition

Suppose we record, for many people, the number of hours they practise a skill and their score on a test of it. For each person, ask two questions: is their practice time above or below the average, and is their score above or below the average?

Covariance turns this into a number. For each observation it multiplies the deviation of xx from its mean by the deviation of yy from its mean. The product is positive when both deviations have the same sign and negative when they differ. Averaging the products tells us which kind of observation dominates.

A scatter plot of 40 points rising from lower left to upper right, split into four quadrants by dashed lines at the two means. Most points lie in the upper-right and lower-left quadrants, where the product of deviations is positive, so the covariance is positive.
The quadrant picture of covariance for 40 simulated points. Dashed lines mark the sample means x̄ and ȳ. Blue circles (26 points) lie where (x − x̄)(y − ȳ) is positive; orange crosses (14 points) lie where it is negative. Positive products dominate, and the sample covariance is about 1.04.

The dashed lines in the figure split the plane into four quadrants. Points in the upper-right and lower-left quadrants contribute positive products; points in the other two contribute negative products. A point far from both means contributes a large product, and a point near either dashed line contributes almost nothing. The covariance is, roughly, the average of all these signed contributions.

Definition

Population covariance

Let XX and YY be two random variables defined on the same population, with means μX=E⁡[X]\mu_X = \E[X] and μY=E⁡[Y]\mu_Y = \E[Y]. Their covariance is the expected value of the product of their deviations from their means:

Cov⁡(X,Y)=E⁡[(X−μX)(Y−μY)].\Cov(X, Y) = \E\big[(X - \mu_X)(Y - \mu_Y)\big].

Expanding the product and using linearity of expectation gives an equivalent formula that is often easier to compute:

Cov⁡(X,Y)=E⁡[XY]−E⁡[X] E⁡[Y].\Cov(X, Y) = \E[XY] - \E[X]\,\E[Y].

To see why, multiply out: (X−μX)(Y−μY)=XY−μYX−μXY+μXμY(X - \mu_X)(Y - \mu_Y) = XY - \mu_Y X - \mu_X Y + \mu_X \mu_Y. Taking expectations, the middle two terms each become μXμY\mu_X \mu_Y, so the right-hand side is E⁡[XY]−μXμY−μXμY+μXμY=E⁡[XY]−μXμY\E[XY] - \mu_X\mu_Y - \mu_X\mu_Y + \mu_X\mu_Y = \E[XY] - \mu_X\mu_Y.

For discrete variables with joint probability mass function p(x,y)=P(X=x,Y=y)p(x, y) = P(X = x, Y = y), the expectation is a sum, Cov⁡(X,Y)=∑x∑y(x−μX)(y−μY) p(x,y)\Cov(X, Y) = \sum_x \sum_y (x - \mu_X)(y - \mu_Y)\, p(x, y). For continuous variables it is the corresponding integral against the joint density.

Sample covariance

From a sample of nn pairs (x1,y1),…,(xn,yn)(x_1, y_1), \dots, (x_n, y_n) with sample means xˉ\bar{x} and yˉ\bar{y}, the sample covariance is

sxy=1n−1∑i=1n(xi−xˉ)(yi−yˉ).s_{xy} = \frac{1}{n - 1} \sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y}).

It estimates Cov⁡(X,Y)\Cov(X, Y). The divisor n−1n - 1 rather than nn plays the same role as in the sample variance: it corrects for measuring deviations from the sample means instead of the unknown population means, and makes sxys_{xy} an unbiased estimator. With n−1n - 1, the sample covariance of xx with itself is exactly the sample variance sx2s_x^2.

Reading the sign

Worked example

Five basketball players report the hours xx they practised free throws last week, and then each takes 10 free throws; yy is the number made.

Step 1: compute the means. xˉ=(1+2+3+4+5)/5=3\bar{x} = (1+2+3+4+5)/5 = 3 hours and yˉ=(2+4+5+4+5)/5=4\bar{y} = (2+4+5+4+5)/5 = 4 throws.

Step 2: tabulate deviations and their products.

Player xix_i yiy_i xi−xˉx_i - \bar{x} yi−yˉy_i - \bar{y} (xi−xˉ)(yi−yˉ)(x_i - \bar{x})(y_i - \bar{y})
1 1 2 −2-2 −2-2 44
2 2 4 −1-1 00 00
3 3 5 00 11 00
4 4 4 11 00 00
5 5 5 22 11 22
Sum 00 00 66

The deviations in each column sum to zero, as they always must. Player 1 (below average on both) and player 5 (above average on both) contribute positive products; nobody contributes a negative one.

Step 3: divide by n−1n - 1.

sxy=65−1=1.5 hour-throws.s_{xy} = \frac{6}{5 - 1} = 1.5 \text{ hour-throws}.

The covariance is positive: in this small sample, more practice goes with more successful throws.

Step 4: notice the units. The unit is “hours × throws”, which has no natural meaning. Worse, if we had recorded practice in minutes, every xi−xˉx_i - \bar{x} would be 60 times larger, and the covariance would be 60×1.5=9060 \times 1.5 = 90, even though the data describe exactly the same relationship. The number 1.5 cannot tell us whether the relationship is strong or weak.

The scale problem and correlation

The worked example shows the main weakness of covariance: its magnitude depends on the units of both variables. In general, for constants aa and cc,

Cov⁡(aX,cY)=ac Cov⁡(X,Y).\Cov(aX, cY) = ac\,\Cov(X, Y).

So covariance tells us the direction of a linear relationship but not, by itself, its strength. Dividing by the two standard deviations removes the units and gives the correlation coefficient, which always lies between −1-1 and 11. For the free-throw data, the sample standard deviations are sx=10/4≈1.58s_x = \sqrt{10/4} \approx 1.58 and sy=6/4≈1.22s_y = \sqrt{6/4} \approx 1.22, and the correlation is 1.5/(1.58×1.22)≈0.771.5 / (1.58 \times 1.22) \approx 0.77, whichever units we use.

Properties

The following hold for any random variables with finite variances, and for any constants a,b,c,da, b, c, d.

Covariance with itself is variance.

Cov⁡(X,X)=E⁡[(X−μX)2]=Var⁡(X).\Cov(X, X) = \E\big[(X - \mu_X)^2\big] = \Var(X).

Symmetry. Cov⁡(X,Y)=Cov⁡(Y,X)\Cov(X, Y) = \Cov(Y, X).

Shifts do not matter; scaling multiplies. Adding a constant moves a variable and its mean by the same amount, so deviations are unchanged:

Cov⁡(aX+b, cY+d)=ac Cov⁡(X,Y).\Cov(aX + b, \, cY + d) = ac\,\Cov(X, Y).

Additivity. Cov⁡(X+Y,Z)=Cov⁡(X,Z)+Cov⁡(Y,Z)\Cov(X + Y, Z) = \Cov(X, Z) + \Cov(Y, Z). Together with the previous rule, this says covariance is linear in each argument (it is bilinear).

Variance of a sum. Applying these rules to Var⁡(X+Y)=Cov⁡(X+Y,X+Y)\Var(X + Y) = \Cov(X + Y, X + Y) gives

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y).\Var(X + Y) = \Var(X) + \Var(Y) + 2\Cov(X, Y).

When XX and YY move together, their sum varies more than the two variances alone would suggest; when they move in opposite directions, they partly cancel and the sum varies less. Only when Cov⁡(X,Y)=0\Cov(X, Y) = 0 do the variances simply add. The same identity holds for sample quantities: for the free-throw data, the values xi+yix_i + y_i are 3,6,8,8,103, 6, 8, 8, 10, whose sample variance is 28/4=728/4 = 7, which equals sx2+sy2+2sxy=2.5+1.5+3s_x^2 + s_y^2 + 2s_{xy} = 2.5 + 1.5 + 3.

Independence and zero covariance

If XX and YY are independent, then E⁡[XY]=E⁡[X] E⁡[Y]\E[XY] = \E[X]\,\E[Y], and so Cov⁡(X,Y)=0\Cov(X, Y) = 0. Independent variables are always uncorrelated.

The converse is false: zero covariance does not imply independence. Covariance only detects linear tendencies, and a relationship can be perfectly strong yet not linear.

A standard example: let XX take the values −1-1, 00, 11, each with probability 1/31/3, and let Y=X2Y = X^2. Then YY is completely determined by XX. Yet

E⁡[X]=0,E⁡[XY]=E⁡[X3]=13(−1+0+1)=0,\E[X] = 0, \qquad \E[XY] = \E[X^3] = \tfrac{1}{3}(-1 + 0 + 1) = 0,

so Cov⁡(X,Y)=0−0⋅E⁡[Y]=0\Cov(X, Y) = 0 - 0 \cdot \E[Y] = 0. The variables are not independent: P(Y=1)=2/3P(Y = 1) = 2/3, but once we know X=1X = 1, we know Y=1Y = 1 for certain. The covariance is zero because the positive products from X=1X = 1 exactly cancel the negative products from X=−1X = -1. The same thing happens with Y=X2Y = X^2 for any XX whose distribution is symmetric about 00 (and has a finite third moment).

Common misunderstandings

“A large covariance means a strong relationship.” Not necessarily. Covariance grows with the spread of each variable and depends on the units. Changing hours to minutes multiplied the covariance by 60 without changing the data. Use correlation to judge strength.

“Zero covariance means the variables are unrelated.” It means there is no linear relationship. Y=X2Y = X^2 with XX symmetric about zero has zero covariance with XX even though YY is a function of XX. Always look at a scatter plot.

“Positive covariance means one variable causes the other.” Covariance describes how variables vary together in the data; it says nothing about why. A third variable can drive both.

“Variances of a sum always add.” Only when the covariance is zero, for example when the variables are independent. Otherwise the 2Cov⁡(X,Y)2\Cov(X, Y) term matters, and it can be large.

Further reading