Covariance
A measure of whether two variables tend to be above and below their means at the same time.
Prerequisites: Expected Value, Variance of a Random Variable, Variance and Standard Deviation.
Covariance measures whether two variables tend to move together. It is positive when large values of one variable tend to go with large values of the other, negative when large values of one tend to go with small values of the other, and near zero when there is no such linear tendency.
Covariance is rarely reported on its own, because its size depends on the units of measurement. But it is the building block of correlation and linear regression, and it appears whenever we add random quantities together, so it is worth understanding well.
Intuition
Suppose we record, for many people, the number of hours they practise a skill and their score on a test of it. For each person, ask two questions: is their practice time above or below the average, and is their score above or below the average?
- If people who practise more than average also tend to score more than average (and those who practise less tend to score less), the two deviations usually have the same sign.
- If the opposite happens, the two deviations usually have opposite signs.
Covariance turns this into a number. For each observation it multiplies the deviation of from its mean by the deviation of from its mean. The product is positive when both deviations have the same sign and negative when they differ. Averaging the products tells us which kind of observation dominates.
The dashed lines in the figure split the plane into four quadrants. Points in the upper-right and lower-left quadrants contribute positive products; points in the other two contribute negative products. A point far from both means contributes a large product, and a point near either dashed line contributes almost nothing. The covariance is, roughly, the average of all these signed contributions.
Definition
Population covariance
Let and be two random variables defined on the same population, with means and . Their covariance is the expected value of the product of their deviations from their means:
Expanding the product and using linearity of expectation gives an equivalent formula that is often easier to compute:
To see why, multiply out: . Taking expectations, the middle two terms each become , so the right-hand side is .
For discrete variables with joint probability mass function , the expectation is a sum, . For continuous variables it is the corresponding integral against the joint density.
Sample covariance
From a sample of pairs with sample means and , the sample covariance is
It estimates . The divisor rather than plays the same role as in the sample variance: it corrects for measuring deviations from the sample means instead of the unknown population means, and makes an unbiased estimator. With , the sample covariance of with itself is exactly the sample variance .
Reading the sign
- : above-average values of tend to occur with above-average values of .
- : above-average values of tend to occur with below-average values of .
- : the variables are uncorrelated. There is no linear tendency, but, as explained below, there may still be a strong relationship of another kind.
Worked example
Five basketball players report the hours they practised free throws last week, and then each takes 10 free throws; is the number made.
Step 1: compute the means. hours and throws.
Step 2: tabulate deviations and their products.
| Player | |||||
|---|---|---|---|---|---|
| 1 | 1 | 2 | |||
| 2 | 2 | 4 | |||
| 3 | 3 | 5 | |||
| 4 | 4 | 4 | |||
| 5 | 5 | 5 | |||
| Sum |
The deviations in each column sum to zero, as they always must. Player 1 (below average on both) and player 5 (above average on both) contribute positive products; nobody contributes a negative one.
Step 3: divide by .
The covariance is positive: in this small sample, more practice goes with more successful throws.
Step 4: notice the units. The unit is “hours × throws”, which has no natural meaning. Worse, if we had recorded practice in minutes, every would be 60 times larger, and the covariance would be , even though the data describe exactly the same relationship. The number 1.5 cannot tell us whether the relationship is strong or weak.
The scale problem and correlation
The worked example shows the main weakness of covariance: its magnitude depends on the units of both variables. In general, for constants and ,
So covariance tells us the direction of a linear relationship but not, by itself, its strength. Dividing by the two standard deviations removes the units and gives the correlation coefficient, which always lies between and . For the free-throw data, the sample standard deviations are and , and the correlation is , whichever units we use.
Properties
The following hold for any random variables with finite variances, and for any constants .
Covariance with itself is variance.
Symmetry. .
Shifts do not matter; scaling multiplies. Adding a constant moves a variable and its mean by the same amount, so deviations are unchanged:
Additivity. . Together with the previous rule, this says covariance is linear in each argument (it is bilinear).
Variance of a sum. Applying these rules to gives
When and move together, their sum varies more than the two variances alone would suggest; when they move in opposite directions, they partly cancel and the sum varies less. Only when do the variances simply add. The same identity holds for sample quantities: for the free-throw data, the values are , whose sample variance is , which equals .
Independence and zero covariance
If and are independent, then , and so . Independent variables are always uncorrelated.
The converse is false: zero covariance does not imply independence. Covariance only detects linear tendencies, and a relationship can be perfectly strong yet not linear.
A standard example: let take the values , , , each with probability , and let . Then is completely determined by . Yet
so . The variables are not independent: , but once we know , we know for certain. The covariance is zero because the positive products from exactly cancel the negative products from . The same thing happens with for any whose distribution is symmetric about (and has a finite third moment).
Common misunderstandings
“A large covariance means a strong relationship.” Not necessarily. Covariance grows with the spread of each variable and depends on the units. Changing hours to minutes multiplied the covariance by 60 without changing the data. Use correlation to judge strength.
“Zero covariance means the variables are unrelated.” It means there is no linear relationship. with symmetric about zero has zero covariance with even though is a function of . Always look at a scatter plot.
“Positive covariance means one variable causes the other.” Covariance describes how variables vary together in the data; it says nothing about why. A third variable can drive both.
“Variances of a sum always add.” Only when the covariance is zero, for example when the variables are independent. Otherwise the term matters, and it can be large.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Covers covariance, correlation, and joint distributions with many examples.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A compact treatment of covariance and its properties.