Variance of a Random Variable
The expected squared distance of a random variable from its mean, measuring how spread out its distribution is.
Prerequisites: Expected Value.
The variance of a random variable measures how spread out its possible values are around its expected value. It is the average squared distance from the mean. Its square root, the standard deviation, measures spread in the same units as the variable itself. Two random variables can have the same mean but very different variances, and that difference is often what matters: in the precision of a measurement, the risk of an investment, or the reliability of an estimate.
This article is about the variance of a probability distribution. The closely related variance of a data set, computed from observed values, is covered in variance and standard deviation.
Intuition
Compare two games. In the first you always win 10 dollars. In the second you win 0 or 20 dollars with equal probability. Both have expected winnings of 10 dollars, but the second is far less predictable. The mean alone cannot tell these apart; we need a measure of how far the outcomes typically land from the mean.
A natural idea is to average the distance from the mean, . But positive and negative deviations cancel: the average of is always . Squaring the deviations removes the signs and makes large deviations count heavily. The variance is the expected value of these squared deviations. For the first game it is ; for the second, every outcome is 10 dollars from the mean, so the variance is (squared dollars) and the standard deviation is 10 dollars.
Definition
Let be a random variable with mean . The variance of is
The standard deviation is its square root,
and the variance is often written .
Spelled out:
- for a discrete variable with PMF : ;
- for a continuous variable with PDF : .
Some immediate facts:
- , because it is an average of squares.
- exactly when is a constant (equal to with probability 1).
- The variance is in squared units: if is measured in metres, is in square metres. The standard deviation is back in metres, which is why it is usually the number reported.
- The variance exists only if is finite. Some heavy-tailed distributions have a mean but infinite variance.
The shortcut formula
Expanding the square and using linearity of expectation gives a formula that is often easier to compute:
So
“the mean of the square minus the square of the mean”. Since the variance is never negative, this also shows that .
Worked example: one die
Roll a fair six-sided die and let be the result. We know .
Using the definition. The squared deviations from 3.5 are for the faces . Each has probability , so
Using the shortcut. First . Then
Both methods agree. The standard deviation is : a typical roll lands roughly 1.7 away from 3.5.
A 0–1 variable. If is with probability and otherwise (a Bernoulli variable), then , so and . The variance is largest, , when , and zero when the outcome is certain.
Properties
Shifting and scaling
For constants and ,
Adding a constant shifts every value and the mean by the same amount, so the deviations , and hence the spread, do not change. Multiplying by multiplies every deviation by , and therefore every squared deviation by .
For a die, has . Converting a temperature from Celsius to Fahrenheit () multiplies its standard deviation by 1.8; the 32 has no effect.
Sums of random variables
For any two random variables and ,
where is the covariance. It measures whether and tend to be above their means at the same time (positive covariance) or on opposite sides (negative covariance).
If and are independent, their covariance is zero and the formula simplifies:
More generally, for independent , the variance of the sum is the sum of the variances. Note what this does not say: standard deviations do not add. It is the squares that add, like the sides of a right triangle.
Two dice. The sum of two independent dice has and standard deviation , not .
The same die twice. Compare this with , the result of one die counted twice. Here the two terms are perfectly dependent, and , twice as large as for two independent dice. Independent errors partly cancel; identical ones do not.
Differences. For independent and , . Subtracting an independent random quantity adds variability, just as adding one does.
This additivity is the basis of a central fact in statistics: the mean of independent observations, each with variance , has variance . Averaging reduces spread, which is why larger samples give more precise estimates; see sampling distributions.
Relationship to the sample variance
The variance of a random variable is a property of a probability distribution, a fixed number usually written . In practice we rarely know the distribution; we have data and estimate with the sample variance
The two follow the same idea, an average squared distance from the mean, but is computed from a sample and changes from sample to sample. Why it divides by rather than is explained in variance and standard deviation.
Common misunderstandings
“Var(X + Y) = Var(X) + Var(Y) always.” Only when the covariance is zero, for example when and are independent. Otherwise the covariance term must be included.
“Var(X − Y) = Var(X) − Var(Y).” For independent variables the variances still add: .
“Var(2X) = 2 Var(X).” Scaling by multiplies the variance by , so . It is the standard deviation that doubles.
“Standard deviations add for independent variables.” Variances add; standard deviations combine as , which is smaller than .
“The variance is in the same units as the data.” It is in squared units. Use the standard deviation when you want a spread on the original scale.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Covers variance, its properties, and the variance of sums.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A concise reference for variance and covariance.