Variance of a Random Variable

The expected squared distance of a random variable from its mean, measuring how spread out its distribution is.

Prerequisites: Expected Value.

The variance of a random variable measures how spread out its possible values are around its expected value. It is the average squared distance from the mean. Its square root, the standard deviation, measures spread in the same units as the variable itself. Two random variables can have the same mean but very different variances, and that difference is often what matters: in the precision of a measurement, the risk of an investment, or the reliability of an estimate.

This article is about the variance of a probability distribution. The closely related variance of a data set, computed from observed values, is covered in variance and standard deviation.

Intuition

Compare two games. In the first you always win 10 dollars. In the second you win 0 or 20 dollars with equal probability. Both have expected winnings of 10 dollars, but the second is far less predictable. The mean alone cannot tell these apart; we need a measure of how far the outcomes typically land from the mean.

A natural idea is to average the distance from the mean, X−μX - \mu. But positive and negative deviations cancel: the average of X−μX - \mu is always 00. Squaring the deviations removes the signs and makes large deviations count heavily. The variance is the expected value of these squared deviations. For the first game it is 00; for the second, every outcome is 10 dollars from the mean, so the variance is 102=10010^2 = 100 (squared dollars) and the standard deviation is 10 dollars.

Definition

Let XX be a random variable with mean μ=E⁡[X]\mu = \E[X]. The variance of XX is

Var⁡(X)=E⁡[(X−μ)2].\Var(X) = \E\big[(X - \mu)^2\big].

The standard deviation is its square root,

σ=Var⁡(X),\sigma = \sqrt{\Var(X)},

and the variance is often written σ2\sigma^2.

Spelled out:

Some immediate facts:

The shortcut formula

Expanding the square and using linearity of expectation gives a formula that is often easier to compute:

Var⁡(X)=E⁡[X2−2μX+μ2]=E⁡[X2]−2μ E⁡[X]+μ2=E⁡[X2]−μ2.\begin{aligned} \Var(X) &= \E\big[X^2 - 2\mu X + \mu^2\big] \\ &= \E[X^2] - 2\mu\,\E[X] + \mu^2 \\ &= \E[X^2] - \mu^2. \end{aligned}

So

Var⁡(X)=E⁡[X2]−(E⁡[X])2,\Var(X) = \E[X^2] - (\E[X])^2,

“the mean of the square minus the square of the mean”. Since the variance is never negative, this also shows that E⁡[X2]≥(E⁡[X])2\E[X^2] \ge (\E[X])^2.

Worked example: one die

Roll a fair six-sided die and let XX be the result. We know E⁡[X]=3.5\E[X] = 3.5.

Using the definition. The squared deviations from 3.5 are 6.25,2.25,0.25,0.25,2.25,6.256.25, 2.25, 0.25, 0.25, 2.25, 6.25 for the faces 1,2,…,61, 2, \dots, 6. Each has probability 1/61/6, so

Var⁡(X)=6.25+2.25+0.25+0.25+2.25+6.256=17.56=3512≈2.92.\Var(X) = \frac{6.25 + 2.25 + 0.25 + 0.25 + 2.25 + 6.25}{6} = \frac{17.5}{6} = \frac{35}{12} \approx 2.92.

Using the shortcut. First E⁡[X2]=(1+4+9+16+25+36)/6=91/6\E[X^2] = (1 + 4 + 9 + 16 + 25 + 36)/6 = 91/6. Then

Var⁡(X)=916−3.52=916−494=182−14712=3512.\Var(X) = \frac{91}{6} - 3.5^2 = \frac{91}{6} - \frac{49}{4} = \frac{182 - 147}{12} = \frac{35}{12}.

Both methods agree. The standard deviation is 35/12≈1.71\sqrt{35/12} \approx 1.71: a typical roll lands roughly 1.7 away from 3.5.

A 0–1 variable. If XX is 11 with probability pp and 00 otherwise (a Bernoulli variable), then X2=XX^2 = X, so E⁡[X2]=p\E[X^2] = p and Var⁡(X)=p−p2=p(1−p)\Var(X) = p - p^2 = p(1 - p). The variance is largest, 0.250.25, when p=0.5p = 0.5, and zero when the outcome is certain.

Properties

Shifting and scaling

For constants aa and bb,

Var⁡(aX+b)=a2 Var⁡(X),so the standard deviation becomes ∣a∣ σ.\Var(aX + b) = a^2\, \Var(X), \qquad \text{so the standard deviation becomes } |a|\,\sigma.

Adding a constant bb shifts every value and the mean by the same amount, so the deviations X−μX - \mu, and hence the spread, do not change. Multiplying by aa multiplies every deviation by aa, and therefore every squared deviation by a2a^2.

For a die, Y=2X+1Y = 2X + 1 has Var⁡(Y)=4×35/12=35/3≈11.67\Var(Y) = 4 \times 35/12 = 35/3 \approx 11.67. Converting a temperature from Celsius to Fahrenheit (F=1.8C+32F = 1.8C + 32) multiplies its standard deviation by 1.8; the 32 has no effect.

Sums of random variables

For any two random variables XX and YY,

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2 Cov⁡(X,Y),\Var(X + Y) = \Var(X) + \Var(Y) + 2\,\Cov(X, Y),

where Cov⁡(X,Y)=E⁡[(X−E⁡[X])(Y−E⁡[Y])]\Cov(X, Y) = \E\big[(X - \E[X])(Y - \E[Y])\big] is the covariance. It measures whether XX and YY tend to be above their means at the same time (positive covariance) or on opposite sides (negative covariance).

If XX and YY are independent, their covariance is zero and the formula simplifies:

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)(independent X,Y).\Var(X + Y) = \Var(X) + \Var(Y) \qquad \text{(independent } X, Y\text{)}.

More generally, for independent X1,…,XnX_1, \dots, X_n, the variance of the sum is the sum of the variances. Note what this does not say: standard deviations do not add. It is the squares that add, like the sides of a right triangle.

Two dice. The sum SS of two independent dice has Var⁡(S)=35/12+35/12=35/6≈5.83\Var(S) = 35/12 + 35/12 = 35/6 \approx 5.83 and standard deviation 35/6≈2.42\sqrt{35/6} \approx 2.42, not 2×1.712 \times 1.71.

The same die twice. Compare this with X+X=2XX + X = 2X, the result of one die counted twice. Here the two terms are perfectly dependent, and Var⁡(2X)=4×35/12≈11.67\Var(2X) = 4 \times 35/12 \approx 11.67, twice as large as for two independent dice. Independent errors partly cancel; identical ones do not.

Differences. For independent XX and YY, Var⁡(X−Y)=Var⁡(X)+(−1)2Var⁡(Y)=Var⁡(X)+Var⁡(Y)\Var(X - Y) = \Var(X) + (-1)^2\Var(Y) = \Var(X) + \Var(Y). Subtracting an independent random quantity adds variability, just as adding one does.

This additivity is the basis of a central fact in statistics: the mean Xˉ\bar{X} of nn independent observations, each with variance σ2\sigma^2, has variance σ2/n\sigma^2 / n. Averaging reduces spread, which is why larger samples give more precise estimates; see sampling distributions.

Relationship to the sample variance

The variance of a random variable is a property of a probability distribution, a fixed number usually written σ2\sigma^2. In practice we rarely know the distribution; we have data x1,…,xnx_1, \dots, x_n and estimate σ2\sigma^2 with the sample variance

s2=1n−1∑i=1n(xi−xˉ)2.s^2 = \frac{1}{n-1} \sum_{i=1}^n (x_i - \bar{x})^2.

The two follow the same idea, an average squared distance from the mean, but s2s^2 is computed from a sample and changes from sample to sample. Why it divides by n−1n - 1 rather than nn is explained in variance and standard deviation.

Common misunderstandings

“Var(X + Y) = Var(X) + Var(Y) always.” Only when the covariance is zero, for example when XX and YY are independent. Otherwise the covariance term must be included.

“Var(X − Y) = Var(X) − Var(Y).” For independent variables the variances still add: Var⁡(X−Y)=Var⁡(X)+Var⁡(Y)\Var(X - Y) = \Var(X) + \Var(Y).

“Var(2X) = 2 Var(X).” Scaling by aa multiplies the variance by a2a^2, so Var⁡(2X)=4Var⁡(X)\Var(2X) = 4\Var(X). It is the standard deviation that doubles.

“Standard deviations add for independent variables.” Variances add; standard deviations combine as σX2+σY2\sqrt{\sigma_X^2 + \sigma_Y^2}, which is smaller than σX+σY\sigma_X + \sigma_Y.

“The variance is in the same units as the data.” It is in squared units. Use the standard deviation when you want a spread on the original scale.

Further reading