Expected Value
The probability-weighted average of a random variable's values, interpreted as its long-run average and as the balance point of its distribution.
Prerequisites: Random Variables.
The expected value (or mean, or expectation) of a random variable is the average of its possible values, each weighted by how probable it is. It is the value you would get, on average, if you could repeat the random experiment many times. The expected value is the most common single-number summary of a distribution, and it is the population counterpart of the familiar sample mean.
Intuition
Roll a fair die many times and average the results. Each face from 1 to 6 comes up about one-sixth of the time, so the average settles down near
That number is the expected value of a die roll. It is a weighted average: every value counts in proportion to its probability. The law of large numbers makes the “settles down” claim precise.
Notice that 3.5 is not a possible result of a roll. The expected value is a long-run average, not a prediction of any single outcome.
A second, physical picture is just as useful. Imagine the possible values marked on a weightless rod, with a weight at each value proportional to its probability. The expected value is the point where the rod would balance.
Definition
Let be a random variable. Its expected value is written , and often when there is no ambiguity.
Discrete case. If has probability mass function , then
where the sum runs over all values can take. Each value is multiplied by its probability, and the results are added.
Continuous case. If has probability density function , the sum becomes an integral:
The density plays the role of the probabilities, and the integral is a continuous weighted average.
When it exists. The expected value is defined only when the sum or integral converges absolutely, that is, when or is finite. For every distribution with finitely many values this is automatic. Some heavy-tailed distributions, such as the Cauchy distribution, have no expected value at all: the averages of repeated samples never settle down.
Expected value of a function
Often we need the expected value of a quantity computed from , such as or a payoff that depends on . For a function ,
We apply to each value but keep the original probabilities. There is no need to work out the distribution of first.
Worked example
A game. You pay 1 dollar to roll a fair die. If it shows a 6 you receive 5 dollars; otherwise you receive nothing. Let be your net winnings in dollars. Then with probability and with probability , so
On average you lose about 17 cents per game. Over 600 games you should expect to be down about 100 dollars, although any particular sequence of games will differ.
The distribution in the figure. Guessing at random on 6 questions, each with 4 options, the number of correct answers has these probabilities:
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | |
|---|---|---|---|---|---|---|---|
| 729 | 1458 | 1215 | 540 | 135 | 18 | 1 |
(These come from the binomial distribution with trials and success probability .) Then
The most likely value is 1, but the small probabilities on 3, 4, 5, and 6 sit far to the right and pull the balance point to 1.5.
A continuous example. If is uniformly distributed on , its density is there, and
the midpoint, as symmetry suggests.
Linearity of expectation
The single most useful property of expected value is that it is linear. For any random variables and and constants and ,
The first rule says that rescaling and shifting a random variable rescales and shifts its mean in the same way. Converting temperatures from Celsius to Fahrenheit, , converts the expected temperature by the same formula.
The second rule says that the expected value of a sum is the sum of the expected values. Crucially, this holds whether or not and are independent. More generally, .
For discrete variables the first rule follows directly from the definition:
using .
Example: two dice. The sum of two dice has the triangular distribution shown in the random variables article. Instead of using that distribution, write , where and are the two dice. Then .
Example: counting with indicators. How many sixes do you expect in 10 rolls of a die? Let be if roll is a six and otherwise. Each has . The number of sixes is , so its expected value is . This trick of writing a count as a sum of 0–1 indicators makes many hard-looking expectations easy, even when the indicators are dependent.
Linearity is also the reason that the sample mean of a random sample has the same expected value as each observation, a fact used throughout point estimation.
Common misunderstandings
“The expected value is the most likely value.” Not in general. For the guessing example, the most likely value is 1 but the mean is 1.5. For a single die every value is equally likely, and the mean, 3.5, is not even a possible value.
“The expected value is what will happen.” It is a long-run average. In a single play of the game above you either win 4 dollars or lose 1 dollar; you never lose exactly 17 cents.
“E[g(X)] = g(E[X]).” This is false unless is linear. For a die, , while . The gap between these two numbers is exactly the variance of a die roll.
“E[XY] = E[X] E[Y].” This product rule needs extra assumptions; it holds when and are independent, but not in general. The difference is the covariance. Linearity for sums needs no such assumption.
“The mean is the middle value.” The middle value is the median. For skewed distributions the mean is pulled toward the long tail, as the figure shows.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Makes extensive use of linearity and indicator variables, with many examples.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A concise treatment of expectation and its properties.