Random Variables

A random variable assigns a number to each outcome of a random experiment; its distribution is described by a PMF, a PDF, or a CDF.

Prerequisites: Introduction to Probability.

A random variable is a numerical quantity whose value depends on the outcome of a random experiment: the number of heads in ten coin tosses, the sum of two dice, a person’s height, the time until the next bus. Random variables let us use ordinary arithmetic and algebra (sums, averages, functions) on uncertain quantities, and they are the objects that statistical models are built from.

Intuition

The outcomes of an experiment are not always numbers. Tossing a coin three times produces sequences like HTH. But we usually care about a number computed from the outcome, such as “how many heads”. That rule (take an outcome, return a number) is a random variable.

The name is a little misleading: a random variable is not a variable in the algebra sense, and the function itself is not random. The randomness is in which outcome occurs; once the outcome is known, the value of the random variable is determined.

Definition

A random variable XX is a function that assigns a real number X(ω)X(\omega) to each outcome ω\omega in the sample space Ω\Omega.

Random variables are written with capital letters (XX, YY, SS), and particular values they can take with lower-case letters (xx, yy, ss). Statements such as “X=2X = 2” or “X≤xX \le x” describe events, namely the set of outcomes for which they are true, so they have probabilities.

Example. Toss a coin three times and let XX be the number of heads. Then X(HHH)=3X(\text{HHH}) = 3, X(HTH)=2X(\text{HTH}) = 2, X(TTT)=0X(\text{TTT}) = 0, and so on. The event {X=2}\{X = 2\} is the set {HHT,HTH,THH}\{\text{HHT}, \text{HTH}, \text{THH}\}, and for a fair coin P(X=2)=3/8P(X = 2) = 3/8.

The distribution of a random variable is the complete description of which values it can take and how probable they are. Two kinds of random variable need slightly different tools.

Discrete random variables and the PMF

A random variable is discrete if it can take only a finite or countable list of values, such as 0,1,2,…0, 1, 2, \dots. Counts are typical examples.

The distribution of a discrete random variable is given by its probability mass function (PMF):

p(x)=P(X=x).p(x) = P(X = x).

Each p(x)p(x) is an actual probability, so a PMF satisfies

p(x)≥0for all x,∑xp(x)=1,p(x) \ge 0 \quad \text{for all } x, \qquad \sum_x p(x) = 1,

where the sum runs over all values XX can take. The probability that XX lands in a set of values is the sum of the masses on those values.

Continuous random variables and the PDF

A random variable is continuous if it can take any value in an interval, and the probability of every single exact value is zero. Measurements such as times, lengths, and weights are typically modelled this way.

Because single values have probability zero, a PMF is useless here. Instead the distribution is described by a probability density function (PDF) f(x)f(x), and probabilities are areas under it:

P(a≤X≤b)=∫abf(x) dx.P(a \le X \le b) = \int_a^b f(x)\, dx.

A PDF satisfies f(x)≥0f(x) \ge 0 and ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1.

The value f(x)f(x) is a density, not a probability. It can be larger than 1; only the area under ff over an interval is a probability. Roughly, for a small interval of width Δx\Delta x around xx, P(x≤X≤x+Δx)≈f(x) ΔxP(x \le X \le x + \Delta x) \approx f(x)\,\Delta x.

Example. If a bus arrives at a uniformly random time within the next 10 minutes, the waiting time XX (in minutes) has density f(x)=1/10f(x) = 1/10 for 0≤x≤100 \le x \le 10 and 00 elsewhere. The probability of waiting at most 3 minutes is the area of a rectangle of width 3 and height 1/101/10, which is 0.30.3. The probability of waiting exactly 3 minutes is 00. This is the uniform distribution.

The cumulative distribution function

One tool works for every random variable, discrete or continuous: the cumulative distribution function (CDF),

F(x)=P(X≤x).F(x) = P(X \le x).

It gives the probability that XX is at most xx. Every CDF starts near 0 for very small xx, never decreases, and approaches 1 for large xx. Probabilities of intervals follow by subtraction: P(a<X≤b)=F(b)−F(a)P(a < X \le b) = F(b) - F(a).

The CDF is connected to the other descriptions:

The article on probability distributions compares PMFs, PDFs, and CDFs side by side.

Worked example: the sum of two dice

Roll two fair six-sided dice and let SS be the sum of the two faces.

Step 1: the sample space. An outcome is an ordered pair (first die, second die), such as (2,5)(2, 5). There are 6×6=366 \times 6 = 36 outcomes, all equally likely, each with probability 1/361/36.

Step 2: the values of S. The sum ranges from 22, from (1,1)(1, 1), to 1212, from (6,6)(6, 6). So SS is a discrete random variable with eleven possible values.

Step 3: count the outcomes for each value. A sum of 7 arises from (1,6),(2,5),(3,4),(4,3),(5,2),(6,1)(1,6), (2,5), (3,4), (4,3), (5,2), (6,1): six outcomes. A sum of 2 arises only from (1,1)(1,1). In general, the number of outcomes giving sum ss is 6−∣s−7∣6 - |s - 7|, so the PMF is

p(s)=P(S=s)=6−∣s−7∣36,s=2,3,…,12.p(s) = P(S = s) = \frac{6 - |s - 7|}{36}, \qquad s = 2, 3, \dots, 12.

Eleven vertical stems at sums 2 through 12. Their heights rise steadily from 1/36 at a sum of 2 to a peak of 6/36 at 7, then fall back to 1/36 at 12, forming a symmetric triangle.
PMF of the sum S of two fair dice. Each stem is a probability P(S = s), shown as a fraction of the 36 equally likely outcomes. The stems add up to 1.

Step 4: use the PMF. Because the stems are probabilities, we can add them:

The example also shows why the counting rule from the introduction to probability must be applied to the right sample space: the 36 ordered pairs are equally likely, but the 11 sums are not.

Summaries of a random variable

A distribution contains all the information about a random variable, but it is often useful to summarize it with a few numbers. The two most important are:

Two random variables can also be related to each other. When knowing the value of one tells us nothing about the other, they are independent.

Common misunderstandings

“A random variable is a random number.” It is a function from outcomes to numbers. The same experiment can carry many random variables: from two dice we can define the sum, the maximum, the difference, and so on.

“The PDF value is the probability of that value.” For a continuous variable, f(x)f(x) is a density and P(X=x)=0P(X = x) = 0. Only areas under ff are probabilities.

“A PMF can be drawn as a smooth curve.” A discrete variable puts probability only on separate points; there is no probability between them. Draw a PMF with stems or bars, not a continuous line.

“X and x are interchangeable.” XX is the random variable; xx is a fixed number. P(X=x)P(X = x) asks how likely the random quantity is to equal that particular number.

Further reading