Random Variables
A random variable assigns a number to each outcome of a random experiment; its distribution is described by a PMF, a PDF, or a CDF.
Prerequisites: Introduction to Probability.
A random variable is a numerical quantity whose value depends on the outcome of a random experiment: the number of heads in ten coin tosses, the sum of two dice, a person’s height, the time until the next bus. Random variables let us use ordinary arithmetic and algebra (sums, averages, functions) on uncertain quantities, and they are the objects that statistical models are built from.
Intuition
The outcomes of an experiment are not always numbers. Tossing a coin three times produces sequences like HTH. But we usually care about a number computed from the outcome, such as “how many heads”. That rule (take an outcome, return a number) is a random variable.
The name is a little misleading: a random variable is not a variable in the algebra sense, and the function itself is not random. The randomness is in which outcome occurs; once the outcome is known, the value of the random variable is determined.
Definition
A random variable is a function that assigns a real number to each outcome in the sample space .
Random variables are written with capital letters (, , ), and particular values they can take with lower-case letters (, , ). Statements such as “” or “” describe events, namely the set of outcomes for which they are true, so they have probabilities.
Example. Toss a coin three times and let be the number of heads. Then , , , and so on. The event is the set , and for a fair coin .
The distribution of a random variable is the complete description of which values it can take and how probable they are. Two kinds of random variable need slightly different tools.
Discrete random variables and the PMF
A random variable is discrete if it can take only a finite or countable list of values, such as . Counts are typical examples.
The distribution of a discrete random variable is given by its probability mass function (PMF):
Each is an actual probability, so a PMF satisfies
where the sum runs over all values can take. The probability that lands in a set of values is the sum of the masses on those values.
Continuous random variables and the PDF
A random variable is continuous if it can take any value in an interval, and the probability of every single exact value is zero. Measurements such as times, lengths, and weights are typically modelled this way.
Because single values have probability zero, a PMF is useless here. Instead the distribution is described by a probability density function (PDF) , and probabilities are areas under it:
A PDF satisfies and .
The value is a density, not a probability. It can be larger than 1; only the area under over an interval is a probability. Roughly, for a small interval of width around , .
Example. If a bus arrives at a uniformly random time within the next 10 minutes, the waiting time (in minutes) has density for and elsewhere. The probability of waiting at most 3 minutes is the area of a rectangle of width 3 and height , which is . The probability of waiting exactly 3 minutes is . This is the uniform distribution.
The cumulative distribution function
One tool works for every random variable, discrete or continuous: the cumulative distribution function (CDF),
It gives the probability that is at most . Every CDF starts near 0 for very small , never decreases, and approaches 1 for large . Probabilities of intervals follow by subtraction: .
The CDF is connected to the other descriptions:
- For a discrete variable, is the sum of over all values , and its graph is a staircase that jumps by at each possible value .
- For a continuous variable, is the area to the left of , and the density is the derivative of the CDF, . In the bus example, for .
The article on probability distributions compares PMFs, PDFs, and CDFs side by side.
Worked example: the sum of two dice
Roll two fair six-sided dice and let be the sum of the two faces.
Step 1: the sample space. An outcome is an ordered pair (first die, second die), such as . There are outcomes, all equally likely, each with probability .
Step 2: the values of S. The sum ranges from , from , to , from . So is a discrete random variable with eleven possible values.
Step 3: count the outcomes for each value. A sum of 7 arises from : six outcomes. A sum of 2 arises only from . In general, the number of outcomes giving sum is , so the PMF is
Step 4: use the PMF. Because the stems are probabilities, we can add them:
- , the most likely sum.
- .
- The CDF at 4 is .
- The masses add to one: .
The example also shows why the counting rule from the introduction to probability must be applied to the right sample space: the 36 ordered pairs are equally likely, but the 11 sums are not.
Summaries of a random variable
A distribution contains all the information about a random variable, but it is often useful to summarize it with a few numbers. The two most important are:
- the expected value , the long-run average value, which for the sum of two dice is ;
- the variance , which measures how spread out the values are around the expected value.
Two random variables can also be related to each other. When knowing the value of one tells us nothing about the other, they are independent.
Common misunderstandings
“A random variable is a random number.” It is a function from outcomes to numbers. The same experiment can carry many random variables: from two dice we can define the sum, the maximum, the difference, and so on.
“The PDF value is the probability of that value.” For a continuous variable, is a density and . Only areas under are probabilities.
“A PMF can be drawn as a smooth curve.” A discrete variable puts probability only on separate points; there is no probability between them. Draw a PMF with stems or bars, not a continuous line.
“X and x are interchangeable.” is the random variable; is a fixed number. asks how likely the random quantity is to equal that particular number.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Introduces random variables carefully as functions, with clear treatments of PMFs, PDFs, and CDFs.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A compact, more mathematical treatment.