Probability Distributions
A description of which values a random variable can take and how probability is spread over them.
Prerequisites: Introduction to Probability, Random Variables.
A probability distribution describes which values a random variable can take and how likely each value, or each range of values, is. It is the complete probabilistic description of an uncertain quantity: once you know the distribution, you can answer any question of the form “what is the probability that the value lands here?”
Distributions are the basic vocabulary of statistics. Statistical models are built by choosing a distribution for the data, and most of inference is about learning the unknown numbers, the parameters, that pin that distribution down.
Intuition
Think of probability as a fixed amount of sand, exactly one unit, that has to be spread over the possible values of a random variable. The distribution says where the sand goes.
- If the variable can only take separate values, such as the number of heads in four coin tosses (0, 1, 2, 3, or 4), the sand is piled up in lumps at those values. Each lump is a probability.
- If the variable can take any value in a range, such as a person’s height, the sand is spread out smoothly, like a layer of varying thickness. No single point holds any sand by itself; only stretches of the line do. The thickness at a point is the density.
These two situations are handled with two different tools, the probability mass function and the probability density function. A third tool, the cumulative distribution function, works for both.
Discrete distributions and the PMF
A random variable is discrete if it can take only a finite or countable list of separate values, such as Counts are the typical example.
A discrete distribution is described by its probability mass function (PMF)
the probability that the random variable takes exactly the value . A PMF must satisfy two conditions:
where the sum runs over all possible values. The probability of a set of values is the sum of their masses. For the number of heads in four fair coin tosses (top-left panel above),
Continuous distributions and the PDF
A random variable is continuous if it can take any value in an interval, and the probability of every single exact value is . Heights, times, and temperatures are usually modelled this way.
A continuous distribution is described by a probability density function (PDF) . Probabilities are areas under the density:
A density must satisfy
In the top-right panel above, the shaded area between and is about , so a standard normal variable lands in that interval with probability about .
Density is not probability
The value is not the probability that ; that probability is for every . A density is probability per unit length. For a small interval of width around ,
So a larger density means values near are more likely, but itself has units of “probability per unit of ”. In particular, a density can be larger than . A variable spread evenly over the interval from to has density on that interval, because the area, , must equal .
A practical consequence: for a continuous variable, and are the same number, because the single point contributes nothing. For a discrete variable they can differ by .
The cumulative distribution function
The cumulative distribution function (CDF) is defined the same way for every random variable:
It answers “what is the probability of a value at most ?” Every CDF starts near on the far left, never decreases, and approaches on the far right.
- For a discrete variable the CDF is a staircase: it is flat between possible values and jumps up by at each possible value (bottom-left panel).
- For a continuous variable the CDF is smooth, and its slope is the density: wherever the derivative exists (bottom-right panel).
The CDF turns interval questions into a subtraction: , whatever the type of variable. It also gives quantiles: the median is the value where first reaches .
Parameters and families
Most named distributions are really families of distributions indexed by a few numbers called parameters. Choosing the parameter values picks out one member of the family.
For example, the normal distribution has two parameters: the mean sets the location and the standard deviation sets the spread. Every choice of and gives a different bell curve, but all of them share the same shape.
Parameters usually have constraints. A probability parameter must lie between and ; a rate or a standard deviation must be positive. Writing reads “ has this distribution”. In statistics the parameters are typically unknown, and estimating them from data is the job of point estimation.
Two summaries of a distribution appear in almost every article: the expected value , its long-run average, and the variance , its spread around that average.
Distributions in this section
| Distribution | Type | Parameters | Possible values | Mean | Variance |
|---|---|---|---|---|---|
| Bernoulli | discrete | ||||
| Binomial | discrete | , | |||
| Poisson | discrete | ||||
| Uniform | continuous | ||||
| Exponential | continuous | ||||
| Normal | continuous | , | all real |
Roughly: the Bernoulli and binomial distributions count successes in yes/no trials; the Poisson distribution counts events in a fixed interval; the uniform distribution spreads probability evenly over an interval; the exponential distribution describes waiting times; and the normal distribution describes measurements that cluster symmetrically around a typical value, as well as sums and averages of many independent pieces.
Common misunderstandings
“The height of a density curve is a probability.” It is a density. Only areas under the curve are probabilities, and a density can exceed ; see above.
“A continuous variable never takes any particular value, so nothing can happen.” Every observation is some particular value; the point is that no single value carries positive probability in advance. Intervals do, and they add up to .
“Drawing a smooth curve through a PMF gives its density.” A discrete variable has no density. Joining the stems of a PMF with a line suggests probability between the integers, where there is none.
“The data follow the distribution exactly.” A named distribution is a model. It is useful when its assumptions are close enough to the truth for the question at hand; checking that is part of the analysis.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Develops random variables, PMFs, PDFs, and CDFs carefully, with many examples.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A compact, more mathematical summary of the standard distributions.