Sampling Distributions
The distribution of a statistic, such as the sample mean, over all the samples that could have been drawn.
Prerequisites: Populations and Samples, Random Variables, Expected Value, Variance of a Random Variable.
The sampling distribution of a statistic is the distribution of values that statistic would take if we drew sample after sample from the same population and recomputed it each time. A sample mean, a sample proportion, or a sample median is not a fixed number: a different random sample would have given a different value. The sampling distribution describes that variation.
It is the central idea of statistical inference. We almost always see just one sample, but to judge how far our one estimate is likely to be from the truth, we need to know how much the estimate would vary across samples. The standard deviation of the sampling distribution, called the standard error, answers exactly that question.
Intuition
Suppose you want to know the average time customers wait at a café. You cannot time every customer, so you time 25 of them and compute their average. A colleague who times a different 25 customers will get a somewhat different average. A third person gets a third value, and so on.
Two things are true about this collection of averages:
- They scatter around the true population average. Some samples happen to contain more long waits, some fewer, but there is no systematic tendency to land too high or too low.
- They scatter much less than individual waiting times do. A single customer might wait 6 minutes, but for a sample of 25 to average 6 minutes, many customers in the sample would have to wait a long time at once, which is unlikely. Averaging cancels out much of the individual variation.
The figure shows this with a simulation. The population on the left is strongly skewed: most waits are short, a few are long. The histogram on the right shows 5,000 sample means, each computed from a fresh sample of 25 waits. They are centred on the population mean, and they are far more tightly concentrated than the individual values. Notice also that the histogram is much more symmetric than the population; that is the central limit theorem at work.
A statistic is a random variable
Here is the key shift in viewpoint. Before the sample is drawn, the observations are unknown, so we model them as random variables . A statistic is any quantity computed from the sample. Because it is a function of random variables, a statistic is itself a random variable, and it has its own distribution: its sampling distribution.
Notation distinguishes the two viewpoints. The sample mean viewed as a random variable is written with a capital letter,
while (lower case) is the particular number computed from the data actually observed. has a sampling distribution; is one draw from it.
Mean and variance of the sample mean
Assume the observations are independent and identically distributed (iid): each is drawn from the same population, with mean and variance , and knowing one tells you nothing about the others. This is a good model for random sampling with replacement, or for sampling without replacement from a population much larger than the sample.
The centre. By linearity of expectation,
On average, the sample mean equals the population mean. It does not systematically over- or underestimate it.
The spread. For independent variables, the variance of a sum is the sum of the variances, and multiplying a variable by a constant multiplies its variance by . So
The variance of the mean shrinks in proportion to . The independence assumption matters here: if the observations are positively correlated (for example, members of the same household), the variance of the mean shrinks more slowly.
These two results hold for any population with a finite variance, whatever its shape. They describe the centre and spread of the sampling distribution, not its shape; for the shape, see the central limit theorem.
Standard error
The standard error (SE) of a statistic is the standard deviation of its sampling distribution. For the sample mean,
The standard deviation describes how much individual values vary. The standard error describes how much the sample mean varies from sample to sample. It is a measure of the precision of as an estimate of .
The square-root law
Because appears under a square root, precision improves more slowly than sample size grows. To halve the standard error you must quadruple the sample size:
Going from 25 to 100 observations halves the standard error; halving it again needs 400. This diminishing return is why large surveys and experiments become expensive quickly.
The estimated standard error
In practice is unknown too, so we replace it with the sample standard deviation computed from the data:
This estimated standard error is what software reports as “standard error” or “SE of the mean”. It is itself an estimate, and it is less reliable when is small, because can then be far from . That extra uncertainty is one reason small-sample methods use the distribution rather than the normal.
Worked example
Waiting times at the café have mean minutes and standard deviation minutes (the population in the figure).
Step 1: standard error for n = 25.
A typical sample mean misses by something on the order of 0.4 minutes, even though a typical individual wait differs from by about 2 minutes. The simulation in the figure agrees: the 5,000 sample means had a standard deviation of about 0.40.
Step 2: quadruple the sample. With , the standard error is minutes: half as large.
Step 3: plan a sample size. How many customers are needed for a standard error of 0.1 minutes? Solve , which gives and .
Step 4: estimate the SE from data. In real life we would not know . Suppose one sample of 25 waits has minutes and minutes. The estimated standard error is
This number, together with the approximate normality of , is what a confidence interval for is built from.
Beyond the sample mean
Every statistic has a sampling distribution: the sample proportion, the median, the sample variance, a regression slope. For a sample proportion from independent yes/no observations with success probability , the same argument gives and , because a proportion is the mean of 0/1 values with variance .
For many statistics there is no simple formula. Then the sampling distribution can be approximated by simulation, as in the figure, or by resampling the observed data (the bootstrap).
Common misunderstandings
“The standard error is the spread of the data.” It is not. The standard deviation describes the spread of individual observations and does not shrink as you collect more data. The standard error describes the uncertainty in the mean and shrinks towards zero as grows. Reporting one when the other is meant can make data look far more (or less) variable than they are.
“The sampling distribution is the distribution of the data in my sample.” The histogram of your observations estimates the population distribution. The sampling distribution is a distribution over hypothetical repeated samples, and you never observe it directly from one sample.
“A bigger population needs a bigger sample.” For a random sample from a large population, the standard error depends on , not on the population size. A random sample of 1,000 is about as precise for a country of 300 million as for a city of 1 million. (When the sample is a sizeable fraction of the population, a finite-population correction makes the SE somewhat smaller.)
“The formulas work for any sample.” assumes independent observations from one population. Clustered, correlated, or non-random samples can have much larger real uncertainty than the formula suggests, and no formula fixes a biased sampling method.
Further reading
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Explains chance variability and the standard error with very little mathematics.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; introduces sampling distributions through simulation.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A compact mathematical treatment.