Central Limit Theorem
The mean of many independent observations is approximately normally distributed, whatever the shape of the population.
Prerequisites: Sampling Distributions, Normal Distribution, Law of Large Numbers.
The central limit theorem (CLT) says that if you add up or average many independent random quantities with the same distribution, the result is approximately normally distributed, even when the individual quantities are far from normal. The larger the number of quantities, the better the approximation.
This is why the normal distribution is everywhere in statistics. We rarely know the shape of a population, but we often need the distribution of a sample mean in order to build a confidence interval or carry out a hypothesis test. The central limit theorem tells us that, for a large enough sample, that distribution is close to a normal one with a known centre and spread.
Intuition
Roll one fair die and every face from 1 to 6 is equally likely: a flat distribution. Roll two dice and add them, and a total of 7 is six times as likely as a total of 2, because many combinations produce middling totals and only one produces each extreme. With more dice, this effect grows: extreme totals need every die to be extreme together, which is rare, while middling totals can arise in an enormous number of ways. The distribution piles up in the middle and takes on a bell shape.
The same happens for skewed variables. The figure below starts with the exponential distribution, which is strongly skewed: it has a peak at zero and a long right tail. Averages of 2 or 5 exponential values are still visibly skewed, but by 30 the histogram is close to the normal curve.
Statement
Let be independent and identically distributed (iid) random variables with mean and finite, positive variance . Let be their sample mean.
From the article on sampling distributions, has mean and standard error . Subtracting the mean and dividing by the standard error gives a standardized mean , which always has mean and variance :
Central limit theorem. As , the distribution of approaches the standard normal distribution: for every real number ,
where is the cumulative distribution function of the standard normal distribution . This is called convergence in distribution.
What this means in practice
For large we use the approximation
read as “the sample mean is approximately normal with mean and variance ”. Equivalently, the sum is approximately .
The centre and spread are not part of the theorem: and hold exactly for every . What the CLT adds is the shape. Without it, we would know how wide the sampling distribution is but not how much probability lies in its tails.
Relation to the law of large numbers
The law of large numbers says gets close to . The central limit theorem zooms in on the remaining error : it is typically of size , and its distribution is approximately normal. Both results describe the same averages at different levels of detail.
Worked example
Waiting times at a service counter are exponentially distributed with mean minutes. For an exponential distribution the standard deviation equals the mean, so minutes. What is the probability that the average wait of 36 customers exceeds 2.5 minutes?
Step 1: centre and spread of the mean. and
Step 2: standardize. The z-score of 2.5 is
Step 3: apply the normal approximation.
So about a 6.7% chance.
Step 4: check against simulation. Simulating 100,000 groups of 36 exponential waiting times gives a proportion of about 0.075 with an average above 2.5 minutes. (Here the exact answer is also available, because a sum of exponential variables has a gamma distribution; it is about 0.074.) The approximation is in the right range but noticeably too small. The reason is the leftover skewness visible in the figure: with an exponential population, even a mean of 36 values has a slightly longer right tail than a normal distribution. In the left tail the error goes the other way: the normal approximation gives , while the exact value is about 0.056.
Why it underlies confidence intervals and tests
Many standard methods have the same structure: compute an estimate, compute its standard error, and use the normal distribution to say how far the estimate is likely to be from the truth. The central limit theorem is what justifies the third step when the data themselves are not normal.
For example, because is approximately normal, it lies within standard errors of in about 95% of samples. Turning that statement around gives the familiar interval for ; see confidence intervals. Similarly, a hypothesis test for a mean compares the observed z-score with the standard normal distribution.
The CLT also explains the normal approximation to the binomial distribution: a binomial count is a sum of independent 0/1 Bernoulli variables, so for many trials it is approximately normal.
Common misunderstandings
“With a large sample, the data become normal.” No. The data have whatever distribution the population has; a large sample of exponential waiting times is still skewed, and its histogram just looks more and more like the exponential density. The CLT is about the sampling distribution of the mean (or sum), the distribution of across repeated samples.
“n ≥ 30 is enough.” This is a rough rule of thumb, not part of the theorem. How large must be depends on the population. For a symmetric population with light tails, a sample of 10 may already be plenty. For a strongly skewed population, or one with occasional extreme values, 30 may be far too few, and the worked example above shows a visible error even at 36 for a moderately skewed one. Tail probabilities, which are what tests and intervals rely on, are usually the slowest part to become accurate.
“It works for any distribution.” It needs a finite variance. For very heavy-tailed distributions, such as the Cauchy distribution, the sample mean does not become normal however large is. In practice, data with rare but enormous values (insurance claims, file sizes, wealth) can make the CLT approximation poor for any realistic sample size.
“It works for any sample.” The basic theorem assumes independent observations from one distribution. Versions exist for some dependent or non-identical data, but strongly dependent observations, such as a time series with long memory, can behave very differently.
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. States and proves the central limit theorem and illustrates it with simulations.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. Covers convergence in probability and in distribution together with the CLT.
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Explains the normal approximation for sums and averages without calculus.