Normal Distribution

The bell-shaped continuous distribution described by a mean and a standard deviation.

Prerequisites: Probability Distributions, Expected Value, Variance of a Random Variable.

The normal distribution (also called the Gaussian distribution) is a continuous probability distribution with a symmetric, bell-shaped density. It is described by two numbers: a mean μ\mu, which says where the bell is centred, and a standard deviation σ\sigma, which says how wide it is.

It is the most important distribution in statistics. Many measurements are approximately normal, and, more importantly, averages of many independent quantities tend to be approximately normal even when the individual quantities are not. That second fact, the central limit theorem, is why the normal distribution appears throughout statistical inference.

Intuition

Think of the heights of adults of one sex in a large population, or the small errors made when repeatedly measuring the same length. Most values sit close to some typical value, values further away become steadily rarer, and deviations above and below the centre are about equally likely. A normal distribution describes exactly this pattern: a single peak at the centre, symmetric tails, and a smooth fall-off.

Two numbers are enough to describe it completely:

Four bell-shaped curves. The three curves centred at zero have standard deviations 0.5, 1 and 2: the smaller the standard deviation, the taller and narrower the curve. A fourth curve with standard deviation 1 is shifted to be centred at 2 and has the same shape as the standard one.
Normal densities for several parameter values. Changing μ shifts the curve; changing σ makes it wider or narrower. Every curve encloses an area of exactly 1, so a narrower curve must also be taller.

Definition

A continuous random variable XX has a normal distribution with mean μ\mu and variance σ2\sigma^2, written X∼N(μ,σ2)X \sim \mathcal{N}(\mu, \sigma^2), if its probability density function is

f(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2),−∞<x<∞.f(x) = \frac{1}{\sigma\sqrt{2\pi}} \exp\!\left( -\frac{(x-\mu)^2}{2\sigma^2} \right), \qquad -\infty < x < \infty.

Here:

Note the convention: the second parameter in N(μ,σ2)\mathcal{N}(\mu, \sigma^2) is the variance, not the standard deviation. Some books and software use σ\sigma instead, so always check which one is meant.

Reading the formula

The formula has two parts.

The density is not a probability

For a continuous variable, f(x)f(x) is a density, not a probability. The probability that XX lands in an interval is the area under the density over that interval:

P(a≤X≤b)=∫abf(x) dx.P(a \le X \le b) = \int_a^b f(x)\, dx.

The probability that XX equals any single exact value is 00. A density value can even exceed 11: the curve with σ=0.5\sigma = 0.5 in the figure above peaks at about 0.800.80, and a curve with σ=0.1\sigma = 0.1 peaks at about 44.

Mean and variance

If X∼N(μ,σ2)X \sim \mathcal{N}(\mu, \sigma^2), then

E⁡[X]=μ,Var⁡(X)=σ2.\E[X] = \mu, \qquad \Var(X) = \sigma^2.

So the parameters are exactly the expected value and the variance of the distribution. Because the density is symmetric about μ\mu, the mean, median, and mode are all equal to μ\mu.

The standard normal distribution

The normal distribution with μ=0\mu = 0 and σ=1\sigma = 1 is called the standard normal distribution. A standard normal variable is usually written ZZ, its density φ(z)\varphi(z) and its cumulative distribution function (CDF) Φ(z)=P(Z≤z)\Phi(z) = P(Z \le z).

Any normal variable can be converted to a standard normal one by subtracting its mean and dividing by its standard deviation:

Z=X−μσ∼N(0,1).Z = \frac{X - \mu}{\sigma} \sim \mathcal{N}(0, 1).

The value z=(x−μ)/σz = (x - \mu)/\sigma is called a z-score. It says how many standard deviations xx lies above (positive) or below (negative) the mean. This standardization is why one table of Φ\Phi, or one software function, is enough to compute probabilities for every normal distribution:

P(X≤x)=Φ ⁣(x−μσ).P(X \le x) = \Phi\!\left(\frac{x - \mu}{\sigma}\right).

There is no formula for Φ\Phi in terms of elementary functions; it is computed numerically. In Python, for example, scipy.stats.norm.cdf(z) returns Φ(z)\Phi(z).

The 68–95–99.7 rule

For every normal distribution, the probability of landing within a fixed number of standard deviations of the mean is the same:

Interval Probability
μ±1σ\mu \pm 1\sigma about 68.3%
μ±2σ\mu \pm 2\sigma about 95.4%
μ±3σ\mu \pm 3\sigma about 99.7%
A standard bell curve with nested shaded bands. About 68.3 percent of the area lies within one standard deviation of the mean, 95.4 percent within two, and 99.7 percent within three.
The 68–95–99.7 rule. Darker bands are closer to the mean. The percentages are areas under the curve and hold for every normal distribution.

This gives a quick sense of scale. A value two standard deviations from the mean is unusual (it happens about 1 time in 22), and a value more than three standard deviations away is rare (about 1 time in 370). The often-quoted “95% within 2 standard deviations” is a rounding: the exact multiplier for 95% is 1.961.96.

Worked example

Scores on a standardized test are approximately normal with mean μ=500\mu = 500 and standard deviation σ=100\sigma = 100. What fraction of test-takers score below 650?

Step 1: standardize. The z-score of 650 is

z=650−500100=1.5.z = \frac{650 - 500}{100} = 1.5.

So 650 is one and a half standard deviations above the mean.

Step 2: look up the standard normal CDF. Φ(1.5)≈0.9332\Phi(1.5) \approx 0.9332. About 93.3% of test-takers score below 650, and about 6.7% score above it.

A second question: what fraction score between 400 and 700? The z-scores are −1-1 and 22, so

P(400≤X≤700)=Φ(2)−Φ(−1)≈0.9772−0.1587=0.8186.P(400 \le X \le 700) = \Phi(2) - \Phi(-1) \approx 0.9772 - 0.1587 = 0.8186.

About 82% of test-takers score in this range. As a check, the answer lies between the 68% for ±1σ\pm 1\sigma and the 95% for ±2σ\pm 2\sigma, which makes sense for an interval running from −1σ-1\sigma to +2σ+2\sigma.

Where the normal distribution appears

Common misunderstandings

“Most data are normal.” Many are not. Incomes, waiting times, and counts of rare events are typically skewed; some data have heavier tails than a normal distribution allows. The normal distribution is a useful model, not a law of nature. Look at a histogram before assuming it.

“The normal distribution can produce any value, so it cannot model heights.” Strictly, a normal variable can take negative values, which a height cannot. But if the mean is many standard deviations above zero, the probability of a negative value is so small that the model is still excellent in practice.

“The density at a point is the probability of that point.” It is not; see above. For a continuous variable, only areas are probabilities.

“Within 2σ means exactly 95%.” It means about 95.4%. The 95% interval is μ±1.96σ\mu \pm 1.96\sigma.

Further reading