Confidence Intervals
A range of plausible values for a parameter, produced by a procedure that captures the true value in a stated fraction of repeated samples.
Prerequisites: Sampling Distributions, Central Limit Theorem, Normal Distribution.
A confidence interval is a range of values, computed from a sample, that is meant to contain an unknown population quantity such as a mean or a proportion. A 95% confidence interval comes from a procedure that, over many repeated samples, produces intervals containing the true value 95% of the time.
A single point estimate, such as a sample mean of 23.4 hours, says nothing about how precise it is. A confidence interval, such as 21.7 to 25.1 hours, reports the estimate together with its uncertainty, which is usually what a reader actually needs.
Intuition
Imagine throwing a ring at a fixed peg that you cannot see. Each throw is one sample; where the ring lands is the interval computed from that sample. A good thrower rings the peg on 95% of throws. After a single throw, the ring either surrounds the peg or it does not; the “95%” describes the thrower’s skill, not that particular throw.
The figure shows this with simulated data. The true mean is known to be , which is never the case in practice. Fifty samples were drawn, and a 95% confidence interval was computed from each.
Every interval has the same width here because is known; only the centre moves from sample to sample. Most intervals cover ; a few, from samples whose means happened to land far from , do not. Nothing in the data of a single sample tells you which kind of interval you have.
The z-interval: a mean with known standard deviation
Let be a random sample from a population with unknown mean and known standard deviation . Assume either that the population is normal or that is large enough for the central limit theorem to apply. Then the sample mean is approximately normal with mean and standard error , so the standardized mean
has a standard normal distribution. Since , we can write
Multiplying through by and rearranging the inequalities to isolate gives the same event in a different form:
Read this carefully: is a fixed number, and the endpoints are random because is. The statement is about the random interval capturing the fixed . Plugging in the observed gives the 95% z-interval
where the critical value for 95% confidence. For other confidence levels, is chosen so that the central area under the standard normal density between and equals the confidence level: for 90% and for 99%. The quantity is called the margin of error.
The t-interval: a mean with unknown standard deviation
In practice is almost never known, so we replace it by the sample standard deviation . This adds uncertainty, because itself varies from sample to sample. To account for it, the critical value comes from the distribution with degrees of freedom instead of the standard normal. The distribution is bell-shaped and centred at zero like the standard normal, but has heavier tails, so its critical values are larger. The t-interval is
where cuts off an area of 2.5% in each tail of the distribution with degrees of freedom (for 95% confidence). For , compared with ; as grows, approaches . The t-interval is exact when the population is normal and is a good approximation for moderately non-normal populations when is not small. For strongly skewed data and small it can be unreliable.
The interval for a proportion
If of independent trials are successes, the sample proportion estimates the population proportion . Its standard error is , which we estimate by plugging in . The resulting Wald interval is
It relies on the normal approximation to the binomial distribution and works well when there are plenty of both successes and failures (a common rule of thumb asks for at least 10 of each). It performs poorly for small or for near or : its actual coverage can be well below the nominal 95%. In the extreme case of successes in trials it gives the interval from to , claiming complete certainty. Better-behaved alternatives such as the Wilson interval are preferred in those situations.
Worked example
A manufacturer tests batteries and records their lifetimes. The sample mean is hours and the sample standard deviation is hours. Lifetimes in this kind of test are roughly normal. Find a 95% confidence interval for the mean lifetime of all batteries of this type.
Step 1: estimated standard error.
Step 2: critical value. With degrees of freedom, the 97.5th percentile of the distribution is (from a table or software, for example scipy.stats.t.ppf(0.975, 15)).
Step 3: margin of error. hours.
Step 4: interval.
Using instead of would give the narrower interval 21.8 to 25.0 hours, which overstates the precision because it ignores the uncertainty in . A 99% interval uses and is wider: about 21.0 to 25.8 hours.
A proportion. In a survey, 540 of 1000 randomly chosen people support a proposal. Then , the estimated standard error is , and the 95% Wald interval is , or about to . Here there are hundreds of successes and failures, so the Wald interval is reliable.
What makes an interval wider or narrower
The margin of error is (critical value) × (standard deviation) / , so three things control the width:
- Confidence level. Higher confidence needs a larger critical value and a wider interval. A 100% interval would have to include every possible value and would be useless.
- Variability in the population. A larger (or ) gives a wider interval.
- Sample size. The width shrinks in proportion to . Quadrupling the sample size halves the width.
Interpreting a confidence interval
This is the part people most often get wrong.
In the frequentist framework used here, the parameter is a fixed, unknown number; it is not random. The 95% describes the procedure: if you repeated the whole study many times, computing a new interval each time, about 95% of those intervals would contain . This long-run fraction is called the coverage of the procedure.
Once you have computed a specific interval, such as 21.7 to 25.1 hours, it either contains or it does not. In this framework it is not correct to say “there is a 95% probability that lies between 21.7 and 25.1”. A correct phrasing is: “we are 95% confident that the interval from 21.7 to 25.1 hours contains the mean lifetime”, where “95% confident” is shorthand for “this interval was produced by a method that succeeds 95% of the time”.
Bayesian credible intervals do make direct probability statements about a parameter, but they require a prior distribution and answer a different question.
Confidence intervals are closely tied to hypothesis testing: for the z- and t-intervals above, the 95% interval contains exactly those values of the mean that the matching two-sided test at the 5% significance level would not reject.
Common misunderstandings
“There is a 95% chance the true mean is in this interval.” In the frequentist sense, no: the 95% refers to the long-run success rate of the procedure, not to one computed interval; see above.
“95% of the data lie inside the interval.” A confidence interval for a mean describes uncertainty about the mean, not the spread of individual values. With batteries, the interval for is about 3.4 hours wide, while individual lifetimes vary with standard deviation 3.2 hours, so many batteries fall outside it. As grows the interval shrinks toward a point; the spread of the data does not.
“The interval accounts for all sources of error.” It accounts only for random sampling variability under the assumed model. A biased sample, measurement error, or a wrong model can make an interval confidently wrong.
“Overlapping intervals mean no significant difference.” Two 95% intervals can overlap somewhat even when a test comparing the two means finds a significant difference. To compare two groups, compute an interval for the difference directly.
Further reading
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; introduces confidence intervals for proportions and means with many examples.
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Particularly clear on what the confidence level does and does not mean.
- NIST/SEMATECH, e-Handbook of Statistical Methods, https://www.itl.nist.gov/div898/handbook/. Practical reference for confidence limits for means and proportions.