Hypothesis Testing
A procedure for deciding whether data are inconsistent enough with a default claim about a population to reject that claim.
Prerequisites: Sampling Distributions, Normal Distribution, Central Limit Theorem.
A hypothesis test is a procedure for deciding whether sample data provide enough evidence against a default claim about a population, called the null hypothesis. If data like ours would be very unusual were the null hypothesis true, we reject it; otherwise we do not.
Hypothesis tests are used to check whether a production process is on target, whether a new treatment differs from a placebo, or whether two groups differ. They give a disciplined answer to the question “could this pattern plausibly be just random variation?”, with a known rate of false alarms.
Intuition
A machine is supposed to fill packages with 500 g of flour. You weigh 25 packages and find an average of 498.2 g. Is the machine underfilling, or is 1.8 g below target just the kind of difference that random variation produces anyway?
A hypothesis test answers by temporarily assuming that nothing is wrong (the true mean fill is exactly 500 g) and asking how surprising the observed average would be under that assumption. If it would be very surprising, the assumption looks doubtful and we reject it. If it would not be surprising, the data are consistent with the machine being on target, though they do not prove it.
The logic resembles a criminal trial. The defendant is presumed innocent (the null hypothesis), and the verdict is “guilty” only if the evidence is strong enough. A verdict of “not guilty” does not mean the defendant was shown to be innocent.
The ingredients of a test
Hypotheses. The null hypothesis is the default claim, usually “no effect” or “no difference”, stated in terms of a population parameter. The alternative hypothesis (also written ) is what we conclude if we reject . For the flour machine, with the true mean fill weight,
Hypotheses are statements about the population, never about the sample: there is no doubt that the sample mean is 498.2 g.
Test statistic. A number computed from the data that measures how far the data depart from . We write for its observed value. For a mean with known population standard deviation , the usual choice is the z statistic
where is the value of under (here 500), is the sample mean, and is its standard error. It counts how many standard errors the sample mean lies from .
Null distribution. The sampling distribution of the test statistic, computed on the assumption that is true. If the data are normal, or is large enough for the central limit theorem, then under the z statistic has a standard normal distribution.
Significance level. A threshold , chosen before looking at the data, for how often we are willing to reject when it is in fact true. The conventional choice is , but there is nothing special about it.
Rejection region. The set of values of the test statistic that lead to rejecting . It is chosen so that, under , the test statistic falls in it with probability . For a two-sided z test at , the rejection region is , because for a standard normal .
Decision. If falls in the rejection region, reject . Otherwise, do not reject .
An equivalent way to make the decision is through the p-value: the probability, computed assuming is true, of a test statistic at least as extreme as the one observed. Reject exactly when the p-value is at most . The p-value adds information the yes-or-no decision hides, namely how far into the tail the observation fell.
Worked example
Return to the flour machine. From long experience, individual fill weights vary with standard deviation g. A sample of packages has mean g. Test against at .
Step 1: standard error.
Step 2: test statistic.
The sample mean is 2.25 standard errors below the target.
Step 3: compare with the rejection region. Since , the statistic falls in the rejection region.
Step 4: decide and state the conclusion in context. Reject at the 5% level. The data give evidence that the machine’s mean fill differs from 500 g; the estimate suggests it is underfilling by roughly 1.8 g. The two-sided p-value is , where is the standard normal CDF.
If σ were unknown. We would replace by the sample standard deviation and compare the statistic with the distribution with degrees of freedom, whose two-sided 5% critical value is about instead of . This is the one-sample t test; it is the version used most often in practice.
Type I and type II errors
A test can be wrong in two ways.
| true | false | |
|---|---|---|
| Reject | Type I error (false alarm), probability | Correct decision, probability (the power) |
| Do not reject | Correct decision, probability | Type II error (missed effect), probability |
- A type I error is rejecting a true null hypothesis. Its probability is the significance level , which we control directly.
- A type II error is failing to reject a false null hypothesis. Its probability is written .
- The power of a test, , is the probability of correctly rejecting when it is false.
Power is not a single number: it depends on how far the truth is from . For the flour test above (, , ), the power is about 0.24 if the true mean is 499 g and about 0.71 if it is 498 g. Small departures are hard to detect. Power increases with the size of the true effect, with the sample size , and with , and decreases with the population standard deviation . Lowering to reduce false alarms, with fixed, raises : there is a trade-off between the two kinds of error. Studies are often planned by choosing large enough to give, say, 80% power for the smallest effect that would matter.
One-sided and two-sided tests
The alternative is two-sided: departures in either direction count as evidence against , and the rejection region is split between both tails.
If only one direction matters, for example if the regulator only cares about underfilling, a one-sided alternative can be used. All of is then placed in one tail: at the rejection region is . A one-sided test has more power in the chosen direction but cannot detect a departure in the other direction.
The choice between one- and two-sided tests must be made before seeing the data, based on the question being asked. Choosing the direction after seeing which way the data point effectively doubles the type I error rate.
Tests and confidence intervals
Two-sided tests and confidence intervals are two views of the same calculation. For the z test above, a two-sided test at level rejects exactly when lies outside the confidence interval .
In the flour example, the 95% interval is , or about 496.6 g to 499.8 g. It does not contain 500, matching the decision to reject. The interval also shows what the test alone does not: which values of are consistent with the data.
This exact correspondence holds when the test and the interval are built from the same statistic and the same standard error, as here. For some other procedures (for example the Wald interval for a proportion paired with a test that uses the null value in its standard error) the two can occasionally disagree near the boundary. A one-sided test corresponds to a one-sided confidence bound rather than to the usual two-sided interval.
Common misunderstandings
“Failing to reject H₀ means H₀ is true.” It means only that the data were not inconsistent enough with to reject it. A small sample may simply lack the power to detect a real effect. “Not guilty” is not “innocent”. The p-values article discusses this further under misinterpretations.
“Rejecting H₀ proves H₁.” A test controls the long-run rate of false rejections; it does not prove anything. About 5% of tests of true null hypotheses at will reject.
“α is the probability that H₀ is true after we reject it.” is the probability of rejecting given that is true. The probability that is true given a rejection is a different conditional probability and cannot be computed from alone.
“Statistically significant means important.” With a large enough sample, even a trivial difference, such as a fill weight 0.05 g below target, produces a significant result. Always report the estimated size of the effect, ideally with a confidence interval.
“The hypotheses can be adjusted after looking at the data.” Choosing the alternative, the significance level, or which test to run after seeing the data invalidates the stated error rates.
Further reading
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; a gentle introduction to the testing framework, errors and power.
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A compact mathematical treatment of tests, power, and the link to confidence sets.
- NIST/SEMATECH, e-Handbook of Statistical Methods, https://www.itl.nist.gov/div898/handbook/. Practical descriptions of common tests with worked industrial examples.