P-values and Statistical Significance
The probability, assuming the null hypothesis and the model are true, of a result at least as extreme as the one observed.
Prerequisites: Hypothesis Testing, Bernoulli and Binomial Distributions.
A p-value is the probability, calculated assuming that the null hypothesis and the rest of the statistical model are true, of getting a test statistic at least as extreme as the one actually observed. A small p-value says the data would be unusual if the null hypothesis were true; a large one says they would not.
P-values are among the most reported numbers in science, and among the most misread. This article defines them, works through an example, and then covers what a p-value does not tell you, following the American Statistical Association’s 2016 statement on p-values.
Intuition
You suspect a coin is biased, so you toss it 100 times and get 60 heads. A fair coin would give about 50 on average, but never exactly 50 every time. Is 60 far enough from 50 to make you doubt the coin is fair?
The p-value answers a precise version of that question: if the coin were fair, how often would 100 tosses give a result at least as lopsided as this one? Here “at least as lopsided” means 60 or more heads, or, by symmetry, 40 or fewer. If such results were common under fairness, 60 heads is not surprising. If they were rare, either something rare happened, or the coin is not fair (or some other part of the model, such as independence of tosses, is wrong).
Definition
Consider a hypothesis test with null hypothesis and test statistic , where values of further from what predicts count as more extreme. Let be the value of calculated from the observed data. The p-value is
where the probability is computed using the distribution of under , the null distribution.
What “at least as extreme” means depends on the alternative hypothesis :
- Two-sided alternatives count departures in both directions. For a statistic whose null distribution is symmetric about the value predicts, the two-sided p-value is twice the one-tail probability beyond the observed value.
- One-sided alternatives count only departures in the direction stated by .
Three things to notice about the definition:
- The p-value is computed assuming is true. It is a probability about data, not about hypotheses.
- It depends on the whole model, not just : for the coin, that the tosses are independent and have the same probability of heads. A small p-value can signal a problem with any of these assumptions.
- It involves results at least as extreme as the one observed, not just the observed result itself. Any single exact outcome may be unlikely; what matters is how far into the tail it lies.
Worked example
Toss a coin times and observe 60 heads. Let be the number of heads. Test
where is the probability of heads on a single toss. (We write rather than the usual to avoid a clash with the p-value.)
Step 1: the null distribution. If is true and the tosses are independent, has a binomial distribution with trials and success probability . It is centred at and symmetric about .
Step 2: which outcomes are at least as extreme? The observed count is away from . Outcomes at least as far from are and .
Step 3: add up their probabilities (exact calculation). Using the binomial PMF,
By symmetry is the same, so the exact two-sided p-value is
If the coin were fair, about 5.7% of experiments of 100 tosses would be at least this lopsided.
A normal approximation. Many textbooks instead approximate the binomial by a normal distribution with the same mean, , and the same standard deviation, . The z statistic is , and the two-sided p-value is , where is the standard normal CDF. With a continuity correction (using in place of , since the count is discrete), and the p-value is about , close to the exact answer.
Notice what just happened. The exact p-value is just above and the uncorrected approximation is just below it. The evidence is the same in both cases; only the arithmetic differs. This is a good reason not to treat as a cliff.
Statistical significance
Before collecting data, the analyst may choose a significance level , conventionally . If , the result is called statistically significant at level , and is rejected. This rule is identical to checking whether the test statistic falls in the rejection region of the corresponding test.
The rule has a useful guarantee. If and the model are true, the probability of getting is at most . So declaring significance whenever produces false alarms in at most about 5% of tests of true null hypotheses. (For a discrete statistic such as a binomial count, the rate is typically a little below .)
“Statistically significant” is a technical term. It means only that the data are hard to reconcile with under the model. It does not mean important, large, or proven.
Common misunderstandings
In 2016 the American Statistical Association published a statement on p-values (Wasserstein and Lazar, listed below) to address widespread misuse. Its six principles say, in brief: p-values can indicate how incompatible data are with a specified model; they do not measure the probability that a hypothesis is true, or the probability that the data were produced by chance alone; conclusions should not rest only on whether a p-value passes a threshold; proper inference requires full reporting and transparency; a p-value does not measure the size or importance of an effect; and by itself a p-value is not a good measure of evidence about a model or hypothesis. The misunderstandings below follow from these principles.
“The p-value is the probability that H₀ is true.” It is not. The p-value is a probability about the data computed assuming is true, roughly . The probability that is true given the data is the reverse conditional probability, and getting it requires Bayes’ theorem and a prior probability for . In the coin example, does not mean there is a 5.7% chance the coin is fair.
“The p-value is the probability that the result is due to chance.” This is the same mistake in different words. The calculation already assumes that chance alone (under ) is at work; it cannot also give the probability of that assumption.
“A small p-value means a large or important effect.” Statistical significance is not practical significance. The p-value depends on the sample size as well as on the effect. Toss a coin a million times and get 50.2% heads: the z statistic is and the two-sided p-value is about , yet a bias of 0.2 percentage points is irrelevant for most purposes. Always report the effect size, the estimated magnitude of the effect, preferably with a confidence interval.
“p > 0.05 shows there is no effect.” A large p-value means the data are reasonably consistent with , but they may be just as consistent with a substantial effect, especially in a small study with low power. Absence of evidence is not evidence of absence. A confidence interval shows which effect sizes remain plausible.
“p = 0.049 and p = 0.051 lead to opposite conclusions.” They represent nearly identical evidence. The worked example above produced p-values on either side of 0.05 from the same data, depending only on the method of calculation. Report the p-value itself rather than only “significant” or “not significant”.
“If I run many tests, a small p-value is still meaningful on its own.” Each test of a true null hypothesis at has a 5% chance of a false alarm, and those chances add up. With 20 independent tests of true null hypotheses, the probability that at least one gives is . Trying many analyses, outcomes, subgroups, or stopping rules and reporting only the ones that “worked” is called p-hacking or data dredging, and it makes reported p-values meaningless. Remedies include deciding the analysis in advance, reporting every test performed, and adjusting for multiple comparisons (for example, the Bonferroni correction, which tests each of hypotheses at level ).
Further reading
- Ronald L. Wasserstein and Nicole A. Lazar, “The ASA Statement on p-Values: Context, Process, and Purpose”, The American Statistician 70(2), 129–133, 2016. Short and readable; the source of the six principles summarized above.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; introduces p-values alongside hypothesis tests with many examples.
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Includes a careful discussion of what significance tests can and cannot show.