Bayes' Theorem

A rule for reversing a conditional probability, turning P(evidence | hypothesis) into P(hypothesis | evidence).

Prerequisites: Conditional Probability.

Bayes’ theorem is a formula for turning one conditional probability around: from the probability of the evidence given a hypothesis, P(B∣A)P(B \mid A), it computes the probability of the hypothesis given the evidence, P(A∣B)P(A \mid B). It is the mathematical rule for updating a belief when new information arrives, and it explains why a positive result from a fairly accurate medical test can still leave the disease unlikely.

Intuition

A test for a disease is described by how it behaves in people whose status is known: how often it is positive for people who have the disease, and how often for people who do not. A patient, however, wants the reverse: given a positive result, how likely is it that I have the disease?

These two questions have different answers, and the gap between them depends on how common the disease is in the first place. If a disease is rare, even a small false-positive rate among the many healthy people can produce more positive results than the disease itself does among the few sick people. Bayes’ theorem keeps track of both sources of positive results and weighs them correctly.

Derivation

Bayes’ theorem follows in two lines from the definition of conditional probability. For events AA and BB with P(A)>0P(A) > 0 and P(B)>0P(B) > 0, the multiplication rule gives the probability of “both AA and BB” in two ways:

P(A∩B)=P(A∣B) P(B)=P(B∣A) P(A).P(A \cap B) = P(A \mid B)\, P(B) = P(B \mid A)\, P(A).

Dividing the two right-hand expressions by P(B)P(B) gives Bayes’ theorem:

P(A∣B)=P(B∣A) P(A)P(B).P(A \mid B) = \frac{P(B \mid A)\, P(A)}{P(B)}.

The denominator P(B)P(B) is usually not known directly. The law of total probability computes it by splitting into the cases “AA” and “not AA”:

P(B)=P(B∣A) P(A)+P(B∣Ac) P(Ac).P(B) = P(B \mid A)\,P(A) + P(B \mid A^c)\,P(A^c).

Substituting gives the form most often used in practice:

P(A∣B)=P(B∣A) P(A)P(B∣A) P(A)+P(B∣Ac) P(Ac).P(A \mid B) = \frac{P(B \mid A)\,P(A)}{P(B \mid A)\,P(A) + P(B \mid A^c)\,P(A^c)}.

More generally, if A1,…,AkA_1, \dots, A_k are mutually exclusive hypotheses, exactly one of which is true, then

P(Aj∣B)=P(B∣Aj) P(Aj)∑i=1kP(B∣Ai) P(Ai).P(A_j \mid B) = \frac{P(B \mid A_j)\,P(A_j)}{\sum_{i=1}^k P(B \mid A_i)\,P(A_i)}.

Prior, likelihood, and posterior

When AA is a hypothesis and BB is observed evidence, each piece of the formula has a name:

Read this way, Bayes’ theorem says: posterior is proportional to likelihood times prior. Evidence raises the probability of hypotheses that predicted it well and lowers the probability of hypotheses that predicted it poorly. The word “likelihood” here is related to, but not the same as, the likelihood function used in estimation.

Worked example: a medical test

A condition affects 1% of a population. A screening test for it has these properties:

A randomly chosen person tests positive. What is the probability that they have the condition?

Let DD be “has the condition” and ++ be “tests positive”. We know

P(D)=0.01,P(+∣D)=0.90,P(+∣Dc)=0.09.P(D) = 0.01, \qquad P(+ \mid D) = 0.90, \qquad P(+ \mid D^c) = 0.09.

Step 1: the probability of a positive test. By the law of total probability,

P(+)=0.90×0.01+0.09×0.99=0.0090+0.0891=0.0981.P(+) = 0.90 \times 0.01 + 0.09 \times 0.99 = 0.0090 + 0.0891 = 0.0981.

Step 2: Bayes’ theorem.

P(D∣+)=P(+∣D) P(D)P(+)=0.00900.0981≈0.092.P(D \mid +) = \frac{P(+ \mid D)\,P(D)}{P(+)} = \frac{0.0090}{0.0981} \approx 0.092.

Despite the test detecting 90% of cases, a positive result means only about a 9.2% chance of having the condition. More than 90% of positive results are false alarms.

The same calculation with natural frequencies

The result is easier to believe when probabilities are replaced by counts of people. Imagine 10,000 people tested.

A tree splitting 10,000 people into 100 with the condition and 9,900 without. Of the 100, 90 test positive; of the 9,900, 891 test positive. So only 90 of the 981 positive tests, about 9 percent, come from people with the condition.
Natural-frequency tree for 10,000 people, with prevalence 1%, sensitivity 90%, and false-positive rate 9%. Positive tests: 90 true positives and 891 false positives.

Altogether 90+891=98190 + 891 = 981 people test positive, and only 90 of them have the condition:

P(D∣+)=90981≈0.092.P(D \mid +) = \frac{90}{981} \approx 0.092.

The reason is now visible: the false-positive rate is small, but it applies to a very large group. Nine percent of 9,900 healthy people is far more than 90% of 100 sick people.

What a negative result means

The same tree answers other questions. Of the 10+9,009=9,01910 + 9{,}009 = 9{,}019 people who test negative, only 10 have the condition, so P(D∣−)=10/9,019≈0.0011P(D \mid -) = 10/9{,}019 \approx 0.0011. A negative result is very reassuring, because the condition was rare to begin with.

Updating twice

Suppose the person who tested positive takes the test again, and assume the two results are independent given the person’s true status. The posterior from the first test, about 0.0920.092, becomes the prior for the second. Of the 981 positives, about 0.90×90=810.90 \times 90 = 81 sick and 0.09×891≈800.09 \times 891 \approx 80 healthy people would test positive again, so

P(D∣two positives)≈8181+80.19≈0.50.P(D \mid \text{two positives}) \approx \frac{81}{81 + 80.19} \approx 0.50.

Two positive results raise the probability to about one half. (In practice, repeated tests on the same person are often not independent, so a different, confirmatory test is usually used.)

Common misunderstandings

The base-rate fallacy. Ignoring the prior P(A)P(A), here the 1% prevalence, and judging the posterior from the test’s accuracy alone. The 90% sensitivity tempts people to answer “about 90%”, but the correct answer is about 9%. Whenever a condition is rare, the base rate matters a great deal.

“P(positive | disease) is the same as P(disease | positive).” These are conditional probabilities in opposite directions. They can be very different, as the example shows. Confusing them is also known as the prosecutor’s fallacy when it appears in court: the probability of the evidence given innocence is not the probability of innocence given the evidence.

“Bayes’ theorem is only for Bayesian statistics.” The theorem itself is an uncontroversial consequence of the definition of conditional probability, accepted by every school of statistics. What is debated is whether it is appropriate to put prior probabilities on unknown parameters, which is the starting point of Bayesian inference.

“A p-value is the probability that the null hypothesis is true.” A p-value is computed assuming the null hypothesis is true, so it is a probability of data given a hypothesis, not of a hypothesis given data. Turning one into the other would require Bayes’ theorem and a prior.

Further reading