Conditional Probability
The probability of an event once we know that another event has occurred.
Prerequisites: Introduction to Probability.
The conditional probability of an event given an event , written , is the probability that happens when we already know that has happened. It is how probabilities are updated in the light of new information, and almost every realistic probability question involves it: the chance of rain given a cloudy morning, the chance of a disease given a positive test, the chance of default given a borrower’s income.
Intuition
Knowing that happened rules out every outcome outside . What remains is a smaller world, and within that world we ask how much of it belongs to .
Draw one card from a well-shuffled standard deck of 52 cards. The probability that it is a king is . Now suppose a friend peeks and tells you the card is a face card (a jack, queen, or king). There are only 12 face cards, and 4 of them are kings, so the probability of a king given this information is . The information shrank the set of possibilities from 52 cards to 12, and kings make up a larger share of the smaller set.
The picture below shows the same idea with areas. Think of the sample space as a rectangle of total area 1, and of each event as a region whose area is its probability. Conditioning on throws away everything outside ; the conditional probability of is the fraction of ’s area that is also inside .
Definition
Let and be events with . The conditional probability of given is
Here is the event “both and happen”. In words: of all the probability sitting on , what fraction also lies in ?
Some remarks:
- The condition is needed because we divide by . Conditioning on an event of probability zero needs a more careful treatment that is beyond this article.
- For a fixed , the function obeys all the axioms of probability. In particular, , and . Conditioning simply produces a new probability on a smaller world.
- When all outcomes are equally likely, the definition reduces to counting inside : , as in the card example.
The multiplication rule
Rearranging the definition gives a rule for the probability that two events both happen:
This is often the natural way to compute a joint probability step by step. For example, draw two cards without replacement. The probability that both are aces is
The second factor is a conditional probability: once an ace has been removed, 3 aces remain among 51 cards.
The law of total probability
Sometimes is hard to find directly but easy to find in each of several separate cases. Suppose the events form a partition of : they are mutually exclusive and together cover every outcome, so exactly one of them happens. Then
The simplest partition is and its complement :
This says that the overall probability of is a weighted average of its conditional probabilities in each case, weighted by how likely each case is. It follows from splitting into the disjoint pieces , adding their probabilities, and applying the multiplication rule to each piece.
Worked example
A survey of 200 households records whether each owns a dog and whether it owns a cat.
| Cat | No cat | Total | |
|---|---|---|---|
| Dog | 30 | 50 | 80 |
| No dog | 40 | 80 | 120 |
| Total | 70 | 130 | 200 |
Choose one of these households at random. Let be “owns a cat” and be “owns a dog”.
Step 1: unconditional probabilities. From the totals, and . Both owning a dog and a cat: . (These are the numbers in the figure above, with and .)
Step 2: condition on owning a dog. Restrict attention to the 80 dog-owning households. Of these, 30 own a cat:
Counting directly in the “Dog” row gives the same answer, .
Step 3: condition the other way. Among the 70 cat-owning households, 30 own a dog:
So . The numerator is the same, households, but the denominators differ: 80 dog owners versus 70 cat owners.
Step 4: check with the law of total probability. Among households without a dog, . Then
which matches Step 1.
Common misunderstandings
“P(A | B) is the same as P(B | A).” It is not, and confusing the two (sometimes called the confusion of the inverse) causes serious errors. The probability that a person who has a disease tests positive can be high while the probability that a person who tests positive has the disease is low. Bayes’ theorem is the rule that converts one into the other.
“P(A | B) is the same as P(A ∩ B).” The joint probability is measured relative to the whole sample space; the conditional probability is measured relative to only. In the example, but .
“Conditioning on B always changes the probability of A.” Not necessarily. If , learning that happened tells us nothing about . Such events are called independent.
“The vertical bar means division.” is read “the probability of given ”. The notation on its own is not an event, and is not .
Further reading
- Joseph K. Blitzstein and Jessica Hwang, Introduction to Probability, 2nd ed., CRC Press, 2019. Covers conditional probability in depth, with many worked examples and common pitfalls.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; introduces conditional probability through contingency tables.