Bernoulli and Binomial Distributions

The distributions of a single yes/no trial and of the number of successes in a fixed number of independent trials.

Prerequisites: Probability Distributions, Statistical Independence, Expected Value.

The Bernoulli distribution describes a single trial with two outcomes, “success” or “failure”, such as one coin toss or one free throw. The binomial distribution describes the number of successes in a fixed number of such trials when the trials are independent and each has the same chance of success.

Together they are the basic model for counting yes/no outcomes: how many patients respond to a treatment, how many of 20 emails are spam, how many of 10 free throws go in. Both are discrete distributions, so they assign probabilities to separate values through a probability mass function.

Intuition

Suppose a basketball player makes each free throw with probability 0.70.7 and takes 10 shots. The number of baskets could be anything from 0 to 10, but not all values are equally likely. Around 7 is typical. Exactly 10 requires every shot to go in, which is rare. Exactly 7 can happen in many ways (any 7 of the 10 shots could be the successful ones), which is why values near the middle collect most of the probability.

The binomial distribution makes this counting precise. It has two ingredients: the probability of one particular sequence of successes and failures, and the number of sequences that give the same total.

The Bernoulli distribution

A random variable XX has a Bernoulli distribution with parameter pp, written X∼Bernoulli(p)X \sim \text{Bernoulli}(p), if it takes the value 11 (success) with probability pp and the value 00 (failure) with probability 1−p1 - p, where 0≤p≤10 \le p \le 1:

P(X=1)=p,P(X=0)=1−p.P(X = 1) = p, \qquad P(X = 0) = 1 - p.

Coding success as 11 and failure as 00 is a convenient choice: the variable then counts successes in one trial, and averages of such variables are proportions.

Its expected value and variance follow directly from the definitions:

E⁡[X]=1⋅p+0⋅(1−p)=p,\E[X] = 1 \cdot p + 0 \cdot (1-p) = p,

and since X2=XX^2 = X when XX is 00 or 11, E⁡[X2]=p\E[X^2] = p, so

Var⁡(X)=E⁡[X2]−(E⁡[X])2=p−p2=p(1−p).\Var(X) = \E[X^2] - (\E[X])^2 = p - p^2 = p(1-p).

The variance is largest at p=0.5p = 0.5, where the outcome is most uncertain, and zero at p=0p = 0 or p=1p = 1, where there is no uncertainty at all.

The binomial distribution

Let XX be the number of successes in nn trials, where

  1. the number of trials nn is fixed in advance,
  2. each trial has two outcomes, success or failure,
  3. every trial has the same probability of success pp, and
  4. the trials are independent.

Then XX has a binomial distribution with parameters nn (a positive integer) and pp (with 0≤p≤10 \le p \le 1), written X∼Binomial(n,p)X \sim \text{Binomial}(n, p). Its probability mass function is

P(X=k)=(nk)pk(1−p)n−k,k=0,1,…,n.P(X = k) = \binom{n}{k} p^k (1-p)^{n-k}, \qquad k = 0, 1, \dots, n.

Reading the formula

Multiplying gives the total probability of all the sequences with kk successes. The Bernoulli distribution is the special case n=1n = 1.

Mean and variance

The cleanest way to find the mean and variance is to write XX as a sum. Let Xi=1X_i = 1 if trial ii is a success and 00 otherwise. Each Xi∼Bernoulli(p)X_i \sim \text{Bernoulli}(p), and

X=X1+X2+⋯+Xn.X = X_1 + X_2 + \dots + X_n.

Expectation is linear, so the expected value of a sum is the sum of the expected values:

E⁡[X]=E⁡[X1]+⋯+E⁡[Xn]=np.\E[X] = \E[X_1] + \dots + \E[X_n] = np.

Because the trials are independent, the variance of the sum is also the sum of the variances:

Var⁡(X)=Var⁡(X1)+⋯+Var⁡(Xn)=np(1−p).\Var(X) = \Var(X_1) + \dots + \Var(X_n) = np(1-p).

The standard deviation is np(1−p)\sqrt{np(1-p)}. As a sanity check, with p=0p = 0 or p=1p = 1 the count is certain (00 or nn) and the variance formula correctly gives 00.

How the shape depends on n and p

Top: three binomial distributions with 10 trials. With p = 0.2 the probabilities pile up near 2 and trail off to the right; with p = 0.5 they are symmetric around 5; with p = 0.8 they pile up near 8 and trail off to the left. Bottom: with 40 trials and p = 0.5 the stems form a symmetric, bell-like shape centred at 20.
Binomial probability mass functions, drawn as stems because only whole numbers are possible. Top: n = 10 with p = 0.2, 0.5 and 0.8 (stems shifted slightly sideways so they do not overlap). Bottom: n = 40, p = 0.5.

Worked example

A player makes each free throw with probability p=0.7p = 0.7, independently of the other shots, and takes n=10n = 10 shots. Let XX be the number made, so X∼Binomial(10,0.7)X \sim \text{Binomial}(10, 0.7).

Step 1: mean and spread. E⁡[X]=10×0.7=7\E[X] = 10 \times 0.7 = 7 and Var⁡(X)=10×0.7×0.3=2.1\Var(X) = 10 \times 0.7 \times 0.3 = 2.1, so the standard deviation is 2.1≈1.45\sqrt{2.1} \approx 1.45 baskets.

Step 2: exactly 8 baskets. The number of ways to choose which 8 of the 10 shots go in is (108)=45\binom{10}{8} = 45. Each such sequence has probability 0.78×0.320.7^8 \times 0.3^2. So

P(X=8)=45×0.78×0.32≈45×0.05765×0.09≈0.2335.P(X = 8) = 45 \times 0.7^8 \times 0.3^2 \approx 45 \times 0.05765 \times 0.09 \approx 0.2335.

Step 3: at least 8 baskets. Add the probabilities of 8, 9, and 10:

P(X=9)=10×0.79×0.3≈0.1211,P(X=10)=0.710≈0.0282.P(X = 9) = 10 \times 0.7^9 \times 0.3 \approx 0.1211, \qquad P(X = 10) = 0.7^{10} \approx 0.0282.

P(X≥8)≈0.2335+0.1211+0.0282=0.3828.P(X \ge 8) \approx 0.2335 + 0.1211 + 0.0282 = 0.3828.

So the player makes 8 or more about 38% of the time. Although 7 is the single most likely value, its probability is only about 0.270.27: no single count is very likely.

When the assumptions fail

The binomial model needs all four conditions above. Common ways they break:

Approximations

Normal approximation. When nn is large and pp is not too close to 00 or 11, the binomial distribution is close to a normal distribution with the same mean npnp and variance np(1−p)np(1-p). This is a consequence of the central limit theorem, since XX is a sum of nn independent Bernoulli variables. A common rule of thumb is to require np≥10np \ge 10 and n(1−p)≥10n(1-p) \ge 10. Because the binomial is discrete, the approximation is better with a continuity correction: treat the integer kk as the interval from k−0.5k - 0.5 to k+0.5k + 0.5. For 100 tosses of a fair coin, P(X≤45)P(X \le 45) is about 0.18410.1841 exactly, and the normal approximation Φ((45.5−50)/5)\Phi\big((45.5 - 50)/5\big) gives about 0.18410.1841 as well.

Poisson approximation. When nn is large and pp is small, so that successes are rare, the binomial is close to a Poisson distribution with mean λ=np\lambda = np.

Common misunderstandings

“The most likely count happens most of the time.” In the example, 7 is the most likely number of baskets, but it happens only about 27% of the time. With larger nn the probability of any single exact count gets smaller still.

“Any count of successes is binomial.” Only if the number of trials is fixed, the trials are independent, and the success probability is the same for each. Check these before using the formulas, particularly the standard deviation, which is too small when the assumptions fail.

“Bernoulli and binomial are unrelated distributions.” A Bernoulli variable is a binomial variable with n=1n = 1, and every binomial variable is a sum of independent Bernoulli variables.

“The variance is np.” That is the Poisson variance. The binomial variance np(1−p)np(1-p) is always smaller than npnp when 0<p0 < p; the two are close only when pp is small.

Further reading