Likelihood and Maximum Likelihood Estimation

The likelihood measures how well each parameter value explains the observed data; the maximum likelihood estimate is the value that explains them best.

Prerequisites: Point Estimation, Random Variables, Bernoulli and Binomial Distributions.

The likelihood of a parameter value is the probability (or probability density) that a model with that parameter value assigns to the data you actually observed. Maximum likelihood estimation picks, as the estimate of the parameter, the value under which the observed data would have been most probable.

It is the most widely used general recipe for building estimators. Given almost any probability model, from a coin toss to logistic regression, maximum likelihood tells you how to turn data into parameter estimates, and in large samples those estimates are usually as precise as any reasonable alternative.

Intuition

You toss a coin 10 times and see 7 heads. What is a sensible guess for pp, the probability of heads?

Try a few candidate values and ask how likely the observed result would be under each:

The data are more than twice as probable under p=0.7p = 0.7 as under a fair coin, and much more probable than under p=0.9p = 0.9. Maximum likelihood formalizes this comparison: compute that probability for every candidate pp, and choose the pp that makes it largest.

Two panels plotted against p from 0 to 1. The likelihood is a single hump that peaks at p = 0.7 with height about 0.27 and is near zero below 0.3 and close to 1. The log-likelihood is a concave curve that also peaks at p = 0.7.
Likelihood (left) and log-likelihood (right) of the probability of heads p after observing 7 heads in 10 tosses. Both are maximized at the same point, p̂ = 0.7, marked with a dot.

The likelihood function

Suppose the data x1,…,xnx_1, \dots, x_n are modelled as observations of random variables whose distribution depends on an unknown parameter θ\theta. Write p(x∣θ)p(x \mid \theta) for the probability mass function of a discrete observation, or f(x∣θ)f(x \mid \theta) for the density of a continuous one. If the observations are independent, the joint probability (or density) of the whole data set is the product of the individual ones.

The likelihood function is that same expression, read as a function of θ\theta with the data held fixed:

L(θ)=∏i=1np(xi∣θ)(discrete),L(θ)=∏i=1nf(xi∣θ)(continuous).L(\theta) = \prod_{i=1}^n p(x_i \mid \theta) \quad \text{(discrete)}, \qquad L(\theta) = \prod_{i=1}^n f(x_i \mid \theta) \quad \text{(continuous)}.

The formula is identical to the probability model. What changes is the question. The model asks, “for a fixed θ\theta, how probable are different data sets?” The likelihood asks, “for these fixed data, how well does each θ\theta explain them?”

The likelihood is not a probability distribution for the parameter

L(θ)L(\theta) is not the probability that θ\theta is the true value. It is a probability of the data, evaluated at many different values of θ\theta. Nothing forces it to add or integrate to 11 over θ\theta. For the coin example,

∫01(107)p7(1−p)3 dp=111≈0.091,\int_0^1 \binom{10}{7} p^7 (1-p)^3 \, dp = \frac{1}{11} \approx 0.091,

not 11. Only ratios of likelihoods are meaningful: L(0.7)/L(0.5)≈2.28L(0.7)/L(0.5) \approx 2.28 says the data are about 2.3 times as probable if p=0.7p = 0.7 as if p=0.5p = 0.5. Turning this into a probability statement about pp itself requires a prior distribution and Bayes’ theorem, which is a different approach to inference.

The log-likelihood

In practice we almost always work with the log-likelihood

ℓ(θ)=log⁡L(θ)=∑i=1nlog⁡p(xi∣θ),\ell(\theta) = \log L(\theta) = \sum_{i=1}^n \log p(x_i \mid \theta),

(with ff in place of pp for continuous data). There are three reasons:

Maximum likelihood estimation

The maximum likelihood estimate (MLE) of θ\theta is the value that maximizes the likelihood:

θ^=arg⁡max⁡θL(θ)=arg⁡max⁡θℓ(θ).\hat{\theta} = \arg\max_{\theta} L(\theta) = \arg\max_{\theta} \ell(\theta).

Here arg⁡max⁡\arg\max means “the value of θ\theta at which the maximum is attained”. When ℓ\ell is smooth and the maximum is not at the edge of the allowed parameter values, it can usually be found by setting the derivative to zero, ℓ′(θ)=0\ell'(\theta) = 0, and checking that this point is a maximum (for example, that ℓ′′(θ^)<0\ell''(\hat\theta) < 0). When no closed-form solution exists, as in logistic regression, the maximum is found numerically.

Worked example: a coin

We observe k=7k = 7 heads in n=10n = 10 independent tosses. The number of heads follows a binomial distribution with parameter pp, where 0<p<10 < p < 1.

Step 1: write the likelihood. The binomial PMF evaluated at the observed kk, viewed as a function of pp:

L(p)=(nk)pk(1−p)n−k=(107)p7(1−p)3.L(p) = \binom{n}{k} p^k (1-p)^{n-k} = \binom{10}{7} p^7 (1-p)^3.

Step 2: take logarithms.

ℓ(p)=log⁡(nk)+klog⁡p+(n−k)log⁡(1−p).\ell(p) = \log\binom{n}{k} + k \log p + (n-k)\log(1-p).

The first term does not depend on pp, so it does not affect where the maximum is.

Step 3: differentiate and set to zero.

ℓ′(p)=kp−n−k1−p=0.\ell'(p) = \frac{k}{p} - \frac{n-k}{1-p} = 0.

Multiplying both sides by p(1−p)p(1-p) gives k(1−p)=(n−k)pk(1-p) = (n-k)p, which simplifies to k=npk = np. So

p^=kn=710=0.7.\hat{p} = \frac{k}{n} = \frac{7}{10} = 0.7.

Step 4: check that it is a maximum. The second derivative ℓ′′(p)=−k/p2−(n−k)/(1−p)2\ell''(p) = -k/p^2 - (n-k)/(1-p)^2 is negative for every pp between 00 and 11, so ℓ\ell is concave and p=0.7p = 0.7 is its single maximum. This matches the peak in the figure.

The answer, “the proportion of heads you observed”, is what intuition suggests. The value of the method is that the same four steps work for models where intuition gives no obvious answer. One edge case: if k=0k = 0 or k=nk = n, the likelihood is largest at the boundary, giving p^=0\hat p = 0 or p^=1\hat p = 1, which is a poor estimate from a small sample.

A continuous example: waiting times

Five waiting times, in minutes, are modelled as independent draws from an exponential distribution with rate λ>0\lambda > 0, whose density is f(x∣λ)=λe−λxf(x \mid \lambda) = \lambda e^{-\lambda x} for x≥0x \ge 0. The data are 2.0,0.5,4.1,1.3,2.62.0, 0.5, 4.1, 1.3, 2.6.

The likelihood is a product of densities:

L(λ)=∏i=1nλe−λxi=λnexp⁡ ⁣(−λ∑i=1nxi),L(\lambda) = \prod_{i=1}^n \lambda e^{-\lambda x_i} = \lambda^n \exp\!\left(-\lambda \sum_{i=1}^n x_i\right),

so ℓ(λ)=nlog⁡λ−λ∑i=1nxi\ell(\lambda) = n \log \lambda - \lambda \sum_{i=1}^n x_i. Setting ℓ′(λ)=n/λ−∑xi=0\ell'(\lambda) = n/\lambda - \sum x_i = 0 gives

λ^=n∑i=1nxi=1xˉ.\hat{\lambda} = \frac{n}{\sum_{i=1}^n x_i} = \frac{1}{\bar{x}}.

Here ∑xi=10.5\sum x_i = 10.5 and xˉ=2.1\bar{x} = 2.1 minutes, so λ^=1/2.1≈0.476\hat{\lambda} = 1/2.1 \approx 0.476 per minute. Since ℓ′′(λ)=−n/λ2<0\ell''(\lambda) = -n/\lambda^2 < 0, this is a maximum.

The same reasoning applied to a normal sample with unknown mean μ\mu shows that the MLE of μ\mu is the sample mean xˉ\bar{x}: the log-likelihood is, apart from constants, −∑(xi−μ)2/(2σ2)-\sum (x_i - \mu)^2 / (2\sigma^2), which is largest when μ\mu minimizes the sum of squared distances to the data, and that is μ=xˉ\mu = \bar{x}.

Properties of maximum likelihood estimators

Under standard regularity conditions (roughly: the model is correct, the number of parameters stays fixed as nn grows, the true value is not on the boundary of the allowed values, and the likelihood is smooth), maximum likelihood estimators have attractive large-sample properties.

These are statements about large samples. In small samples, an MLE can be noticeably biased. The standard example is the variance of a normal distribution: its MLE is

σ^2=1n∑i=1n(xi−xˉ)2,\hat{\sigma}^2 = \frac{1}{n} \sum_{i=1}^n (x_i - \bar{x})^2,

with nn rather than n−1n - 1 in the denominator. On average it underestimates σ2\sigma^2 by the factor (n−1)/n(n-1)/n, for example by 20% when n=5n = 5. That is why the usual sample variance divides by n−1n - 1. The bias vanishes as nn grows, consistent with the large-sample properties above.

Common misunderstandings

“The likelihood is the probability that the parameter is correct.” It is the probability (or density) of the observed data under that parameter value. It does not sum or integrate to 11 over the parameter, and it says nothing about which parameter values were plausible before seeing the data.

“Likelihood and probability are the same word.” In everyday speech they are. In statistics, probability describes possible data for a fixed parameter; likelihood compares parameter values for fixed data.

“The MLE is the true value.” It is an estimate, and it varies from sample to sample like any other statistic. With 7 heads in 10 tosses, p^=0.7\hat p = 0.7, but a fair coin produces 7 or more heads in 10 tosses about 17% of the time.

“Maximum likelihood always gives the best estimator.” Its optimality is a large-sample result that assumes the model is correct. With small samples, many parameters, or a misspecified model, it can be biased or unstable, and other estimators may do better.

Further reading