Likelihood and Maximum Likelihood Estimation
The likelihood measures how well each parameter value explains the observed data; the maximum likelihood estimate is the value that explains them best.
Prerequisites: Point Estimation, Random Variables, Bernoulli and Binomial Distributions.
The likelihood of a parameter value is the probability (or probability density) that a model with that parameter value assigns to the data you actually observed. Maximum likelihood estimation picks, as the estimate of the parameter, the value under which the observed data would have been most probable.
It is the most widely used general recipe for building estimators. Given almost any probability model, from a coin toss to logistic regression, maximum likelihood tells you how to turn data into parameter estimates, and in large samples those estimates are usually as precise as any reasonable alternative.
Intuition
You toss a coin 10 times and see 7 heads. What is a sensible guess for , the probability of heads?
Try a few candidate values and ask how likely the observed result would be under each:
- if , getting exactly 7 heads has probability about ;
- if , it has probability about ;
- if , it has probability about .
The data are more than twice as probable under as under a fair coin, and much more probable than under . Maximum likelihood formalizes this comparison: compute that probability for every candidate , and choose the that makes it largest.
The likelihood function
Suppose the data are modelled as observations of random variables whose distribution depends on an unknown parameter . Write for the probability mass function of a discrete observation, or for the density of a continuous one. If the observations are independent, the joint probability (or density) of the whole data set is the product of the individual ones.
The likelihood function is that same expression, read as a function of with the data held fixed:
The formula is identical to the probability model. What changes is the question. The model asks, “for a fixed , how probable are different data sets?” The likelihood asks, “for these fixed data, how well does each explain them?”
The likelihood is not a probability distribution for the parameter
is not the probability that is the true value. It is a probability of the data, evaluated at many different values of . Nothing forces it to add or integrate to over . For the coin example,
not . Only ratios of likelihoods are meaningful: says the data are about 2.3 times as probable if as if . Turning this into a probability statement about itself requires a prior distribution and Bayes’ theorem, which is a different approach to inference.
The log-likelihood
In practice we almost always work with the log-likelihood
(with in place of for continuous data). There are three reasons:
- the logarithm turns a product into a sum, which is much easier to differentiate;
- a product of many probabilities smaller than quickly becomes too small for a computer to store, while a sum of logarithms does not;
- is strictly increasing, so and are largest at exactly the same . Nothing is lost by switching.
Maximum likelihood estimation
The maximum likelihood estimate (MLE) of is the value that maximizes the likelihood:
Here means “the value of at which the maximum is attained”. When is smooth and the maximum is not at the edge of the allowed parameter values, it can usually be found by setting the derivative to zero, , and checking that this point is a maximum (for example, that ). When no closed-form solution exists, as in logistic regression, the maximum is found numerically.
Worked example: a coin
We observe heads in independent tosses. The number of heads follows a binomial distribution with parameter , where .
Step 1: write the likelihood. The binomial PMF evaluated at the observed , viewed as a function of :
Step 2: take logarithms.
The first term does not depend on , so it does not affect where the maximum is.
Step 3: differentiate and set to zero.
Multiplying both sides by gives , which simplifies to . So
Step 4: check that it is a maximum. The second derivative is negative for every between and , so is concave and is its single maximum. This matches the peak in the figure.
The answer, “the proportion of heads you observed”, is what intuition suggests. The value of the method is that the same four steps work for models where intuition gives no obvious answer. One edge case: if or , the likelihood is largest at the boundary, giving or , which is a poor estimate from a small sample.
A continuous example: waiting times
Five waiting times, in minutes, are modelled as independent draws from an exponential distribution with rate , whose density is for . The data are .
The likelihood is a product of densities:
so . Setting gives
Here and minutes, so per minute. Since , this is a maximum.
The same reasoning applied to a normal sample with unknown mean shows that the MLE of is the sample mean : the log-likelihood is, apart from constants, , which is largest when minimizes the sum of squared distances to the data, and that is .
Properties of maximum likelihood estimators
Under standard regularity conditions (roughly: the model is correct, the number of parameters stays fixed as grows, the true value is not on the boundary of the allowed values, and the likelihood is smooth), maximum likelihood estimators have attractive large-sample properties.
- Consistency. As the sample size grows, gets arbitrarily close to the true with probability approaching .
- Asymptotic normality. For large , the sampling distribution of is approximately normal, centred at the true . Its standard deviation can be estimated from how sharply curved the log-likelihood is at its peak: a sharp peak means a precise estimate. For the coin, this gives the familiar standard error , used to build confidence intervals.
- Efficiency. In large samples, no other well-behaved estimator has a smaller variance.
- Invariance. If is the MLE of , then is the MLE of for any function . In the waiting-time example, the MLE of the mean waiting time is minutes.
These are statements about large samples. In small samples, an MLE can be noticeably biased. The standard example is the variance of a normal distribution: its MLE is
with rather than in the denominator. On average it underestimates by the factor , for example by 20% when . That is why the usual sample variance divides by . The bias vanishes as grows, consistent with the large-sample properties above.
Common misunderstandings
“The likelihood is the probability that the parameter is correct.” It is the probability (or density) of the observed data under that parameter value. It does not sum or integrate to over the parameter, and it says nothing about which parameter values were plausible before seeing the data.
“Likelihood and probability are the same word.” In everyday speech they are. In statistics, probability describes possible data for a fixed parameter; likelihood compares parameter values for fixed data.
“The MLE is the true value.” It is an estimate, and it varies from sample to sample like any other statistic. With 7 heads in 10 tosses, , but a fair coin produces 7 or more heads in 10 tosses about 17% of the time.
“Maximum likelihood always gives the best estimator.” Its optimality is a large-sample result that assumes the model is correct. With small samples, many parameters, or a misspecified model, it can be biased or unstable, and other estimators may do better.
Further reading
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. Covers maximum likelihood and its large-sample theory concisely.
- George Casella and Roger L. Berger, Statistical Inference, 2nd ed., Duxbury, 2002. A rigorous treatment of likelihood, including the invariance property and regularity conditions.