Point Estimation
Using a single number computed from a sample to estimate an unknown population parameter, and judging how good that rule is.
Prerequisites: Sampling Distributions, Expected Value, Variance of a Random Variable.
Point estimation is the use of sample data to produce a single best guess, a point estimate, for an unknown quantity describing a population, such as its mean, a proportion, or a variance. The sample mean as a guess for the population mean is the most familiar example.
Since any guess from a random sample will be somewhat off, the important question is not whether one estimate is right, but whether the rule that produced it tends to land close to the truth. The ideas of bias, variance, and mean squared error make that question precise, and they let us compare different ways of estimating the same thing.
Parameters, estimators, and estimates
A parameter, written generically as (theta), is a fixed but unknown number describing the population: the mean , the variance , the proportion of voters supporting a candidate.
An estimator (“theta hat”) is a rule for computing a guess of from the data . Because the data are random, an estimator is a random variable with a sampling distribution.
An estimate is the number you get when you apply the rule to the data you actually observed.
For example, the sample mean is an estimator of . If five measured waiting times average 2.3 minutes, then 2.3 is an estimate. Asking “is 2.3 biased?” makes no sense; asking “is biased?” does, because it is a question about how the rule behaves over all the samples that could have occurred.
Bias, variance, and mean squared error
Picture the sampling distribution of an estimator. Two things determine whether it is a good one: where its distribution is centred, and how spread out it is.
Bias measures the centre. It is the difference between the estimator’s average value and the true parameter:
An estimator with zero bias for every possible value of is called unbiased: on average, over many samples, it hits the target.
Variance measures the spread: says how much the estimate changes from sample to sample. Its square root is the standard error of the estimator.
Mean squared error (MSE) combines the two into a single measure of how far the estimate typically lands from the truth:
The bias–variance decomposition
The MSE splits exactly into a variance part and a bias part:
To see why, write and add and subtract it inside the square:
Now take expectations term by term. The first term gives . In the middle term, is a constant and , so it vanishes. The last term is the constant .
The decomposition says there are two ways to be wrong: systematically (bias) and randomly (variance). Reducing one often increases the other.
The bias–variance tradeoff
An unbiased estimator is not automatically the best one. Accepting a small bias can sometimes buy a large reduction in variance, and so a smaller MSE.
In the figure, the population mean is , and each sample has observations from a normal population with . Estimator A is the sample mean. Estimator B, , averages the sample mean with a prior guess of 4. Then:
- A is unbiased, with variance , so its MSE is .
- B has mean , so its bias is . Its variance is . Its MSE is .
Here the biased estimator is better on average. But it is better only because the guess of 4 happened to be close to the truth; if the true mean were 10, B’s bias would be and its MSE about 9.45, far worse than A. Shrinking towards a guess helps when the guess is good and the sample is small. This tradeoff appears throughout statistics and machine learning, for example in regularized regression.
Consistency
An estimator is consistent if it gets arbitrarily close to as the sample size grows: for every ,
where is the estimator computed from observations. A convenient sufficient condition is that both the bias and the variance go to zero as grows, because then the MSE goes to zero.
The sample mean is consistent for : this is the law of large numbers. Consistency is a minimal requirement. An estimator that does not home in on the truth even with unlimited data is rarely worth using. Bias and consistency are different properties: an estimator can be biased for every finite and still consistent, as the next example shows.
Example: estimating a variance
Let be iid with mean and variance . A natural estimator of averages the squared deviations from the sample mean:
Its expected value. Expanding the square shows that . Since and ,
So : the estimator is biased downwards, by . The deviations are measured from , which is fitted to the same data and so sits closer to them than does. Dividing by instead gives the unbiased sample variance ; see variance and standard deviation. Since the bias shrinks to zero, is still consistent.
Worked example. Take observations from a normal population with . For normal data, the variance of is .
- Unbiased : bias , variance , so MSE .
- Divide-by- estimator : bias , variance , so MSE .
The biased estimator has the smaller MSE. Even so, is the standard choice, because unbiasedness is convenient when estimates are combined and because inference methods such as the test are built around it. Neither is “the correct” estimator; they are optimal by different criteria.
How estimators are constructed
Bias, variance, and MSE judge an estimator once you have one. Two general recipes produce them:
- Method of moments: set sample averages equal to their population counterparts and solve. For example, set equal to , written in terms of the parameters.
- Maximum likelihood: choose the parameter value under which the observed data are most probable. This is the most widely used general method, and it has good large-sample properties; see maximum likelihood estimation.
A point estimate on its own says nothing about its own uncertainty. Reporting it with a standard error, or as a confidence interval, tells the reader how far off it could plausibly be.
Common misunderstandings
“An unbiased estimator gives the right answer.” It gives the right answer on average over many samples. Any single estimate can be far off, especially if the variance is large.
“Unbiased is always better than biased.” Not if the goal is to be close to the truth. As both examples above show, a slightly biased estimator can have a smaller mean squared error.
“If is unbiased for , then is unbiased for .” It is not: taking a square root does not commute with averaging. For normal observations, . The bias shrinks as grows, and is usually ignored.
“Bias means the data were collected badly.” Statistical bias is a property of an estimator under an assumed model. Sampling bias, where the sample is not representative of the population, is a different and often more serious problem that no choice of estimator can fix.
Further reading
- Larry Wasserman, All of Statistics: A Concise Course in Statistical Inference, Springer, 2004. A concise treatment of bias, standard error, MSE and consistency.
- George Casella and Roger L. Berger, Statistical Inference, 2nd ed., Duxbury, 2002. The standard graduate-level reference on point estimation, including the method of moments and maximum likelihood.