Point Estimation

Using a single number computed from a sample to estimate an unknown population parameter, and judging how good that rule is.

Prerequisites: Sampling Distributions, Expected Value, Variance of a Random Variable.

Point estimation is the use of sample data to produce a single best guess, a point estimate, for an unknown quantity describing a population, such as its mean, a proportion, or a variance. The sample mean as a guess for the population mean is the most familiar example.

Since any guess from a random sample will be somewhat off, the important question is not whether one estimate is right, but whether the rule that produced it tends to land close to the truth. The ideas of bias, variance, and mean squared error make that question precise, and they let us compare different ways of estimating the same thing.

Parameters, estimators, and estimates

A parameter, written generically as θ\theta (theta), is a fixed but unknown number describing the population: the mean μ\mu, the variance σ2\sigma^2, the proportion pp of voters supporting a candidate.

An estimator θ^\hat{\theta} (“theta hat”) is a rule for computing a guess of θ\theta from the data X1,…,XnX_1, \dots, X_n. Because the data are random, an estimator is a random variable with a sampling distribution.

An estimate is the number you get when you apply the rule to the data you actually observed.

For example, the sample mean Xˉ=1n∑i=1nXi\bar{X} = \frac{1}{n}\sum_{i=1}^n X_i is an estimator of μ\mu. If five measured waiting times average 2.3 minutes, then 2.3 is an estimate. Asking “is 2.3 biased?” makes no sense; asking “is Xˉ\bar{X} biased?” does, because it is a question about how the rule behaves over all the samples that could have occurred.

Bias, variance, and mean squared error

Picture the sampling distribution of an estimator. Two things determine whether it is a good one: where its distribution is centred, and how spread out it is.

Bias measures the centre. It is the difference between the estimator’s average value and the true parameter:

Bias⁡(θ^)=E⁡[θ^]−θ.\operatorname{Bias}(\hat{\theta}) = \E[\hat{\theta}] - \theta.

An estimator with zero bias for every possible value of θ\theta is called unbiased: on average, over many samples, it hits the target.

Variance measures the spread: Var⁡(θ^)\Var(\hat{\theta}) says how much the estimate changes from sample to sample. Its square root is the standard error of the estimator.

Mean squared error (MSE) combines the two into a single measure of how far the estimate typically lands from the truth:

MSE⁡(θ^)=E⁡[(θ^−θ)2].\operatorname{MSE}(\hat{\theta}) = \E\big[(\hat{\theta} - \theta)^2\big].

The bias–variance decomposition

The MSE splits exactly into a variance part and a bias part:

MSE⁡(θ^)=Var⁡(θ^)+Bias⁡(θ^)2.\operatorname{MSE}(\hat{\theta}) = \Var(\hat{\theta}) + \operatorname{Bias}(\hat{\theta})^2.

To see why, write m=E⁡[θ^]m = \E[\hat{\theta}] and add and subtract it inside the square:

(θ^−θ)2=(θ^−m)2+2(θ^−m)(m−θ)+(m−θ)2.(\hat{\theta} - \theta)^2 = (\hat{\theta} - m)^2 + 2(\hat{\theta} - m)(m - \theta) + (m - \theta)^2.

Now take expectations term by term. The first term gives E⁡[(θ^−m)2]=Var⁡(θ^)\E[(\hat{\theta} - m)^2] = \Var(\hat{\theta}). In the middle term, m−θm - \theta is a constant and E⁡[θ^−m]=0\E[\hat{\theta} - m] = 0, so it vanishes. The last term is the constant (m−θ)2=Bias⁡(θ^)2(m - \theta)^2 = \operatorname{Bias}(\hat{\theta})^2.

The decomposition says there are two ways to be wrong: systematically (bias) and randomly (variance). Reducing one often increases the other.

The bias–variance tradeoff

An unbiased estimator is not automatically the best one. Accepting a small bias can sometimes buy a large reduction in variance, and so a smaller MSE.

Two bell-shaped sampling distributions over the value of the estimate, with a dotted vertical line at the true mean of 5. The sample mean is centred exactly on 5 but is wide. The shrunk mean is centred slightly to the left, at 4.5, but is much narrower, and on average it lands closer to 5.
Sampling distributions of two estimators of a population mean μ = 5, from n = 5 independent normal observations with σ = 3. A: the sample mean X̄, unbiased with variance σ²/n = 1.80, so MSE 1.80. B: the shrunk mean (X̄ + 4)/2, which pulls X̄ halfway towards a prior guess of 4; its bias is −0.5 and its variance is 0.45, so MSE = 0.45 + 0.25 = 0.70.

In the figure, the population mean is μ=5\mu = 5, and each sample has n=5n = 5 observations from a normal population with σ=3\sigma = 3. Estimator A is the sample mean. Estimator B, (Xˉ+4)/2(\bar{X} + 4)/2, averages the sample mean with a prior guess of 4. Then:

Here the biased estimator is better on average. But it is better only because the guess of 4 happened to be close to the truth; if the true mean were 10, B’s bias would be −3-3 and its MSE about 9.45, far worse than A. Shrinking towards a guess helps when the guess is good and the sample is small. This tradeoff appears throughout statistics and machine learning, for example in regularized regression.

Consistency

An estimator is consistent if it gets arbitrarily close to θ\theta as the sample size grows: for every ε>0\varepsilon > 0,

P(∣θ^n−θ∣>ε)→0as n→∞,P\big(|\hat{\theta}_n - \theta| > \varepsilon\big) \to 0 \quad \text{as } n \to \infty,

where θ^n\hat{\theta}_n is the estimator computed from nn observations. A convenient sufficient condition is that both the bias and the variance go to zero as nn grows, because then the MSE goes to zero.

The sample mean is consistent for μ\mu: this is the law of large numbers. Consistency is a minimal requirement. An estimator that does not home in on the truth even with unlimited data is rarely worth using. Bias and consistency are different properties: an estimator can be biased for every finite nn and still consistent, as the next example shows.

Example: estimating a variance

Let X1,…,XnX_1, \dots, X_n be iid with mean μ\mu and variance σ2\sigma^2. A natural estimator of σ2\sigma^2 averages the squared deviations from the sample mean:

σ^2=1n∑i=1n(Xi−Xˉ)2.\hat{\sigma}^2 = \frac{1}{n} \sum_{i=1}^n (X_i - \bar{X})^2.

Its expected value. Expanding the square shows that ∑i(Xi−Xˉ)2=∑iXi2−nXˉ2\sum_{i}(X_i - \bar{X})^2 = \sum_i X_i^2 - n\bar{X}^2. Since E⁡[Xi2]=σ2+μ2\E[X_i^2] = \sigma^2 + \mu^2 and E⁡[Xˉ2]=Var⁡(Xˉ)+μ2=σ2/n+μ2\E[\bar{X}^2] = \Var(\bar{X}) + \mu^2 = \sigma^2/n + \mu^2,

E⁡[∑i=1n(Xi−Xˉ)2]=n(σ2+μ2)−n(σ2n+μ2)=(n−1)σ2.\E\Big[\sum_{i=1}^n (X_i - \bar{X})^2\Big] = n(\sigma^2 + \mu^2) - n\Big(\frac{\sigma^2}{n} + \mu^2\Big) = (n-1)\sigma^2.

So E⁡[σ^2]=n−1nσ2\E[\hat{\sigma}^2] = \frac{n-1}{n}\sigma^2: the estimator is biased downwards, by −σ2/n-\sigma^2/n. The deviations are measured from Xˉ\bar{X}, which is fitted to the same data and so sits closer to them than μ\mu does. Dividing by n−1n - 1 instead gives the unbiased sample variance s2s^2; see variance and standard deviation. Since the bias −σ2/n-\sigma^2/n shrinks to zero, σ^2\hat{\sigma}^2 is still consistent.

Worked example. Take n=5n = 5 observations from a normal population with σ2=1\sigma^2 = 1. For normal data, the variance of s2s^2 is 2σ4/(n−1)2\sigma^4/(n-1).

  1. Unbiased s2s^2: bias 00, variance 2/4=0.52/4 = 0.5, so MSE =0.5= 0.5.
  2. Divide-by-nn estimator σ^2=n−1ns2=0.8 s2\hat{\sigma}^2 = \frac{n-1}{n}s^2 = 0.8\,s^2: bias 0.8−1=−0.20.8 - 1 = -0.2, variance 0.82×0.5=0.320.8^2 \times 0.5 = 0.32, so MSE =0.32+0.04=0.36= 0.32 + 0.04 = 0.36.

The biased estimator has the smaller MSE. Even so, s2s^2 is the standard choice, because unbiasedness is convenient when estimates are combined and because inference methods such as the tt test are built around it. Neither is “the correct” estimator; they are optimal by different criteria.

How estimators are constructed

Bias, variance, and MSE judge an estimator once you have one. Two general recipes produce them:

A point estimate on its own says nothing about its own uncertainty. Reporting it with a standard error, or as a confidence interval, tells the reader how far off it could plausibly be.

Common misunderstandings

“An unbiased estimator gives the right answer.” It gives the right answer on average over many samples. Any single estimate can be far off, especially if the variance is large.

“Unbiased is always better than biased.” Not if the goal is to be close to the truth. As both examples above show, a slightly biased estimator can have a smaller mean squared error.

“If s2s^2 is unbiased for σ2\sigma^2, then ss is unbiased for σ\sigma.” It is not: taking a square root does not commute with averaging. For n=5n = 5 normal observations, E⁡[s]≈0.94 σ\E[s] \approx 0.94\,\sigma. The bias shrinks as nn grows, and is usually ignored.

“Bias means the data were collected badly.” Statistical bias is a property of an estimator under an assumed model. Sampling bias, where the sample is not representative of the population, is a different and often more serious problem that no choice of estimator can fix.

Further reading