Posts
Why the sample variance divides by n − 1
A short proof that dividing by n − 1 removes the bias, and a reminder that the unbiased estimator is not the most accurate one.
Every introductory course defines the sample variance with in the denominator, and most students accept it on trust. The reason fits in three lines, and so does the caveat that usually goes unmentioned.
The bias, and its correction
#Split each deviation from the true mean into a deviation from the sample mean and the error of the sample mean. The cross terms sum to zero, which gives the identity . Taking expectations,
because independence gives .
The intuition is that is fitted to the same data it is measured against. The sample mean is the point that minimizes the sum of squared deviations, so the data always look a little less spread out around than around the unknown . Estimating the mean uses up one degree of freedom, and dividing by instead of restores the expected value exactly.1
Unbiased is not the same as accurate
#The correction fixes the mean of the estimator, not its spread, and for small samples the spread dominates.
For normal data, follows a distribution, so . The mean squared error of is
which is minimized at . Dividing by gives the only unbiased estimator of the three, and the largest error.
| Size | Mean of | MSE of | MSE of | MSE of |
|---|---|---|---|---|
| 2 | 0.500 (0.500) | 2.000 (2.009) | 0.750 (0.752) | 0.667 (0.667) |
| 3 | 0.667 (0.666) | 1.000 (1.004) | 0.556 (0.558) | 0.500 (0.502) |
| 5 | 0.800 (0.801) | 0.500 (0.501) | 0.360 (0.360) | 0.333 (0.333) |
| 10 | 0.900 (0.900) | 0.222 (0.223) | 0.190 (0.191) | 0.182 (0.182) |
| 30 | 0.967 (0.968) | 0.069 (0.069) | 0.066 (0.066) | 0.065 (0.065) |
With two observations the unbiased estimator has three times the mean squared error of . By thirty observations their errors differ by less than ten per cent, and the choice hardly matters.
Which one to report
#Report when the variance feeds into something that relies on unbiasedness: pooled estimates, statistics, and the standard formulas that expect it. Use when it is the maximum-likelihood estimate in a larger model and consistency with that model matters more. And when a single small-sample variance must be as close as possible to the truth, remember that neither is the most accurate choice.
The correction is named after the astronomer Friedrich Wilhelm Bessel. ↩︎