Skip to content
DataLeaf

Posts

Why the sample variance divides by n − 1

A short proof that dividing by n − 1 removes the bias, and a reminder that the unbiased estimator is not the most accurate one.

Every introductory course defines the sample variance with n−1n - 1 in the denominator, and most students accept it on trust. The reason fits in three lines, and so does the caveat that usually goes unmentioned.

Definition — Sample variance
For observations x1,…,xnx_1, \dots, x_n with mean xˉ=1n∑ixi\bar x = \frac{1}{n} \sum_{i} x_i, let S=∑i=1n(xi−xˉ)2S = \sum_{i=1}^{n} (x_i - \bar x)^2. The sample variance is s2=S/(n−1)s^2 = S / (n - 1); the plug-in variance is σ^2=S/n\hat\sigma^2 = S / n.

The bias, and its correction

#
Theorem — Unbiasedness
If X1,…,XnX_1, \dots, X_n are independent with mean μ\mu and variance σ2<∞\sigma^2 < \infty, then 𝔼[S]=(n−1) σ2\mathbb{E}[S] = (n - 1)\,\sigma^2. Consequently 𝔼[s2]=σ2\mathbb{E}[s^2] = \sigma^2, while 𝔼[σ^2]=n−1n σ2\mathbb{E}[\hat\sigma^2] = \frac{n-1}{n}\,\sigma^2.
Proof.

Split each deviation from the true mean into a deviation from the sample mean and the error of the sample mean. The cross terms sum to zero, which gives the identity ∑i(Xi−μ)2=S+n(Xˉ−μ)2\sum_i (X_i - \mu)^2 = S + n (\bar X - \mu)^2. Taking expectations,

𝔼[S]=∑i=1n𝔼[(Xi−μ)2]−n 𝔼[(Xˉ−μ)2]=nσ2−nVar⁡(Xˉ)=nσ2−σ2=(n−1) σ2, \begin{aligned} \mathbb{E}[S] &= \sum_{i=1}^{n} \mathbb{E}\bigl[(X_i - \mu)^2\bigr] - n\, \mathbb{E}\bigl[(\bar X - \mu)^2\bigr] \\ &= n \sigma^2 - n \operatorname{Var}(\bar X) = n\sigma^2 - \sigma^2 = (n - 1)\, \sigma^2, \end{aligned}

because independence gives Var⁡(Xˉ)=σ2/n\operatorname{Var}(\bar X) = \sigma^2 / n.

End of proof

The intuition is that xˉ\bar x is fitted to the same data it is measured against. The sample mean is the point that minimizes the sum of squared deviations, so the data always look a little less spread out around xˉ\bar x than around the unknown μ\mu. Estimating the mean uses up one degree of freedom, and dividing by n−1n - 1 instead of nn restores the expected value exactly.1

Unbiased is not the same as accurate

#

The correction fixes the mean of the estimator, not its spread, and for small samples the spread dominates.

Remark

For normal data, S/σ2S / \sigma^2 follows a χn−12\chi^2_{n-1} distribution, so Var⁡(S)=2(n−1) σ4\operatorname{Var}(S) = 2 (n - 1)\, \sigma^4. The mean squared error of S/cS / c is

MSE⁡(S/c)=σ4[2(n−1)c2+(n−1c−1)2], \operatorname{MSE}(S / c) = \sigma^4 \left[ \frac{2(n-1)}{c^2} + \left( \frac{n-1}{c} - 1 \right)^{2} \right],

which is minimized at c=n+1c = n + 1. Dividing by n−1n - 1 gives the only unbiased estimator of the three, and the largest error.

Example — A simulation
Draw 200,000 samples of size nn from a standard normal distribution, so that σ2=1\sigma^2 = 1, and compute all three estimators for each sample. The table gives the exact values from the remark, with the simulated values in parentheses.
Size nnMean of S/nS/nMSE of S/(n−1)S/(n-1)MSE of S/nS/nMSE of S/(n+1)S/(n+1)
20.500 (0.500)2.000 (2.009)0.750 (0.752)0.667 (0.667)
30.667 (0.666)1.000 (1.004)0.556 (0.558)0.500 (0.502)
50.800 (0.801)0.500 (0.501)0.360 (0.360)0.333 (0.333)
100.900 (0.900)0.222 (0.223)0.190 (0.191)0.182 (0.182)
300.967 (0.968)0.069 (0.069)0.066 (0.066)0.065 (0.065)

With two observations the unbiased estimator has three times the mean squared error of S/(n+1)S / (n + 1). By thirty observations their errors differ by less than ten per cent, and the choice hardly matters.

Which one to report

#

Report s2s^2 when the variance feeds into something that relies on unbiasedness: pooled estimates, tt statistics, and the standard formulas that expect it. Use σ^2\hat\sigma^2 when it is the maximum-likelihood estimate in a larger model and consistency with that model matters more. And when a single small-sample variance must be as close as possible to the truth, remember that neither is the most accurate choice.


  1. The correction is named after the astronomer Friedrich Wilhelm Bessel. ↩︎