An unbiased estimator of a parameter satisfies for every permissible value of . Since the sample mean is linear,
so is unbiased.
Use the identity
The two expectations on the right are and , respectively. Hence
Thus
The unbiased sample variance instead divides the same sum of squares by .
The sample mean is
Write the Gaussian sample vector as its orthogonal projection onto the span of plus its projection onto the orthogonal complement. These two Gaussian projections are independent. Their squared standardized lengths give Cochran's theorem:
Because every column of has sample mean zero, every column of also has sample mean zero. The spectral theorem for real symmetric matrices gives
Hence for the sample covariance is , while the sample variance of is
Thus the are pairwise uncorrelated sample principal components, ordered by decreasing sample variance.
A statistic is sufficient for if the conditional distribution of the full sample given does not depend on . It is minimal sufficient if it is a function of every sufficient statistic.
The likelihood factors as
so the Fisher-Neyman factorization theorem shows that , and hence its one-to-one transform , is sufficient. Moreover, for two samples and , the ratio is independent of exactly when their maxima agree. The likelihood-ratio criterion for minimal sufficiency therefore shows that and are minimal sufficient.
For , the sample mean is not sufficient: two samples can have the same mean but different maxima, and their likelihood ratio then depends on through the support indicators. It is consequently not minimal sufficient. For the degenerate special case , the sample mean and maximum coincide and both conclusions reverse.
Let have regular density , let
be its score function, and let
be its Fisher information. Regularity gives the mean-zero score identity .
Suppose an estimator has mean . Differentiating under the integral gives
The Cauchy-Schwarz inequality therefore yields
and hence the Cramer-Rao bound
In particular, an unbiased estimator of has variance at least . For an independent sample, the information is the sum of the individual informations.
In a decision problem with risk of a decision rule , a rule is minimax when
Now let be independent variables with , under squared-error loss. The sample mean is unbiased with variance , so
for every . Thus the minimax value is at most .
For the matching lower bound, choose a continuously differentiable density on that vanishes at both endpoints and has finite prior information
for example, . For , define the prior
It is supported inside the parameter space, and its information is .
Here the likelihood score is
whose Fisher information is . For any decision rule , combine it with the prior score to form the joint score
Integration by parts in , with no boundary term because vanishes there, gives
The likelihood score has conditional mean zero, so its cross term with the prior score vanishes and
Cauchy--Schwarz now proves the Van Trees inequality in this case:
The worst-case risk dominates every integrated risk, and therefore
Letting shows that every rule has worst-case risk at least . Since attains that value, the result on the minimax sample mean for a nonnegative normal location is