For independent observations with a normal distribution , fixed, every estimator has . Compare and : the joint Kullback-Leibler divergence is , so the total variation–Hellinger–relative entropy inequality bounds total variation distance by . Apply the Le Cam lower bound under absolute-error loss. Degenerate zero-noise laws do not satisfy a positive lower bound.
For two observation laws at real parameters separated by , every estimator satisfies . With common densities, the sum of the two risks is at least , at least by the triangle inequality. Divide by two. The same proof works for any metric loss.
A sufficient statistic has a conditional distribution of the full data given its value that is independent of the unknown parameter. A minimal sufficient statistic is sufficient and is a function of every other sufficient statistic, up to null sets. For a dominated family with common positive support, the likelihood-ratio criterion for minimal sufficiency states that is minimal sufficient if holds exactly when is independent of .
For , the joint probability mass function is
The Fisher-Neyman factorization theorem proves sufficiency of . For two samples, the likelihood ratio is , independent of exactly when their sums agree. Thus is minimal sufficient. Any one-to-one transformation of it is also minimal sufficient, so choose
This chosen estimator is an unbiased estimator, since each observation has mean . If instead the statistic is reported as , its expectation is , and as an estimator of it has bias (zero only when ). Being a minimal sufficient statistic is unchanged by this rescaling; being an unbiased estimator is not.
For a target , the bias of an integrable estimator is ; it is an unbiased estimator if this is zero at every allowed parameter value. Here has a zero-truncated Poisson distribution, , and . Write an estimator based only on as , assuming a finite expectation for every . Unbiasedness requires
Absolute integrability at every positive parameter ensures that the power series on the left converges absolutely on every complex disc. Uniqueness of coefficients in a power series therefore gives for odd and for positive even . Conversely these values give the displayed identity, so
is the unique unbiased estimator of this form. The uniqueness claim concerns nonrandomized functions of ; allowing external randomness would permit addition of independent mean-zero noise.
For this parity estimator for a zero-truncated Poisson count, usefulness depends on the loss, but under squared-error loss it performs poorly despite being an unbiased estimator. Since , and
It takes the inadmissible value with positive probability, and its variance tends to one rather than zero as . Clipping to the parameter interval gives . This has bias , but
Thus a simple biased estimator strictly improves its mean squared error at every parameter value; uniqueness among unbiased estimators does not make it optimal for this loss.
The Rao-Blackwell theorem applies to an estimator with finite second moment and a sufficient statistic . The conditional estimator can be chosen as a function of the observed that does not depend on the unknown parameter: this uses the parameter-independent conditional data distribution in the definition of sufficient statistic. The tower property of conditional expectation gives , so their bias agrees. The law of total variance gives
For any target it follows that
Equality holds exactly when almost surely under that parameter value. This proves variance reduction and reduction of mean squared error, whether or not the original estimator is an unbiased estimator. More generally, conditional Jensen inequality proves the corresponding inequality for any convex loss in the estimate, whenever the expectations exist.
Let and be Radon-Nikodym derivatives. Use the following normalizations of the Kullback-Leibler divergence, total variation distance and unnormalized Hellinger distance:
We use and assign infinite Kullback-Leibler divergence if , where denotes absolute continuity of measures. Common domination alone also suffices for the inequalities with the extended-value convention. The negative part of the logarithmic integral is integrable: on , . Thus the extended-value integral defining the Kullback-Leibler divergence is well-defined. These expressions do not depend on which common dominating measure is chosen. The Kullback-Leibler divergence is generally asymmetric and therefore fails a defining property of a metric.
For completeness, the integral formula for total variation distance follows by taking : since , the positive and negative parts of have the same integral, and this set attains the supremum.
Put , by the Cauchy-Schwarz inequality. A second application gives
The logarithmic inequality in the hint implies for . Apply it to on :
Multiplying by and integrating proves
If on a set of positive -measure, the right side is finite but the Kullback-Leibler divergence is infinite, so the inequality is immediate. Thus
This total variation–Hellinger–relative entropy inequality requires the stated Hellinger distance normalization for its first constant. With the alternative convention , the correct comparison is and . In particular, and have .
Here is a testing form of the Le Cam two-point lemma. Suppose two parameter values in a metric space have separation , and let be the corresponding observation laws. Every estimator satisfies
To prove it, classify by the closer of the two parameters, breaking a tie in favour of one. Let . On , the triangle inequality implies , while on it implies . Hence the sum of the two probabilities in the display is at least
At least one is at least half the sum, proving the claim. It also gives a lower bound for the maximum expected metric loss.
For absolute-error loss, or any loss given by a metric, there is a stronger expectation form. Take densities relative to and use the triangle inequality under their overlap:
Therefore the Le Cam lower bound under absolute-error loss is
Both arguments are valid for randomized estimators as well, by adjoining the independent randomization to the observation; doing so does not change total variation distance.
For the normal distribution model, assume the usual nondegenerate convention and use the full product observation laws. Choose
For one observation, its Gaussian log likelihood ratio is
Under the first law, and . Thus the Kullback-Leibler divergence between normal distributions is
For the independent and identically distributed random variables, the joint log likelihood ratio is the sum of the individual ones, so the Kullback-Leibler divergence of the two full statistical samples is . The comparison just proved implies . Applying the expectation form of the Le Cam two-point lemma gives
This is a Gaussian location minimax lower bound under absolute-error loss. The constant may depend on the fixed known , but neither on the estimator nor on . If a signed nonzero scale were used, the same result holds with in place of . If zero noise were allowed, would estimate exactly, so a positive lower bound would be false; positive variance is necessary.