The total variation distance is
If , the Kullback-Leibler divergence is
Pinsker's inequality states
The squared-loss Le Cam two-point lemma says that for two experiments with scalar parameters ,
Indeed, classify the data as when is closer to and as otherwise. On a classification error, the estimation error is at least . The sum of the two testing error probabilities is at least . Averaging the two risks and then bounding their maximum proves the result.
Let . The Hölder class on consists of functions with derivatives through order whose th derivative is Hölder of order with constant , with the standard integer-order convention.
Fix any . First compare the constant regression functions and with . Both belong to every Hölder class under the seminorm convention, and the normal-product divergence is
Pinsker and Le Cam therefore give a lower bound .
For the smoothness-dependent term, use the supplied smooth bump , translated one-sidedly near a boundary when necessary, and compare
Choose its fixed normalization so that whenever . The Gaussian divergence satisfies
Take
with the bandwidth and amplitude truncated at constants when this expression leaves . Then the divergence remains bounded and
The same construction can be placed at every , with a one-sided bump at the endpoints. Pinsker and Le Cam, combined with the constant alternatives, prove
where depends only on .
Put
The degree- local polynomial estimator is , where
Define the local polynomial Gram matrix
Assume that it is positive definite, and write
The normal equations then give
Thus the estimator is a linear estimator in nonparametric regression, and the displayed are its effective kernel weights.
If is a polynomial of degree at most , then for some . Feeding into the weighted least-squares problem gives the exact zero-residual fit , whose intercept is . Hence the polynomial reproduction property of local polynomial regression is
Let . The Hölder class consists of functions with derivatives through order bounded by and
The subclass additionally makes every derivative of order below -Lipschitz.
Suppose . Taylor's theorem and that additional Lipschitz condition give, for the degree- Taylor polynomial at ,
On the support of , , so
For the regular design , at most points satisfy when . Polynomial reproduction cancels , and therefore
Thus the universal exponent in the question is , with the displayed choice of .
Put and . The density Hölder class consists of nonnegative functions integrating to one, with derivatives through order , such that
This is the density version of the Hölder class. A kernel for density estimation is an integrable function with . It has order when for and .
Choose a bounded kernel of order at least with , and use the kernel density estimator
Taylor's theorem at and the vanishing kernel moments cancel every polynomial term below the remainder. Hence, for a constant depending only on and the fixed kernel,
We also need a uniform density bound. The standard Hölder interpolation argument combines nonnegativity, , and the Hölder constraint to give
Indeed, near a point where is close to its maximum , Taylor's theorem and the derivative bounds implied by the Hölder constraint keep of order on an interval of length comparable to ; integrating over that interval gives .
Using this bound and independence,
Thus
Choose
Both terms then have order . Since the infimum over all measurable estimators is no larger than the risk of this particular estimator,
This is the pointwise minimax rate for Hölder density estimation upper bound.
Write , , and define
The degree- local polynomial regression fit minimizes
Assume its local polynomial Gram matrix
is positive definite. The normal equations then give
Because the polynomial is written in the scaled coordinate, the local polynomial derivative estimator is . If , then
For a polynomial of degree at most , the local least-squares fit to is exactly the Taylor polynomial , so its scaled linear coefficient is . Therefore the polynomial reproduction property of local polynomial regression gives
Positive definiteness and the fixed finite-dimensional basis provide a number such that
Since outside ,
and
The regular design has at most points in this window when . Since the errors are independent with variance at most ,
Finally let and take the degree- Taylor polynomial of at . The Hölder class remainder satisfies
Polynomial reproduction removes from the bias. On the kernel window the remainder is at most , so the weight-sum bound gives
Let . The Hölder class consists of functions with continuous derivatives whose th derivative satisfies
with the equivalent Lipschitz convention when is an integer.
The degree- local polynomial regression estimator minimizes
and takes . Put , , and rescale by . If is positive definite, weighted least squares gives
Define
Then . If has degree at most , fitting the noiseless response reproduces that polynomial exactly, and hence
For nonzero , the polynomial cannot vanish throughout . Therefore
and compactness of the unit sphere makes its minimum eigenvalue positive. Choose and so that makes the supplied lower bound on at least . On the kernel support, is bounded by a constant depending only on , and . It follows that only weights are nonzero and
Polynomial reproduction cancels the Taylor polynomial of degree . The Hölder remainder on is at most , so the squared bias is at most . Independence and bound the variance by . Combining them uniformly in and proves