A centered random variable is sub-Gaussian with parameter when
For , the Chernoff bound gives
Minimizing at yields . Applying the same argument to proves the left-tail bound.
The moment-generating-function inequality and its version at imply
Comparing the second-order terms as gives . Since ,
A centered is Sub-Gamma random variable in the right tail with variance factor and scale factor when
The corresponding Bernstein's inequality is
A standard squared-sub-Gaussian lemma, obtained by integrating the sub-Gaussian tail or expanding exponential moments, states that
Scaling a sub-Gamma variable by multiplies its variance factor by and its scale by ; independent sums add variance factors and take the largest scale. Decompose
Their sum is therefore sub-Gamma on the right with variance factor
and scale factor
Apply Bernstein to the sum at threshold to obtain
A kernel for density estimation is an integrable function with . The kernel density estimator is
It has order when for and .
Differentiation gives the derivative kernel density estimator
Dropping the negative square of its mean from the variance and using Tonelli and the substitution ,
Thus and .
For , the Nikolsky class consists of functions with square-integrable derivatives through order and
using the equivalent finite-difference definition when is an integer. Integration by parts shows
Taylor expansion in , cancellation of the moments through order for , and the generalized Minkowski inequality give
Consequently
where one admissible choice is
Hence .
The mean integrated squared error is
the sum of integrated variance and squared bias. The preceding bounds give
Taking proves the MISE rate for derivative kernel density estimation
Thus . Estimating a density of the same smoothness has variance order and rate ; estimating its derivative is harder because differentiation amplifies high-frequency noise, changing to .
At a target , the degree- local polynomial regression estimate minimizes
and . Let have row
and let be diagonal with the kernel weights. Assume the local polynomial Gram matrix is positive definite. The weighted normal equations give
so it is a linear estimator in nonparametric regression.
Define
Inverting the two-by-two Gram matrix for yields
Now put . At and for the uniform kernel,
For the usual bandwidth range , . The elementary power-sum formulas and show, for , that
The limiting moment determinant is
For , the displayed errors can consume at most a fixed fraction of this value, so
for a universal and .
As printed, the claim “for all ” cannot hold for this design: if , all points receive weight , each , and the determinant is , not bounded below by a positive multiple of . The standard bandwidth restriction is therefore necessary for that intermediate assertion. The final bias bound remains harmless for , since and the finite design response means are uniformly bounded.
The polynomial reproduction property of local polynomial regression makes the local-linear weights reproduce both and . Write
for . The reproduced terms have zero bias. Using , the numerator contributed by the remainders is bounded by
Division by the determinant lower bound gives
The degree-zero Nadaraya–Watson estimator reproduces constants but not linear functions. The term therefore remains, and its boundary bias is bounded by ; this first-order boundary bias is the improvement that local linear fitting removes.
The total variation distance is
If , the Kullback-Leibler divergence is
Pinsker's inequality states
The squared-loss Le Cam two-point lemma says that for two experiments with scalar parameters ,
Indeed, classify the data as when is closer to and as otherwise. On a classification error, the estimation error is at least . The sum of the two testing error probabilities is at least . Averaging the two risks and then bounding their maximum proves the result.
Let . The Hölder class on consists of functions with derivatives through order whose th derivative is Hölder of order with constant , with the standard integer-order convention.
Fix any . First compare the constant regression functions and with . Both belong to every Hölder class under the seminorm convention, and the normal-product divergence is
Pinsker and Le Cam therefore give a lower bound .
For the smoothness-dependent term, use the supplied smooth bump , translated one-sidedly near a boundary when necessary, and compare
Choose its fixed normalization so that whenever . The Gaussian divergence satisfies
Take
with the bandwidth and amplitude truncated at constants when this expression leaves . Then the divergence remains bounded and
The same construction can be placed at every , with a one-sided bump at the endpoints. Pinsker and Le Cam, combined with the constant alternatives, prove
where depends only on .

Articles by others on the same topic (0)

There are currently no matching articles.