A centered random variable is sub-Gaussian with parameter whenFor , the Chernoff bound givesMinimizing at yields . Applying the same argument to proves the left-tail bound.
The moment-generating-function inequality and its version at implyComparing the second-order terms as gives . Since ,
A centered is Sub-Gamma random variable in the right tail with variance factor and scale factor whenThe corresponding Bernstein's inequality is
A standard squared-sub-Gaussian lemma, obtained by integrating the sub-Gaussian tail or expanding exponential moments, states thatScaling a sub-Gamma variable by multiplies its variance factor by and its scale by ; independent sums add variance factors and take the largest scale. DecomposeTheir sum is therefore sub-Gamma on the right with variance factorand scale factorApply Bernstein to the sum at threshold to obtain
A kernel for density estimation is an integrable function with . The kernel density estimator isIt has order when for and .
Differentiation gives the derivative kernel density estimatorDropping the negative square of its mean from the variance and using Tonelli and the substitution ,Thus and .
For , the Nikolsky class consists of functions with square-integrable derivatives through order andusing the equivalent finite-difference definition when is an integer. Integration by parts showsTaylor expansion in , cancellation of the moments through order for , and the generalized Minkowski inequality giveConsequentlywhere one admissible choice isHence .
The mean integrated squared error isthe sum of integrated variance and squared bias. The preceding bounds giveTaking proves the MISE rate for derivative kernel density estimationThus . Estimating a density of the same smoothness has variance order and rate ; estimating its derivative is harder because differentiation amplifies high-frequency noise, changing to .
At a target , the degree- local polynomial regression estimate minimizesand . Let have row
and let be diagonal with the kernel weights. Assume the local polynomial Gram matrix is positive definite. The weighted normal equations giveso it is a linear estimator in nonparametric regression.
and let be diagonal with the kernel weights. Assume the local polynomial Gram matrix is positive definite. The weighted normal equations giveso it is a linear estimator in nonparametric regression.
Now put . At and for the uniform kernel,For the usual bandwidth range , . The elementary power-sum formulas and show, for , thatThe limiting moment determinant isFor , the displayed errors can consume at most a fixed fraction of this value, sofor a universal and .
As printed, the claim “for all ” cannot hold for this design: if , all points receive weight , each , and the determinant is , not bounded below by a positive multiple of . The standard bandwidth restriction is therefore necessary for that intermediate assertion. The final bias bound remains harmless for , since and the finite design response means are uniformly bounded.
The polynomial reproduction property of local polynomial regression makes the local-linear weights reproduce both and . Writefor . The reproduced terms have zero bias. Using , the numerator contributed by the remainders is bounded byDivision by the determinant lower bound givesThe degree-zero Nadaraya–Watson estimator reproduces constants but not linear functions. The term therefore remains, and its boundary bias is bounded by ; this first-order boundary bias is the improvement that local linear fitting removes.
The squared-loss Le Cam two-point lemma says that for two experiments with scalar parameters ,Indeed, classify the data as when is closer to and as otherwise. On a classification error, the estimation error is at least . The sum of the two testing error probabilities is at least . Averaging the two risks and then bounding their maximum proves the result.
Let . The Hölder class on consists of functions with derivatives through order whose th derivative is Hölder of order with constant , with the standard integer-order convention.
Fix any . First compare the constant regression functions and with . Both belong to every Hölder class under the seminorm convention, and the normal-product divergence isPinsker and Le Cam therefore give a lower bound .
For the smoothness-dependent term, use the supplied smooth bump , translated one-sidedly near a boundary when necessary, and compareChoose its fixed normalization so that whenever . The Gaussian divergence satisfiesTakewith the bandwidth and amplitude truncated at constants when this expression leaves . Then the divergence remains bounded andThe same construction can be placed at every , with a one-sided bump at the endpoints. Pinsker and Le Cam, combined with the constant alternatives, provewhere depends only on .
Articles by others on the same topic
There are currently no matching articles.