For a density with finitely many jumps at separated points and the unit-width box kernel, take smaller than all consecutive jump spacings. The smoothing bias is supported in disjoint length- neighbourhoods of the jumps. At each jump its two sides form triangles of height , and exact integration gives
Combined with the integrated variance of a kernel density estimator, this gives expected error at .
Use the usual probability-kernel convention: , , and , in addition to symmetry and compact support. These conditions are needed for the stated finite bound; symmetry and compact support by themselves do not impose normalization or nonnegativity. Interpret the differentiability assumption in the usual Sobolev space sense , or assume classical regularity sufficient for the integral Taylor remainder. The kernel density estimator is
The bias-variance decomposition of mean squared error, followed by Tonelli theorem, yields
Independence and a change of variables give the exact integrated variance of a kernel density estimator
All changes of integration order are justified either by nonnegativity or by the displayed finite integrals.
For the bias of a kernel density estimator, symmetry gives . Taylor's integral remainder, with translations understood in , gives
The Minkowski inequality and translation invariance of the L2 norm imply
where . Consequently we obtain the slightly stronger conclusion
Since , this proves the requested bound. For signed kernels, the same argument has in place of ; the printed bound is not generally justified merely by signed-kernel symmetry. For example, let and . This symmetric compactly supported signed kernel has integral one and , but its fourth moment is nonzero. For a normal density its convolution has nonzero bias at any fixed positive bandwidth. Sending would therefore contradict a bound containing only the integrated variance term.
Writing the requested bound as , differentiation gives when . Thus gives mean integrated squared error of order . With the sharper constant one instead has ; the rate is identical.
A direct heuristic for is to choose a twice differentiable pilot kernel , estimate by the second derivative kernel density estimator
and integrate its square. Its integrated variance contains a diagonal term , which should not be mistaken for the target. A useful corrected U-statistic removes that diagonal:
Its expectation is . Thus smoothing bias of an estimator, rather than the diagonal contribution, remains. For a smooth compactly supported probability pilot kernel, the same variance calculation gives
The second term tends to zero because convolution with a shrinking probability kernel approximates every function. Thus and make the raw squared pilot estimate consistent for the quadratic density derivative functional. If , then
so the diagonal-corrected U-statistic is consistent under the same choice. Consistency already uses the assumed ; the additional third-derivative hypothesis is relevant to more ambitious rate heuristics.
Three bounded continuous derivatives do not justify a general root- estimation rate for . For a smooth perturbation with compact support, the directional derivative is
where the last expression requires a fourth derivative in the appropriate weak sense. A regular root- argument would need an influence function proportional to , with finite variance under . Three derivatives do not supply that condition.
For a more quantitative heuristic, impose the additional assumption , use a smooth Fourier cutoff of width , and estimate the quadratic functional by its off-diagonal U-statistic. The tail bias of an estimator is because the Plancherel theorem weights this functional by frequency to the fourth power. The degenerate part of its standard deviation is : the squared norm of the quadratic kernel scales as . At , these two terms balance at and have size , already slower than ; any additional first-order variance cannot improve that balance. To make both terms requires and , which are compatible only when . This is an illustrative smoothness calculation under extra Sobolev and moment assumptions, not a rate theorem asserted solely from bounded third derivatives. Particular much smoother probability density functions may of course permit root- estimation.
For the unit-width box kernel, the convolution and kernel density estimator are
and
In particular . The scaling gives . Even if is not square-integrable, Young's convolution inequality gives because and .
Using independence to eliminate the cross terms in the centred estimator, and Tonelli theorem to integrate the nonnegative variance, gives the exact integrated variance of a kernel density estimator:
The Cauchy-Schwarz inequality now yields
This constant comes from the unit norm of the unscaled unit-width box kernel.
For the piecewise constant probability density function, integrability forces the constants on both unbounded outer intervals to be zero. Thus is bounded, has compact support, and has finitely many jumps. Let denote the jump at . For smaller than the least gap between consecutive breakpoints, the smoothing regions do not overlap. Outside these regions the convolution equals .
Within a region, put . For , the bias is ; for , it is . Endpoint values do not affect an norm. Integrating these two triangular errors gives
This is the box-kernel bias of a piecewise constant density. By the triangle inequality, with ,
Balancing this bias-variance tradeoff with gives, for all sufficiently large ,
Differentiating a twice differentiable kernel density estimator twice estimates . Its integrated variance of a kernel density estimator analogue is at most . This explains why derivative estimation requires more smoothing than ordinary density estimation.