For a density with finitely many jumps at separated points and the unit-width box kernel, take smaller than all consecutive jump spacings. The smoothing bias is supported in disjoint length- neighbourhoods of the jumps. At each jump its two sides form triangles of height , and exact integration givesCombined with the integrated variance of a kernel density estimator, this gives expected error at .
Past exam of the mathematics course of the University of Cambridge 2012 iii Paper 39 3 Solution Created 2026-10-03 Updated 2026-10-07
Use the usual probability-kernel convention: , , and , in addition to symmetry and compact support. These conditions are needed for the stated finite bound; symmetry and compact support by themselves do not impose normalization or nonnegativity. Interpret the differentiability assumption in the usual Sobolev space sense , or assume classical regularity sufficient for the integral Taylor remainder. The kernel density estimator isThe bias-variance decomposition of mean squared error, followed by Tonelli theorem, yieldsIndependence and a change of variables give the exact integrated variance of a kernel density estimatorAll changes of integration order are justified either by nonnegativity or by the displayed finite integrals.
For the bias of a kernel density estimator, symmetry gives . Taylor's integral remainder, with translations understood in , givesThe Minkowski inequality and translation invariance of the L2 norm implywhere . Consequently we obtain the slightly stronger conclusionSince , this proves the requested bound. For signed kernels, the same argument has in place of ; the printed bound is not generally justified merely by signed-kernel symmetry. For example, let and . This symmetric compactly supported signed kernel has integral one and , but its fourth moment is nonzero. For a normal density its convolution has nonzero bias at any fixed positive bandwidth. Sending would therefore contradict a bound containing only the integrated variance term.
Writing the requested bound as , differentiation gives when . Thus gives mean integrated squared error of order . With the sharper constant one instead has ; the rate is identical.
A direct heuristic for is to choose a twice differentiable pilot kernel , estimate by the second derivative kernel density estimatorand integrate its square. Its integrated variance contains a diagonal term , which should not be mistaken for the target. A useful corrected U-statistic removes that diagonal:Its expectation is . Thus smoothing bias of an estimator, rather than the diagonal contribution, remains. For a smooth compactly supported probability pilot kernel, the same variance calculation givesThe second term tends to zero because convolution with a shrinking probability kernel approximates every function. Thus and make the raw squared pilot estimate consistent for the quadratic density derivative functional. If , thenso the diagonal-corrected U-statistic is consistent under the same choice. Consistency already uses the assumed ; the additional third-derivative hypothesis is relevant to more ambitious rate heuristics.
Three bounded continuous derivatives do not justify a general root- estimation rate for . For a smooth perturbation with compact support, the directional derivative iswhere the last expression requires a fourth derivative in the appropriate weak sense. A regular root- argument would need an influence function proportional to , with finite variance under . Three derivatives do not supply that condition.
For a more quantitative heuristic, impose the additional assumption , use a smooth Fourier cutoff of width , and estimate the quadratic functional by its off-diagonal U-statistic. The tail bias of an estimator is because the Plancherel theorem weights this functional by frequency to the fourth power. The degenerate part of its standard deviation is : the squared norm of the quadratic kernel scales as . At , these two terms balance at and have size , already slower than ; any additional first-order variance cannot improve that balance. To make both terms requires and , which are compatible only when . This is an illustrative smoothness calculation under extra Sobolev and moment assumptions, not a rate theorem asserted solely from bounded third derivatives. Particular much smoother probability density functions may of course permit root- estimation.
Past exam of the mathematics course of the University of Cambridge 2013 iii Paper 33 2 Solution Created 2026-10-03 Updated 2026-10-07
For the unit-width box kernel, the convolution and kernel density estimator areandIn particular . The scaling gives . Even if is not square-integrable, Young's convolution inequality gives because and .
Using independence to eliminate the cross terms in the centred estimator, and Tonelli theorem to integrate the nonnegative variance, gives the exact integrated variance of a kernel density estimator:The Cauchy-Schwarz inequality now yieldsThis constant comes from the unit norm of the unscaled unit-width box kernel.
For the piecewise constant probability density function, integrability forces the constants on both unbounded outer intervals to be zero. Thus is bounded, has compact support, and has finitely many jumps. Let denote the jump at . For smaller than the least gap between consecutive breakpoints, the smoothing regions do not overlap. Outside these regions the convolution equals .
Within a region, put . For , the bias is ; for , it is . Endpoint values do not affect an norm. Integrating these two triangular errors givesThis is the box-kernel bias of a piecewise constant density. By the triangle inequality, with ,Balancing this bias-variance tradeoff with gives, for all sufficiently large ,
Second derivative kernel density estimator 2026-10-07
Differentiating a twice differentiable kernel density estimator twice estimates . Its integrated variance of a kernel density estimator analogue is at most . This explains why derivative estimation requires more smoothing than ordinary density estimation.