Integrated mean squared error Created 2026-10-05 Updated 2026-10-06
The integrated mean squared error is over a specified domain . The Tonelli theorem permits integration of the pointwise bias-variance decomposition of mean squared error.
An uncorrected nonnegative symmetric kernel density estimator can leak an order-one amount across a jump in a probability density function. If its transition region has width proportional to the smoothing bandwidth , the integrated mean squared error includes squared bias of an estimator of order , rather than the order familiar for twice-smooth interior densities. For the unit-width uniform window and the exponential distribution of rate one, the integrated squared bias of an estimator is .
Put and , where is a nonnegative integer and is the smoothing bandwidth. The local polynomial estimator is the intercept of the weighted least squares fit
When the local polynomial Gram matrix is invertible, the normal equations give
A singular local polynomial Gram matrix requires a selection convention and may leave the intercept unidentified; an arbitrary nonnegative regression kernel alone does not guarantee uniqueness. At , provided , the Nadaraya–Watson estimator is
Multiplying every weight by changes neither fit.
For the uniform kernel, let and . Then . For and , the interval is inside , and the regular design has
The same lower bound holds when the interval starts at zero, despite the absence of the design point zero. In particular the Nadaraya–Watson estimator is defined throughout the integration interval.
For every Lipschitz continuous mean function in , the bias of an estimator and variance satisfy
Here the variance uses the independence of the errors in the fixed-design nonparametric regression model. The bias-variance decomposition of mean squared error and Tonelli theorem give a slightly stronger integrated mean squared error bound than required:
Thus
The printed integration interval represents a nonnegative risk function for . For , the written integral has reversed limits and is nonpositive, making the displayed upper bound trivial; it should not be interpreted as an integrated mean squared error in that range. For , evaluation outside would additionally require a specified extension of the mean function. The subsequent optimization restricts to .
Minimizing the stated right side gives
For sufficiently large , and . For example take . Therefore the upper bound holds with
For the lower bound, the PDF has the mean squared error ; the TeX erroneously moves the square outside the expected value. The PDF's suggested smooth bump function is , with value zero at the endpoints, rather than the corrupted TeX formula. A triangular alternative suffices because only Lipschitz continuity is needed.
The Le Cam two-point lemma says that for two data laws with parameter separation , every estimator satisfies
Indeed, classify according to the nearer parameter value. Misclassification incurs squared-error loss at least , and the sum of the two error probabilities is at least one minus the total variation distance. Averaging the two risk functions proves the stated form.
Let , , and use the two-point lower bound with a triangular bump
Both functions are in , and . Assume first that and . At most design points meet the bump's support. The Kullback-Leibler divergence between normal distributions and chain rule for relative entropy give
The Pinsker's inequality therefore gives . The Le Cam two-point lemma yields the pointwise minimax rate for Lipschitz regression
This universal constant is valid for .
If the printed lower bound is read as applying to all positive integers , it also holds with a smaller constant depending on the fixed . For , use the constant mean functions and . They have Kullback-Leibler divergence , so Pinsker's inequality and the Le Cam two-point lemma give minimax risk at least . Combining the two ranges gives the explicit all- choice
The universal asymptotic constant and this parameter-dependent all- constant are distinct conventions; the latter supplies the assertion without inserting an unstated large- restriction.