For fixed-design nonparametric regression, let and . The degree- local polynomial estimator takes the fitted intercept , whereUse a nonnegative regression kernel and an invertible local polynomial Gram matrix to obtain a unique fit. For , solving the single weighted least squares normal equation gives the local constant estimator:At the interior point , use the standard unit-integral regression kernel convention . Otherwise the leading variance below is ; unlike a kernel density estimator, the local constant estimator itself is unchanged by rescaling .
The local constant estimator is a linear estimator in nonparametric regression with weights . The independent unit-variance errors implyTo control the bias of an estimator, differentiability at this one point is enough. WriteSymmetry makes , and the supplied moment approximation gives . Therefore the linear contribution to the bias of an estimator is . The remaining contribution has absolute value at most , because only receive positive weight. This is the interior first-order bias cancellation for local constant regression, and gives squared bias of an estimator . Combining with the variance provesFor fixed , take . It satisfies both asymptotic smoothing bandwidth conditions, and the displayed remainder, multiplied by , tends to zero for this fixed . ThusLetting proves the optimized scaled pointwise error tends to zero. The order of these limits matters: no remainder uniform in has been assumed.
For the lower bound over the whole function class, the absence of a uniform bound on derivatives permits a particularly direct argument. It also avoids assuming a normal distribution for errors when the model only supplies their first two moments. Fix and . Choose a smooth bump function on the interval, with , whose support contains no design point except possibly . Such a bump exists because the design is finite; at an endpoint use the restriction of a smooth bump on the real line. Every , , belongs to the differentiability class.
If is not a design point, all observation means are zero for every . The Le Cam two-point lemma applied to has total variation distance zero and gives a lower bound tending to infinity with . In fact, the supremum of the mean squared error is infinite at every such .
If is a design point, only its response depends on . The resulting location family is the single observation ; all remaining observations are independent of and give no extra information. The single-observation location minimax bound isHere is a proof covering discrete as well as continuous errors. Give the prior distribution , independently of the noise. Put and . Consider an easier experiment in which an oracle also reveals whether , and reveals itself on the complementary event. On the event and , every possible truncated noise value places within . The flat prior distribution then makes the Bayesian posterior noise law exactly its truncated original law. The smallest conditional squared-error loss is its conditional variance , attained by the posterior mean.
The probability of that event is at least for , by restricting further to . Thus every rule's supremum risk function is at least the original Bayes risk, which is at least the easier experiment's Bayes risk, and therefore at least . Let and then . Finite second moment and zero mean give . The rule has constant mean squared error one, proving the equality.
Consequently, at every , the unrestricted differentiability class satisfies the stronger conclusionThus works for every . There is no conflict with the preceding fixed-function limit: the worst functions, including arbitrarily narrow bumps, can vary with . This is the pointwise versus uniform risk distinction. A smoothness ball with a common derivative bound would pose a different minimax risk problem.
Articles by others on the same topic
There are currently no matching articles.