Write and . A function bracket contains the measurable functions satisfying for every . Require its endpoints to be integrable and call its width. A sufficient condition is that, for every , finitely many brackets of width at most cover the whole class . Under this condition the uniform strong law from finite L1 bracketing statesThe conclusion holds outside a common measurable null set. If the supremum is not initially known to be measurable, this formulation means pathwise convergence on a measurable probability-one event; a pointwise separable function class, including the application below, has a measurable supremum. Pointwise brackets also ensure the sample inequalities hold simultaneously over the class.
To prove the result, choose a finite -cover , . If belongs to bracket , monotonicity of the empirical measure and of expectation givesandConsequentlyApply the strong law of large numbers to these finitely many integrable endpoints. On a probability-one event, the maximum tends to zero. Repeat with , , and intersect the countably many probability-one events. The limiting supremum is bounded by for every , hence is zero. This proves the uniform law of large numbers without a boundedness assumption on the class itself.
For the moment-generating function, use the empirical measure estimatorIf , this estimator and are both identically one. Otherwise, makes increasing in , and provides an integrable envelope. The dominated convergence theorem shows that is continuous on , hence uniformly continuous.
For any , choose a partition so that for every . If , then for every . These endpoint functions form finitely many integrable function brackets with the required widths. The just-proved uniform law of large numbers therefore givesBoth functions of are continuous; their supremum equals the supremum over a countable dense subset, so it is measurable. This proves uniform consistency of an empirical moment-generating function using only the observed sample.
For the unit-width box kernel, the convolution and kernel density estimator areandIn particular . The scaling gives . Even if is not square-integrable, Young's convolution inequality gives because and .
Using independence to eliminate the cross terms in the centred estimator, and Tonelli theorem to integrate the nonnegative variance, gives the exact integrated variance of a kernel density estimator:The Cauchy-Schwarz inequality now yieldsThis constant comes from the unit norm of the unscaled unit-width box kernel.
For the piecewise constant probability density function, integrability forces the constants on both unbounded outer intervals to be zero. Thus is bounded, has compact support, and has finitely many jumps. Let denote the jump at . For smaller than the least gap between consecutive breakpoints, the smoothing regions do not overlap. Outside these regions the convolution equals .
Within a region, put . For , the bias is ; for , it is . Endpoint values do not affect an norm. Integrating these two triangular errors givesThis is the box-kernel bias of a piecewise constant density. By the triangle inequality, with ,Balancing this bias-variance tradeoff with gives, for all sufficiently large ,
Use half-open intervals and setThe Haar scaling functions and Haar wavelets areFor fixed , each scaling function has squared integral , and different have disjoint supports, proving that this family is orthonormal. Each wavelet also has norm one and integral zero. At a fixed level, distinct wavelets have disjoint supports. At different levels, their dyadic supports are either disjoint or the finer support lies in a half on which the coarser wavelet is constant. The finer wavelet has integral zero, so their inner product vanishes. Finally, a wavelet with is supported inside a unit interval; it is orthogonal to its containing because its integral is zero, and to all other unit scaling functions because their supports are disjoint. Thus both the scaling family at each fixed level and the complete stated Haar family are orthonormal sets.
The Haar refinement identity isIt shows that the closed scaling spaces satisfy , where is the closed span of the level- wavelets. The union of the is dense in : continuous compactly supported functions can be approximated in by their dyadic cell averages, and such continuous functions are dense in . Hence the orthonormal Haar family is also a complete orthonormal basis.
For a locally integrable function, define the compact-support coefficient integralsThese exist without requiring . The two formulas for the Haar approximation areFor the wavelet sum is empty. For every fixed , these sums are locally finite, so they make sense pointwise and locally in . The refinement identities and their orthogonal two-by-two coefficient transformation establish equality by induction, even for merely locally integrable . If , this is also the orthogonal projection onto .
In particular, for and ,Now impose the symmetry and monotonicity conditions. The limit zero and decrease on imply for . For , the cell average and every value inside lie between and . ThereforeSumming over the positive half-line telescopes and gives a bound . Reflection maps each dyadic cell to another dyadic cell up to endpoints. Since is even, its Haar approximation is even almost everywhere as well, so the negative half-line has the same error. ConsequentlyThe argument proves integrability of the difference, even if itself is not integrable on the whole line. It is the symmetric monotone case of the Haar approximation error for a function of bounded variation.
For the unit-width box kernel, the order-zero local polynomial estimator minimizes, over constants ,A common factor such as in these weights does not change the minimizer. Put and . If , differentiating the quadratic gives its unique minimizer:The last expression is the Nadaraya–Watson estimator. Every is within of some design point . Since gives , that point is in the closed kernel window. Thus and the equality holds throughout the domain, including and .
For the error bound, first take the usual bandwidth range . The clipped window has length at least . Any closed interval of length in contains at least points of the design . Consequentlywhere the last inequality uses . This is a window occupancy for an equally spaced regression design bound that remains valid at the boundary.
Let , the Lipschitz constant supplied by the bounded-derivative assumption. The bias is bounded byIndependence and the variance bound giveBy the Cauchy-Schwarz inequality and the triangle inequality,Thus works uniformly in and the design size in this bandwidth range. The same clipped-length proof in fact works for .
There is a genuine omission in the unrestricted formulation: an upper restriction on bandwidth is necessary for the displayed variance scale. Take , , and . The window contains the sole observation for every , and holds. Its mean absolute error is , whereas the proposed variance term is . No fixed can satisfy this for unbounded .
A valid statement for every satisfying uses . The clipped window has length at least , so , while the preceding coverage argument gives . If , the first bound gives ; if , the second does. Therefore the same bias and variance calculation proves the unrestricted correctionThis mean absolute error of local constant regression bound reduces to the requested rate for ordinary small bandwidths and saturates its stochastic term when the window covers the whole design.
Articles by others on the same topic
There are currently no matching articles.