For a cumulative distribution function , its quantile function is the generalized inverse
with the infimum allowed to be . For observations , the empirical distribution function is
If are the order statistics, then
The Bennett inequality says that if are independent, , almost surely, and , then
To prove it, convexity of on , followed by the power-series bound for centered , gives
Independence and the Chernoff bound therefore yield
The minimizing value satisfies , namely . Substitution gives the stated exponent.
For independent , the uniform order statistic satisfies
Set . The event means that at least sample points lie in . If this occurs, some -element subset consists entirely of such points. The union bound gives
This is exactly the required inequality for .
Put and . The density Hölder class consists of nonnegative functions integrating to one, with derivatives through order , such that
This is the density version of the Hölder class. A kernel for density estimation is an integrable function with . It has order when for and .
Choose a bounded kernel of order at least with , and use the kernel density estimator
Taylor's theorem at and the vanishing kernel moments cancel every polynomial term below the remainder. Hence, for a constant depending only on and the fixed kernel,
We also need a uniform density bound. The standard Hölder interpolation argument combines nonnegativity, , and the Hölder constraint to give
Indeed, near a point where is close to its maximum , Taylor's theorem and the derivative bounds implied by the Hölder constraint keep of order on an interval of length comparable to ; integrating over that interval gives .
Using this bound and independence,
Thus
Choose
Both terms then have order . Since the infimum over all measurable estimators is no larger than the risk of this particular estimator,
This is the pointwise minimax rate for Hölder density estimation upper bound.
Write , , and define
The degree- local polynomial regression fit minimizes
Assume its local polynomial Gram matrix
is positive definite. The normal equations then give
Because the polynomial is written in the scaled coordinate, the local polynomial derivative estimator is . If , then
For a polynomial of degree at most , the local least-squares fit to is exactly the Taylor polynomial , so its scaled linear coefficient is . Therefore the polynomial reproduction property of local polynomial regression gives
Positive definiteness and the fixed finite-dimensional basis provide a number such that
Since outside ,
and
The regular design has at most points in this window when . Since the errors are independent with variance at most ,
Finally let and take the degree- Taylor polynomial of at . The Hölder class remainder satisfies
Polynomial reproduction removes from the bias. On the kernel window the remainder is at most , so the weight-sum bound gives
The total variation distance and Kullback-Leibler divergence are
with the divergence when is not absolutely continuous with respect to .
For and ,
If , direct integration on the intervals cut by and gives
Consequently
One squared-loss form of Assouad's lemma is as follows. Suppose is an Assouad hypercube such that
and every pair of neighboring vertices satisfies . Then
To prove it, decode from a nearest hypercube vertex. The separation condition converts estimation loss into a constant multiple of its Hamming error. Average this error under the uniform prior on and sum coordinatewise. For each coordinate, the two equally weighted mixtures with that bit equal to zero or one have total variation at most by convexity. The minimum average error of any binary test is , which gives the displayed bound after the nearest-vertex factor.
We now build such a hypercube inside the monotone cone. Let and . Divide the first coordinates into consecutive blocks. For , set on block
and extend the remaining coordinates at the last level. If are sufficiently small universal constants, every signal is nondecreasing and belongs to .
Neighboring vertices differ by on one block, so their squared Euclidean separation is . Their product experiments differ only on that block and have
By Pinsker's inequality, choosing small makes every neighboring total variation distance at most, say, .
Apply Assouad's lemma with normalized squared loss . Here , and hence
This proves the lower half of the minimax rate for isotonic sequence estimation.

Articles by others on the same topic (0)

There are currently no matching articles.