For a cumulative distribution function , its quantile function is the generalized inversewith the infimum allowed to be . For observations , the empirical distribution function isIf are the order statistics, then
The Bennett inequality says that if are independent, , almost surely, and , thenTo prove it, convexity of on , followed by the power-series bound for centered , givesIndependence and the Chernoff bound therefore yieldThe minimizing value satisfies , namely . Substitution gives the stated exponent.
For independent , the uniform order statistic satisfiesSet . The event means that at least sample points lie in . If this occurs, some -element subset consists entirely of such points. The union bound givesThis is exactly the required inequality for .
Put and . The density Hölder class consists of nonnegative functions integrating to one, with derivatives through order , such thatThis is the density version of the Hölder class. A kernel for density estimation is an integrable function with . It has order when for and .
Choose a bounded kernel of order at least with , and use the kernel density estimatorTaylor's theorem at and the vanishing kernel moments cancel every polynomial term below the remainder. Hence, for a constant depending only on and the fixed kernel,
We also need a uniform density bound. The standard Hölder interpolation argument combines nonnegativity, , and the Hölder constraint to giveIndeed, near a point where is close to its maximum , Taylor's theorem and the derivative bounds implied by the Hölder constraint keep of order on an interval of length comparable to ; integrating over that interval gives .
Using this bound and independence,ThusChooseBoth terms then have order . Since the infimum over all measurable estimators is no larger than the risk of this particular estimator,This is the pointwise minimax rate for Hölder density estimation upper bound.
Write , , and defineThe degree- local polynomial regression fit minimizesAssume its local polynomial Gram matrixis positive definite. The normal equations then give
Because the polynomial is written in the scaled coordinate, the local polynomial derivative estimator is . If , thenFor a polynomial of degree at most , the local least-squares fit to is exactly the Taylor polynomial , so its scaled linear coefficient is . Therefore the polynomial reproduction property of local polynomial regression gives
Positive definiteness and the fixed finite-dimensional basis provide a number such thatSince outside ,and
The regular design has at most points in this window when . Since the errors are independent with variance at most ,
Finally let and take the degree- Taylor polynomial of at . The Hölder class remainder satisfiesPolynomial reproduction removes from the bias. On the kernel window the remainder is at most , so the weight-sum bound gives
The total variation distance and Kullback-Leibler divergence arewith the divergence when is not absolutely continuous with respect to .
One squared-loss form of Assouad's lemma is as follows. Suppose is an Assouad hypercube such thatand every pair of neighboring vertices satisfies . ThenTo prove it, decode from a nearest hypercube vertex. The separation condition converts estimation loss into a constant multiple of its Hamming error. Average this error under the uniform prior on and sum coordinatewise. For each coordinate, the two equally weighted mixtures with that bit equal to zero or one have total variation at most by convexity. The minimum average error of any binary test is , which gives the displayed bound after the nearest-vertex factor.
We now build such a hypercube inside the monotone cone. Let and . Divide the first coordinates into consecutive blocks. For , set on block and extend the remaining coordinates at the last level. If are sufficiently small universal constants, every signal is nondecreasing and belongs to .
Neighboring vertices differ by on one block, so their squared Euclidean separation is . Their product experiments differ only on that block and haveBy Pinsker's inequality, choosing small makes every neighboring total variation distance at most, say, .
Apply Assouad's lemma with normalized squared loss . Here , and henceThis proves the lower half of the minimax rate for isotonic sequence estimation.
Articles by others on the same topic
There are currently no matching articles.