Put for . This is the quantile function, rather than an ordinary inverse requiring strict increase. The limits of the distribution function at infinity make finite, and the fact that is a right-continuous function ensures . ConsequentlyIncreasing shrinks the set in the infimum, so the quantile function is a monotone function. To prove the left continuity of the quantile function, fix and let . If , choose . Then , so some satisfies , a contradiction. Thus as . At a flat stretch of the distribution function, the quantile function may jump immediately to the right, consistent with this left-continuous convention.
For the uniform order statistic, count the observations at most . That count has the binomial distribution with parameters , givingDifferentiating makes adjacent terms telescope: use and . The probability density function is thereforeIn particular, has the Beta distribution with parameters .
Write for the specified median. Continuity of the distribution function gives and no mass at , even if the distribution function has a flat stretch there. Thus has the binomial distribution with parameters . Except on a null event, the order-statistic confidence interval for a median covers exactly when . Symmetry of the binomial distribution makes its two failure probabilities equal. Alternatively, the probability integral transform givesHence the coverage is exactly , withThe asymmetric open/closed endpoint convention does not change the coverage because the continuous probability distribution assigns no mass to the median.
For independent and identically distributed random variables sampled from a probability density function , a kernel density estimator with kernel for density estimation and smoothing bandwidth isThe usual kernel for density estimation is nonnegative; finite makes the integrated variance finite. The mean integrated squared error and its bias-variance decomposition of mean squared error areThe interchange is justified by the Tonelli theorem.
For the exponential distribution and the specified unit-width uniform kernel, put . The expected value is . Intersecting this interval with the positive half-line givesOn the true probability density function vanishes, whereasHere for . This leakage across the support boundary proves the integrated squared bias from a density jump:In particular, and suffice. The phenomenon differs from a smooth interior bias of a kernel density estimator: the order-one boundary error persists over a region of width proportional to .
By independence, . Integrating the first term and using yieldsAlso : the kernel density estimator mean is an average of translates of , so the Jensen inequality followed by the Tonelli theorem bounds its squared integral.
To optimize over every , we must establish that useful smoothing bandwidths approach zero, rather than optimize only a formal small- expression. An exact calculation supplies that justification. The autocorrelation integral of the exponential distribution density isAveraging over the two uniform windows, or integrating the piecewise expression for , givesThese formulas show that is continuous on , is strictly positive there because on , and tends to as . Consequently, is bounded away from zero on for every . Its expansion agrees with the supplied .
The choice givesChoose smoothing bandwidths within of the infimum. Their mean integrated squared error tends to zero, and nonnegativity of the integrated variance gives , hence . For any , eventuallyThe last step is the arithmetic-geometric mean inequality. Taking the lower limit and then matches the upper bound. ThusThe original PDF has this constant; the TeX aid's is a transcription error.
For fixed-design nonparametric regression, let and . The degree- local polynomial estimator takes the fitted intercept , whereUse a nonnegative regression kernel and an invertible local polynomial Gram matrix to obtain a unique fit. For , solving the single weighted least squares normal equation gives the local constant estimator:At the interior point , use the standard unit-integral regression kernel convention . Otherwise the leading variance below is ; unlike a kernel density estimator, the local constant estimator itself is unchanged by rescaling .
The local constant estimator is a linear estimator in nonparametric regression with weights . The independent unit-variance errors implyTo control the bias of an estimator, differentiability at this one point is enough. WriteSymmetry makes , and the supplied moment approximation gives . Therefore the linear contribution to the bias of an estimator is . The remaining contribution has absolute value at most , because only receive positive weight. This is the interior first-order bias cancellation for local constant regression, and gives squared bias of an estimator . Combining with the variance provesFor fixed , take . It satisfies both asymptotic smoothing bandwidth conditions, and the displayed remainder, multiplied by , tends to zero for this fixed . ThusLetting proves the optimized scaled pointwise error tends to zero. The order of these limits matters: no remainder uniform in has been assumed.
For the lower bound over the whole function class, the absence of a uniform bound on derivatives permits a particularly direct argument. It also avoids assuming a normal distribution for errors when the model only supplies their first two moments. Fix and . Choose a smooth bump function on the interval, with , whose support contains no design point except possibly . Such a bump exists because the design is finite; at an endpoint use the restriction of a smooth bump on the real line. Every , , belongs to the differentiability class.
If is not a design point, all observation means are zero for every . The Le Cam two-point lemma applied to has total variation distance zero and gives a lower bound tending to infinity with . In fact, the supremum of the mean squared error is infinite at every such .
If is a design point, only its response depends on . The resulting location family is the single observation ; all remaining observations are independent of and give no extra information. The single-observation location minimax bound isHere is a proof covering discrete as well as continuous errors. Give the prior distribution , independently of the noise. Put and . Consider an easier experiment in which an oracle also reveals whether , and reveals itself on the complementary event. On the event and , every possible truncated noise value places within . The flat prior distribution then makes the Bayesian posterior noise law exactly its truncated original law. The smallest conditional squared-error loss is its conditional variance , attained by the posterior mean.
The probability of that event is at least for , by restricting further to . Thus every rule's supremum risk function is at least the original Bayes risk, which is at least the easier experiment's Bayes risk, and therefore at least . Let and then . Finite second moment and zero mean give . The rule has constant mean squared error one, proving the equality.
Consequently, at every , the unrestricted differentiability class satisfies the stronger conclusionThus works for every . There is no conflict with the preceding fixed-function limit: the worst functions, including arbitrarily narrow bumps, can vary with . This is the pointwise versus uniform risk distinction. A smoothness ball with a common derivative bound would pose a different minimax risk problem.
A non-degenerate probability distribution is not concentrated at a single point. In extreme value theory, a non-degenerate distribution function is max-stable if, for each integer , constants satisfy for every . This says that a normalized sample maximum has the same law as one observation. A distribution function belongs to the maximum domain of attraction of a non-degenerate if there are such thatat every continuity point of . Equivalently, has convergence in distribution to .
The extremal types theorem says that any such non-degenerate limit is max-stable and, up to a positive affine change of variable, is exactly one of the following:These are respectively the Gumbel distribution, Fréchet distribution, and negative Weibull distribution. The theorem is also known as the Fisher–Tippett–Gnedenko theorem.
Useful sufficient conditions can be expressed using the survival function and the right endpoint of a distribution . An infinite endpoint with for every gives the Fréchet distribution domain. A finite endpoint with as gives the negative Weibull distribution domain. For the Gumbel distribution, it suffices that a positive Gumbel auxiliary function gives for every real as . These are regular variation and exponential tail-ratio conditions; they need no proof here.
In case (i), on , so the finite-endpoint ratio is exactly . Therefore the domain is negative Weibull with shape . With and ,and the limit is one for .
In case (ii), has exponential tail ratio with the constant Gumbel auxiliary function . Therefore the domain is Gumbel. Choosing and givesIn case (iii), has regular variation of index , so the domain is Fréchet with shape . The convenient choices and givewith limit zero for .
Finally, define the empirical distribution function . On a sample with positive threshold and exactly observations strictly above , its survival function has . Using the Tonelli theorem to integrate the finite nonnegative sum givesThis is the Hill estimator as an empirical plug-in version of the tail-integral limit. It does not prove consistency for fixed ; that is a separate issue.
There is a finite-sample qualification because the question permits ties and does not assume positive observations. The formula requires and . If because the threshold is tied, direct substitution giveswhere terms equal to the threshold contribute zero. This is generally times the displayed Hill estimator, rather than the same estimator. If , the empirical plug-in denominator is zero. For a continuous probability distribution the no-tie condition holds almost surely; for a Fréchet distribution domain and fixed , the threshold is positive with probability tending to one. Thus the stated plug-in identity holds at a positive untied threshold, and the literal claim for all samples from an arbitrary distribution function needs this qualification.
Articles by others on the same topic
There are currently no matching articles.