Nearest neighbour density estimation 2026-10-05
In one dimension, a nearest neighbour density estimator is . It uses the nearest neighbour distance to select an interval containing a fixed number of observations. The factor corresponds to the exact expected value of the associated uniform order statistic.
Nearest neighbour distance 2026-10-05
For observations , the nearest neighbour distance of order at is the th order statistic of the distances . An atomless distance distribution function ensures ties occur with probability zero.
Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 210 1 Solution Created 2026-10-03 Updated 2026-10-05
For , the nearest neighbour distance is the th order statistic of . Each has continuous distribution function on . The probability integral transform makes independent and identically distributed random variables with the uniform distribution on . Monotonicity of gives , a uniform order statistic. Strict monotonicity is not needed at this stage.
For , counting observations below givesDifferentiating and cancelling consecutive terms gives the Beta distribution:The Gamma function identity identifies this with the stated normalization. Ratios of Beta function integrals yield the momentsThis is the probability content of a nearest neighbour ball identity, specialized to intervals on the real line.
Now write . The Lipschitz bound givesThusStrict positivity of the probability density function makes continuous and strictly increasing from zero to one, so its inverse function is defined for every .
For the requested conditional calculation, suppose . At the preceding bound givesTherefore , and substituting in the approximation proves the inverse probability content bound for a Lipschitz density:The same argument works, without , on the meaningful local range with .
Since the nearest neighbour density estimator obeys , its conditional bound follows from the Beta distribution moments:Hence, under the printed location condition,The conditional inverse function bound also bounds , so the expected value in this calculation is finite under those formal hypotheses.
There is, however, a genuine flaw in the original PDF: the printed condition has no admissible points when is strictly positive on all of . A probability density function on all of cannot have Lipschitz constant zero, so . Its Lipschitz bound forcesIntegrating this triangular lower envelope gives the Lipschitz density height boundThe inequality is strict because the probability density function has positive mass outside the finite interval. Thus everywhere. The last two conditional conclusions above are valid implications but vacuous as printed. Dropping strict positivity gives only , with equality forcing the entire probability density function to be the normalized triangular envelope. The local inverse bound does not justify replacing the printed condition by its reverse in the final expected value bound, since ranges over all of .