For , the risk is
The Bayes classifier is , with risk .
The K-nearest neighbors algorithm takes the majority label among the training features closest to the query, with a stated tie rule. Its data-dependent risk is the conditional test error given the training sample, and denotes its expectation over that sample.
For one nearest neighbour, condition on a feature value and couple the coincident nearest feature as . The two labels are conditionally independent Bernoulli, so their mismatch probability is
Integration over gives
For each , the call to knn.cv performs leave-one-out nearest-neighbour classification and the third line stores the fraction of omitted observations misclassified. Choose a minimizer of ls, then fit the nearest-neighbour classifier with that to the full data.
A weighted nearest-neighbours classifier predicts one when
where neighbours are distance ordered. Under the usual smooth-density and smooth-regression assumptions, asymptotically optimal weights downweight distant neighbours, for example normalized positive parts of . The optimal weighted-nearest-neighbour theorem gives smaller leading asymptotic regret than equal weights.

Articles by others on the same topic (0)

There are currently no matching articles.