Let index the closest training covariates to . The K-nearest neighbors algorithm estimates
and predicts one when this average is at least .
Small gives low smoothing bias but high sampling variance and a jagged decision boundary. Larger averages more labels, reducing variance and producing a smoother boundary, but it mixes increasingly distant covariates and raises bias. The optimal balance depends on sample size, dimension, and smoothness of , and is commonly selected by cross-validation.

Articles by others on the same topic (0)

There are currently no matching articles.