Solution (source code)

= Solution

Let $N_k(x)$ index the $k$ closest training covariates to $x$. The <K-nearest neighbors algorithm> estimates
$$
\widehat\eta_1(x)=\frac1k\sum_{i\in N_k(x)}Y_i
$$
and predicts one when this average is at least $1/2$.

Small $k$ gives low smoothing bias but high sampling variance and a jagged <decision boundary>. Larger $k$ averages more labels, reducing variance and producing a smoother boundary, but it mixes increasingly distant covariates and raises bias. The optimal balance depends on sample size, dimension, and smoothness of $\eta_1$, and is commonly selected by <cross-validation>.