Solution (source code)

= Solution

<Principal component analysis> centers the word-count vectors, diagonalizes their sample <covariance matrix>, and projects onto the eigenvectors with the largest eigenvalues. Fitting logistic regression to $d<n$ principal-component scores removes exact collinearity and yields a lower-dimensional design.

Each principal component is an unsupervised linear combination of many words chosen to explain predictor variance; it is not a word selected for association with spam. The dimension can be chosen from a scree plot or cumulative explained variance, but cross-validating the downstream classification loss better targets prediction.