Principal component analysis centers the word-count vectors, diagonalizes their sample covariance matrix, and projects onto the eigenvectors with the largest eigenvalues. Fitting logistic regression to principal-component scores removes exact collinearity and yields a lower-dimensional design.
Each principal component is an unsupervised linear combination of many words chosen to explain predictor variance; it is not a word selected for association with spam. The dimension can be chosen from a scree plot or cumulative explained variance, but cross-validating the downstream classification loss better targets prediction.
Articles by others on the same topic
There are currently no matching articles.