LDA assumes with prior probabilities and
where the class means may differ but the nonsingular covariance matrix is common. Bayes' rule chooses the class maximizing . Cancelling terms common to every class gives the discriminant
The boundary between classes and is , an affine hyperplane with normal . Hence the Bayes rule is a linear classifier.
The fitted LDA rule replaces by class proportions, class sample means, and the pooled within-class covariance, then maximizes the resulting . Means and covariances have unbounded sensitivity, so a gross outlier can substantially move every LDA boundary. A soft-margin linear support vector machine uses hinge loss; observations beyond the correctly classified margin cease contributing, though mislabeled or extreme points can still matter according to the penalty . Thus the SVM is generally more robust to well-classified extremes, while neither method is automatically robust to adversarial outliers.
When , the pooled within-class covariance has rank at most , so it is singular and ordinary LDA is not defined. For every ,
is positive definite because . Replacing the covariance by this matrix gives well-defined regularized LDA. Choose for predictive performance by cross-validation.
Let be a positive-definite kernel with feature map into a reproducing-kernel Hilbert space. Perform regularized LDA on , replacing the within-class covariance operator by with , which is invertible. The representer property expresses all required inner products and discriminants through the Gram matrix and vectors . This is Kernel LDA. For a nonlinear kernel such as the Gaussian radial-basis kernel, its affine boundaries in feature space pull back to nonlinear boundaries in the original input space.

Articles by others on the same topic (0)

There are currently no matching articles.