For and centered data, the objective is the Rayleigh quotient
This is the first-direction optimization in principal component analysis. Therefore is any unit eigenvector of the sample covariance matrix corresponding to its largest eigenvalue.
Write . By the definition of the kernel matrix,
and
The irrelevant positive factor gives exactly the stated constrained optimization.
Let , where is the largest eigenvalue and . The generalized Rayleigh quotient is maximized by
up to sign and addition of a vector in , which does not change . Repeated leading eigenvectors give further principal directions.
Kernel principal component analysis diagonalizes the centered by kernel matrix instead of an explicit covariance operator in a possibly infinite-dimensional feature space. A new point has coordinate
so the leading coordinates require only evaluations of the positive-semidefinite kernel.
Running the K-nearest neighbors algorithm in a truncated collection of these coordinates can remove low-variance noise, reduce effective dimension, and allow a nonlinear boundary in the original covariates. This is the kernel trick: every feature-space inner product needed for fitting and projection is replaced by without constructing explicitly.

Articles by others on the same topic (0)

There are currently no matching articles.