= Solution
<Principal component analysis> replaces the original coordinates by orthogonal linear combinations, ordered so that the first few retain as much <variance> as possible. It gives a lower-dimensional summary for visualization or compression and reveals the main directions of variation. A direction of large <variance> need not be the direction most useful for classification.
Write $\mu=\mathbb EX$. For a unit vector $u$, the centered score $Y=u^T(X-\mu)$ has <variance>
$$
\operatorname{Var}(Y)=u^T\Sigma u,\qquad u^Tu=1.
$$
Maximizing this <Rayleigh quotient> with a Lagrange multiplier gives
$$
\nabla_u\bigl[u^T\Sigma u-\lambda(u^Tu-1)\bigr]=2\Sigma u-2\lambda u=0.
$$
Thus a maximizing direction must be an <eigenvector> of the <covariance matrix>. To identify the maximum, use the <spectral theorem for real symmetric matrices> to choose an <orthonormal basis> of <eigenvectors> $v_1,\ldots,v_p$, with <eigenvalues> $\lambda_1\geq\cdots\geq\lambda_p\geq0$. Writing $u=\sum_jc_jv_j$ gives
$$
u^T\Sigma u=\sum_j\lambda_jc_j^2\leq\lambda_1\sum_jc_j^2=\lambda_1.
$$
Equality is attained at $u=v_1$. Consequently the first <population principal component> is
$$
\boxed{Y_1=v_1^T(X-\mu),\qquad\operatorname{Var}(Y_1)=\lambda_1.}
$$
For each subsequent <principal component>, maximize the same <variance> over unit coefficient vectors orthogonal to all the previously chosen ones. Restricting the expansion above to the remaining <eigenvectors> gives $v_2$, then $v_3$, and so on. Thus
$$
\boxed{Y_j=v_j^T(X-\mu),\qquad\operatorname{Cov}(Y_i,Y_j)=v_i^T\Sigma v_j=\lambda_j\mathbf1_{\{i=j\}}.}
$$
The <principal components> are uncorrelated; they need not be independent unless additional distributional assumptions, such as joint Gaussianity, hold. Their total <variance> is $\operatorname{tr}\Sigma=\sum_j\lambda_j$. Signs of <eigenvectors> are arbitrary, and a repeated <eigenvalue> permits any <orthonormal basis> in its eigenspace. In <sample principal component> calculations, replace $\mu$ by the <sample mean> and $\Sigma$ by the <sample covariance matrix>.
Back to article page