Solution (source code)

= Solution

<LDA> assumes $Y\in\{1,\ldots,L\}$ with prior probabilities $\pi_l>0$ and
$$
X\mid Y=l\sim N_p(\mu_l,\Sigma),
$$
where the class means may differ but the nonsingular covariance matrix is common. Bayes' rule chooses the class maximizing $\log\pi_l+\log f_l(x)$. Cancelling terms common to every class gives the discriminant
$$
\delta_l(x)=x^T\Sigma^{-1}\mu_l
-\frac12\mu_l^T\Sigma^{-1}\mu_l+\log\pi_l.
$$
The boundary between classes $l$ and $m$ is $\delta_l(x)=\delta_m(x)$, an affine hyperplane with normal $\Sigma^{-1}(\mu_l-\mu_m)$. Hence the Bayes rule $\arg\max_l\delta_l(x)$ is a linear classifier.