= Solution
Write $x_r,x_t$ for two rows of the data matrix and $\delta=x_r-x_t$. Three possible dissimilarities for continuous variables are the following.
The <Euclidean distance> is
$$
\boxed{d_E(x_r,x_t)=\sqrt{\sum_{j=1}^p\delta_j^2}.}
$$
It is simple and has the usual geometric interpretation when coordinates have comparable units. Its disadvantage is sensitivity to the numerical scales: a change of units or a high-variance variable can dominate it. Strongly correlated coordinates can also give repeated weight to essentially the same information. Squared <Euclidean distance> is useful in some methods, but is not itself a metric because it need not satisfy the triangle inequality.
The <standardized Euclidean distance> is
$$
\boxed{d_S(x_r,x_t)=\sqrt{\sum_{j=1}^p\frac{\delta_j^2}{s_j^2}},}
$$
where $s_j^2$ is the <sample variance> of coordinate $j$ across individuals. It measures differences in standard-deviation units and is invariant under positive coordinate rescalings when the scales are estimated consistently. It is useful for differently measured variables, but requires nonzero <sample variances>, ignores correlations between coordinates and can give excessive influence to a low-variance noisy variable. With fixed positive scales it is a <Euclidean distance> after a diagonal change of coordinates, so it is a metric.
The <Mahalanobis distance> is
$$
\boxed{d_M(x_r,x_t)=\sqrt{\delta^TS^{-1}\delta},}
$$
where $S$ is the pooled <sample covariance matrix> of all individuals. If $S=CC^T$ is its <Cholesky decomposition>, then $d_M=\|C^{-1}\delta\|_2$: it is <Euclidean distance> in whitened coordinates. It adjusts for both scales and correlations, avoiding duplicate weight for highly correlated variables, and is invariant under nonsingular affine transformations when the <sample covariance matrix> is transformed accordingly. No <multivariate normal distribution> assumption is required merely to define the distance. However, $S$ must be a <positive-definite matrix>; when there are too few observations or dependent coordinates, its inverse is unavailable. Estimation of $S^{-1}$ can also be unstable and sensitive to outliers. These trade-offs determine whether covariance adjustment is preferable to the simpler distances.
Back to article page