Solution (source code)

= Solution

<Agglomerative hierarchical clustering> starts with each observation as a singleton <cluster in cluster analysis>, repeatedly joins the closest pair of current clusters, and continues until one remains. The input is a <dissimilarity matrix>, so choosing variable scales and a scientifically sensible dissimilarity is part of the analysis.

Different linkages define closeness differently:
$$
\begin{aligned}
d_{\mathrm{single}}(A,B)&=\min_{i\in A,j\in B}d_{ij},\\
d_{\mathrm{complete}}(A,B)&=\max_{i\in A,j\in B}d_{ij},\\
d_{\mathrm{average}}(A,B)&=\frac1{|A||B|}\sum_{i\in A,j\in B}d_{ij}.
\end{aligned}
$$
<Single-linkage clustering> can connect elongated chains through nearest neighbors; <complete-linkage clustering> emphasizes compact clusters; <average-linkage clustering> averages all cross-pair distances. For Euclidean observations, <Ward minimum-variance clustering> instead chooses the smallest increase in within-cluster squared error,
$$
\Delta(A,B)=\frac{|A||B|}{|A|+|B|}\|\bar x_A-\bar x_B\|^2.
$$
This follows by expanding squared deviations around the merged mean and using the zero sums of within-cluster residuals.

A <dendrogram> plots merge heights. In the lower-left sketch, two close pairs first merge at small heights, then join across a much larger separation; a horizontal cut between these heights yields two groups. \b[The hierarchy gives nested partitions, not a uniquely established number of populations.] Examine sensitivity to scaling, linkage and ties; once an agglomerative merge is made, the method does not undo it. The absence of a probability model also means that a visually attractive <dendrogram> is not itself a significance test.