Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2015/iii/paper-33/6/c/solution

The three-component finite Gaussian mixture with a common variance has the largest displayed maximized log-likelihood, but it also has more fitted parameters. For components there are free weights, means and one common variance, making residual-distribution parameters. Applying the nominal Akaike information criterion and Bayesian information criterion to the three residual fits gives
Both criteria favour the two-component model. It improves the likelihood substantially over one normal component; the third component gains only in log-likelihood at a cost of two further parameters. The Akaike information criterion difference between two and three components is small, about , so that criterion alone gives only a slight preference; the Bayesian information criterion preference for two is stronger. Their nearly indistinguishable overlaid curves also support retaining the simpler two-component fit.
These calculations are model-selection aids, not an exact test. Ordinary chi-squared calibration of a likelihood-ratio test for the number of mixture components is invalid in general: under a smaller-component null, some weights lie on the boundary and extra-component parameters are unidentified. This is nonregular mixture model selection. A parametric bootstrap or predictive cross-validation is preferable for a formal comparison, and should account for the mean-adjustment stage. The table uses the provided residual likelihoods and their nominal parameter counts; first-stage regression uncertainty and possible local EM maxima remain qualifications. On the supplied evidence, select two components without claiming that two distinct biological groups have been established.

New to topics? Read the docs here!