Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 6 c Solution Created 2026-10-03 Updated 2026-10-06
The three-component finite Gaussian mixture with a common variance has the largest displayed maximized log-likelihood, but it also has more fitted parameters. For components there are free weights, means and one common variance, making residual-distribution parameters. Applying the nominal Akaike information criterion and Bayesian information criterion to the three residual fits givesBoth criteria favour the two-component model. It improves the likelihood substantially over one normal component; the third component gains only in log-likelihood at a cost of two further parameters. The Akaike information criterion difference between two and three components is small, about , so that criterion alone gives only a slight preference; the Bayesian information criterion preference for two is stronger. Their nearly indistinguishable overlaid curves also support retaining the simpler two-component fit.
These calculations are model-selection aids, not an exact test. Ordinary chi-squared calibration of a likelihood-ratio test for the number of mixture components is invalid in general: under a smaller-component null, some weights lie on the boundary and extra-component parameters are unidentified. This is nonregular mixture model selection. A parametric bootstrap or predictive cross-validation is preferable for a formal comparison, and should account for the mean-adjustment stage. The table uses the provided residual likelihoods and their nominal parameter counts; first-stage regression uncertainty and possible local EM maxima remain qualifications. On the supplied evidence, select two components without claiming that two distinct biological groups have been established.