For a finite Gaussian mixture with a common variance, introduce independent latent labels with probabilities , , and conditional responses , with common . The observed density and log-likelihood are
The expectation-maximization algorithm replaces the difficult log of sums by an expected complete-data objective. Choose positive initial weights and variance and separated initial means. At iteration , the E-step computes the mixture responsibilities
The common normalizing factor cancels. For numerical stability the probabilities can be evaluated by subtracting the largest log weight before exponentiating.
With the old responsibilities held fixed, the expected complete-data log-likelihood, up to irrelevant constants, is
Let . A Lagrange multiplier for , weighted least squares for the means, and differentiation in the common variance give the M-step:
The variance update uses the new means and old responsibilities, with denominator , not a residual degrees-of-freedom adjustment: this is maximum likelihood estimation. Repeat the E- and M-steps until the observed log-likelihood and parameters stabilize.
By EM likelihood monotonicity, exact updates do not decrease the observed likelihood. They need not reach its global maximum, so use several starting configurations and keep the best converged fit, checking for empty or nearly empty components and vanishing variance. The labels are interchangeable; sorting means after fitting supplies an interpretable labeling. Distinct starting means do not guarantee that all fitted components remain distinct. With a common variance and the usual fixed small relative to distinct observations the model avoids the individual-component variance-collapse pathology of unrestricted Gaussian mixtures, but degenerate data or too many components still require attention.
The three-component finite Gaussian mixture with a common variance has the largest displayed maximized log-likelihood, but it also has more fitted parameters. For components there are free weights, means and one common variance, making residual-distribution parameters. Applying the nominal Akaike information criterion and Bayesian information criterion to the three residual fits gives
Both criteria favour the two-component model. It improves the likelihood substantially over one normal component; the third component gains only in log-likelihood at a cost of two further parameters. The Akaike information criterion difference between two and three components is small, about , so that criterion alone gives only a slight preference; the Bayesian information criterion preference for two is stronger. Their nearly indistinguishable overlaid curves also support retaining the simpler two-component fit.
These calculations are model-selection aids, not an exact test. Ordinary chi-squared calibration of a likelihood-ratio test for the number of mixture components is invalid in general: under a smaller-component null, some weights lie on the boundary and extra-component parameters are unidentified. This is nonregular mixture model selection. A parametric bootstrap or predictive cross-validation is preferable for a formal comparison, and should account for the mean-adjustment stage. The table uses the provided residual likelihoods and their nominal parameter counts; first-stage regression uncertainty and possible local EM maxima remain qualifications. On the supplied evidence, select two components without claiming that two distinct biological groups have been established.