For a finite Gaussian mixture with a common variance, introduce independent latent labels with probabilities , , and conditional responses , with common . The observed density and log-likelihood areThe expectation-maximization algorithm replaces the difficult log of sums by an expected complete-data objective. Choose positive initial weights and variance and separated initial means. At iteration , the E-step computes the mixture responsibilitiesThe common normalizing factor cancels. For numerical stability the probabilities can be evaluated by subtracting the largest log weight before exponentiating.
With the old responsibilities held fixed, the expected complete-data log-likelihood, up to irrelevant constants, isLet . A Lagrange multiplier for , weighted least squares for the means, and differentiation in the common variance give the M-step:The variance update uses the new means and old responsibilities, with denominator , not a residual degrees-of-freedom adjustment: this is maximum likelihood estimation. Repeat the E- and M-steps until the observed log-likelihood and parameters stabilize.
By EM likelihood monotonicity, exact updates do not decrease the observed likelihood. They need not reach its global maximum, so use several starting configurations and keep the best converged fit, checking for empty or nearly empty components and vanishing variance. The labels are interchangeable; sorting means after fitting supplies an interpretable labeling. Distinct starting means do not guarantee that all fitted components remain distinct. With a common variance and the usual fixed small relative to distinct observations the model avoids the individual-component variance-collapse pathology of unrestricted Gaussian mixtures, but degenerate data or too many components still require attention.
The statistician first models the age- and gender-dependent mean by linear regression. Its fitted value is . At fixed gender, a year of age is associated with a decrease of BMI units; the group coded gender has a fitted mean units above the group coded at the same age. This question does not state which gender receives which code. The intercept refers to age zero in the reference group, outside a typical adult-patient range, so its substantive interpretation is limited. Both slope Student t-tests have small p-values, and the overall F-test supports an age/gender-dependent mean. The coefficient of determination is only , leaving considerable individual variation to investigate.
The residual diagnostics check whether mean adjustment is plausible and whether a single normal error law is adequate. The residual-versus-fitted plot has no strong curved mean trend, although its smooth curve is not perfectly flat. The scale-location plot suggests some decline in spread as fitted BMI increases, so common residual variance is a working approximation rather than an established fact. The normal Q-Q plot has systematic departures, notably shorter tails than the normal reference, suggesting the residual law is not exactly normal. The regression leverage plot does not display an obviously extreme leverage point, but the marked observations and Cook's distance should still be checked for influence. Independence between patients cannot be diagnosed from these four plots alone.
Next the statistician extracts the regression residuals and fits an intercept-only normal model as a baseline residual distribution. With an intercept in the original ordinary least squares model, the residuals sum to zero by the normal equation; the near-zero fitted residual mean and its are therefore automatic, not new evidence of good fit. The standard error in this refit uses degrees of freedom, whereas the original uses after estimating three regression coefficients. The reported log-likelihood for the single-normal residual model is .
The histogram suggests a shape worth exploring beyond one normal density, with broad shoulders and some asymmetry, without showing unambiguous separated clusters. The statistician fits two- and three-component finite Gaussian mixtures with a common variance by the expectation-maximization algorithm, starting from separated means and positive weights. Successive likelihoods increase and the displayed final iterations stabilize, as expected from EM likelihood monotonicity. This is evidence of numerical convergence from those starts, not proof of a global maximum.
The two-component fit assigns weights , residual means and common variance . The three-component fit assigns weights , residual means and common variance . The code correctly passes the square roots of those variances to
dnorm and overlays the resulting weighted densities. Both curves broadly follow the histogram and are very similar. In particular, a two-component mixture need not have two distinct visible modes.These fits explore residual mixture clustering: groups differ in BMI relative to the same age/gender-adjusted mean, rather than simply in raw BMI. Membership can be summarized by the fitted mixture responsibilities, retaining uncertainty instead of asserting a certain label for each patient. The common-variance assumption is economical, but should be checked; the common regression slope assumption and constancy of mixture proportions across age and gender are also substantive. A residual mixture alone does not establish genuine biological subpopulations; a flexible continuous distribution, missing predictors, nonlinear mean effects or heteroscedasticity could explain a similar marginal shape.
There is also a two-stage residual mixture fitting limitation. Fitted regression residuals are not exactly independent identically distributed observations: even under a normal regression, their covariance matrix is , and estimating the mean adjustment introduces uncertainty. With subjects and three initial regression coefficients, treating them as independent is a useful approximation for exploring shape, but ordinary mixture likelihoods omit the first-stage uncertainty. A joint mixture regression with shared slopes, or a suitable bootstrap of the whole analysis, would give a firmer basis for parameter uncertainty and cluster selection. Examine component membership versus age/gender and repeat the algorithm from multiple starts before making substantive clustering claims.
The three-component finite Gaussian mixture with a common variance has the largest displayed maximized log-likelihood, but it also has more fitted parameters. For components there are free weights, means and one common variance, making residual-distribution parameters. Applying the nominal Akaike information criterion and Bayesian information criterion to the three residual fits givesBoth criteria favour the two-component model. It improves the likelihood substantially over one normal component; the third component gains only in log-likelihood at a cost of two further parameters. The Akaike information criterion difference between two and three components is small, about , so that criterion alone gives only a slight preference; the Bayesian information criterion preference for two is stronger. Their nearly indistinguishable overlaid curves also support retaining the simpler two-component fit.
These calculations are model-selection aids, not an exact test. Ordinary chi-squared calibration of a likelihood-ratio test for the number of mixture components is invalid in general: under a smaller-component null, some weights lie on the boundary and extra-component parameters are unidentified. This is nonregular mixture model selection. A parametric bootstrap or predictive cross-validation is preferable for a formal comparison, and should account for the mean-adjustment stage. The table uses the provided residual likelihoods and their nominal parameter counts; first-stage regression uncertainty and possible local EM maxima remain qualifications. On the supplied evidence, select two components without claiming that two distinct biological groups have been established.
Articles by others on the same topic
There are currently no matching articles.