The E-step computes mixture responsibilities from the old parameters. Put . The M-step updates , and . The variance uses new means and old responsibilities. EM likelihood monotonicity guarantees nondecrease of the observed likelihood, not a global optimum.
Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 6 a Solution Created 2026-10-03 Updated 2026-10-06
For a finite Gaussian mixture with a common variance, introduce independent latent labels with probabilities , , and conditional responses , with common . The observed density and log-likelihood areThe expectation-maximization algorithm replaces the difficult log of sums by an expected complete-data objective. Choose positive initial weights and variance and separated initial means. At iteration , the E-step computes the mixture responsibilitiesThe common normalizing factor cancels. For numerical stability the probabilities can be evaluated by subtracting the largest log weight before exponentiating.
With the old responsibilities held fixed, the expected complete-data log-likelihood, up to irrelevant constants, isLet . A Lagrange multiplier for , weighted least squares for the means, and differentiation in the common variance give the M-step:The variance update uses the new means and old responsibilities, with denominator , not a residual degrees-of-freedom adjustment: this is maximum likelihood estimation. Repeat the E- and M-steps until the observed log-likelihood and parameters stabilize.
By EM likelihood monotonicity, exact updates do not decrease the observed likelihood. They need not reach its global maximum, so use several starting configurations and keep the best converged fit, checking for empty or nearly empty components and vanishing variance. The labels are interchangeable; sorting means after fitting supplies an interpretable labeling. Distinct starting means do not guarantee that all fitted components remain distinct. With a common variance and the usual fixed small relative to distinct observations the model avoids the individual-component variance-collapse pathology of unrestricted Gaussian mixtures, but degenerate data or too many components still require attention.
Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 6 b Solution Created 2026-10-03 Updated 2026-10-06
The statistician first models the age- and gender-dependent mean by linear regression. Its fitted value is . At fixed gender, a year of age is associated with a decrease of BMI units; the group coded gender has a fitted mean units above the group coded at the same age. This question does not state which gender receives which code. The intercept refers to age zero in the reference group, outside a typical adult-patient range, so its substantive interpretation is limited. Both slope Student t-tests have small p-values, and the overall F-test supports an age/gender-dependent mean. The coefficient of determination is only , leaving considerable individual variation to investigate.
The residual diagnostics check whether mean adjustment is plausible and whether a single normal error law is adequate. The residual-versus-fitted plot has no strong curved mean trend, although its smooth curve is not perfectly flat. The scale-location plot suggests some decline in spread as fitted BMI increases, so common residual variance is a working approximation rather than an established fact. The normal Q-Q plot has systematic departures, notably shorter tails than the normal reference, suggesting the residual law is not exactly normal. The regression leverage plot does not display an obviously extreme leverage point, but the marked observations and Cook's distance should still be checked for influence. Independence between patients cannot be diagnosed from these four plots alone.
Next the statistician extracts the regression residuals and fits an intercept-only normal model as a baseline residual distribution. With an intercept in the original ordinary least squares model, the residuals sum to zero by the normal equation; the near-zero fitted residual mean and its are therefore automatic, not new evidence of good fit. The standard error in this refit uses degrees of freedom, whereas the original uses after estimating three regression coefficients. The reported log-likelihood for the single-normal residual model is .
The histogram suggests a shape worth exploring beyond one normal density, with broad shoulders and some asymmetry, without showing unambiguous separated clusters. The statistician fits two- and three-component finite Gaussian mixtures with a common variance by the expectation-maximization algorithm, starting from separated means and positive weights. Successive likelihoods increase and the displayed final iterations stabilize, as expected from EM likelihood monotonicity. This is evidence of numerical convergence from those starts, not proof of a global maximum.
The two-component fit assigns weights , residual means and common variance . The three-component fit assigns weights , residual means and common variance . The code correctly passes the square roots of those variances to
dnorm and overlays the resulting weighted densities. Both curves broadly follow the histogram and are very similar. In particular, a two-component mixture need not have two distinct visible modes.These fits explore residual mixture clustering: groups differ in BMI relative to the same age/gender-adjusted mean, rather than simply in raw BMI. Membership can be summarized by the fitted mixture responsibilities, retaining uncertainty instead of asserting a certain label for each patient. The common-variance assumption is economical, but should be checked; the common regression slope assumption and constancy of mixture proportions across age and gender are also substantive. A residual mixture alone does not establish genuine biological subpopulations; a flexible continuous distribution, missing predictors, nonlinear mean effects or heteroscedasticity could explain a similar marginal shape.
There is also a two-stage residual mixture fitting limitation. Fitted regression residuals are not exactly independent identically distributed observations: even under a normal regression, their covariance matrix is , and estimating the mean adjustment introduces uncertainty. With subjects and three initial regression coefficients, treating them as independent is a useful approximation for exploring shape, but ordinary mixture likelihoods omit the first-stage uncertainty. A joint mixture regression with shared slopes, or a suitable bootstrap of the whole analysis, would give a firmer basis for parameter uncertainty and cluster selection. Examine component membership versus age/gender and repeat the algorithm from multiple starts before making substantive clustering claims.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 216 2 Solution Created 2026-10-03 Updated 2026-10-06
Write for the directed edges from parent to child and for the state at site . The global Markov property for an undirected graph on this tree, together with the specified parent transitions, gives the complete-data likelihood functionOne can obtain the factorization by successively removing terminal subtrees: after conditioning on their parent, each subtree is independent of the rest. The root factor and transition products also show that different sites are independent. Only ancestral states are latent variables.
At the current Markov kernel , define expected transition countsThe E-step of the expectation-maximization algorithm isMaximize separately for each row under and . For , the Lagrange multiplier equation is , with zero-count entries set to zero at a boundary maximum. Thus the EM transition-count update on a tree isIf , that row does not occur in ; any row probability distribution, including its previous value, maximizes the objective.
The expected counts can be computed exactly by the Felsenstein pruning algorithm and an outward belief propagation pass. For one site, let be the conditional probability of observations in the subtree below , given state . For a leaf with observed state , . For an internal vertex,Define as the joint mass of state at and all observations outside its subtree. The root has . For a child of , putThe edge posterior probability required above isMultiply the inside and outside factors to obtain the joint mass of the observations and endpoint states; division by the site likelihood function proves this expression. Sum over edges and sites to form . The two passes take operations on a binary tree. Scaling messages or computing them logarithmically avoids underflow.
A strictly positive initial Markov kernel ensures for every observed leaf pattern, so every E-step is defined. At a later boundary iterate, require positive likelihood for the actual data and restrict the conditional distribution to its support. EM likelihood monotonicity proves that the update cannot decrease observed-data likelihood function. It is a maximum-likelihood iteration, not a guarantee of finding the global maximum-likelihood estimate; different initial kernels may lead to different stationary points.
Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 216 4 c Solution Created 2026-10-03 Updated 2026-10-05
Fix the observed , let , and let . The E-step formsFor finite expected log-likelihoods and , the Jensen inequality givesThe first inequality allows additional mass outside the old complete-data support; it is equality when the relevant supports coincide. Ratios are taken on the support of . If a candidate assigns zero density on positive -mass, its is minus infinity and it cannot be an improving finite M-step.
An M-step of the expectation-maximization algorithm chooses with , soThis EM likelihood monotonicity also holds for a generalized M-step that merely increases , rather than maximizing it exactly. It is a statement about the observed-data likelihood function.