Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 5 b ii Solution Created 2026-10-03 Updated 2026-10-06
The first model's residual deviance is on residual degrees of freedom, a ratio of about . This is far above the scale expected from a well-fitting dispersion-one grouped binomial regression. It suggests an inadequate mean function, overdispersion, or both. The linear age and exposure restrictions may miss nonlinear associations; the generalized additive model allows those shapes to be estimated through penalized cubic regression splines rather than assumed.
The very small exposure Wald p-value in the first fit suggests an association, but does not establish that a linear logit effect is adequate. Nor does the nonsignificant linear age term exclude every possible nonlinear age effect. Allowing estimated scale also addresses the excessive residual variation, which can arise from clustered data because several births belong to the same woman. The second model's estimated scale confirms that permitting smooth means has not removed all extra variation. The reason to proceed is poor initial fit and possible nonlinear effects, with dispersion-aware inference, not merely the wish to obtain more significant tests. These observational fits alone do not establish causation by the disaster.
Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 5 b i Solution Created 2026-10-03 Updated 2026-10-06
For woman , write for the number of births, for the number with defects, for baseline age, and for exposure. Let be the common marginal defect probability for her births. The first fit is a grouped grouped-binomial logistic regression:It assumes independent women and, conditionally on their covariates, independent births with the same probability within each woman. The supplied birth totals are the binomial denominators; fitting the proportions with weights is equivalent to this grouped-binomial specification. The fitted linear predictor is .
The second fit is a generalized additive model with the same logit link function and the mean specificationThe functions are penalized cubic regression splines with natural boundary conditions. Centering constraints such as separate them from the intercept; the fitted centered intercept is . Their second derivative roughness penalties control complexity.
The call estimates a common scale rather than fixing it at one. Its working variance specification for the counts isEquivalently the variance of is . This is the working moment interpretation of an overdispersed binomial generalized additive model: the code uses binomial deviance for fitting and estimated scale for inference. With , it is not an exact independent-binomial sampling model. Women are still treated as independent groups, but within-woman dependence or unobserved heterogeneity can motivate the extra dispersion. The quasibinomial regression interpretation states what the scaled analysis assumes without inventing a full probability distribution for its counts.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 3 e Solution Created 2026-10-03 Updated 2026-10-06
The residual smooth rises from negative values, arches above zero, then falls at high fitted values. This suggests a nonlinear altitude effect on the log mean. Since the fitted log mean in the linear model increases with altitude, the horizontal fitted-value axis is also an increasing rescaling of altitude. A generalized additive model with Quasi-Poisson regression is therefore a reasonable candidate:The centering constraint identifies the intercept. Here
bs="cr" uses a penalized cubic regression spline. The plotted count residuals are deviance residuals, with the Q-Q panel using leverage-standardized versions. The quantile-quantile plot also shows tail discrepancies, but normality of count residuals is not a distributional assumption of this model: its main requirements are an adequate mean, variance relation, and independence. Investigate the labelled islands and possible excess variation as well.The estimated smooth in the second plot increases fairly rapidly at low altitude, flattens around altitude 20 to 25, and is nearly level at the upper end. The pointwise bands widen where there are fewer observations. A penalized spline can describe a rise followed by a plateau without forcing a global parabola. A quadratic log mean has slope , so negative curvature eventually forces a decline and its curvature is constant everywhere. A cubic regression spline allows different curvature over different ranges, with a penalty controlling unnecessary wiggles and natural linear tails outside its boundary knots. The displayed smooth supports this greater flexibility, but it is not by itself a formal rejection of every quadratic model. Select complexity and compare predictive performance using cross-validation, then reassess residuals and the Pearson dispersion estimator.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 4 d Solution Created 2026-10-03 Updated 2026-10-06
The proposed generalized additive model retains the binary response and logit link, but replaces the linear age effect by a smooth function:A penalized cubic regression spline represents ; frailty remains a parametric coefficient. The plotted age smooth is centered, so its vertical values represent an age contribution to log odds, not probabilities or the entire fitted predictor.
The plot gives no evidence that the nonlinear model improves the mean structure. The fitted curve is essentially a straight descending line, and the label
s(age,1) reports effective degrees of freedom close to one, consistent with a linear effect. The widening bands near the ends reflect sparse data and increasing uncertainty, not detected curvature. A linear-age logistic regression is therefore adequate on this evidence and simpler to interpret. A formal comparative claim would need a suitable test or predictive/model-selection criterion, but the figure supplies no reason to retain extra nonlinear flexibility. Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 6 e Solution Created 2026-10-03 Updated 2026-10-06
The generalized additive mixed model fitswhere denotes the furnace random intercept and is a centered penalized cubic regression spline. Its wiggly components have a mixed-model representation; there is no explicit furnace-specific random slope in this formula.
The plotted smooth gives no evidence that a nonparametric fixed effect is needed. It is essentially a straight rising line, and
s(stir,1) indicates effective degrees of freedom close to one. The wider bands at extreme stirring rates reflect uncertainty, not detected nonlinearity. A linear fixed stirring effect is adequate on this evidence.If the three supplied Akaike information criterion values are on a common valid likelihood basis, choose
strength_model3, the fixed linear slope plus random-intercept model: its AIC is smallest, versus for the additional random slope and for the smooth. The differences are and , respectively, so especially the random-slope comparison represents modest support rather than decisive evidence. Retaining the population slope is consistent with part (c); this choice removes only its random variation.There is a likelihood-basis caveat in the literal printed commands. The supplied ranking favours the simpler random-intercept model; the displayed original calls alone do not certify that ranking as a consistent ML comparison. This distinction is particularly relevant when the smooth changes the fixed-effect space. The Gaussian
lmer fits by REML by default, whereas the Gaussian gamm call defaults to ML; direct AIC on these objects need not compare the same type of likelihood. Moreover the earlier ML ANOVA prints for strength_model, whereas this table prints , so these numbers should not silently be treated as one identical ML fit. To obtain a defensible common comparison, fit all candidates to the same observations by ML:fit_slope <- update(strength_model, REML = FALSE)
fit_intercept <- update(strength_model3, REML = FALSE)
fit_smooth <- mgcv::gamm(strength ~ s(stir, bs = "cr"),
random = list(furnace = ~ 1), method = "ML")
AIC(fit_slope, fit_intercept, fit_smooth$lme)gamm ML default and mixed-model representation are documented at stat.ethz.ch/R-manual/R-devel/library/mgcv/html/gamm.html.