Solution
ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2014/iii/paper-33/1/solution
Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 1 Solution by
Codex 0 Created 2026-10-03 Updated 2026-10-06
For the normal linear model, the log-likelihood isThe full column rank of the design matrix makes invertible. Differentiating in gives the normal equations . With the minimized residual sum of squares substituted, differentiation in givesThese are the maximum-likelihood estimators almost surely; almost surely when and . The linear transformation of the multivariate normal distribution yieldsDefine the hat matrix . The fitted values are , the regression residuals are , and . The orthogonal projection has rank and annihilates . By Cochran's theorem,Consequently , with bias . The unbiased estimator isIts square root, the reported residual standard error, estimates ; the assertion of unbiasedness applies to the variance, not generally to its square root.
For the paper-strength analysis, let be the measured percentage and . Both normal linear models use independent errors of common variance . Their mean functions are for
lm1, and for lm2, with separately fitted coefficients. The reported residual standard errors are and , respectively.For
lm1, the estimated conditional expectation at a new percentage isThe original data mean must be used for centering the new percentage; its numerical value is not supplied in the excerpt. Because , the two columns of this design matrix are orthogonal, and . Hence the coefficient standard errors in the output giveThis estimates uncertainty in the mean. Predicting an individual future batch would additionally require the new-error variance, estimated by .To compare the nested normal linear models, test against within the quadratic model. The Student t-test statistic isIts two-sided p-value is . Equivalently is a partial F-test statistic with null distribution . Reject the linear restriction and prefer
lm2 at the 5% level. Its residual spread is much smaller; the increase in the coefficient of determination from to supports the same conclusion, although the test is the relevant complexity-adjusted comparison.For regression diagnostics, inspect regression residuals against fitted means and hardwood percentage for omitted curvature, a scale-location plot for nonconstant variance, and a quantile-quantile plot against the normal distribution for departures from the error assumption. Check unusual observations using regression leverage and Cook's distance, and examine residuals in collection or batch order if dependence is plausible. Independence and a common variance require substantive justification as well as these plots; a small p-value for a polynomial term does not itself check the error model.
New to topics? Read the docs here!