Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 1 Solution Created 2026-10-03 Updated 2026-10-06
For the normal linear model, the log-likelihood isThe full column rank of the design matrix makes invertible. Differentiating in gives the normal equations . With the minimized residual sum of squares substituted, differentiation in givesThese are the maximum-likelihood estimators almost surely; almost surely when and . The linear transformation of the multivariate normal distribution yieldsDefine the hat matrix . The fitted values are , the regression residuals are , and . The orthogonal projection has rank and annihilates . By Cochran's theorem,Consequently , with bias . The unbiased estimator isIts square root, the reported residual standard error, estimates ; the assertion of unbiasedness applies to the variance, not generally to its square root.
For the paper-strength analysis, let be the measured percentage and . Both normal linear models use independent errors of common variance . Their mean functions are for
lm1, and for lm2, with separately fitted coefficients. The reported residual standard errors are and , respectively.For
lm1, the estimated conditional expectation at a new percentage isThe original data mean must be used for centering the new percentage; its numerical value is not supplied in the excerpt. Because , the two columns of this design matrix are orthogonal, and . Hence the coefficient standard errors in the output giveThis estimates uncertainty in the mean. Predicting an individual future batch would additionally require the new-error variance, estimated by .To compare the nested normal linear models, test against within the quadratic model. The Student t-test statistic isIts two-sided p-value is . Equivalently is a partial F-test statistic with null distribution . Reject the linear restriction and prefer
lm2 at the 5% level. Its residual spread is much smaller; the increase in the coefficient of determination from to supports the same conclusion, although the test is the relevant complexity-adjusted comparison.For regression diagnostics, inspect regression residuals against fitted means and hardwood percentage for omitted curvature, a scale-location plot for nonconstant variance, and a quantile-quantile plot against the normal distribution for departures from the error assumption. Check unusual observations using regression leverage and Cook's distance, and examine residuals in collection or batch order if dependence is plausible. Independence and a common variance require substantive justification as well as these plots; a small p-value for a polynomial term does not itself check the error model.
Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 2 Solution Created 2026-10-03 Updated 2026-10-06
Let be speed and the tool type. Each fit is a normal linear model with independent errors. The three mean specifications areThe constraints are treatment coding with type 1 as reference. The parameter counts are , respectively. Model 2 gives parallel lines; model 3 allows interaction terms and distinct slopes.
In the analysis of variance interaction row, adding three independent slope differences costs three degrees of freedom. Its extra sum of squares is the reduction in residual sum of squares, . Its mean square is , and its F-test statistic isThus the four missing entries are , , , and , to the precision allowed by the rounded output. This tests against the alternative that at least one slope difference is nonzero. Under and the normal linear model assumptions, the statistic has distribution . Its p-value gives no reason to reject at 5%.
The simpler common-line model is inadequate compared with the parallel-line model. Testing against at least one nonzero type effect, while retaining speed, gives the partial F-testIts p-value is far below , so reject the common-line restriction. This reduction differs from the sequential type sum of squares in the displayed table, because that table adds type before speed. Recommend model 2: different intercepts and a common decreasing slope.
The selected coefficient estimates giveThe intercept estimates the expected lifetime of type 1 at speed zero; if zero speed is outside the data range, it is only an extrapolated intercept. Its standard error is , and its Student t-test ratio is , with two-sided p-value . At any fixed speed, type 2 differs from type 1 by hours, with standard error , and p-value , giving little evidence of a difference. Types 3 and 4 exceed type 1 by and hours, with standard errors and ; their test ratios and and p-values and support positive differences. Each of these individual coefficient tests has null value zero, alternative nonzero, and null distribution .
The speed coefficient means a reduction of hours per extra revolution per minute, or hours per additional 100 rpm, for every type. Its standard error is ; the ratio has null distribution when the common slope is zero, and two-sided p-value .
The residual standard error estimates the common noise standard deviation in hours using 15 residual degrees of freedom, since . The coefficient of determination means that about of the corrected lifetime variation is explained by this fit. The overall F-test statistic tests the simultaneous restriction against at least one nonzero coefficient; its null distribution is and its p-value rejects an intercept-only mean. These conclusions remain conditional on suitable regression diagnostics for independent, homoscedastic, approximately normal errors.
The required sketch plots the four fitted lines. Its speed interval is illustrative because the observed speeds are not supplied; the ordering and vertical gaps are determined by the fitted coefficients.
Parallel fitted tool-lifetime lines for the four tool types
. Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 1 Solution Created 2026-10-03 Updated 2026-10-06
The ordinary least squares objective has gradient . The normal equation therefore givesThe Gram matrix is a positive-definite matrix because has full column rank. More explicitly, for any ,since . This proves that the displayed solution is the unique global minimum.
The fitted values are , where the hat matrix isThe inverse of the symmetric matrix is symmetric, so . Multiplication gives . Thus is the orthogonal projection onto the column space of the design matrix.
The regression residual vector is with . Consequently , , and the residual sum of squares isBy fitted-residual orthogonality, . Using , the cross-covariance matrix isBoth vectors are jointly normal, so they are also independent by independence of uncorrelated jointly normal variables. The zero covariance calculation itself only needs the common error variance and lack of error correlations.
For the first wind fit, let be wind velocity and electrical output. The simple linear regression is , with independent . It has two fitted mean parameters and residual degrees of freedom. The missing analysis of variance entries are therefore velocity degrees of freedom , velocity mean square , and residual degrees of freedom . The residual mean square is , and , consistent with the printed value after rounding.
Because the model includes a regression intercept, the coefficient of determination is the explained sum of squares divided by the total centered sum of squares:A large coefficient of determination does not rule out a wrong mean function. In the PDF's first residual-versus-fitted plot, residuals are negative at both ends and positive in the middle. This curved pattern agrees with the visibly flattening output-versus-velocity relationship and motivates a polynomial regression with a quadratic term.
The second wind model is , again with independent common-variance normal errors. Its fitted mean is . The line marked (A) is a two-sided Student t-test of against , conditional on retaining the intercept and linear term. Under the null hypothesis,Its two-sided p-value is . This is far below , so reject a purely linear mean in favour of the quadratic fit. Equivalently, the extra-term nested-model F-test has and null law .
The third fit is a reciprocal-predictor regression, , with estimated mean . It has two mean parameters rather than the quadratic model's three. Its residual standard error is smaller, versus , and its coefficient of determination is larger, versus . The corresponding residual sums of squares are approximately and . The reciprocal fit is preferable on both these fit measures and parsimony. The two models are not nested, so an ordinary extra-term F-test between them is inappropriate. On the common normal-error likelihood, the difference also favours the third model.
Neither lower residual sum of squares nor a higher coefficient of determination establishes adequate assumptions. The quadratic residual-versus-fitted plot removes the original pronounced curvature; the reciprocal plot also has no comparably obvious mean trend. The low-output residuals appear somewhat more spread out, so check scale-location plots and residuals against velocity for heteroscedasticity. Q-Q plots assess normality; residuals against observation order assess serial dependence; regression leverage and Cook's distance identify influential observations. Further cross-validation, replicate observations at comparable velocities, and prediction errors would help choose between their extrapolation behaviours. Both fitted shapes should be judged principally over the observed positive-velocity range.
Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 2 Solution Created 2026-10-03 Updated 2026-10-06
Let be recall for age level , processing level , and replicate . Write for Younger and for Older. With Older and A as the reference levels in a regression factor, the additive two-factor normal linear model under treatment coding iswhere the errors are independent , with common positive variance. The design matrix is fixed; subjects supply independent observations; the specified conditional mean is correct; and the absence of an interaction term means the age difference is constant across processing levels. These are model assumptions, not consequences of random assignment. The treatment allocation supports independence between processing assignment and background characteristics, but age itself was not randomized.
An older subject in A has and , so the estimated mean is words. Its standard error is the intercept's . There are residual degrees of freedom, giving the 95% confidence interval for this mean:This is a mean-response confidence interval, not a prediction interval for a new individual's recall; the latter also includes the new observation's error variance.
The second two-factor normal linear model adds age-by-processing interaction terms:Its ten free mean parameters are equivalent to one mean for each of the ten cells. To compare the fits, the null hypothesis is . The full residual sum of squares is on degrees of freedom. Adding interaction reduces the additive model's residual sum of squares by , so the reduced value is . The nested-model F-test statistic isIts p-value is . Reject additivity and retain the interaction model. The significant interaction means that a single age effect averaged over all processing methods does not adequately summarize the data.
The regression intercept is the older-A mean. The age coefficient is the younger-minus-older contrast specifically in A. The processing coefficients compare B, C, D and E with A specifically among older subjects. The interaction coefficients are differences between the age contrast in each of those methods and its value in A. Adding the relevant coefficients givesThus A, C and especially D favour younger subjects in their estimated means, whereas B and E have little estimated age difference. In the younger group, D has the largest estimated recall, followed by C and A; B and E are much lower. Among older subjects C is largest, D and A are intermediate, and B and E are lowest. Comparing these orders is a description of estimates, not a claim that every pairwise difference is significant.
Each coefficient's printed standard error estimates its uncertainty; its statistic divides the estimate by that error, and its two-sided p-value uses under the relevant zero-contrast null hypothesis. For example, the older B-versus-A and E-versus-A contrasts have and . The C-versus-A and D-versus-A older contrasts have and . The B interaction has , while the D interaction is borderline at and the other two individual interaction tests are not significant at . These individual tests do not override the joint interaction F-test, and multiple comparisons require care. A simple younger-versus-older contrast outside A combines coefficients and must use their estimated covariance, rather than adding their marginal errors.
The residual standard error estimates . The coefficient of determination is the fraction of centered sample recall variation explained by the ten-cell model. The overall with null law tests all nine non-intercept coefficients jointly against zero, giving extremely strong evidence that the cell means are not all equal. The sequential analysis of variance separates age, process, and their interaction; the balanced design makes the main-effect sums orthogonal, whereas the coefficient table uses the specified reference-level contrasts.
To assess pooling into three processing types, retain age-by-type interaction because the previous test rejects additivity. The reduced model has six cell means. Its null hypothesis says A and C have equal means within each age, and B and E have equal means within each age: four restrictions. Compare it with the ten-cell model usingIn fact the supplied fitted cell means and equal cell sizes determine the increase without raw observations. Pooling two ten-person cells with means adds to the residual sum of squares. Here the four differences are , soTherefore and . At , the data do not reject the three-type reduction, though the result is borderline and is not proof that the pooled means are identical. This factor-level pooling test keeps the previously supported interaction structure; reducing the categories and removing all interaction simultaneously would test a different set of restrictions.
Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 6 b Solution Created 2026-10-03 Updated 2026-10-06
The statistician first models the age- and gender-dependent mean by linear regression. Its fitted value is . At fixed gender, a year of age is associated with a decrease of BMI units; the group coded gender has a fitted mean units above the group coded at the same age. This question does not state which gender receives which code. The intercept refers to age zero in the reference group, outside a typical adult-patient range, so its substantive interpretation is limited. Both slope Student t-tests have small p-values, and the overall F-test supports an age/gender-dependent mean. The coefficient of determination is only , leaving considerable individual variation to investigate.
The residual diagnostics check whether mean adjustment is plausible and whether a single normal error law is adequate. The residual-versus-fitted plot has no strong curved mean trend, although its smooth curve is not perfectly flat. The scale-location plot suggests some decline in spread as fitted BMI increases, so common residual variance is a working approximation rather than an established fact. The normal Q-Q plot has systematic departures, notably shorter tails than the normal reference, suggesting the residual law is not exactly normal. The regression leverage plot does not display an obviously extreme leverage point, but the marked observations and Cook's distance should still be checked for influence. Independence between patients cannot be diagnosed from these four plots alone.
Next the statistician extracts the regression residuals and fits an intercept-only normal model as a baseline residual distribution. With an intercept in the original ordinary least squares model, the residuals sum to zero by the normal equation; the near-zero fitted residual mean and its are therefore automatic, not new evidence of good fit. The standard error in this refit uses degrees of freedom, whereas the original uses after estimating three regression coefficients. The reported log-likelihood for the single-normal residual model is .
The histogram suggests a shape worth exploring beyond one normal density, with broad shoulders and some asymmetry, without showing unambiguous separated clusters. The statistician fits two- and three-component finite Gaussian mixtures with a common variance by the expectation-maximization algorithm, starting from separated means and positive weights. Successive likelihoods increase and the displayed final iterations stabilize, as expected from EM likelihood monotonicity. This is evidence of numerical convergence from those starts, not proof of a global maximum.
The two-component fit assigns weights , residual means and common variance . The three-component fit assigns weights , residual means and common variance . The code correctly passes the square roots of those variances to
dnorm and overlays the resulting weighted densities. Both curves broadly follow the histogram and are very similar. In particular, a two-component mixture need not have two distinct visible modes.These fits explore residual mixture clustering: groups differ in BMI relative to the same age/gender-adjusted mean, rather than simply in raw BMI. Membership can be summarized by the fitted mixture responsibilities, retaining uncertainty instead of asserting a certain label for each patient. The common-variance assumption is economical, but should be checked; the common regression slope assumption and constancy of mixture proportions across age and gender are also substantive. A residual mixture alone does not establish genuine biological subpopulations; a flexible continuous distribution, missing predictors, nonlinear mean effects or heteroscedasticity could explain a similar marginal shape.
There is also a two-stage residual mixture fitting limitation. Fitted regression residuals are not exactly independent identically distributed observations: even under a normal regression, their covariance matrix is , and estimating the mean adjustment introduces uncertainty. With subjects and three initial regression coefficients, treating them as independent is a useful approximation for exploring shape, but ordinary mixture likelihoods omit the first-stage uncertainty. A joint mixture regression with shared slopes, or a suitable bootstrap of the whole analysis, would give a firmer basis for parameter uncertainty and cluster selection. Examine component membership versus age/gender and repeat the algorithm from multiple starts before making substantive clustering claims.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 1 b Solution Created 2026-10-03 Updated 2026-10-06
Let , , , and . The horizontal axis of the residual-versus-fitted plot is , the fitted values, and the vertical axis is , the regression residual, in millimetres. Under the normal linear model, , where is the regression leverage.
In the quantile-quantile plot, the vertical values are the ordered standardized regression residualsand the horizontal values are corresponding theoretical quantiles of the standard normal distribution, with plotting positions such as . The exact plotting-position convention has little practical effect here. Under Gaussian errors these points should approximately follow a straight line; they are not independent because fitting induces residual correlations.
The main visible concern is curvature in the conditional mean. The red smooth in the residual-versus-fitted plot descends from positive residuals at low fitted values, becomes negative in the middle, and rises again at high fitted values. This suggests the strictly linear time effects may be inadequate. There is no clear monotone widening of the residual scatter, so strong heteroscedasticity is not apparent. The quantile-quantile plot has modest tail deviations and a few labelled observations, but does not show a dramatic departure from normality. These plots cannot establish independence across days or laboratories; check residuals against day, laboratory, and sampling order as well. A large coefficient of determination does not remove the visible mean-model concern.
