Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 33 5 b iii Solution Created 2026-10-03 Updated 2026-10-06
The trace of the smoother's influence matrix is its total effective degrees of freedom. Include the unpenalized intercept as well as both centered smooth terms:The
Ref.df entries are for approximate significance calibration and must not be summed to obtain this trace. A consistency check using the generalized cross-validation form and scale estimate gives ; the small difference is explained by rounding in the displayed quantities. The required trace is about . Past exam of the mathematics course of the University of Cambridge 2015 iii Paper 33 4 b Solution Created 2026-10-03 Updated 2026-10-06
The generalized cross-validation curve reaches its minimum at among the supplied grid values. This balances improved conditioning against shrinkage bias using an estimate of prediction error, rather than choosing the penalty with the smallest training residual sum of squares. Nearby values have similar errors, so the plot does not establish a highly precise optimal penalty.
Reading the PDF table at gives the fitted regression intercept and slopes:These table entries are necessary because the TeX stores the table only inside a figure. The ridge regression slopes are shrunk relative to the zero-penalty fit; their small nonzero values do not represent variable exclusion.
There is a minor source inconsistency: the prose says the predictors are centered, but the table's intercept varies with . With exactly centered predictor columns and an unpenalized intercept it would remain . The numbers above faithfully report the printed table, while part (a) gives the centered formula and the general uncentered conversion. The table therefore reflects a different or incompletely described preprocessing convention.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 b Solution Created 2026-10-03 Updated 2026-10-06
Assume initially that the design points are distinct and ordered on . The cubic smoothing spline minimizesover functions with an absolutely continuous first derivative and square-integrable second derivative. Dividing the residual term by simply changes the scaling assigned to . The minimum roughness property of the natural cubic spline interpolant reduces the problem to natural cubic splines with knots at the design points: among functions with any prescribed fitted values, that interpolant has the smallest integrated squared curvature. The minimizer is therefore a natural cubic smoothing spline, with natural second-derivative boundary conditions and linear tails when extended.
Choose a full basis for this natural-spline space. Put and . Writing changes the objective toDifferentiation gives . Although the curvature penalty has an affine nullspace, the full basis at distinct design points makes the combined matrix positive definite. HenceEquivalently, use the cardinal basis that interpolates the observations, making and . These are smoothing matrices, symmetric but generally not idempotent. Their effective degrees of freedom are . As the fit interpolates; as it approaches ordinary least squares on the affine functions. With replicated predictors, use a design matrix for distinct knots and observation weights; the basis form remains valid when identifiable.
Useful choices include Leave-one-out cross-validation, generalized cross-validation, and estimation of smoothness by restricted maximum likelihood through a mixed-model representation. For fixed , the leave-one-out residual identity for a linear smoother givesMinimize these over a suitable range. Choose smoothing by predictive performance or a variance-component likelihood, not by arbitrarily forcing a visually appealing curve.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 c Solution Created 2026-10-03 Updated 2026-10-06
Both fits use the Gaussian generalized additive modelwhere is experience and education. Centering makes the additive decomposition identifiable. No experience-education interaction term is included.
The first fit uses local linear regression for each term.
span=1 includes all observations in each nearest-neighbour neighbourhood, but distance weights and a moving target still vary: it does not make the fit exactly a global straight line. The second uses cubic smoothing splines, with separate tuning values spar=0.8 and spar=1.2. Increasing spar increases the curvature penalty and smooths more; spar is a monotone tuning parametrization, not the numerical in part (b). The relation depends on the predictor scaling and design.The plotted experience effect is increasing in both fits. The local fit gives a broadly smooth, nearly linear rise; the spline fit follows more of the steep low-experience rise and subsequent flattening, especially near the lower boundary. Education has a weaker increasing effect. The local estimate shows some curvature, while the spline fit at
spar=1.2 is closer to a straight line. The clearest difference is the amount of smoothing and the experience curvature, rather than opposite effect directions. These are centered partial-effect plots with partial residuals, not scatterplots of wage against each predictor alone.A smaller local span would allow the first fit to capture more experience curvature. A moderately smaller education
spar could reveal nonlinear detail suppressed in the second fit; its small but significant nonparametric component in part (e) supports checking that possibility. Do not tune solely to follow every point, especially at sparse boundaries. Compare candidate spans or separate spline penalties using cross-validation or generalized cross-validation; increase smoothing if flexible fits introduce unsupported wiggles. See the documented monotone spar–penalty convention at stat.ethz.ch/R-manual/R-patched/library/stats/html/smooth.spline.html.