Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 206 3 b Solution Created 2026-10-03 Updated 2026-10-05
Choose knots independently of the eventual validation responses, and let be a natural cubic spline basis as in part (a). A regression spline writes . Define the design matrix . Under the stated normal linear model, maximum likelihood estimation for is the same as ordinary least squares:Differentiation gives the normal equations . If the design has full column rank,If it does not, a Moore-Penrose inverse gives the fitted values, but not every coefficient or off-sample prediction is identifiable. Merely choosing distinct knots does not guarantee full rank for an arbitrary set of observed predictor values.
More knots allow greater curvature and reduce approximation bias at the cost of larger estimation variance. Too many or poorly located knots can cause unstable fits. Equally spaced knots are simple; empirical predictor quantiles put more flexibility where observations are dense; subject knowledge can concentrate knots near likely changes. Both the number and positions can be compared using cross-validation or an appropriate penalized model-selection criterion. Adaptive knot selection must be included within validation, since treating a data-selected basis as prespecified understates its flexibility. Boundary knots should cover the region where inference is intended, and natural linear tails do not by themselves justify far-out extrapolation.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 206 3 d Solution Created 2026-10-03 Updated 2026-10-05
Because no family is specified, the command fits a Gaussian identity-link generalized additive model:The are centered penalized regression splines with cubic regression bases; centering separates them from the intercept. The age curve in the original figure falls to a minimum near age 25, then rises markedly, particularly from about 40 onwards. Its effective degrees of freedom are , reflecting curvature. The white-cell curve has effective degrees of freedom one and is nearly flat with a slight positive slope; the pointwise uncertainty bands are consistent with no substantial effect. These plots display centered partial mean contributions, not numbers of prescriptions by themselves.
A simpler mean model keeps the age regression spline and replaces the white-cell smooth by a linear term. A separate issue is that the response is a nonnegative count: a Poisson regression with logarithmic link function respects that support and positivity of the mean. A reasonable candidate isIf diagnostics show overdispersion, one can fitusing a negative binomial regression. Alternatively, if a Gaussian approximation is adequate, the minimal simplification is
model1p <- gam(npres ~ s(age, bs="cr") + wbc, family=poisson(link="log"))
gam.check(model1p)model1nb <- gam(npres ~ s(age, bs="cr") + wbc, family=nb(link="log"))gam(npres ~ s(age, bs="cr") + wbc). These are candidates to compare by diagnostics and validation, rather than guaranteed improvements from the partial plots alone. Retain the curved age effect; consider a linear or removable white-cell effect, and assess an appropriate count family. Refit before interpreting the old curves on a new link scale. Penalized regression spline 2026-10-05
A penalized regression spline estimates basis coefficients by minimizing a residual or likelihood criterion plus a curvature penalty. Basis size limits the available complexity; the penalty controls how much of that complexity is used.
Regression spline 2026-10-05
A regression spline fits a regression function in a finite-dimensional spline approximation space with selected knots. Evaluating its basis at predictor values gives a design matrix, reducing fitting to regression on basis coefficients.