Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 a Solution Created 2026-10-03 Updated 2026-10-06
At a target , approximate the unknown mean locally by and solve the local linear regression weighted least-squares problemwhere is a nonnegative regression kernel and a smoothing bandwidth. A common factor in all weights makes no difference. Estimate by the fitted intercept , and optionally by . Repeat at each target rather than fitting a single global straight line.
For an explicit expression put and . The normal equations yieldThe denominator is positive provided there are at least two distinct positively weighted predictor values. The weights sum to one and their first centered moment is zero, so local linear regression exactly reproduces affine functions and reduces boundary bias compared with a local constant fit. Smaller uses more local detail and increases variance; larger smooths more and can increase bias. Choose the bandwidth by cross-validation or another justified criterion. A nearest-neighbour span provides an adaptive bandwidth, as in the local fits used later.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 c Solution Created 2026-10-03 Updated 2026-10-06
Both fits use the Gaussian generalized additive modelwhere is experience and education. Centering makes the additive decomposition identifiable. No experience-education interaction term is included.
The first fit uses local linear regression for each term.
span=1 includes all observations in each nearest-neighbour neighbourhood, but distance weights and a moving target still vary: it does not make the fit exactly a global straight line. The second uses cubic smoothing splines, with separate tuning values spar=0.8 and spar=1.2. Increasing spar increases the curvature penalty and smooths more; spar is a monotone tuning parametrization, not the numerical in part (b). The relation depends on the predictor scaling and design.The plotted experience effect is increasing in both fits. The local fit gives a broadly smooth, nearly linear rise; the spline fit follows more of the steep low-experience rise and subsequent flattening, especially near the lower boundary. Education has a weaker increasing effect. The local estimate shows some curvature, while the spline fit at
spar=1.2 is closer to a straight line. The clearest difference is the amount of smoothing and the experience curvature, rather than opposite effect directions. These are centered partial-effect plots with partial residuals, not scatterplots of wage against each predictor alone.A smaller local span would allow the first fit to capture more experience curvature. A moderately smaller education
spar could reveal nonlinear detail suppressed in the second fit; its small but significant nonparametric component in part (e) supports checking that possibility. Do not tune solely to follow every point, especially at sparse boundaries. Compare candidate spans or separate spline penalties using cross-validation or generalized cross-validation; increase smoothing if flexible fits introduce unsupported wiggles. See the documented monotone spar–penalty convention at stat.ethz.ch/R-manual/R-patched/library/stats/html/smooth.spline.html.