Backfitting algorithm 2026-10-06
The backfitting algorithm estimates each additive mean function by smoothing its partial residual, obtained after subtracting the current fits of all other terms, and cycling through terms until convergence. Centering each smooth separates its constant part from the intercept. For fixed quadratic smoothness penalties it is block coordinate minimization of a penalized least-squares objective; identifiability and positive definiteness of the constrained problem ensure a unique converged fit. Non-Gaussian generalized additive models can use weighted backfitting inside iteratively reweighted least squares.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 c Solution Created 2026-10-03 Updated 2026-10-06
Both fits use the Gaussian generalized additive modelwhere is experience and education. Centering makes the additive decomposition identifiable. No experience-education interaction term is included.
The first fit uses local linear regression for each term.
span=1 includes all observations in each nearest-neighbour neighbourhood, but distance weights and a moving target still vary: it does not make the fit exactly a global straight line. The second uses cubic smoothing splines, with separate tuning values spar=0.8 and spar=1.2. Increasing spar increases the curvature penalty and smooths more; spar is a monotone tuning parametrization, not the numerical in part (b). The relation depends on the predictor scaling and design.The plotted experience effect is increasing in both fits. The local fit gives a broadly smooth, nearly linear rise; the spline fit follows more of the steep low-experience rise and subsequent flattening, especially near the lower boundary. Education has a weaker increasing effect. The local estimate shows some curvature, while the spline fit at
spar=1.2 is closer to a straight line. The clearest difference is the amount of smoothing and the experience curvature, rather than opposite effect directions. These are centered partial-effect plots with partial residuals, not scatterplots of wage against each predictor alone.A smaller local span would allow the first fit to capture more experience curvature. A moderately smaller education
spar could reveal nonlinear detail suppressed in the second fit; its small but significant nonparametric component in part (e) supports checking that possibility. Do not tune solely to follow every point, especially at sparse boundaries. Compare candidate spans or separate spline penalties using cross-validation or generalized cross-validation; increase smoothing if flexible fits introduce unsupported wiggles. See the documented monotone spar–penalty convention at stat.ethz.ch/R-manual/R-patched/library/stats/html/smooth.spline.html. Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 206 5 d Solution Created 2026-10-03 Updated 2026-10-06
Use the backfitting algorithm to alternate conditional updates of the two additive functions. Initialize and both centered function vectors at zero. With the smoothing matrices for the chosen experience and education spline penalties, one cycle isfollowed byThe centering transfers any constant component to the intercept; update to the mean of if necessary, which is under the stated constraints. Repeat until changes in the functions or objective are negligible. Each function is estimated by smoothing its partial residual, namely the response after subtracting all other currently fitted terms.
For fixed penalties these updates are block minimizations ofsubject to the centering constraints. Each step cannot increase the objective; with identifiable additive terms and a positive-definite constrained quadratic, backfitting converges to its unique minimizer. Highly dependent predictors can make the decomposition poorly identified or slow convergence. The Gaussian identity-link model needs these least-squares updates directly; non-Gaussian generalized additive models use weighted backfitting inside iteratively reweighted least squares. The package implementation also separates each spline's unpenalized linear component from its nonlinear component, giving the two ANOVA tables displayed in part (e).