The trace of the smoother's influence matrix is its total effective degrees of freedom. Include the unpenalized intercept as well as both centered smooth terms:
The Ref.df entries are for approximate significance calibration and must not be summed to obtain this trace. A consistency check using the generalized cross-validation form and scale estimate gives ; the small difference is explained by rounding in the displayed quantities. The required trace is about .
The generalized cross-validation curve reaches its minimum at among the supplied grid values. This balances improved conditioning against shrinkage bias using an estimate of prediction error, rather than choosing the penalty with the smallest training residual sum of squares. Nearby values have similar errors, so the plot does not establish a highly precise optimal penalty.
Reading the PDF table at gives the fitted regression intercept and slopes:
These table entries are necessary because the TeX stores the table only inside a figure. The ridge regression slopes are shrunk relative to the zero-penalty fit; their small nonzero values do not represent variable exclusion.
There is a minor source inconsistency: the prose says the predictors are centered, but the table's intercept varies with . With exactly centered predictor columns and an unpenalized intercept it would remain . The numbers above faithfully report the printed table, while part (a) gives the centered formula and the general uncentered conversion. The table therefore reflects a different or incompletely described preprocessing convention.
Assume initially that the design points are distinct and ordered on . The cubic smoothing spline minimizes
over functions with an absolutely continuous first derivative and square-integrable second derivative. Dividing the residual term by simply changes the scaling assigned to . The minimum roughness property of the natural cubic spline interpolant reduces the problem to natural cubic splines with knots at the design points: among functions with any prescribed fitted values, that interpolant has the smallest integrated squared curvature. The minimizer is therefore a natural cubic smoothing spline, with natural second-derivative boundary conditions and linear tails when extended.
Choose a full basis for this natural-spline space. Put and . Writing changes the objective to
Differentiation gives . Although the curvature penalty has an affine nullspace, the full basis at distinct design points makes the combined matrix positive definite. Hence
Equivalently, use the cardinal basis that interpolates the observations, making and . These are smoothing matrices, symmetric but generally not idempotent. Their effective degrees of freedom are . As the fit interpolates; as it approaches ordinary least squares on the affine functions. With replicated predictors, use a design matrix for distinct knots and observation weights; the basis form remains valid when identifiable.
Useful choices include Leave-one-out cross-validation, generalized cross-validation, and estimation of smoothness by restricted maximum likelihood through a mixed-model representation. For fixed , the leave-one-out residual identity for a linear smoother gives
Minimize these over a suitable range. Choose smoothing by predictive performance or a variance-component likelihood, not by arbitrarily forcing a visually appealing curve.
Both fits use the Gaussian generalized additive model
where is experience and education. Centering makes the additive decomposition identifiable. No experience-education interaction term is included.
The first fit uses local linear regression for each term. span=1 includes all observations in each nearest-neighbour neighbourhood, but distance weights and a moving target still vary: it does not make the fit exactly a global straight line. The second uses cubic smoothing splines, with separate tuning values spar=0.8 and spar=1.2. Increasing spar increases the curvature penalty and smooths more; spar is a monotone tuning parametrization, not the numerical in part (b). The relation depends on the predictor scaling and design.
The plotted experience effect is increasing in both fits. The local fit gives a broadly smooth, nearly linear rise; the spline fit follows more of the steep low-experience rise and subsequent flattening, especially near the lower boundary. Education has a weaker increasing effect. The local estimate shows some curvature, while the spline fit at spar=1.2 is closer to a straight line. The clearest difference is the amount of smoothing and the experience curvature, rather than opposite effect directions. These are centered partial-effect plots with partial residuals, not scatterplots of wage against each predictor alone.
A smaller local span would allow the first fit to capture more experience curvature. A moderately smaller education spar could reveal nonlinear detail suppressed in the second fit; its small but significant nonparametric component in part (e) supports checking that possibility. Do not tune solely to follow every point, especially at sparse boundaries. Compare candidate spans or separate spline penalties using cross-validation or generalized cross-validation; increase smoothing if flexible fits introduce unsupported wiggles. See the documented monotone spar–penalty convention at stat.ethz.ch/R-manual/R-patched/library/stats/html/smooth.spline.html.