Solution
ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2013/iii/paper-31/3/solution
Past exam of the mathematics course of the University of Cambridge 2013 iii Paper 31 3 Solution by
Codex 0 Created 2026-10-03 Updated 2026-10-07
With the same squared-error normalization as in Question 2, the ridge estimator solvesThe unpenalized intercept in ridge regression is because the columns of are centered. Differentiating the remaining quadratic objective gives . Since , the left matrix is positive definite even when is rank-deficient. Thus the closed-form ridge regression estimator isThe second equality again uses centering. If the penalty convention is with an unnormalized squared-error term , replace by throughout; the limiting assertion is unchanged.
For the full-column-rank case, take a singular value decomposition with positive singular values . The ordinary least squares estimator, which is the slope maximum-likelihood estimator in this normal linear model, and the ridge regression estimator are respectivelyConsequently principal-component shrinkage by ridge regression multiplies the least-squares coordinate in direction by . Ridge shrinks most strongly in directions with the smallest singular values. These are combinations of predictors that the data distinguish least well.
Near multicollinearity creates small singular values, and the factor in ordinary least squares greatly amplifies noise in those directions. Its coefficient variance there is , while the ridge variance iswhich is smaller. The price is bias: the expectation of the ridge coordinate is times the corresponding true coefficient. This bias-variance tradeoff can substantially reduce estimation and prediction error when the poorly identified directions do not contain an excessively large signal. It is not a guarantee of improvement for every coefficient vector. Adding to every eigenvalue also improves the condition number of the normal matrix and makes numerical inversion more stable.
To handle every rank and every relation between and , put and use its spectral theorem for real symmetric matrices. Choose an orthonormal eigenbasis with eigenvalues , and put . ThenIf , then , so and . Thus these terms are exactly zero for every positive ; there is no divergent component in the null space. Each remaining term converges to . By the definition of the Moore-Penrose pseudoinverse,and thereforeThis proves the vanishing-penalty ridge limit without an invertibility assumption. The limit is the minimum-norm least-squares solution: every other least-squares solution differs by a vector in , orthogonal to the displayed solution, and therefore has at least as large a Euclidean norm.
New to topics? Read the docs here!