The ridge regression estimator minimizes
Its gradient vanishes exactly when . Since makes this matrix positive definite,
Use a singular value decomposition . Then
Each nonzero singular value contributes , while each zero one contributes zero. Thus
using the Moore-Penrose inverse.
Put . Its ridgeless minimum-norm fit is , so its first coordinates are
The strong law of large numbers applied entrywise gives almost surely as . Continuity of inversion then yields
almost surely, where the middle equality is the push-through identity.
Because is independent, conditional prediction risk equals . With and ,
The noise term has conditional mean zero and covariance . Taking squared norms proves
Apply the stated deterministic equivalent for the squared resolvent with . The bias term tends to . For the variance use
Normalized traces and therefore give
Thus the conclusion printed in the paper also has both signs involving reversed. The hypotheses as printed imply the formula above; they cannot imply the requested one because is positive definite while for a resolvent limit.

Articles by others on the same topic (0)

There are currently no matching articles.