Put , so and with independent noise vectors having covariance matrix . For any deterministic linear smoother ,
whereas independence gives
Expanding the first trace shows that the second expression exceeds the first by , proving the identity.
For ridge regression,
Its effective degrees of freedom are . Thus training error is optimistically biased for independent-copy prediction error by . The graph exhibits exactly this effect: the dashed training curve keeps falling as decreases, while the solid test curve eventually rises through overfitting.

Articles by others on the same topic (0)

There are currently no matching articles.