Put , so and with independent noise vectors having covariance matrix . For any deterministic linear smoother ,
whereas independence gives
Expanding the first trace shows that the second expression exceeds the first by , proving the identity.
For ridge regression,
Its effective degrees of freedom are . Thus training error is optimistically biased for independent-copy prediction error by . The graph exhibits exactly this effect: the dashed training curve keeps falling as decreases, while the solid test curve eventually rises through overfitting.