= Solution
Use an <unpenalized intercept in ridge regression>. For centered predictor columns, the <ridge regression> optimization is
$$
\min_{\alpha,\beta}\ \|Y-\alpha\mathbf1-X\beta\|_2^2+\lambda\|\beta\|_2^2,\qquad\lambda\geq0.
$$
Differentiating with respect to $\alpha$ gives $\widehat\alpha=\overline Y$. Put $Y_c=Y-\overline Y\mathbf1$. Differentiation with respect to $\beta$ gives the ridge <normal equation>
$$
(X^TX+\lambda I_6)\widehat\beta_\lambda=X^TY_c.
$$
Consequently the <closed-form ridge regression estimator> is
$$
\boxed{\widehat\beta_\lambda=(X^TX+\lambda I_6)^{-1}X^TY_c,\qquad\widehat\alpha=\overline Y.}
$$
Because $X^T\mathbf1=0$, the numerator may also be written $X^TY$. For $\lambda>0$ the inverse exists even with <multicollinearity> or deficient column rank; at $\lambda=0$ a unique <ordinary least squares> coefficient vector requires full rank. The quadratic penalty stabilizes nearly singular directions but generally does not set individual coefficients exactly to zero.
For uncentered original predictors, use $X_c=X-\mathbf1\overline x^T$, apply the displayed estimator to $X_c,Y_c$, and recover $\widehat\alpha=\overline Y-\overline x^T\widehat\beta$. Any internal scaling used by `lm.ridge` must be undone to report coefficients on the original predictor scale. Multiplying the objective by a constant changes the numerical penalty convention unless $\lambda$ is rescaled as well.
Back to article page