For fixed , letting gives ridge regression, including ordinary least squares when ; letting forces every coefficient to zero. For fixed , letting gives the Lasso. Letting both penalties vanish gives an ordinary least squares solution, unique when has full column rank and otherwise potentially nonunique or path-dependent.
When , the term is strictly convex. Its sum with the convex squared loss and penalty is strictly convex and coercive, so the elastic net solution exists and is unique.
For coordinate , hold all other coefficients fixed and form the partial residual
The one-coordinate coordinate descent problem is
Its exact update is
Cyclically update coordinates and their residuals until the objective or coefficients converge. Convexity makes every limit point a global minimizer, and makes it the unique minimizer.
When , expanding the objective separates it by coordinates:
The soft thresholding solution is therefore
Model m1 is ridge regression, while m2 is the Lasso. Ridge shrinks but normally retains every coefficient; the Lasso's penalty sets many coefficients exactly to zero. Model m3 combines sparsity with the grouping effect of the elastic net: correlated predictors tend to enter together and receive more similar coefficients.
Accordingly, m3 keeps the weak variables age, lcp, and gleason at zero as m2 does, but retains lweight, lbph, svi, and pgg45 as the ridge fit does. The correlation between svi and lcavol explains why m3 keeps both with substantial coefficients, whereas m2 selects lcavol and discards svi. This lies between the dense ridge behavior and the more aggressively sparse Lasso behavior.
Choose on a grid by K-fold cross-validation, comparing the same held-out prediction loss and optionally applying the one-standard-error rule for a simpler model. A genuinely untouched test set can then estimate final prediction error.
Ordinary model-based intervals after selecting nonzero coefficients ignore selection and are generally invalid. Valid approaches include Debiased Lasso or a selective-inference procedure under its assumptions, sample splitting followed by an unpenalized refit and inference on the independent half, or a bootstrap that repeats both tuning and fitting and is interpreted with care near the nonsmooth zero threshold.

Articles by others on the same topic (0)

There are currently no matching articles.