Use the following normalization for the Lasso estimator, leaving the intercept unpenalized:
A different convention for the factor in front of squared error rescales the regularization parameter. With this convention, centering of the columns gives . Put and . Then and . The Lasso minimizes . A minimum exists since makes this continuous objective coercive; the bound below holds for every minimizer.
We first prove the needed sharp two-sided Gaussian tail bound. If is standard normal and , substituting in its density integral yields
The last integral without is one after multiplication by . In particular this proves the two-sided bound with no extra factor of two.
Since , each has a standard normal distribution. The union bound, which does not require independent columns or independent scores, gives the Gaussian score event for the Lasso
Here the positive tuning choice presupposes and . When the displayed lower bound is negative it is simply a valid, vacuous lower bound. For the prescribed parameter is zero and the probability lower bound is also zero, so no positive-parameter assertion is supplied by that prescription.
Let and . Comparing the Lasso objective at its minimizer and at and expanding the squares gives the Basic inequality for the Lasso
On , the stochastic term is at most . Because , the triangle inequality gives
It follows that
In particular , which is the Lasso cone condition. The given Compatibility condition for the Lasso is therefore applicable and gives .
Adding to the preceding inequality and using compatibility now yields
The last inequality is just the nonnegativity of . Thus , and dropping one nonnegative coefficient-error term gives
This prediction and coefficient error bound for compatible Lasso holds on , and hence with the required probability. If , the earlier basic inequality directly forces on , so the zero right side is also correct.