Past exam of the mathematics course of the University of Cambridge 2018 iii Paper 205 4 Solution Created 2026-10-03 Updated 2026-10-05
For a convex function , its subdifferential at is the set of supporting slopesThe subgradient optimality condition isIndeed, the defining inequality with is exactly the global minimum inequality. For the subdifferential of the L1 norm, the coordinates vary independently:
Let . The Karush-Kuhn-Tucker conditions for Lasso give a vector such that . If , the unsquared residual norm is differentiable at this fit, with gradientUsing the subdifferential sum rule, choose to obtainConvexity now proves the tuning relation between Lasso and square-root Lasso:The positive residual assumption is needed both for this gradient and for the quotient defining the tuning parameter.
For the final score calculation, use the usual fixed-design interpretation of the normal linear model: regard and the tuning parameter in the regression of as fixed, and choose that regression's minimizer using only those inputs. Its nonzero residual then has the Square-root Lasso optimality conditionExpanding the response residual givesSince is a fixed unit vector and , a linear image of a multivariate normal vector gives . Finally, Holder inequality and the residual score bound yieldThis is the residual-score decomposition from square-root Lasso. The normality statement also holds conditionally on a design independent of the noise. If were chosen from the response noise, for example through the preceding formula, the direction would require a separate independence argument; the final calculation treats as fixed.