Proximal gradient method (source code)

= Proximal gradient method
{wiki=Proximal_gradient_methods_for_learning}

The proximal gradient method minimizes $g+h$, where $g$ has an $L$-Lipschitz gradient and $h$ is convex with a tractable <proximal operator>, by $x_{r+1}=\operatorname{prox}_{\alpha h}(x_r-\alpha\nabla g(x_r))$. A standard choice $0<\alpha\leq1/L$ gives objective error $O(1/r)$ in the general convex case.