The objective is strictly convex, so the minimizer is unique. The Slater condition makes the Karush-Kuhn-Tucker conditions necessary and sufficient. Absorb the box constraints into the Euclidean projection onto a convex set and attach a scalar multiplier to . Stationarity over the box is equivalent to
while primal feasibility requires . Coordinatewise, these conditions are
They are also sufficient because they minimize the Lagrangian over the box and satisfy the equality constraint. Thus the projection onto a box-constrained hyperplane reduces to solving the displayed one-dimensional continuous, nonincreasing equation for . The multiplier need not be unique on a flat interval, but the projected vector is unique.
For a proper lower-semicontinuous convex function , its proximal operator is
The squared norm is strongly convex, so the minimizer is unique. The subdifferential sum rule gives the necessary and sufficient condition
More generally,
The subgradient inversion rule for the convex conjugate says exactly when . Hence
which is precisely the proximal optimality condition
Since , this proves the generalized Moreau decomposition
The function is the support function . For a nonempty compact convex set,
so its convex conjugate is the indicator function . Applying the Moreau decomposition,
Multiplication of an indicator function by a positive scalar does not change it, and its proximal operator is the Euclidean projection onto a convex set. Therefore
Take and , so
is the capped simplex. A linear objective over this convex polytope attains its maximum at a zero-one extreme point. Choosing the coordinates at which is largest gives
Equivalently, an exchange of weight from a smaller component to a larger one never decreases the objective. Thus the sum of the largest components is the support function .
Part c now gives
By the projection onto a box-constrained hyperplane, has
Consequently the proximal operator is evaluated by solving this one-dimensional equation for , then substituting the resulting projection.
Set
The gradient has Lipschitz continuity with constant
where the norm is the spectral norm. The proximal gradient method is therefore
For , part d makes the second step explicit:
where is the capped simplex; when , this proximal step is the identity.
A standard fixed choice is ; the wider interval also gives convergence under the usual forward-backward conditions. For a general convex objective, the function-value error is . If has full column rank, the quadratic term is strongly convex and an appropriate fixed step gives a linear convergence rate.

Articles by others on the same topic (0)

There are currently no matching articles.