Use the convention
This proximal map is also the resolvent of a monotone operator . A proper lower semicontinuous convex function has an affine minorant, so the quadratic term makes this minimization coercive and strongly convex. A unique minimizer exists for every . Its subgradient optimality condition is
The Moreau–Yosida regularisation is . The conjugate of an infimal convolution and the quadratic conjugate give
The factor here is essential. Apply subgradient inversion under convex conjugacy, followed by the subdifferential sum rule with the everywhere differentiable quadratic:
The last equivalence is precisely the unique proximal minimization condition. Thus
This proves both existence and uniqueness of the subgradient, rather than only identifying a possible element. The finite convex function is therefore differentiable, with , the gradient of a Moreau envelope.
For completeness, monotonicity of applied to the two proximal conditions gives
Hence is firmly nonexpansive. Expanding the same inequality shows that is firmly nonexpansive too. In particular, is -Lipschitz continuous. None of this requires a bounded effective domain; the result applies to the next example as well.
For the Euclidean norm, rotational symmetry and minimizing the radial objective give the radial soft thresholding formula
Indeed, if , an optimal is a nonnegative scalar multiple of ; minimizing over gives . The gradient of a Moreau envelope is consequently
At this is , and the two expressions agree when . The Moreau envelope of the Euclidean norm itself is
Thus a quadratic core replaces the nondifferentiable tip while the outer gradient remains the normalized radial direction. The Euclidean norm has an unbounded effective domain, unlike the earlier bounded-domain hypothesis; the preceding argument explicitly shows that this restriction is unnecessary here.
Figure 1.
The absolute value and its Moreau envelope with tau equal to one, together with the continuous clipped gradient replacing the jump at the origin
.
For convex optimization, this Moreau–Yosida regularisation permits gradient descent and other methods for functions with a Lipschitz gradient, using a gradient Lipschitz constant . It approximates the norm uniformly: . Smaller improves approximation but makes the permitted gradient steps smaller. Minimizing alone preserves the minimizers and minimum value of ; replacing one term in a larger objective can shift the minimizer, so the smoothing parameter controls that approximation error.