= Solution
The <conjugate of an infimal convolution> is the sum of the conjugates:
$$
f^*(p)=\sup_{y,z}\{\langle p,y+z\rangle-g(y)-h(z)\}=g^*(p)+h^*(p).
$$
Here $g^*(p)=\frac12\|p\|_2^2$, and $h^*(p)=\delta_{[-1,1]^n}(p)$, the <indicator functional> of the unit infinity-norm ball. Thus
$$
f(x)=\sup_{\|p\|_\infty\leq1}\left(\langle p,x\rangle-\frac12\|p\|_2^2\right).
$$
The <infimal convolution> is finite convex and continuous, so the <Fenchel-Moreau theorem> applies without a closure defect. By equality in the <Fenchel–Young inequality> and the <subdifferential sum rule>,
$$
p\in\partial f(x)\quad\Longleftrightarrow\quad x\in\partial(g^*+h^*)(p)
=p+N_{[-1,1]^n}(p).
$$
This is precisely the variational characterization of projecting $x$ onto the cube. Consequently
$$
\boxed{\partial f(x)=\{\operatorname{clip}(x,-1,1)\},\qquad
f(x)=\sum_{i=1}^n\begin{cases}\frac12x_i^2,&|x_i|\leq1,\\|x_i|-\frac12,&|x_i|>1.\end{cases}}
$$
For the primal split, $y_i=\operatorname{clip}(x_i,-1,1)$ and $z_i=\operatorname{sign}(x_i)(|x_i|-1)_+$. These directly minimize the two scalar terms. The <Huber loss> is continuously differentiable, including at $x_i=\pm1$, but its second derivative changes there. The one-dimensional sketch shows a quadratic center joined tangentially to linear tails:
\Image[/past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2014/iii/paper-65-huber-function.png]
{title=The scalar Huber function with quadratic center and linear tails joined at minus one and one}
{height=400}
Applied to a discrete gradient, the <Huber gradient regularizer> penalizes small slopes quadratically and large slopes linearly. Compared with pure squared-gradient smoothing it preserves large edges better; compared with pure <total variation denoising> it encourages small smooth variations and reduces the strong preference for piecewise-constant plateaus. It can therefore be useful for denoising signals or images containing both smooth regions and sharp transitions. It still penalizes edges and can bias their amplitude, and it does not guarantee complete elimination of <staircasing in total variation denoising>. The unit threshold must be scaled appropriately for data units and grid spacing.
Back to article page