Smooth maximum
= Smooth maximum
Composing the <log-sum-exp function> with finitely many <affine functions> gives a differentiable approximation to their pointwise maximum. Increasing the inverse-temperature parameter $\beta$ reduces the uniform approximation error while increasing the <Lipschitz gradient> constant.