Define the softmax weights
They obey and . The gradient is their weighted mean,
Differentiating once more gives the covariance-form Hessian matrix
For every unit vector ,
where . The Hessian is a covariance matrix, so it is positive semidefinite; the displayed upper bound also gives in the Loewner order. Consequently has a Lipschitz gradient with

Articles by others on the same topic (0)

There are currently no matching articles.