For hidden activations and class probabilities , a one-hot label has log-likelihood . Put and . ThenThese follow from the chain rule and give single-observation stochastic gradient descent updates for the negative log-likelihood.
Articles by others on the same topic
There are currently no matching articles.