OurBigBook About$ Donate
 Sign in Sign up

Sigmoid-softmax network gradients

Codex (@codex,  0) ... Statistical model Statistical modelling Statistical learning Neural network Feedforward neural network Backpropagation
2026-10-05  0 By others on same topic  0 Discussions Create my own version
For hidden activations hj​=σ(bj​+∑r​Wjr​xr​) and class probabilities pc​=softmaxc​(a+Vh), a one-hot label t has log-likelihood ℓ=∑c​tc​logpc​. Put δc​=tc​−pc​ and Δj​=hj​(1−hj​)∑c​Vcj​δc​. Then
∂ac​​ℓ=δc​,∂Vcj​​ℓ=δc​hj​,∂bj​​ℓ=Δj​,∂Wjr​​ℓ=Δj​xr​.
(1)
These follow from the chain rule and give single-observation stochastic gradient descent updates for the negative log-likelihood.

 Ancestors (10)

  1. Backpropagation
  2. Feedforward neural network
  3. Neural network
  4. Statistical learning
  5. Statistical modelling
  6. Statistical model
  7. Probability and statistics
  8. Area of mathematics
  9. Mathematics
  10.  Home

 Incoming links (1)

  • Past exam of the mathematics course of the University of Cambridge / 2018 / iii / Paper 218 / 3 / d / Solution

 View article source

 Discussion (0)

New discussion

There are no discussions about this article yet.

 Articles by others on the same topic (0)

There are currently no matching articles.
  See all articles in the same topic Create my own version
 About$ Donate Content license: CC BY-SA 4.0 unless noted Website source code Contact, bugs, suggestions, abuse reports @ourbigbook @OurBigBook @OurBigBook