Past exam of the mathematics course of the University of Cambridge 2019 iii Paper 218 5 c Solution 2026-10-03
For one-hot labels and predicted probabilities , the categorical cross-entropy loss isStochastic gradient descent initializes , randomly orders the observations in each of five epochs, and for each single-observation batch computes a forward pass, the sample loss, and its gradient, then updatesBackpropagation is used after the forward loss evaluation to compute this gradient from the output layer back through the hidden layers.