Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2019/iii/paper-218/5/c/solution

For one-hot labels and predicted probabilities , the categorical cross-entropy loss is
Stochastic gradient descent initializes , randomly orders the observations in each of five epochs, and for each single-observation batch computes a forward pass, the sample loss, and its gradient, then updates
Backpropagation is used after the forward loss evaluation to compute this gradient from the output layer back through the hidden layers.

New to topics? Read the docs here!