Solution (source code)

= Solution

For one-hot labels $y_{il}$ and predicted probabilities $p_{il}(\theta)$, the <categorical cross-entropy loss> is
$$
L(\theta)=-\sum_{i=1}^n\sum_{l=1}^{26}y_{il}\log p_{il}(\theta).
$$
Stochastic gradient descent initializes $\theta$, randomly orders the observations in each of five epochs, and for each single-observation batch computes a forward pass, the sample loss, and its gradient, then updates
$$
\theta\leftarrow\theta-\eta\nabla_\theta L_i(\theta).
$$
<Backpropagation> is used after the forward loss evaluation to compute this gradient from the output layer back through the hidden layers.