Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2023/iii/paper-218/4/d/solution

The first observation has and one-hot label in the output order . If every kernel weight and bias initially equals one, each hidden preactivation is two, so . Both logits equal five and .
For stochastic gradient descent on one cross-entropy observation,
Hence the output-weight gradient has first row and second row , while the output-bias gradient is . With learning rate one,
Using the old output weights for backpropagation gives
so every entry of and remains equal to one after this batch.

New to topics? Read the docs here!