Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2019/iii/paper-218/5/d/solution

With identity hidden activations and no biases, the pre-softmax map is a product of weight matrices. For any nonzero scalar , multiplying the incoming weights of one hidden neuron by and dividing its outgoing weights by leaves that product, every output probability, and the cross-entropy unchanged. Every minimizer therefore belongs to an infinite continuum of equivalent parameterizations; more generally, invertible changes of hidden coordinates and their inverse in the adjacent layer give the same network function.
I would use batch_size=1. The resulting stochastic-gradient noise helps move along flat nonidentifiable directions and escape saddle regions, whereas full-batch gradient descent is deterministic and can stagnate in this highly nonconvex, singular parameterization.

New to topics? Read the docs here!