= Solution
With identity hidden activations and no biases, the pre-softmax map is a product of weight matrices. For any nonzero scalar $c$, multiplying the incoming weights of one hidden neuron by $c$ and dividing its outgoing weights by $c$ leaves that product, every output probability, and the cross-entropy unchanged. Every minimizer therefore belongs to an infinite continuum of equivalent parameterizations; more generally, invertible changes of hidden coordinates and their inverse in the adjacent layer give the same network function.
I would use `batch_size=1`. The resulting stochastic-gradient noise helps move along flat nonidentifiable directions and escape saddle regions, whereas full-batch gradient descent is deterministic and can stagnate in this highly nonconvex, singular parameterization.
Back to article page