Past exam of the mathematics course of the University of Cambridge 2019 iii Paper 218 5 b Solution 2026-10-03
The loop divides each feature by its sample maximum, putting features with nonnegative values on comparable scales and improving numerical conditioning for gradient optimization. If exactly one hidden layer must use the sigmoid function, use layer 3. Sigmoid derivatives can become small in saturated regions; placing it latest minimizes the number of subsequent gradient multiplications affected by this saturation, while the earlier ReLU layers retain efficient gradient propagation.