Past exam of the mathematics course of the University of Cambridge 2026 iii Paper 218 5 b Solution Created 2026-09-24 Updated 2026-09-24
Stochastic gradient descent replaces the full empirical-loss gradient by the gradient on a randomly ordered observation or mini-batch, then updates . The training half contains observations, so batches of 16 give updates per epoch. Over 100 epochs every parameter is updated times.