Solution

ID: past-exam-of-the-mathematics-course-of-the-university-of-cambridge/2021/iii/paper-218/4/c/solution

Deep sigmoid networks suffer from vanishing gradients because derivatives become small in saturated units and multiply across layers. The ReLU has derivative one on its positive half-line and therefore propagates gradients more effectively there.

New to topics? Read the docs here!