Solution
= Solution
Deep sigmoid networks suffer from vanishing gradients because derivatives become small in saturated units and multiply across layers. The <ReLU> has derivative one on its positive half-line and therefore propagates gradients more effectively there.