Deep sigmoid networks suffer from vanishing gradients because derivatives become small in saturated units and multiply across layers. The ReLU has derivative one on its positive half-line and therefore propagates gradients more effectively there.
Articles by others on the same topic
There are currently no matching articles.