A dead ReLU unit has nonpositive preactivation for every training input, so its output and gradient are always zero. Replace by the leaky ReLUIts negative-side derivative permits recovery while it remains nonsaturating, preserving the gradient advantage over sigmoid activation.
Articles by others on the same topic
There are currently no matching articles.