Solution
= Solution
A dead ReLU unit has nonpositive preactivation for every training input, so its output and gradient are always zero. Replace $\max(0,z)$ by the leaky ReLU
$$
\phi_a(z)=\max(z,az),\qquad0<a<1.
$$
Its negative-side derivative $a$ permits recovery while it remains nonsaturating, preserving the gradient advantage over sigmoid activation.