= Solution
The first observation has $x=(1,0,0)^T$ and one-hot label $y=(0,1)^T$ in the output order $(\mathrm{no},\mathrm{yes})$. If every kernel weight and bias initially equals one, each hidden preactivation is two, so $h=(2,2)^T$. Both logits equal five and $p=(1/2,1/2)^T$.
For <stochastic gradient descent> on one cross-entropy observation,
$$
\frac{\partial L}{\partial z}=p-y=(1/2,-1/2)^T.
$$
Hence the output-weight gradient has first row $(1,1)$ and second row $(-1,-1)$, while the output-bias gradient is $(1/2,-1/2)$. With learning rate one,
$$
V^{\mathrm{new}}=
\begin{pmatrix}0&0\\2&2\end{pmatrix},
\qquad
c^{\mathrm{new}}=(1/2,3/2)^T.
$$
Using the old output weights for backpropagation gives
$$
V^T(p-y)=(0,0)^T,
$$
so every entry of $W$ and $b$ remains equal to one after this batch.
Back to article page