Solution (source code)

= Solution

The fitted classifier is
$$
\widehat C(x)=\operatorname*{argmax}_{k\in\{0,1\}}\widehat p_k(x),
$$
and training minimizes the empirical <categorical cross-entropy loss>
$$
-\sum_i\sum_{k=0}^1y_{ik}\log p_{ik}.
$$

The <softmax non-identifiability> already proves that the coefficient vector is not unique. For any $a\in\mathbb R^2$ and $r\in\mathbb R$, replace both rows of $V$ by $V_{k\cdot}+a^T$ and both output biases by $c_k+r$. Every logit then gains the same value $a^Th+r$, so all softmax probabilities, classifications, and losses remain unchanged. Permuting the two hidden units supplies another non-uniqueness.