The fitted classifier is
and training minimizes the empirical categorical cross-entropy loss
The softmax non-identifiability already proves that the coefficient vector is not unique. For any and , replace both rows of by and both output biases by . Every logit then gains the same value , so all softmax probabilities, classifications, and losses remain unchanged. Permuting the two hidden units supplies another non-uniqueness.

Articles by others on the same topic (0)

There are currently no matching articles.