With and , the model-matrix input isThe feedforward neural network has three inputs, a fully connected layer of two ReLU units, and a fully connected two-class softmax output. Algebraically,where , , , and . The coding is for no click and for a click. There areparameters.
The softmax non-identifiability already proves that the coefficient vector is not unique. For any and , replace both rows of by and both output biases by . Every logit then gains the same value , so all softmax probabilities, classifications, and losses remain unchanged. Permuting the two hidden units supplies another non-uniqueness.
For the one-layer softmax fit, write its class logits as . ThenThus an equivalent logistic regression classifier predicts a click exactly whenUsing one reference class removes the common-logit-shift non-identifiability. Subject to the usual full-rank and no-separation conditions, the logistic parameter is identifiable. It uses four effective parameters rather than an eight-parameter redundant softmax representation, has a convex loss, and gives directly interpretable log-odds coefficients, so it is preferable for this binary linear classifier.
The first observation has and one-hot label in the output order . If every kernel weight and bias initially equals one, each hidden preactivation is two, so . Both logits equal five and .
For stochastic gradient descent on one cross-entropy observation,Hence the output-weight gradient has first row and second row , while the output-bias gradient is . With learning rate one,Using the old output weights for backpropagation givesso every entry of and remains equal to one after this batch.
Articles by others on the same topic
There are currently no matching articles.