The first observation has and one-hot label in the output order . If every kernel weight and bias initially equals one, each hidden preactivation is two, so . Both logits equal five and .
For stochastic gradient descent on one cross-entropy observation,Hence the output-weight gradient has first row and second row , while the output-bias gradient is . With learning rate one,Using the old output weights for backpropagation givesso every entry of and remains equal to one after this batch.
Articles by others on the same topic
There are currently no matching articles.