For the one-layer softmax fit, write its class logits as . Then
Thus an equivalent logistic regression classifier predicts a click exactly when
Using one reference class removes the common-logit-shift non-identifiability. Subject to the usual full-rank and no-separation conditions, the logistic parameter is identifiable. It uses four effective parameters rather than an eight-parameter redundant softmax representation, has a convex loss, and gives directly interpretable log-odds coefficients, so it is preferable for this binary linear classifier.

Articles by others on the same topic (0)

There are currently no matching articles.