For the one-layer softmax fit, write its class logits as . ThenThus an equivalent logistic regression classifier predicts a click exactly whenUsing one reference class removes the common-logit-shift non-identifiability. Subject to the usual full-rank and no-separation conditions, the logistic parameter is identifiable. It uses four effective parameters rather than an eight-parameter redundant softmax representation, has a convex loss, and gives directly interpretable log-odds coefficients, so it is preferable for this binary linear classifier.
Articles by others on the same topic
There are currently no matching articles.