= Solution
For the one-layer softmax fit, write its class logits as $z_k=a_k+b_k^Tx$. Then
$$
\log\frac{p_1(x)}{p_0(x)}
=(a_1-a_0)+(b_1-b_0)^Tx.
$$
Thus an equivalent <logistic regression> classifier predicts a click exactly when
$$
\delta_0+\delta^Tx\geq0,
\qquad
\delta_0=a_1-a_0,
\quad \delta=b_1-b_0.
$$
Using one reference class removes the common-logit-shift non-identifiability. Subject to the usual full-rank and no-separation conditions, the logistic parameter is identifiable. It uses four effective parameters rather than an eight-parameter redundant softmax representation, has a convex loss, and gives directly interpretable log-odds coefficients, so it is preferable for this binary linear classifier.
Back to article page