Solution (source code)

= Solution

Writing $\eta_i=\beta_0+x_i^T\beta$, the <log-likelihood> is
$$
\ell(\beta_0,\beta)
=\sum_{i=1}^n
\left\{y_i\eta_i-\log(1+e^{\eta_i})\right\}.
$$
With 500 observations and 5000 word predictors, the augmented design matrix cannot have full column rank. There are nonzero coefficient directions that leave every $\eta_i$ unchanged, so any maximizer belongs to an affine family and is not unique. In addition, the high-dimensional features may completely separate the classes, in which case no finite maximizer exists at all.