Put and . The probit regression likelihood is , so with the stated normal prior,
Differentiating gives the probit posterior score
Equivalently, the th likelihood term is . The prior's gradient is essential.
This posterior density is smooth, positive everywhere, and bounded above by a constant times the Gaussian prior, since the likelihood is at most one and has positive normalizing constant. Its derivatives are integrable: differentiating a likelihood factor gives a bounded normal density, and differentiating the prior gives a linear factor times a Gaussian. Integration by parts therefore yields
Thus each component has a known zero expected value and can be used as a posterior score control variate. This identity differentiates the posterior density in the random parameter; it is distinct from the usual mean-zero score identity for a likelihood under repeated sampling.
The requisite second moments are finite. For , is bounded. For ,
Hence the score grows at most exponentially in , and the Gaussian upper bound on the posterior gives finite second moments. For a fixed vector , averaging is therefore an unbiased Monte Carlo estimator of the same mean; suitable coefficients can reduce its variance.