A maximum marginal likelihood estimator chooses a hyperparameter by maximizing the data density obtained after integrating the model parameter against its prior distribution.
After integrating out the multivariate normal distribution , the marginal distribution is
Up to terms independent of , the log-likelihood is
Differentiating and setting the result to zero gives the maximum marginal likelihood estimator
This is an Empirical Bayes method because the estimated hyperparameter is then inserted into the prior and posterior distributions.
Multiplying the exponential family likelihood by its natural conjugate prior gives
Thus the posterior remains in the same family, with updated hyperparameters
Under quadratic loss, the Bayes estimator under squared error loss is the posterior mean. Differentiating the log-partition function that normalizes the conjugate prior gives