Bayesian deviance 2026-10-06
A Bayesian deviance is minus twice the log-likelihood, with any chosen additive data-only constant held consistent across the models being compared. Its Bayesian posterior expectation measures average fit. In the deviance information criterion, evaluating it at a parameter's posterior mean also enters the effective complexity penalty. This likelihood-based convention differs by a data-only constant from a saturated-model exponential-family deviance when a common saturated model exists.
Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 35 4 g Solution Created 2026-10-03 Updated 2026-10-06
For the common Bayesian deviance convention used in all three models,Here
Dhat is , an at-posterior-mean fit measure, and is an effective parameter count. The independent model's Dhat of 53.1 is almost identical to the exchangeable model's 53.2; both improve on the common model's 57.8. The common model's corresponds to six intercepts plus one shared effect. Independence uses roughly twelve effective parameters. Partial pooling reduces the exchangeable model's effective complexity to about 8.7 while retaining nearly the same fitted likelihood function as independence.The exchangeable model has the lowest reported DIC, but the common model is competitive. Their difference is only about 1.3, whereas independence is worse by about 6.3. The deviance information criterion measures penalized fit for a predictive comparison, not model posterior probabilities, and these numbers do not establish overwhelming evidence for heterogeneity. The displayed exchangeable is , rather than the printed 70.5; rounding of the underlying values can account for a tenth and does not change this interpretation.
Past exam of the mathematics course of the University of Cambridge 2014 iii Paper 35 4 j Solution Created 2026-10-03 Updated 2026-10-06
Fit both the normal and Student t random-effect model with comparable proper prior distributions. Compare priors on the same spread measure: a normal distribution scale is a standard deviation, whereas the standard deviation is .
Use a posterior predictive check: draw study effects and binomial counts from each fitted hierarchy and compare replicated dispersion and extreme study contrasts with the observations. For predicting a new study, generate a new effect from the hierarchy rather than reusing an existing fitted effect. A Leave-one-out cross-validation with entire studies held out can compare integrated predictive probabilities for both arms of each omitted trial, averaging over hyperparameters and its unobserved study effect. Leave-one-study-out influence analysis also reveals whether the difference is driven by a single trial.
Prefer the heavier-tailed hierarchy if it improves the relevant predictive checks and held-out study predictions robustly to reasonable prior choices. The deviance information criterion can supplement the comparison, but its effective parameter count can depend on the latent-variable representation, and six studies give limited information about tail shape. A small numerical criterion difference alone is insufficient evidence.