Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 207 2 b iii Solution Created 2026-10-03 Updated 2026-10-05
Let contain independent complete observations of the -dimensional vector, and let be a Directed acyclic graph. A Gaussian Bayesian network uses, for node ,Here is an intercept, the vector of parent coefficients, and the conditional variance; the vector collects these statistical parameters. Its likelihood function factors asChoose a proper statistical parameter prior distribution and a graph prior distribution . Bayes theorem gives the Bayesian network structure score, up to the normalization common to graphs,The integral is Bayesian model evidence: it averages over nuisance parameters rather than substituting their best-fitting values. A graph prior can favor sparse graphs, while the evidence balances fit against the amount of prior statistical parameter space that predicts the observations well. Compare these scores over admissible Directed acyclic graphs, using enumeration when feasible or a search procedure otherwise; search need not find a global maximum, and observational data need not identify a unique causal orientation.
A conjugate prior makes the integral analytic. For example, for take , , where is positive-definite matrix and ; these are independent local normal-inverse-gamma priors. Under prior independence between nodes, the evidence is a product of local regression evidences. Each posterior has the same family, so the local integral is the ratio of prior and posterior normalization constants, with the likelihood's constants included. This avoids costly numerical integration and makes local graph updates inexpensive. Hyperparameters must be specified coherently if score equivalence between observationally equivalent graphs is desired; arbitrary local priors do not automatically have that property.