For observations and graph , Bayes theorem gives . Here is the graph prior distribution and the local distribution statistical parameters. The integral is Bayesian model evidence; it averages over nuisance parameters with a proper statistical parameter prior distribution. For independent complete observations, the node-factorized likelihood function and independence of local statistical parameter prior distributions make the evidence factor over nodes; suitable conjugate priors can make the local integrals analytic. With missing node observations, integrating out unobserved values can couple the local parameters, so independence of local prior distributions alone does not guarantee this factorization. Proper priors and coherent hyperparameters matter for comparing different graphs.
Nuisance parameter 2026-10-05
A nuisance parameter is a statistical parameter required to specify the distribution of observations but not itself the target of inference. A profile likelihood maximizes over it; Bayesian model evidence integrates over it using a proper prior distribution. These operations differ and need not give the same inference. A baseline hazard in a Cox proportional-hazards model is an infinite-dimensional example.
Let contain independent complete observations of the -dimensional vector, and let be a Directed acyclic graph. A Gaussian Bayesian network uses, for node ,
Here is an intercept, the vector of parent coefficients, and the conditional variance; the vector collects these statistical parameters. Its likelihood function factors as
Choose a proper statistical parameter prior distribution and a graph prior distribution . Bayes theorem gives the Bayesian network structure score, up to the normalization common to graphs,
The integral is Bayesian model evidence: it averages over nuisance parameters rather than substituting their best-fitting values. A graph prior can favor sparse graphs, while the evidence balances fit against the amount of prior statistical parameter space that predicts the observations well. Compare these scores over admissible Directed acyclic graphs, using enumeration when feasible or a search procedure otherwise; search need not find a global maximum, and observational data need not identify a unique causal orientation.
A conjugate prior makes the integral analytic. For example, for take , , where is positive-definite matrix and ; these are independent local normal-inverse-gamma priors. Under prior independence between nodes, the evidence is a product of local regression evidences. Each posterior has the same family, so the local integral is the ratio of prior and posterior normalization constants, with the likelihood's constants included. This avoids costly numerical integration and makes local graph updates inexpensive. Hyperparameters must be specified coherently if score equivalence between observationally equivalent graphs is desired; arbitrary local priors do not automatically have that property.