For observations and graph , Bayes theorem gives . Here is the graph prior distribution and the local distribution statistical parameters. The integral is Bayesian model evidence; it averages over nuisance parameters with a proper statistical parameter prior distribution. For independent complete observations, the node-factorized likelihood function and independence of local statistical parameter prior distributions make the evidence factor over nodes; suitable conjugate priors can make the local integrals analytic. With missing node observations, integrating out unobserved values can couple the local parameters, so independence of local prior distributions alone does not guarantee this factorization. Proper priors and coherent hyperparameters matter for comparing different graphs.
Nuisance parameter 2026-10-05
A nuisance parameter is a statistical parameter required to specify the distribution of observations but not itself the target of inference. A profile likelihood maximizes over it; Bayesian model evidence integrates over it using a proper prior distribution. These operations differ and need not give the same inference. A baseline hazard in a Cox proportional-hazards model is an infinite-dimensional example.
Let contain independent complete observations of the -dimensional vector, and let be a Directed acyclic graph. A Gaussian Bayesian network uses, for node ,
Here is an intercept, the vector of parent coefficients, and the conditional variance; the vector collects these statistical parameters. Its likelihood function factors as
Choose a proper statistical parameter prior distribution and a graph prior distribution . Bayes theorem gives the Bayesian network structure score, up to the normalization common to graphs,
The integral is Bayesian model evidence: it averages over nuisance parameters rather than substituting their best-fitting values. A graph prior can favor sparse graphs, while the evidence balances fit against the amount of prior statistical parameter space that predicts the observations well. Compare these scores over admissible Directed acyclic graphs, using enumeration when feasible or a search procedure otherwise; search need not find a global maximum, and observational data need not identify a unique causal orientation.
A conjugate prior makes the integral analytic. For example, for take , , where is positive-definite matrix and ; these are independent local normal-inverse-gamma priors. Under prior independence between nodes, the evidence is a product of local regression evidences. Each posterior has the same family, so the local integral is the ratio of prior and posterior normalization constants, with the likelihood's constants included. This avoids costly numerical integration and makes local graph updates inexpensive. Hyperparameters must be specified coherently if score equivalence between observationally equivalent graphs is desired; arbitrary local priors do not automatically have that property.
The original PDF shows the chain , not a collider. Its Bayesian network factorization is . For any with positive marginal probability,
Integrating or summing over also gives the same conditional marginal . Thus
This is conditional independence; the graph does not generally imply marginal independence, since summing over can transmit dependence from to . Special statistical parameter choices can make marginal independence hold as well, so the claim is that only the conditional statement is guaranteed by the graph. Conditional distributions on null values of are immaterial to this assertion.
Reading the arrows in the original diagram gives parents , and . The Bayesian network factorization is
For binary variables, each conditional probability table has one free probability per parent configuration, since its two entries sum to one. There are one each for and , four for , and two for , giving
This is the dimension of the unrestricted binary Bayesian network family; extra statistical parameter equalities or deterministic relationships would describe a smaller model and are not imposed by the diagram.
Set , and . With equal positive rates, the one-year probability of a change is , so the conditional likelihood function is proportional to
For , its unrestricted binomial maximizer is . A finite interior positive-rate maximum exists exactly when , or
Indeed the binomial log-likelihood is strictly concave at this interior maximum and is strictly increasing. If and , the likelihood increases towards its supremum as , requiring ; no finite maximum exists. Equality also has only this limiting maximum.
If , allowing the boundary gives a finite maximum there; insisting on gives only a supremum as . Thus if “valid” includes zero rates, the condition is with ; if rates must be strictly positive, it is . If , every rate has the same empty-data likelihood and the statistical parameter is not identifiable. These boundary qualifications prevent a negative or infinite plug-in rate from being called a valid interior estimate.
Use one joint hypothesis test for . Fit the unrestricted model with statistical parameters , where and for treatment indicator . Refit under , estimating the two common baseline intensities; do not fix them at their unrestricted estimates. The likelihood-ratio test statistic is
Under regular identifiable positive-rate models, sufficient independent patient trajectories, and the null hypothesis, Wilks theorem gives two degrees of freedom because two coefficients are constrained. Reject if exceeds the chosen upper chi-squared distribution quantile. Alternatively a two-dimensional Wald test uses , including the off-diagonal covariance. Two separate tests do not provide this single calibrated assessment. The point estimates alone cannot produce a numerical p-value without fitted likelihoods or a coefficient covariance matrix.
Statistical parameter 2026-10-05
A statistical parameter indexes a family of probability distributions in a statistical model. It is fixed in a frequentist sampling model; a prior distribution makes it random for Bayesian statistics. Different statistical parameter values should produce different observable distributions when identifiability is claimed.