At zero a centered normal distribution with standard deviation has density . The two prior standard deviations are and , giving
The wider prior distribution has a lower density at its center.
Let , and . These are the two prior precision parameters. Completing the square in gives the Normal-normal conjugacy update
Thus the posterior mean is a precision-weighted average of the observed mean and the prior center zero. The posterior variance is the reciprocal of the total precision.
Write , with independent and . The convolution of independent random variables is again a normal distribution, so the prior predictive laws are
These are predictive distributions before observing , hence the Bayesian model evidence for the observed mean. The residual information in the original observations is common to both models and cancels in their Bayes factor.
Divide the two normal distribution predictive densities to obtain the Bayes factor
At this becomes
The wider prior distribution spreads its predictive mass over more possible means, giving the narrower model more Bayesian model evidence for observations very near zero. Away from zero, the exponential term opposes that factor.
Under the point hypothesis , the test statistic has a standard normal distribution. The observed statistic is three, giving a two-sided p-value . This is conventionally strong evidence against that point hypothesis. The narrow model is a continuous prior distribution around zero, rather than a point hypothesis.
For and , Normal-normal conjugacy gives
The narrow-model posterior mean is strongly pulled toward zero and its standard deviation is about . The wide-model posterior mean is about , very close to the observation, with standard deviation about . Each model produces a markedly different posterior, so choosing between them requires their predictive evidence.
The requested conclusion is false for the printed parameter values. They give and . Substituting in the Bayes factor yields
Thus the Bayes factor favours by about to one, not . Nor does the conclusion hold for arbitrary large : both parameters enter the expression explicitly. With , taking a much wider alternative, for example , would instead give . That is a different prior assumption and cannot repair the printed calculation silently.
With equal model prior probabilities, Bayes factor updating gives and . The Bayesian model averaging posterior density is therefore
This is a two-component mixture model of the normal distributions already derived. For the numerical observation above, and . Consequently it is predominantly the wide-model posterior, not predominantly the narrow one.
Put . The posterior mean of the Gaussian practical-null mixture is
Near zero, , so the narrow component has high posterior probability and is small when . This pulls the Bayesian model averaging posterior toward zero. For the given ratios, writing gives
Thus the order statement alone does not guarantee strong pull toward zero: already has the opposite model preference.
As grows, makes , and the mixture model approaches . Its center is exactly , and approximately only for a sufficiently diffuse wide prior, meaning . Here , so the relative displacement is about one percent. Large observations alone do not remove this finite-prior shrinkage.
In a normal linear model, attach a continuous spike-and-slab prior to each regression coefficient. With suitably scaled predictors, introduce indicators and set
The narrow component describes practically negligible effects; the wide component permits substantial ones. Fit the joint Bayesian posterior of coefficients, indicators and any unknown residual variance. Bayesian model averaging gives shrinkage toward zero for poorly supported effects, while quantifies wide-component support. A shared Beta distribution prior on can represent uncertainty about how many effects are substantial.
Select on scientifically meaningful effect size, for example a high for a prechosen threshold in meaningful predictor units. Membership in the wide component alone does not imply a large realized effect: its normal distribution still permits values near zero. Correlated predictors also require interpretation of the joint Bayesian posterior, rather than treating each coefficient as an isolated test.

Articles by others on the same topic (0)

There are currently no matching articles.