Jeffreys prior 2026-10-05
The Jeffreys prior uses the square root of the determinant of the Fisher information matrix as a parameter-density kernel. It is invariant under smooth one-to-one reparameterization, since the information and density Jacobians transform compatibly. It may be an improper prior; invariance does not guarantee posterior propriety. With nuisance parameters, a scalar conditional Jeffreys prior and the joint Jeffreys prior need not coincide.
Several chains from dispersed initial states, trace plots, autocorrelation, effective sample sizes, and the rank-normalized split potential scale reduction factor assess mixing and disagreement between chains. An close to one is useful evidence but cannot prove convergence or prove posterior propriety. Comparing chains that explore different modes is especially important; monitor each scientifically important observable and its Monte Carlo uncertainty. A long apparent plateau can also occur when an improper-target sampler is drifting toward a boundary.
For a regular one-parameter sampling distribution, the Fisher information and Jeffreys prior are
Under the usual differentiation and integrability conditions, . This prior distribution transforms as a density under smooth one-to-one reparameterizations, so the rule is coordinate invariant. Its integral need not be finite; posterior propriety must still be established if it is an improper prior.
Let . Bayes theorem and the stated prior factorization give the posterior density
The normalizing constant, when finite and positive, is the integral of this product over its admissible parameter domain. An improper prior requires checking posterior propriety before interpreting this expression as a probability density.
A log-flat improper prior is not appropriate when all known are positive: the integrated likelihood in (iii) has a positive finite limit as , so its integral against diverges. One defensible noninformative choice is the scalar Jeffreys prior for an additive variance component, calculated holding the mean parameters fixed:
Indeed the normal variance score has information . This is a scalar conditional-information prior, not a claim that its product with the flat mean prior is the joint Jeffreys prior. It is bounded near zero and is at infinity. The integrated likelihood is uniformly bounded near zero and is bounded by a constant times at infinity, independently of the mean-shape parameters, since . Thus and a proper external prior on the physically admissible ensure posterior propriety. Proper weak scale priors are another option. If known errors vanish, the boundary argument and appropriate prior need separate reconsideration.
For an efficient parameterization, use with and . A chain on this joint reduced parameter space targets
The factor is the Jacobian determinant. At each iteration propose , with for a fixed nonsingular proposal covariance, and accept with probability . Proposals outside the physical prior domain have target zero. Tune during warmup and freeze it for the retained Random-walk Metropolis algorithm. To obtain samples of the full original parameter vector, independently draw from its retained positive prior and reconstruct ; the resulting vector is . Analytically marginalizing using (iii) would also be valid.
Use several dispersed chains, trace plots, rank-normalized split , and the effective sample size of a Markov chain for each parameter and for . These Markov chain Monte Carlo convergence diagnostics reveal poor mixing and disagreement but do not prove convergence. Estimate the integrated autocorrelation time ; with retained draws, . Default to no thinning of a Markov chain, so the thinning factor is . If storage requires thinning, choose a spacing after inspecting the autocorrelation and verify the retained-chain effective sample size; no finite spacing guarantees independent draws.
Using the retained samples, compute
These are posterior summaries, provided the corresponding moments exist, not uncertainties of the numerical estimates. A proper prior alone does not ensure finite second moments; choose or check the external prior accordingly. The Monte Carlo standard error of the posterior mean is approximately .
Write , and . A flat density in means with respect to ; likewise . The Jacobian determinant must therefore appear when using the original scale variables. The joint hierarchical Bayesian model kernel, with respect to , is
This is a kernel, not a normalized joint probability distribution, because the stated priors are improper. Conditioning on would still require posterior propriety. In fact that check fails here, as shown in part (iii)(c); proper formal full conditionals alone do not repair it. The Gaussian–exponential hierarchical colour model distinguishes the intrinsic normal colour, positive interstellar dust reddening, and measurement noise.