Normalizing constant 2026-10-06
A normalizing constant makes a nonnegative integrable function into a probability density function: if , then is a probability density function.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 208 6 a Solution Created 2026-10-03 Updated 2026-10-06
Let be a nonnegative target kernel and a proposal distribution. At state , propose , draw an independent uniform , and move to ifOtherwise keep . This is the Metropolis–Hastings algorithm. Start on the positive target support, with the usual zero-ratio conventions. The unknown normalizing constant cancels.
An unnormalized kernel is sufficient, provided . A genuinely improper target with infinite integral is not a probability density and cannot supply a stationary target probability distribution. The acceptance formula can still be written formally, but cannot be said to generate samples from that nonexistent probability law.
Because rejected proposals leave the state unchanged, the Markov kernel includes an atom:Detailed balance with a probability measure means the measure identityWhere both sides have ordinary densities, this reads . The measure formulation also includes the rejection atom. Integrating it shows that is invariant.
Past exam of the mathematics course of the University of Cambridge 2016 iii Paper 208 6 d Solution Created 2026-10-03 Updated 2026-10-06
Use the ergodic theorem for a positive Harris recurrent Markov chain: for an invariant distribution and an integrable function , the time average converges almost surely to under the usual ergodicity conditions. This is the relevant theorem for correlated Metropolis–Hastings algorithm output, rather than the strong law of large numbers. Here is integrable. Indeed, puttingwe have , which bounds the unnormalized density by a constant times an integrable Gaussian density and also gives finite first moments.
For a proposal scale , initialize anywhere in . At each step draw independent random variables with the standard normal distribution and an independent uniform , propose , and accept whenOtherwise set . The normal proposal is symmetric, so its density cancels; the target normalizing constant also cancels. Include every state in the average, including repeated states after rejection. The ergodic theorem for a positive Harris recurrent Markov chain then givesA fixed discarded initial segment does not change this limit.
Small proposal variance gives high acceptance but tiny moves and strong serial dependence. Large variance gives more ambitious moves, but many proposals enter very low-density regions and are rejected, creating long runs at one state. An intermediate scale should be assessed by exploration and effective sample size of a Markov chain, not acceptance rate alone.
This density has two modes, at and their interchange. To see this, stationary points satisfy , implying either or . The unequal solutions have and are local minima of ; the equal solution is a saddle. Mode switching is therefore a material part of proposal-scale selection: a chain confined to one mode can have high acceptance and a misleading finite-run estimate.
Two modes of the quartic target density, with the diagonal saddle between them, explaining slow mode switching in random-walk Metropolis-Hastings
. Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 216 4 a Solution Created 2026-10-03 Updated 2026-10-06
Write . Mean-field variational inference minimizes the reverse Kullback-Leibler divergenceEach is a probability density function; without further parametric restrictions, the optimization is over all such product probability distributions. Equivalently it maximizes the evidence lower bound for any unnormalized posterior density . Its difference from is exactly the Kullback-Leibler divergence.
Fix and writeAssume and that the displayed expected values and objective decomposition are well defined. The optimal factor isIndeed, the part of the objective depending on isThe remaining term is fixed. Gibbs inequality gives nonnegativity of the Kullback-Leibler divergence, with equality precisely at almost everywhere. This proves global optimality of that coordinate update; it does not assert global optimality of a sequence of coordinate updates for the full nonconvex product-family problem.
The printed upper index in the list of remaining coordinates is inconsistent with the dimension ; the natural interpretation is . Also, if the exponential expression has zero or infinite normalizing constant, or the expected values are undefined, the usual coordinate formula needs additional support or integrability hypotheses. A restricted parametric factor family need not contain this unrestricted optimal factor.
Past exam of the mathematics course of the University of Cambridge 2017 iii Paper 216 6 Solution Created 2026-10-03 Updated 2026-10-06
Use the extended target probability density functionThe desired stationary distribution for positions is ; the full position-momentum joint probability distribution is , not alone. Gaussian momentum refreshment preserves , since it redraws its momentum marginal distribution independently while keeping the position fixed.
Let and be the surrogate leapfrog integration map. The proposal in the paper is : its notation includes the final momentum flip. Reversibility gives , so . Both and preserve volume. Thus is an involutive Metropolis proposal, and the required acceptance probability isEquivalently, in terms of the surrogate Hamiltonian ,Only evaluations of the true target probability density function are needed at endpoints; the trajectory uses the surrogate gradient. Target and surrogate normalizing constants cancel.
For , the accepted flux satisfiesChanging variables has unit absolute Jacobian determinant, so this identity proves detailed balance for accepted moves. The rejection mass stays at the same point and is also reversible. The accept/reject step therefore preserves . Since momentum refreshment also preserves , their composition and either phase of the alternating process preserve . Projecting onto positions proves the position stationary distribution is exactly the original target. Stationarity alone does not establish uniqueness or convergence from every start; those require additional irreducible Markov chain and aperiodic Markov chain hypotheses.
There is a genuine defect in the printed smoothing claim. For the positive, nondifferentiable Laplace distribution on the real line, direct minimization gives, for every ,For the minimizer is ; for it is ; at zero both signs minimize. Hence normalization leaves , still nondifferentiable at zero. The displayed construction does not in general produce the asserted smooth surrogate. For example, the also nondifferentiable target makes that infimum for every , since the negative quartic term dominates the quadratic penalty. The invariance proof above is valid conditional on actually having a usable smooth surrogate, as the subsequent algorithm assumes.
A corrected sufficient construction is Moreau smoothing of a negative log-density. For a proper convex function with sequential lower semicontinuity , setThe Moreau envelope is differentiable, with ; require to be integrable to normalize it. The signs differ from the printed formula. For the Laplace distribution, this givesa genuinely differentiable potential with integrable exponential tails. Using that surrogate with the boxed acceptance probability still targets the original Laplace distribution.
