The likelihood function and prior distribution give the posterior densityCompleting the square in the quadratic form givesThus Gaussian conjugacy for a normal linear model yields
After integrating out the multivariate normal distribution , the marginal distribution isUp to terms independent of , the log-likelihood isDifferentiating and setting the result to zero gives the maximum marginal likelihood estimatorThis is an Empirical Bayes method because the estimated hyperparameter is then inserted into the prior and posterior distributions.
For a decision in the unit ball, conditional expectation under the posterior givesThe symmetric positive semidefinite matrix has a Rayleigh quotient maximized over by any unit eigenvector corresponding to its largest eigenvalue. Bayes decision rule therefore chooses any such eigenvector for the stated utility.
For the first observations, writeBy Gaussian conjugacy for a normal linear model, the prefix posterior isCompute a Cholesky decomposition of once. If is row of the design matrix, thenThe Rank-one Cholesky update obtains a triangular factor from in operations. Two triangular solves give , and, for , another solve givesThe initial factorization costs and all updates and samples cost . This is within the requested bound.
Put . The Weighted graph Laplacian of the tree is defined byMultiplication of the Gaussian likelihood by the prior shows that, conditionally on the precision parameter ,Completing the square therefore givesAs a function of , the posterior density isso, in shape-rate notation,The precision matrix has the sparsity pattern of a tree. A sparse Cholesky decomposition and its triangular solves have cost and storage on this graph, while the gamma update also costs . Hence each systematic-scan Gibbs sampler iteration costs .
A Markov kernel with stationary distribution is geometrically ergodic if there are and a finite function such thatfor every and almost every starting point , where is total variation distance.
A measurable set is a small set with minorisation constant if some integer and some probability measure satisfyfor every and every measurable .
A standard drift-minorisation condition is that the chain be irreducible and aperiodic, and that there exist a measurable , a small set , constants and such thatThese conditions imply geometric ergodicity.
Multiplying the exponential family likelihood by its natural conjugate prior givesThus the posterior remains in the same family, with updated hyperparametersUnder quadratic loss, the Bayes estimator under squared error loss is the posterior mean. Differentiating the log-partition function that normalizes the conjugate prior gives
Let be the importance weight. The two normalized densities giveSince , there is a finite constant such that everywhere. The Independence Metropolis–Hastings algorithm has an accepted transition density satisfyingConsequently the whole state space is a small set, with the one-step minorization condition . Iterating this Doeblin condition gives uniform geometric convergence in total variation distance, so the chain is geometrically ergodic.
Markov chain Monte Carlo asymptotic variance for a stationary Markov chain and iswhenever the limit and series exist. If a reversible Markov chain has positive spectral gap , then the spectral theorem for normal operators on a separable Hilbert space gives
For , detailed balance says that the measure is invariant under exchanging and . ThereforeThis is precisely the defining identity for a self-adjoint operator.
By the stationary distribution property, and have the same marginal distribution . Expanding the square givesThis is the probabilistic representation of the Dirichlet form of a Markov chain.
Subtracting the expected value of does not alter either side, so suppose . Since reversibility makes a self-adjoint operator and stationarity makes it a contraction,The Discrete-time Poincaré inequality for a Markov kernel is therefore equivalent toApplying this inequality successively to yieldsConversely, the asserted variance contraction with rearranges to the Poincaré inequality. Hence the two statements are equivalent.
Part c gives the integral representationThe integrand vanishes on the diagonal . Thus the assumed off-diagonal Peskun ordering impliesfor every .
On the mean-zero subspace, the variational characterization of the spectral gap of a positive reversible kernel isIt follows immediately thatEquivalently, the energy inequality says in the Löwner order on . Positivity permits the operator monotonicity of the square root and hence ; the spectral representations in the question identify the top spectral values and give the same gap inequality.
Hamiltonian Monte Carlo augments the position by an independent momentum and uses the Hamiltonian functionFrom the current , draw a fresh , apply a fixed number of leapfrog steps to approximate Hamiltonian flow, and obtain . The Metropolis–Hastings acceptance probability isotherwise retain . Momentum negation may be included to make the proposal explicitly reversible. The leapfrog map is volume preserving and reversible, while the acceptance step corrects its discretization error, leaving invariant after the momentum is discarded.
For a sufficiently regular Itô diffusion with , a density is stationary if and only if it solves the stationary Fokker-Planck equationwith integrable Fokker-Planck probability current and boundary conditions that make its outward flux vanish. Here .
For the displayed parametrization, assumewhere , that is symmetric positive semidefinite, and that is antisymmetric. Substituting and the stated into the probability current cancels all terms. The remaining divergence isbecause the second derivatives are symmetric in whereas . Thus these conditions, together with the boundary and regularity assumptions, imply stationarity. More generally, the divergence equation above is the exact necessary and sufficient condition; within this construction, and antisymmetric are the standard way to satisfy it.
Invariant distribution of an Itô diffusion, specialized to Underdamped Langevin dynamics, has densityThus has density proportional to , and conditionally and marginally . The Hamiltonian transport between and preserves this density, while the Ornstein-Uhlenbeck process in momentum has exactly that Gaussian invariant law.
Articles by others on the same topic
There are currently no matching articles.