A convenient general-state-space ergodic theorem is the following. For a positive Harris recurrent Markov chain with invariant probability measure , and a measurable function with ,
Here is the stationary expected value. Harris recurrence means that every set of positive irreducibility measure is visited almost surely from every state; positive recurrence supplies an invariant probability rather than only an infinite invariant measure. The ergodic theorem for a positive Harris recurrent Markov chain holds from any starting state under these Harris hypotheses. aperiodicity is not necessary just for averages.
A useful sufficient form of the central limit theorem for a geometrically ergodic Markov chain adds aperiodicity, geometric ergodicity, and for some . It gives
These are sufficient hypotheses, not a claim that irreducibility alone ensures a central limit theorem. The Markov chain Monte Carlo asymptotic variance uses stationary covariances , with . The series is absolutely convergent under the stated sufficient assumptions. If , write and , where the integrated autocorrelation time is . When , the effective sample size of a Markov chain is approximately . Negative correlations can reduce the asymptotic variance; a zero asymptotic variance gives a degenerate normal limit. For constant , variance is zero and the autocorrelation normalization is undefined.
Apply the Monte Carlo estimator to the observable :
The theoretical justification is the ergodic theorem for a positive Harris recurrent Markov chain, not an independent-sample law of large numbers. The Markov chain has finite state space, is irreducible, and has the strictly positive posterior distribution as invariant law, so it is positive recurrent and Harris recurrent with counting measure. The observable is bounded: . Thus the estimator converges almost surely to the posterior expected value from any initial configuration.
Each state also has positive self-transition probability, so the chain is an aperiodic Markov chain. Finiteness then gives geometric convergence and the relevant central limit theorem; serial dependence determines the Markov chain Monte Carlo asymptotic variance if error bars are required. A fixed finite burn-in can be discarded without changing consistency, but independence of the retained states should not be assumed.
Use the ergodic theorem for a positive Harris recurrent Markov chain: for an invariant distribution and an integrable function , the time average converges almost surely to under the usual ergodicity conditions. This is the relevant theorem for correlated Metropolis–Hastings algorithm output, rather than the strong law of large numbers. Here is integrable. Indeed, putting
we have , which bounds the unnormalized density by a constant times an integrable Gaussian density and also gives finite first moments.
For a proposal scale , initialize anywhere in . At each step draw independent random variables with the standard normal distribution and an independent uniform , propose , and accept when
Otherwise set . The normal proposal is symmetric, so its density cancels; the target normalizing constant also cancels. Include every state in the average, including repeated states after rejection. The ergodic theorem for a positive Harris recurrent Markov chain then gives
A fixed discarded initial segment does not change this limit.
Small proposal variance gives high acceptance but tiny moves and strong serial dependence. Large variance gives more ambitious moves, but many proposals enter very low-density regions and are rejected, creating long runs at one state. An intermediate scale should be assessed by exploration and effective sample size of a Markov chain, not acceptance rate alone.
This density has two modes, at and their interchange. To see this, stationary points satisfy , implying either or . The unequal solutions have and are local minima of ; the equal solution is a saddle. Mode switching is therefore a material part of proposal-scale selection: a chain confined to one mode can have high acceptance and a misleading finite-run estimate.
Figure 1.
Two modes of the quartic target density, with the diagonal saddle between them, explaining slow mode switching in random-walk Metropolis-Hastings
.
A Harris recurrent chain visits every measurable set of positive irreducibility measure almost surely. Positive Harris recurrence additionally supplies an invariant probability measure. The ergodic theorem for a positive Harris recurrent Markov chain gives consistency of time averages of integrable functions.