Variational inference approximates a posterior distribution by minimizing a Kullback-Leibler divergence over a chosen family of probability distributions. For an unnormalized posterior density and a trial probability density function , minimizing is equivalent to minimizing .
A reparameterization gradient differentiates an expected value by representing the simulated random variable as a differentiable transformation of parameter-independent noise. For , the identity requires justified differentiation under the integral sign. For a normal distribution, use with drawn from the standard normal distribution.
Markov chain variational inference optimizes an initial parametric probability distribution by minimizing after a fixed number of Markov kernel updates. The best initial probability distribution can depend on ; Kullback-Leibler divergence contraction does not imply that optimization commutes with applying the Markov kernel.
Mean-field variational inference restricts the trial probability density function to . With the other factors fixed, , provided this expression has a finite positive normalization constant. Subtracting the objective at leaves , proving the update is optimal when the terms are well defined.

Articles by others on the same topic (0)

There are currently no matching articles.