For a convex function satisfying , the f-divergence between probability distributions and is
where , , and the integrand uses the lower-semicontinuous perspective convention at .
For probability densities and , the Kullback-Leibler divergence from to is
For probability distributions and on the same finite set,
with equality exactly when on the support of .
For probability distributions and on a product space, the Kullback-Leibler divergence decomposes into the divergence of a marginal and the expected divergence of the corresponding conditional distributions. Iteration gives
Passing two probability distributions through the same Markov kernel cannot increase their relative entropy. For a kernel ,
For nonnegative , with and , the log-sum inequality is
It proves both Gibbs inequality and the data processing inequality for relative entropy.
For a probability mass function of full support on a finite alphabet and a real function ,
where uses natural logarithms. The unique maximizing probability mass function is the exponential tilt .
For univariate normal distributions,
For probability mass functions with absolutely continuous with respect to ,
For a measurable function , every f-divergence contracts under pushforward:
This follows from conditional Jensen inequality applied to the jointly convex perspective .
The squared Hellinger distance is
where and are densities with respect to any common dominating measure .

Articles by others on the same topic (0)

There are currently no matching articles.