For a convex function satisfying , the f-divergence between probability distributions and iswhere , , and the integrand uses the lower-semicontinuous perspective convention at .
For probability distributions and on the same finite set,with equality exactly when on the support of .
For probability distributions and on a product space, the Kullback-Leibler divergence decomposes into the divergence of a marginal and the expected divergence of the corresponding conditional distributions. Iteration gives
Passing two probability distributions through the same Markov kernel cannot increase their relative entropy. For a kernel ,
For nonnegative , with and , the log-sum inequality isIt proves both Gibbs inequality and the data processing inequality for relative entropy.
For a probability mass function of full support on a finite alphabet and a real function ,where uses natural logarithms. The unique maximizing probability mass function is the exponential tilt .
For a measurable function , every f-divergence contracts under pushforward:This follows from conditional Jensen inequality applied to the jointly convex perspective .
The squared Hellinger distance iswhere and are densities with respect to any common dominating measure .
Articles by others on the same topic
There are currently no matching articles.