For nonnegative numbers , put and . The log-sum inequality iswith and the usual extended-value convention when a denominator vanishes. Equality holds precisely when is constant over the indices with .
Let be a Markov kernel, and let , be the output probability distributions. Applying the log-sum inequality for each to and givesSumming over and using yields the data processing inequality for relative entropy
The elementary logarithm inequality givesSince ,This proves , relating relative entropy to chi-squared divergence.
Writing for the laws of , respectively, and using Jensen inequality for the concave natural logarithm,This is the lower-bound half of the Gibbs variational principle for relative entropy.
Let and define the exponentially tilted probability mass functionFor this choice every ratio equals , so equality holds in the Jensen inequality used in part (d). Henceand is the maximizer. This is the finite-alphabet Gibbs variational principle for relative entropy.
Articles by others on the same topic
There are currently no matching articles.