For nonnegative , with and , the log-sum inequality is
with the usual extended-value conventions. Equality holds when is constant wherever .
Let and put and . For each alphabet symbol , apply the log-sum inequality to , and the corresponding values. Summing over gives
which is joint convexity in .
For nonnegative numbers , put and . The log-sum inequality is
with and the usual extended-value convention when a denominator vanishes. Equality holds precisely when is constant over the indices with .
Let be a Markov kernel, and let , be the output probability distributions. Applying the log-sum inequality for each to and gives
Summing over and using yields the data processing inequality for relative entropy