Data processing inequality for relative entropy Created 2026-09-24 Updated 2026-09-24
Passing two probability distributions through the same Markov kernel cannot increase their relative entropy. For a kernel ,
Past exam of the mathematics course of the University of Cambridge 2026 iii Paper 224 1 a Solution Created 2026-09-24 Updated 2026-09-24
Suppose is a Markov chain, so . The chain rule for mutual information givesand alsobecause conditional mutual information is nonnegative. Therefore . Similarly,while , so . These are the two data processing inequalities. In particular, applying any deterministic function or Markov kernel to either argument cannot increase mutual information.
Past exam of the mathematics course of the University of Cambridge 2026 iii Paper 224 3 b Solution Created 2026-09-24 Updated 2026-09-24
Let be a Markov kernel, and let , be the output probability distributions. Applying the log-sum inequality for each to and givesSumming over and using yields the data processing inequality for relative entropy