Fantastic KL Divergence and How to (Actually) Compute It
Jia-Bin Huang · 11:45
KL divergence measures the *extra* bits you waste when you encode data drawn from a true distribution P using a code optimized for a wrong distribution Q — and because computing it exactly is infeasible for large voca...