Sharp Inequalities between Total Variation and Hellinger Distances for Gaussian Mixtures
Abstract
Lay Summary
Many machine learning methods rely on comparing probability distributions, but different ways of measuring how different two distributions are can behave very differently. In this paper, we study two widely used measures — total variation distance and Hellinger distance — for Gaussian mixture models, which are fundamental tools in statistics, clustering, and Bayesian learning. We prove a sharp mathematical relationship between these two distances. We show that if two Gaussian mixtures are close in total variation distance, then they must also be close in Hellinger distance, up to a correction term that is essentially the best possible. This resolves an open problem posed in earlier work. Beyond the theoretical result itself, our work has practical implications for robust machine learning. In settings where data may contain corrupted observations or outliers, our inequality leads to optimal guarantees for estimating Gaussian mixtures under contamination. These results also improve the theoretical understanding of empirical Bayes methods, which are widely used for large-scale statistical inference and denoising problems. Overall, the paper strengthens the theoretical foundations of robust learning and statistical estimation for Gaussian mixture models.