Taiji Suzuki: Statistical optimality theory and adaptation method of diffusion models
Abstract
This talk explores recent theoretical and methodological advancements in diffusion models. On the theoretical side, we demonstrate that diffusion models can circumvent the curse of dimensionality by discovering intrinsic low-dimensional structures within data distributions. We discuss this capability for both continuous and discrete variables and establish their theoretical optimality, showing that these advantages are driven by their feature learning abilities.
Methodologically, we introduce novel inference-time-adaptation/post-training techniques for discrete diffusion models. Leveraging Doob's h-transform as our primary technical tool, we demonstrate that integrating this transform with the efficient sampling capabilities of diffusion models facilitates effective inference time adaptation without explicit parameter updates. Specifically, our approach enables reinforcement learning on the unmasking order distribution without requiring model updates, yielding substantial performance improvements, and we show the unmasking order adaptation has a significant impact on the performance improvement.