Why DDIM Hallucinates More Than DDPM: A Theoretical Analysis of Reverse Dynamics
Abstract
Lay Summary
Diffusion models are state-of-the-art AI image and video generators, owing to their impressive synthesized samples. Yet, hallucinations, generations that violate structural or semantic constraints, remain a critical problem for these models. For example, a diffusion model may accidentally produce a hand with six fingers, even though a hand with five was requested. We describe when and where hallucinations arise in diffusion models in a simplified setting (i.e., mixture of Gaussians, which is a standard theoretical setting). Specifically, we focus on the types of hallucinations that emerge when the sample lands in between two modes of the distribution. We also examine why introducing randomness in diffusion models helps reduce hallucination rate. This will help to design better, safer diffusion models for more reliable generative AI.