Stable Velocity: A Variance Perspective on Flow Matching
Abstract
Lay Summary
Modern generative models can create high-quality images and videos, but training them efficiently remains challenging. One reason is that the learning signals they rely on can be very noisy, especially in the early stages of generation, which makes training unstable and slow. In this work, we analyze where this noise comes from and show that it is much higher when the model starts from random inputs, but becomes much lower as the model gets closer to real data. Based on this observation, we propose a new method called Stable Velocity that improves both training and generation. During training, our method reduces noise in the learning signals and focuses more on the parts of the process that are easier to learn. During generation, we take advantage of simpler dynamics in the low-noise region to speed up sampling without additional training. Our approach consistently makes training more efficient and can generate images and videos more than twice as fast, while maintaining the same quality.