Is your Flow Matching Model Really Generalising? A Path-Length Diagnostic
Abstract
Flow-matching models learn a time-dependent velocity field that transports a Gaussian source to the data distribution along ODE trajectories. We propose a simple geometric diagnostic for assessing whether such models have truly learnt to generate from the data manifold: the path length, defined as the integrated velocity norm along a trajectory. Studying three flow-matching models trained on MNIST under different computational budgets, we find a striking pattern. Path-length distributions on training and test data, computed via the reverse ODE, are nearly indistinguishable and shrink monotonically as training progresses. Path-length distributions on freshly generated samples, by contrast, remain essentially constant across training budgets. As a result, three regimes emerge: at high FID (approximately 130) generated paths are shorter than data paths, at moderate FID (approximately 13) the two distributions match, and at low FID (approximately 8) generated paths exceed data paths. This pattern is not explained by memorisation, nor by numerical integration error, nor by velocity-magnitude artefacts at the trajectory endpoints. We interpret matching path-length distributions as a signal that generation has landed on the data manifold, and we relate the observed asymmetry to recent results on the implicit regularisation and trajectory stability of flow matching. Path length thus provides a cheap, model-internal probe that complements FID and memorisation metrics.