Measuring and Reducing Train--Inference Mismatch in Discrete Diffusion Language Models
Julien Coquet ⋅ Tiago Pimentel ⋅ Dimitri von Rütte ⋅ Yuhui Ding ⋅ Thomas Hofmann
Abstract
Discrete diffusion language models are trained on states sampled from a forward corruption process, but they generate by following states induced by a reverse sampler driven by the learned denoiser. This creates a train–inference mismatch: the denoiser may be queried on states unlike those seen during training. We formalize this mismatch as the Mirror Gap, a time-indexed discrepancy between the forward marginal at each denoising time and the sampler-induced reverse marginal. We estimate projected versions of this gap using lightweight classifiers on frozen denoiser hidden states. The resulting signal is both predictive and actionable: early scores predict final sample quality, and using the same signal online substantially improves the quality–compute frontier. Notably, once the sampling trajectory is controlled directly, common token-level heuristics such as top-$p$ sampling can become unnecessary or even counterproductive. These findings recast diffusion text degeneration as trajectory-level drift, and show that this drift can be measured and reduced.
Chat is not available.
Successful Page Load