Task-Dependent Inference-Compute Scaling Frontiers: Diffusion vs. Autoregressive Language Models
Khurram Khalil ⋅ Ripan K Kundu
Abstract
Scaling test-time compute is a critical driver of language model performance, yet current empirical scaling laws primarily characterize discrete autoregressive (AR) architectures. Diffusion language models (DLMs) construct outputs via high-dimensional, continuous iterative denoising, presenting a fundamentally different generation dynamic. However, it is unknown how their quality-vs.-compute scaling frontiers compare to AR models at matched inference FLOPs. We systematically benchmark inference-compute scaling curves of AR (Qwen2.5) and DLMs (Dream, LLaDA) across sequential reasoning and constraint-satisfaction workloads, moving beyond parameter-count comparisons to rigorously track per-candidate FLOPs. Our results reveal a sharp, task-dependent double dissociation. On constraint-satisfaction tasks (evaluated via Pass@$N$), AR generation quickly saturates at $\sim$31\% accuracy, whereas DLMs' holistic state-space traversal avoids early-commitment traps, scaling smoothly to 98.7\%. Conversely, on sequential chain-of-thought reasoning (GSM8K via majority vote), AR scales highly efficiently to 97.1\%, maintaining $>$10$\times$ per-FLOP efficiency advantage over DLMs. % Broader Impact These findings demonstrate that test-time scaling limits are inherently dictated by interplay between task structure and high-dimensional generation dynamics. This provides first unified scaling perspective across AR and DLM paradigms and motivates task-aware inference routing.
Chat is not available.
Successful Page Load