Temporal Backtracking Search for Test-time Generative Video Reasoning
Sejoon Jun ⋅ Zheng Ding ⋅ Huangyuan Su ⋅ Weirui Ye ⋅ Yilun Du
Abstract
While test-time scaling has revolutionized reasoning in large language models, generative video reasoning remains bottlenecked by a single-shot paradigm. We demonstrate that searching over denoising steps cannot rescue logically flawed rollouts because spatial trajectories commit early in the diffusion process. Root-level Best-of-$N$ (BoN) sampling is similarly inefficient: reasoning errors cluster early in the temporal axis, and resampling blindly discards verified upstream progress. To unlock effective test-time scaling for video models, we introduce \textbf{Temporal Backtracking Search (TBS)}, which shifts the search space to the temporal axis. TBS transforms video generation into an iterative generate--verify--restart loop via three core mechanisms: (1) \textit{variable-$K$ conditioning} to resume generation from arbitrary clean prefixes; (2) \textit{temporal process verification} to localize failures and extract valid restart anchors; and (3) \textit{prefix-based search} to reallocate compute toward extending correct trajectories rather than root resampling. Across algorithmic, navigation, and robotics domains, TBS Pareto-dominates matched-budget BoN. In a strict out-of-distribution setting where one-shot generation collapses ($0.7\%$ for BoN), TBS achieves $22.7\%$, with every solved episode stemming from a restarted branch. Ultimately, TBS reveals that the local reasoning competence of video models far exceeds what single-shot rollouts indicate, providing a scalable test-time framework to unlock it.
Chat is not available.
Successful Page Load