Scalable Robot Policy Evaluation via Autoregressive Video World Models
Byeongguk Jeon ⋅ Seonghyeon Ye ⋅ JaeHyeok Doo ⋅ Sungdong Kim ⋅ Minjoon Seo ⋅ Hyungmok Son ⋅ Kimin Lee
Abstract
Video world models offer a scalable alternative to real-world and simulation-based robot policy evaluation, serving as neural simulators. However, video world models suffer from compounding errors over long rollouts and slow inference. To address this gap, we propose WorldArena, a scalable evaluation pipeline that combines a few-step autoregressive video world model with a rubric-guided vision-language model (VLM) judge to prevent world-model errors from propagating into evaluation outcomes. We introduce Step Forcing, which closes the train--test gap by matching noise schedules, contexts, and priors between training and inference. Step Forcing enables stable long-horizon autoregressive rollouts without teacher distillation or costly self-rollouts at training time. We evaluate eight generalist robot policies on WorldArena over 4,188 rollouts, achieving a Pearson correlation of $r = 0.989$ and a Spearman correlation of $\rho = 0.970$ with the real-world RoboArena leaderboard.
Chat is not available.
Successful Page Load