Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
Abstract
What if a world simulation model could render not an imagined environment but a city that actually exists? Prior generative world models synthesize visually plausible yet artificial environments by imagining all content. We present Seoul World Model (SWM), a city-scale world model grounded in the real city of Seoul. SWM anchors autoregressive video generation through retrieval-augmented conditioning on nearby street-view images. However, this design introduces challenges including temporal misalignment between retrieved references and the dynamic target scene, limited trajectory diversity and data sparsity from vehicle-mounted captures, and long-horizon error accumulation. We address these challenges through cross-temporal pairing, a synthetic urban dataset and view interpolation pipeline, and a Virtual Lookahead Sink that continuously re-grounds each chunk to a retrieved future reference. SWM outperforms recent video world models across Seoul, Busan, and Ann Arbor while supporting diverse camera movements and text-prompted scenario variations.