LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
Abstract
Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present LoCoT2V-Bench, a benchmark for long video generation (LVG) featuring multi-scene prompts with hierarchical metadata (e.g., character settings and camera behaviors), constructed from collected real-world videos. We further propose LoCoT2V-Eval, a multi-dimensional framework covering perceptual quality, text-video alignment, temporal quality, dynamic quality, and Human Expectation Realization Degree (HERD), with an emphasis on aspects such as fine-grained text-video alignment and temporal character consistency. Experiments on 17 representative LVG models reveal pronounced capability disparities across evaluation dimensions, with strong perceptual quality and background consistency but markedly weaker fine-grained text-video alignment and character consistency. These findings suggest that improving prompt faithfulness and identity preservation remains a key challenge for long-form video generation. Our code and data are released at https://github.com/XqZeppelinhead0702/LoCoT2V-Bench.
Lay Summary
While AI can generate impressive short videos, it struggles with the complex, multi-scene scripts needed for professional filmmaking, and researchers lack tools to accurately evaluate these longer videos. To solve this, we created LoCoT2V-Bench, a new testing framework featuring highly detailed text prompts based on real-world videos. We also developed an automated evaluator that acts like a film critic, assessing visual quality, narrative flow, and whether characters remain consistent across different scenes. Testing 17 popular AI video models revealed a major flaw: while they create stunning backgrounds, they consistently fail to maintain exact character details and follow precise instructions over time. Our benchmark highlights the gap between casual AI video and professional needs, providing a clear roadmap to help researchers build more reliable, controllable AI tools for storytellers.