From Synopsis to Storyboard: Enhancing Prompt Expressiveness for Multi-shot Video Generation
Abstract
AI-generated content (AIGC) has enabled highquality video generation from textual prompts. While modern video models possess strong inherent controllability, their performance is often limited by the quality of user-provided prompts. Professional creators can naturally embed cinematic information, such as shot scale, composition, character, and layout, whereas amateur users typically provide vague story descriptions, resulting in inconsistent and incoherent videos. To address this, we propose QwenCine, a large language model fine-tuned from Qwen-based via supervised learning to automatically convert simple story synopses into structured, storyboard-level prompts. QwenCine produces model-readable shot description captions and infers shot scale, viewpoint, spatial layout, and key visual details, enabling users without filmmaking expertise to generate near-professional videos. We further construct CineFlow-Storyboard, a high-quality movie storyboard multi-shot description dataset derived from an open source dataset, containing story synopses, keyframe sequences, and shot-level prompts. Experiments demonstrate that QwenCine significantly improves shot accuracy, narrative coherence, and video controllability, effectively bridging the gap between amateur and professional video creation.