OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
Abstract
In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. We present OmniShow, the first all-in-one model tailored for this practical yet challenging task, capable of harmonizing multimodal conditions and delivering industry-grade performance. To overcome the trade-off between controllability and quality, we introduce Unified Channel-wise Conditioning for efficient image and pose injection, and Gated Local-Context Attention to ensure precise audio-visual synchronization. To effectively address data scarcity, we develop a Decoupled-Then-Joint Training strategy that leverages a multi-stage training process with model merging to efficiently harness heterogeneous sub-task datasets. Furthermore, to fill the evaluation gap in this field, we establish HOIVG-Bench, a dedicated and comprehensive benchmark for HOIVG. Extensive experiments demonstrate that OmniShow achieves overall state-of-the-art performance across various multimodal conditioning settings, setting a solid standard for the emerging HOIVG task.
Lay Summary
Creating realistic videos of people interacting with objects is difficult because many pieces of information must work together, such as text instructions, character appearance, body motion, and audio. Existing video generation systems usually support only part of this information at a time, which limits their usefulness in real applications like digital content creation, avatar animation, and video editing. We introduce OmniShow, a unified AI system that can generate human-object interaction videos while taking all of these conditions together. Our method is designed to keep the character and object appearance consistent, follow the desired motion, and stay synchronized with audio. To make this possible, we also develop a training strategy that learns from several different types of data instead of requiring one perfect dataset for every condition. In addition, we build HOIVG-Bench, a benchmark for evaluating this task in a more complete way. Experiments show that OmniShow performs better than strong existing methods across multiple settings. We hope this work helps make controllable video generation more practical and reliable for real-world creative applications.