HECTOR: Hybrid Editable Compositional Object References for Video Generation
Abstract
Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holistically and therefore lack mechanisms for explicit compositional manipulation. To address this limitation, we propose HECTOR, a generative pipeline that enables fine-grained compositional control. In contrast to prior methods, HECTOR supports hybrid reference conditioning, allowing generation to be simultaneously guided by static images and/or dynamic videos. Moreover, users can explicitly specify the trajectory of each referenced element, precisely controlling its location, scale, and speed (see Figure1). This design allows the model to synthesize coherent videos that satisfy complex spatiotemporal constraints while preserving high-fidelity adherence to references. Extensive experiments demonstrate that HECTOR achieves superior visual quality, stronger reference preservation, and improved motion controllability compared with existing approaches.
Lay Summary
Most current AI video generation models synthesize scenes holistically, preventing users from precisely controlling the motion and interactions of individual objects. We introduce HECTOR, a compositional framework that allows users to guide video generation using a mix of static images and dynamic video references. By explicitly specifying exact trajectories for each element, users can independently control an object's location, scale, and speed. This approach synthesizes coherent videos that satisfy complex spatial and temporal constraints while preserving high-fidelity reference details. Ultimately, HECTOR unlocks powerful editing capabilities like precise object replacement, addition, and background modification for advanced creative workflows.