AQUA: Aligned Query Fusion for Reference-Unbiased and Temporally Consistent Video Motion Transfer
Abstract
Diffusion models have enabled motion transfer that reflects motion from a reference video while aligning with a given target prompt. While prior methods often require costly model training or fine-tuning, training-free alternatives utilizing self-attention query features have emerged as a flexible solution. However, direct use of reference query features often generates videos with unwanted visual details from the reference and temporal inconsistency. In this paper, we propose AQUA, an optimization-free framework that modulates query features to address these challenges. Specifically, AQUA adaptively balances reference and target query features to remain faithful to the target prompt. Furthermore, AQUA employs a multi-frame guidance to ensure temporal consistency across the generated video.