Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses
Abstract
Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing defense strategies struggle against LipSync forgeries, as state-of-the-art generative models not only optimize for the lip synchronization but also significantly eliminate visual artifacts, resulting in the lack of key detection signals. Inspired by the inherent biological coupling between lip movements and head poses in natural speech, we observe that generative models fundamentally disrupt this global coordination when optimizing for local lip motion. In this paper, we propose LipDA, a novel framework for joint LipSync Detection and Attribution, which takes advantage of the inconsistency between head and lip. For detection, the framework learns to quantify this discrepancy by contrasting lip and pose features from authentic versus forged videos. For attribution, our method is designed to capture the unique temporal dynamics and audio-visual synchronization patterns that act as generative fingerprints, enabling source tracing. To validate our approach, we conduct extensive experiments on two challenging LipSync benchmarks as well as on our own proposed large-scale and multi-generator dataset, LipSyncBench-A. LipDA achieves over 97% AUC in detection and 97.5% accuracy in model attribution, significantly outperforming existing methods.
Lay Summary
Modern artificial intelligence systems can now modify a real video so that a person appears to say words they never spoke. Such lip-synchronization forgeries have already been exploited for financial fraud, including a recent incident involving a loss of 25 million dollars, and existing detection methods are becoming increasingly ineffective because the latest generators leave almost no visible artifacts. We observe that authentic human speech exhibits a tight coordination between lip movements and subtle head motion, arising from a shared physiological control system. Current AI generators, in contrast, synthesize lip motion and head pose as two independent components, which inadvertently disrupts this natural coupling. Our framework, LipDA, learns to quantify this discrepancy and uses it as a reliable signal for distinguishing authentic videos from forged ones. We further observe that each family of generators imprints a distinctive temporal pattern on the synthesized motion, which enables LipDA to identify the specific model responsible for a given forgery. Experimental results demonstrate that LipDA detects forgeries that elude existing methods and remains accurate under common video degradations such as blurring and compression, providing a more dependable safeguard against video-based misinformation and fraud.