Expressively Controllable Talking Face Generation via Identity-Aware AU Conditioning and Neutral-Reference Debiasing
Abstract
Precise control over expressive dynamics in audio-driven talking face generation remains challenging due to a lack of granular control signals. Existing methods typically rely on coarse categorical emotion labels applied across an entire video, resulting in rigid, unchanging facial poses that lack temporal diversity. To address this, we propose IUFace, a framework that utilizes Action Units (AUs) as a continuous, multimodal conditioning signal for anatomically grounded, frame-level expression control. To ensure these movements translate naturally across diverse subjects, we introduce Identity-Aware AU Conditioning to personalize AU activations based on specific facial structures. Furthermore, our Neutral-Reference Debiasing paradigm explicitly decouples static identity from dynamic expression during training, effectively eliminating expression leakage and reference-pose tethering. Extensive evaluations demonstrate that IUFace achieves state-of-the-art performance in fine-grained expression alignment, temporal coherence, and reliable generation.