STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
Abstract
This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and dynamic viewpoint control, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusion models depend heavily on reference guidance, leading to limited motion diversity. At the same time, 3D-aware animation typically relies on inversion through pretrained tri-plane generators, which often leads to imperfect reconstructions and identity drift. We rethink reference- and geometry-based paradigms in two ways. First, we deviate from strict reference conditioning at pretraining by introducing softer identity constraints. Second, we address 3D awareness implicitly within the 2D video domain by leveraging the inherent multi-view nature of video data. STARCaster adopts a compositional approach progressing from ID-aware motion modeling, to audio-visual synchronization via lip reading-based supervision, and finally to novel view animation through temporal-to-spatial adaptation. To overcome the scarcity of 4D audio-visual data, we propose a decoupled learning approach in which view consistency and temporal coherence are trained independently. Comprehensive evaluations demonstrate that STARCaster generalizes effectively across tasks and identities, consistently surpassing prior approaches in different benchmarks.
Lay Summary
Creating realistic AI-generated videos of talking people is a major challenge in computer science. Current artificial intelligence tools struggle to make a digital character speak realistically while simultaneously shifting the camera angle. Often, these tools either produce rigid, unnatural facial expressions or accidentally change the person's unique facial features entirely when the camera moves. To solve this, we introduced STARCaster, a new AI system that separates the learning of facial identity, speech, and movement. Instead of forcing the AI to strictly copy a single reference image, we allowed it more flexibility to naturally learn how faces move. We also taught the AI to understand 3D space by training it on how real videos naturally change perspective, allowing for seamless camera angle adjustments. Our framework allows users to generate highly realistic, moving videos of a person speaking just from a single picture and a voice clip. This breakthrough has major implications for creating high-quality digital avatars, improving virtual assistant interactions, and making filmmaking and video editing tools much more accessible.