EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition
Abstract
Dynamic Facial Expression Recognition (DFER) models the temporal evolution of facial expressions in videos. In real-world scenarios, changing scenarios distort expression trajectories, challenging existing methods. Most current approaches address this via passive feature alignment or domain-incremental learning but do not explicitly model scenario evolution, limiting their ability to capture expression dynamics under scenario-incremental changes. To address this, we propose EmWorld, an emotion world model for DFER that explicitly models latent emotion state evolution under scenario variations. Specifically, EmWorld formulates scenario-incremental DFER as a progressive Bayesian inference problem over latent world states with dual temporal scales. Slow-timescale component (STS) models scenario evolution using stochastic evolutionary priors, capturing long-term scenario effects and providing proactive guidance in new scenarios. Fast-timescale component (FTS) models frame-level expression dynamics with temporally consistent latent transitions, decoupling expression dynamics from scenario influences. By jointly inferring latent states at both timescales, EmWorld shifts DFER from a passive feature discrimination to active probabilistic state inference under evolving scenarios. Experiments on FERV39k, DFEW, and MAFW demonstrate that EmWorld consistently outperforms state-of-the-art methods, achieving up to 3.84\% improvement while exhibiting strong cross-scenario stability and long-term robustness.
Lay Summary
We teach computers to recognize emotions from people’s facial expressions in videos, but in the real world, changing scenarios can distort how expressions evolve over time, making it hard for existing methods to remain accurate. To address this, we developed EmWorld, an emotion world model. It treats emotions as “world states” that evolve over time and captures expression dynamics on two timescales: a slow timescale that learns long-term scenario changes to help the model adapt to new environments, and a fast timescale that tracks frame-by-frame expression changes to ensure consistent recognition. By combining these two perspectives, the model can separate actual facial expressions from distracting scenario influences, making emotion recognition more robust. Experiments on multiple large-scale facial expression video datasets show that EmWorld outperforms previous methods, achieving up to 3.84% higher accuracy while maintaining stability across different scenarios. This approach provides a reliable new way for computers to understand emotions in real-world conditions.