VIBE: Disentangling Social Dynamics via Kinematics-Informed Variational Inference for Behavioral Emotion
Abstract
Group Emotion Recognition (GER) is crucial for understanding social dynamics, ranging from interpreting intimate conversations to evaluating crowd behavior in large-scale surveillance scenarios. While current AI models can analyze these scenes, they often act as black boxes that take shortcuts. Instead of focusing on how people are actually behaving, these models often get distracted by the background environment, leading to inaccurate results. To bridge this gap, we introduce VIBE (Variational Inference for Behavioral Emotion), a kinematics-aware framework that integrates audio, video, and text modalities through causal structuring. Unlike standard models that simply mix data together, VIBE utilizes mathematical constraints to filter out background noise and isolate the genuine emotions of the people involved. This purified representation enables our model to focus exclusively on the sociological mechanics of the crowd, dynamically modulating neural attention based on raw physical synchrony. Simultaneously, we align visual dynamics with human interpretability by projecting latent representations into a semantically structured space informed by textual descriptions. Comprehensive experiments demonstrate that VIBE consistently outperforms state-of-the-art methods. Code is available at GitHub.
Lay Summary
Understanding a crowd's collective mood is important for applications such as social robotics and large-scale event surveillance. However, current AI models often act as "black boxes" that take hidden shortcuts. Instead of analyzing how people actually behave and interact, these models are often distracted by the background environment, leading to inaccurate and biased predictions. To fix this, we created VIBE, a new AI framework that uses constraints to actively filter out background noise and isolate genuine human emotion. Rather than simply mixing all visual data together, VIBE specifically focuses on physical synchrony, measuring how individuals physically move and coordinate as a group. It also links these visual movements with text descriptions to ensure the model understands the real-world meaning behind the actions. VIBE consistently outperforms existing state-of-the-art methods in recognizing group emotions. By forcing the model to base its decisions on authentic human interactions rather than environmental distractions, this research paves the way for fairer, more transparent, and highly reliable tools for understanding real-world social dynamics.