Attend to Anything: Foundation Model for Unified Human Attention Modeling
Abstract
Lay Summary
Human attention naturally focuses on the most important regions when people view images and videos or perceive sounds. However, existing attention and saliency models are often limited to specific scenarios or task settings, such as static images, particular types of videos, or individual audio-visual contexts. As a result, they struggle to generalize flexibly in real-world applications. We propose the Attend to Anything Model (AAM), a unified model for understanding attention across images, videos, and audio-visual scenes. AAM formulates attention as a cognitive entailment relationship, enabling the model to reason not only about what is salient, but also about hierarchical task relationships from general to specific through language prompts. To bridge static image attention and dynamic video attention, we further draw inspiration from fluid dynamics and model the temporal evolution of attention in videos as a diffusion process over time. Across 16 public benchmarks, AAM consistently outperforms existing state-of-the-art methods by an average of approximately 6% on diverse attention and saliency tasks, while achieving about a 4× speedup in video inference. These results suggest that AAM can serve as a more general and efficient foundation model for attention modeling, supporting future research on predicting where humans attend across image, video, and audio-visual tasks.