OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
Abstract
While humans perceive the world through diverse modalities that operate synergistically to support a holistic understanding of their surroundings, existing omnivideo models still face substantial challenges on audio-visual understanding tasks. In this paper, we propose OmniVideo-R1, a novel reinforced framework that improves mixed-modality reasoning. OmniVideo-R1 empowers models to "think with omnimodal cues" by two key strategies: (1) query‑intensive grounding based on self‑supervised learning paradigms; and (2) modality‑attentive fusion built upon contrastive learning paradigms. Extensive experiments on multiple benchmarks demonstrate that OmniVideo-R1 consistently outperforms strong baselines, highlighting its effectiveness and robust generalization capabilities.
Lay Summary
While humans naturally combine what they see and hear to understand the real world, current AI models often struggle with this balance. Many models suffer from "modality bias", where they rely too heavily on visual cues and ignore crucial audio information, leading to errors in complex reasoning tasks. To bridge this gap, we developed OmniVideo-R1, a training framework designed to enhance how AI "thinks" across different senses. Our approach introduces two core strategies: first, it encourages the model to identify key audio-visual cues before answering. Second, it uses a specialized learning technique to ensure the model integrates both sound and vision rather than relying on just one. Our results show that OmniVideo-R1 significantly improves performance on challenging benchmarks. By helping AI better synchronize its "eyes" and "ears", this research paves the way for more reliable assistants that can truly understand the multisensory world we live in.