T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
Abstract
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts. To address this limitation, we present T2AV-Compass, a unified benchmark for comprehensive evaluation of T2AV systems, consisting of 500 diverse and complex prompts constructed via a taxonomy-driven pipeline to ensure semantic richness and physical plausibility. Besides, T2AV-Compass introduces a dual-level evaluation framework that integrates objective signal-level metrics for video quality, audio quality, and cross-modal alignment with a subjective MLLM-as-a-Judge protocol for instruction following and realism assessment. Extensive evaluation of 15 representative T2AV systems reveals that even the strongest models fall substantially short of human-level realism and cross-modal consistency, with persistent failures in audio realism, fine-grained synchronization, instruction following, etc. These results indicate significant improvement room for future models and highlight the value of T2AV-Compass as a challenging and diagnostic testbed for advancing text-to-audio-video generation.
Lay Summary
AI systems can now generate videos with sound from written prompts, but it is hard to tell whether the sound truly matches the scene and whether the result follows the user's request. We introduce T2AV-Compass, a test suite with 500 detailed prompts for evaluating text-to-audio-video generation. It checks visual quality, audio quality, audio-video matching, timing, realism, and instruction following. We use T2AV-Compass to compare 15 current systems. The results show that even strong models still struggle with realistic audio, precise synchronization, and complex multi-event prompts. We hope this benchmark helps researchers find clear weaknesses and build more faithful and reliable audio-video generation systems.