AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
Abstract
Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity, failing to capture fine-grained joint correctness required by realistic prompts. We introduce AVGen-Bench, a task-driven benchmark for T2AV generation, featuring high-quality prompts across 11 real-world categories. To support comprehensive assessment, we propose a multi-granular evaluation framework that combines lightweight specialist models with Multimodal Large Language Models (MLLMs), enabling evaluation from perceptual quality to fine-grained semantic controllability. Our evaluation reveals a pronounced gap between strong audio-visual aesthetics and weak semantic reliability, including persistent failures in text rendering, speech coherence, physical reasoning, and universal breakdown in musical pitch control.
Lay Summary
Text-to-video systems are starting to generate videos with sound, including speech, music, and environmental audio. These systems often look and sound impressive at first glance, but they can still fail in ways that matter for real use: text in the video may be unreadable, speech may not match the scene, music may play the wrong notes, faces may change across shots, or physical events may not make sense. This paper introduces AVGen-Bench, a benchmark for testing whether audio-video generation systems can follow detailed user instructions. The benchmark contains realistic prompts from areas such as advertisements, movie trailers, tutorials, sports, and physical or chemical demonstrations. It evaluates not only general visual and audio quality, but also whether the generated video satisfies specific details in the prompt. Our results show that current systems are much better at producing attractive videos than at reliably following fine-grained instructions. We hope this benchmark helps researchers and developers identify these weaknesses and build more controllable, reliable, and responsible audio-video generation systems.