AVENUE: Audio-Video EditiNg Understanding and Evaluation
Abstract
Audio-video (AV) editing requires models to infer a modality-selective edit scope from the prompt alone—determining not only what should change, but also which modality should be preserved. However, existing benchmarks provide limited edit-type and modality coverage, while current evaluation systems are modality-blind and sample-agnostic. We introduce AVENUE (Audio-Video EditiNg Understanding and Evaluation), comprising: (1) a benchmark of 1,291 source clips and 7,957 editing instructions across audio-targeted, video-targeted, and AV-coupled edit types, curated and human-verified from VGGSound; and (2) a sample-specific, modality-aware evaluation framework. Our findings reveal that when editing one modality, existing models frequently induce unintended changes in the other, regardless of paradigm.