Testing Audio Captioning Metrics with Controlled Semantic Perturbations
Abstract
There has been rapid progress in audio captioning models and evaluation metrics in recent years. However, existing captioning metrics are often evaluated only through aggregate benchmark scores, with limited analysis of their robustness to semantic variations, paraphrases, omissions, and hallucinated details. In this work, we introduce StressCaps, a diagnostic challenge set designed to systematically evaluate captioning metrics under controlled semantic perturbations. The benchmark contains both meaning-preserving transformations and semantically corrupted captions. Using StressCaps, we evaluate a broad set of commonly used audio captioning metrics and analyze their strengths and limitations across different perturbation categories. Our experiments show that many metrics remain overly sensitive to surface-level textual changes despite preserving semantic meaning, while semantic similarity metrics such as FENSE and SBERT demonstrate stronger robustness to paraphrasing but remain vulnerable to unsupported or hallucinated additions. These findings highlight significant limitations of current automatic caption evaluation methods and motivate the development of more semantically reliable metrics for long-form and open-ended caption generation. The dataset and StressCaps generation pipeline are available at https://github.com/stresscaps/stresscaps.