Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models
Abstract
The integration of audio modality into Large Audio Language Models (LALMs) significantly expands their attack surface. Existing jailbreaks predominantly treat audio as a carrier for malicious payloads, relying on semantic optimization, acoustic parameter control, or additive perturbation to embed harmful content into the audio signal. In this work, we challenge this necessity and propose a new paradigm in which the role of audio shifts from content injection to safety alignment interference. We reveal that LALM safety alignment can be compromised solely by specific Acoustic Latent Semantics (ALS), the underlying paralinguistic features intrinsic to the priors of audio generative models. Distinct from previous works that leverage explicit acoustic parameters to merely style malicious audio, we demonstrate that interference audio, benign in content but infused with specific ALS, can serve as a universal jailbreak trigger. Leveraging this insight, we propose Acoustic Interference Attack (AIA), which decouples the attack payload from the audio. It employs a set of universal, instruction-neutral interference audio, enabling standard malicious text queries to bypass safety alignment without instance-specific optimization. Experiments on 10 LALMs across five datasets demonstrate that AIA achieves the state-of-the-art attack success rate. Furthermore, our interpretability analysis uncovers the inference path drift induced by AIA and identifies the inherent effective patterns within ALS, revealing the fundamental vulnerability of cross-modal alignment in LALMs.
Lay Summary
Advanced AI models processing both text and audio face unique security risks. Existing attacks bypass safety alignment by embedding malicious content into complex audio signals. We discovered a simpler attack called "Acoustic Interference". By pairing harmful text with harmless-sounding audio containing specific natural voice traits, like a distinct pitch or speed, we can completely disrupt the safety guardrails. Our method successfully compromised 10 popular AI models. Exposing this critical blind spot provides developers with the knowledge needed to build more secure multimodal AI systems.