Hearing Without Noticing? Attention-Aware Stealthy Black-Box Adversarial Audio Attacks
Abstract
Automatic Speech Recognition (ASR) systems, such as those in intelligent assistants, are vulnerable to adversarial examples (AEs). Benign audio clips like music, when embedded with small perturbations, can trick ASR models into recognizing attacker-specified commands. Prior studies focus on minimizing perturbation magnitude to craft AEs. However, they fails to achieve high attack stealthiness against black-box ASR systems in the physical world. In this paper, we introduce the first music carrier selection algorithm and an attention-aware stealthiness loss function to generate stealthy AEs. Extensive evaluations on five commercial ASR APIs and three widely-used voice assistants demonstrate that our method significantly outperforms state-of-the-art techniques in both effectiveness and stealthiness. Notably, in a user study involving 200 participants, 55.6\% of participants perceived our physical adversarial examples as benign audio, which is an improvement of over 20\% compared to existing methods.
Lay Summary
Hackers can trick voice assistants using regular-sounding music secretly embedded with malicious commands. However, earlier methods usually created weird static or faint whispers that people easily noticed, making the attacks easy to catch. We developed a stealthy attack that exploits the psychology of how human hearing and attention work. Our method automatically picks complex music (like electronic beats) that naturally hides extra sounds. It then weaves the secret command into the music so smoothly that the human brain simply ignores it, while the voice assistant still "hears" the command. In real-world tests, popular voice assistants executed these hidden commands with a 100% success rate, yet over half of the participants thought they were just listening to regular music. This reveals a major security flaw in everyday voice-controlled devices and highlights the urgent need for stronger defenses.