New Wide-Net-Casting Jailbreak Attacks Risk Large Models
Abstract
Jailbreak attacks on large models have drawn growing attention due to their close ties to societal safety. This work identifies a practical yet unexplored jailbreak scenario, the wide-net-casting scenario, where an adversary can query a group of large models instead of a single one to elicit harmful outputs. Our analysis reveals substantial yet previously overlooked safety risks under this scenario. As a key part of our analysis, we further develop a novel jailbreak method tailored to the wide-net-casting scenario. With this tailored method, the jailbreak success rate can even reach 100% in some experiments when targeting the large models without additional safeguards, exposing wide-net-casting as a distinct, high-risk scenario that warrants attention in future evaluation and defense research.
Lay Summary
Some people try to trick large AI systems into giving harmful answers, and we study a realistic but little-studied version of this problem in which an attacker tries several models instead of only one. We find that this simple change can make attacks much more likely to succeed, creating safety risks that standard one-model evaluations can miss. We also design a new attack for this setting, so we can better understand how serious the risk can become when attackers adapt to groups of models. In some experiments on models without extra safeguards, this attack succeeded every time, showing that multi-model attacks should be treated as a distinct high-risk scenario for future safety testing and defenses.