Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Alexander Panfilov ⋅ Peter Romov ⋅ Igor Shilov ⋅ Yves-Alexandre de Montjoye ⋅ Jonas Geiping ⋅ Maksym Andriushchenko
Abstract
We show that an autoresearch pipeline powered by Claude Code discovers novel white-box adversarial attack algorithms that significantly outperform all existing methods in jailbreaking and prompt injection evaluations. Starting from existing attack implementations, the agent iterates to produce new algorithms achieving up to 40\% attack success rate on CBRN queries against GPT-OSS-Safeguard-20B, compared to $\leq$10\% for existing methods. In a separate experiment, autoresearch on random-token targets yields attack algorithms that transfer directly to a prompt injection task, reaching 100\% ASR against Meta-SecAlign-70B versus 56\% for the best baseline. We release all discovered attacks alongside baseline implementations and evaluation code.
Chat is not available.
Successful Page Load