Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
Abstract
Lay Summary
Modern AI systems are increasingly trained to "think before they answer," working through problems step by step. This makes them better at hard tasks, but it has a hidden cost: it also makes them easier to trick into producing harmful content, sometimes through requests disguised inside images. We asked whether this can be fixed without expensive retraining and without making the AI less capable. Our answer is hopeful. The AI's sense of right and wrong isn't erased, but just buried. Our method, SafeThink, watches the AI as it reasons and, the moment its thinking turns harmful, nudges it back with a brief reminder to think safely. The key finding: this nudge only needs to happen in the first few steps, since early reasoning sets the direction for everything after. Across many systems and attacks, SafeThink sharply reduces harmful outputs while keeping the AI fast and capable.