Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Abstract
LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly shapes subsequent actions. Small deviations in these thoughts can therefore propagate into unsafe behaviors, yet existing guardrails typically operate only on final outputs or require intrusive model modifications. We introduce Thought-Aligner, a lightweight plug-in safety model that performs causal correction on unsafe thoughts before action execution, without altering the underlying agent. The corrected thoughts are fed back into the agent, steering its decision process and tool use toward safer trajectories. Because it operates solely at the thought level, Thought-Aligner is model-agnostic and can be integrated into diverse agent frameworks. We train Thought-Aligner via two-stage contrastive learning on paired safe and unsafe thoughts generated across ten risk scenarios. Experiments on diverse agent-safety benchmarks and six LLMs show that Thought-Aligner increases behavioral safety from about 50% without protection to around 90% on average, exceeding state-of-the-art guardrails by roughly 23%, while also improving helpfulness by about 5%. The method incurs low per-step latency and minimal overhead, enabling scalable and practical deployment. We publicly release Thought-Aligner-7B at https://huggingface.co/WhitzardAgent/Thought-Aligner-7B.
Lay Summary
AI agents can now use tools and take multi-step actions, such as managing files, sending emails, or operating online services. These abilities make them useful, but also create risks: a small mistake in an agent’s intermediate reasoning can lead to unsafe actions, such as leaking private information, spending money without permission, or deleting important data. Our paper introduces Thought-Aligner, a lightweight safety module that checks and corrects an agent’s intermediate “thoughts” before the agent acts. Instead of simply blocking a final response or retraining the whole agent, Thought-Aligner guides the agent toward safer actions while still allowing it to continue the task. We train Thought-Aligner using paired examples of unsafe and safer agent thoughts across common risk scenarios. Across several agent-safety benchmarks and six language models, it substantially improves behavioral safety while preserving helpfulness. This suggests that correcting reasoning before action is a practical way to make AI agents safer.