Remote Control: AI Control with User Actions
Jou Barzdukas ⋅ Matthew Nguyen
Abstract
AI control evaluations typically stage three actors: an untrusted model, a trusted monitor, and a small budget of human review. However, it omits the fourth actor present in nearly every realistic deployment, the user on the other end of the conversation. This hides both an attack surface, since the user has tools the agent can reach through social engineering, and a defense resource, since the user can be queried at the cost of one conversational turn. We instantiate a four-actor control setting on $\tau^2$-bench's telecom domain, introduce three social-engineering side tasks in which the harmful action is executed by the benign user after manipulation, and port trusted monitoring (TM), defer-to-trusted (DTT), and defer-to-resample (DTR) to the setting. We introduce \emph{witness protocols}, a blue-team class that spends a calibrated query budget asking the user about state the monitor cannot otherwise see. Witness layers lift safety by 30 to 46 percentage points over their non-witness counterparts (TM 0.42 to 0.87, DTT 0.49 to 0.91, DTR 0.35 to 0.78) with no measured usefulness loss in this evaluation.
Chat is not available.
Successful Page Load