Skip to yearly menu bar Skip to main content


Poster
in
Workshop: Trustworthy AI for Good Workshop

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

Catherine Ge-Wang ⋅ Tyler Crosse ⋅ Benjamin Hadad ⋅ Joachim Schaeffer ⋅ Ram Potham ⋅ Tyler Tracy

Abstract

Chat is not available.