Skip to yearly menu bar Skip to main content


Poster
in
Workshop: 2nd Workshop on Compositional Learning: Safety, Interpretability, and Agents
Fri, Jul 10, 2026 • 7:30 PM – 8:30 PM PDT

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

Shunchang Liu ⋅ Xin Chen ⋅ Belen Martin Urcelay ⋅ Francesco Croce

Abstract

Chat is not available.