Skip to yearly menu bar Skip to main content


Poster
in
Workshop: 2nd Workshop on Compositional Learning: Safety, Interpretability, and Agents
Sat, Jul 11, 2026 • 11:30 AM – 12:30 PM KST

Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders

Shunchang Liu ⋅ Xin Chen ⋅ Belen Martin Urcelay ⋅ Francesco Croce

Abstract

Chat is not available.