What do Uncertainty Lens tell about Emergent Misalignment?
Abstract
Emergent misalignment (EM) is a phenomenon in which language models display broad generalization of undesirable behavior after training on a narrow dataset of harmful examples. In this paper, we study emergent misaligned models through the lens of uncertainty quantification. While EM models demonstrate significantly higher uncertainty compared to the original models, a control model finetuned on benign data samples is similarly uncertain – suggesting that magnitude of uncertainty is more related to domain shift rather than to alignment properties. However, we find that EM model have a characteristic dynamic pattern in uncertainty: increased lag-1 autocorrelation across token-wise entropies. We interpret this as a hallmark of high-level uncertainty about aligned/misaligned behavior, rather than e.g. choosing a particular wording among semantically identical paraphrases.