Model Incrimination: Investigating Whether Concerning Behavior Reflects Misalignment
Abstract
A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior, but behavior alone is not sufficient to establish misalignment: a concerning action can arise from benign causes such as confusion. This raises the problem of determining whether malign intent underlies such behavior, a process we term model incrimination. The goal of this paper is to develop effective methods for doing so. To enable this, we create a suite of six agentic environments where models exhibit concerning behavior as practice grounds, and follow a simple two-step protocol for investigating the causes behind model behavior: hypothesis generation via reading the chain of thought -- which, while not always faithful, is a rich source of hypotheses about what drives model behavior -- followed by hypothesis validation via environment interventions and additional methods as appropriate. As we do not have access to ground truth about why a model takes an action, we rely on convergent findings across independent experiments as our standard of evidence. By following our protocol, we learn effective methods for concretely determining motivations: for example, we use predictions to back out latent properties of model behavior, like the fact that Kimi K2 Thinking takes shortcuts due to a legitimate disposition towards less effortful courses of action, and that reward hacking in frontier models is strategic, while in weaker models it is not. However, some unanswered questions require further methodological development: we construct an absence-of-evidence case against Kimi K2 Thinking believing it is going against the user's wishes while taking shortcuts, but run into confounds that limit our confidence in its absence. Overall, a key takeaway is that simple methods like reading the CoT and environment interventions are highly effective. More broadly, our work shows model incrimination is a tractable empirical problem with significant room for progress, and establishes a baseline for future work.