The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
Andrzej Szablewski ⋅ Raffaello Fornasiere ⋅ Gabriel Konar-Steenberg ⋅ Nikita Menon ⋅ Stefan Heimersheim
Abstract
Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural transcripts or synthetic documents. Prior research has shown that interpretability methods can easily identify hidden behaviours in these MOs. However, recent work suggests that such post-hoc training methods may make interpretability unrealistically easy. We investigate this claim by constructing a suite of 54 $\verb|OLMo2-1B|$- and $\verb|Gemma-3-1B|$-based MOs trained with seven different techniques, including standard post-hoc SFT methods, post-hoc DPO, and more realistic integration of MO data into the OLMo post-training DPO phase. We use these MO variants to benchmark activation oracles, activation steering, logit lens, and sparse autoencoders. Our findings suggest that (i) current interpretability methods are not as capable as suggested by prior work; (ii) MO interpretability depends strongly on training methodology, target behaviour, interpretability technique, and model architecture; and (iii) substantial variance remains even after controlling for differences in the strength of target behaviour expression. Our results cast substantial doubt on the validity of current MOs as interpretability proxies. Our code is available here: https://github.com/anonsubmissionneurips2026/model-organism-lottery.
Chat is not available.
Successful Page Load