Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
Abstract
Training modern large language models (LLMs) has become a veritable smorgasbord of algorithms and datasets designed to elicit particular behaviors, making it critical to develop techniques to understand the effects of datasets on the model's properties. This is exacerbated by recent experiments that show datasets can transmit signals that are not directly observable from individual datapoints (Halawi et al., 2024; Betley et al., 2025b;Cloud et al., 2025; Betley et al., 2025a), posing a conceptual challenge for dataset-centric understandings of LLM training and suggesting a missing fundamental account of such phenomena. Towards understanding such effects, inspired by recent work on the linear structure of LLMs (Park et al., 2024; Golowich et al., 2025b), we uncover a general mechanism through which hidden subtexts can arise in generic datasets. We introduce LOGIT-LINEAR SELECTION (LLS), a method that prescribes how to select subsets of a generic preference dataset to elicit a wide range of hidden effects. We apply LLS to discover subsets of real-world datasets so that models trained on them exhibit behaviors ranging from having specific preferences, to responding to prompts in a different language not present in the dataset, to taking on a different persona. Crucially, the effect persists for the selected subset, across models with varying architectures, supporting its generality and universality.
Lay Summary
Modern AI language models are trained on a huge mix of datasets chosen to give them particular skills and behaviors, so it really matters to understand how a dataset shapes the model you get. That's gotten harder because recent work shows datasets can carry hidden signals you can't spot by just looking at the individual examples — which means our usual way of thinking about training data is missing something fundamental. Building on recent findings that AI models have a simple underlying mathematical structure, we figured out a general explanation for how these hidden messages can be buried in everyday datasets. We introduce a method called Logit-Linear Selection (LLS) that tells you which examples to pick out of an ordinary dataset to plant a chosen hidden effect. Using it on real datasets, we made models take on specific preferences, reply in a language that never appears in the data, or adopt a different persona — and the same selected subset keeps working across different kinds of models, showing the effect is general and not a fluke.