Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting
Abstract
Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outputs without considering the underlying mechanisms. Using election forecasting as a test case, we introduce mechanistic forecasting, a method that demonstrates that probing internal model representations offers a fundamentally different---and sometimes more effective--- approach to preference prediction. Examining over 24 million configurations across 7 models, 6 national elections, multiple persona attributes, and prompt variations, we systematically analyze how demographic and ideological information activates latent party-encoding components within the respective models. We find that leveraging this internal knowledge via mechanistic forecasting, opposed to solely relying on surface-level predictions, can improve prediction accuracy. The effects vary across demographic versus opinion-based attributes, political parties, national contexts, and models. Our findings demonstrate that the latent representational structure of LLMs contains systematic, exploitable information about human preferences, establishing a new paradigm for using language models in social science prediction tasks.
Lay Summary
Large language models are increasingly used to predict human preferences, but existing approaches only look at what models say — not what’s happening inside them. We introduce mechanistic forecasting, which reads a model’s internal representations instead of its outputs. Using election forecasting as a test case, we study over 24 million configurations across 7 models and 6 national elections, and show that internal model structure contains systematic, exploitable information about political preferences — information that surface-level predictions sometimes miss entirely. This opens a new path for using LLMs in social science.