What Do Historical Language Models Model?
Abstract
Historical language models are increasingly used to infer attitudes, beliefs, or viewpoints from past societies. This paper argues that such uses rest on fragile epistemic assumptions. We show that historical language models do not simulate past minds or populations, but instead model the structure of surviving textual archives, which are shaped by systematic biases of literacy, genre, and preservation. By introducing a validity ladder that distinguishes textual, discursive, and population-level claims, we provide a framework for evaluating what kinds of historical inferences these models can legitimately support. This perspective clarifies how historical language models can contribute to research in the social sciences and humanities without encouraging overinterpretation.