Explanation in an Emerging Science of Large Language Models
Abstract
LLM research increasingly takes on the character of scientific inquiry, but the field draws on heterogeneous explanatory approaches, including mechanistic interpretability, analyses of training dynamics, and statistical learning theory, without a clear account of how their explanatory strategies differ. Drawing on philosophy of science, and on the philosophy of cognitive neuroscience in particular, this paper develops a taxonomy of explanatory forms in LLM research along four axes: level of organization, mode of analysis, temporal scope, and scope of generalization. It uses this taxonomy to map the explanatory landscape of current work on in-context learning and grokking. The framework is both descriptive and normative. It clarifies the forms of explanation the field already produces, and articulates what a more mature science of LLMs would require.