A Bayesian epistemology for LLM evaluation
Abstract
Evaluating the capacities of large language models (LLMs) requires inferring latent abilities from observable task performance. We propose to formalize such inferences within a Bayesian framework, drawing on the epistemology of comparative cognitive science. The framework makes explicit how posterior credences about a system's abilities depend on prior probabilities over hypotheses about computational strategies, likelihood functions shaped by task demands, and marginalization over auxiliary factors that influence performance independently of the target ability. Within this framework, we argue that the relative evidential weight of behavioral versus mechanistic evidence for algorithmic-level hypotheses concerning the representations and computations underlying task performance depends systematically on background knowledge about the target system. For human cognition, decades of research in cognitive science, neuroscience, and evolutionary biology provide strong priors that constrain the space of plausible computational strategies, while well-characterized auxiliary factors (attention, motivation, working memory limits) allow informative marginalization. Behavioral paradigms consequently provide substantial evidential leverage for discriminating between algorithmic hypotheses. For LLMs, this evidential hierarchy shifts: weak priors about learned computational strategies and poorly characterized auxiliary factors (e.g., sensitivity to prompt formatting, tokenization artifacts, training distribution) make behavioral evidence systematically underdetermining across competing hypotheses. Mechanistic interpretability, by contrast, provides more direct access to representations and computations, offering evidence that can discriminate where behavioral evidence cannot.
Speaker