Ted Underwood
Abstract
Can Language Models Speak for the Past?
The concrete center of this talk is a benchmark that evaluates models' representation of English-language writing from 1875 to 1924. But why build this? Evaluating representation of history has become important for two reasons. 1) Humanists and social scientists have proposed that models trained on historical sources could provide a new kind of evidence about the past, and models to fill that need are beginning to emerge. 2) The AI community more broadly is concerned to improve models' representation of cultural difference — and history makes a good test case. I'll share benchmark results, which suggest historical models have several distinct purposes that may need to be evaluated separately. In most cases, I’ll suggest, “context” is a better name than “culture” for the thing being measured.