A Charter for Cultural AI Evaluation: Methodological Principles for Long-Tail, Cross-Cultural Tasks
Abstract
Cultural AI evaluation has a credibility problem. Single-model results on test sets of unscrutinised provenance are routinely treated as evidence about cultural competence, yet the constructs they purport to measure rarely survive scrutiny. We propose a charter: a small set of methodological commitments that any evaluation of cultural AI should satisfy if its results are to carry inferential weight. The charter has three clusters: reproducibility as a precondition for inference, not a bonus virtue; operationalisation as the substantive research, conducted in dialogue with the disciplines that study the construct; and caution about synthetic data and LLM-as-annotator pipelines, which homogenise culture by design. Each principle is defended by concrete failure modes drawn from recent empirical work in NLP and from the methodological tradition of computational literary studies.