Consensus Is Not Enough: Disagreement-Preserving Evaluation for Cultural AI
Abstract
Many evaluations of generative AI treat agreement as success: one preferred response, one majority label, one reward score, one consensus answer. This reduction is useful when disagreement is noise. In cultural and value-laden tasks, however, disagreement can be part of the object being evaluated. It may reflect ambiguity, taste, expertise, stakeholder position, moral conflict, or power asymmetry. We call the loss of this structure consensus collapse: the compression of meaningful disagreement into a single answer or score whose presentation conceals what was compressed. We introduce Disagreement-Preserving Evaluation. Prior work preserves disagreement at the label level, the training-target level, and the preference-distribution level; our focus is downstream of all of these: whether disagreement remains visible in the final user-facing response and headline evaluation report. The framework has two modes. In distributional settings, it evaluates whether systems preserve the shape of annotator disagreement rather than only matching a majority label. In open-ended cultural settings, it scores responses using a five-dimensional vector: Visibility, Tradeoffs, Calibration, Boundaries, and Actionability. We anchor the framework in a single domain (education) so the rubric is concrete and falsifiable. We contribute a taxonomy of disagreement types; a response-level rubric; an ordinary-metric-versus-DP-EVAL contrast; an explicit boundary-calibration subprotocol with a 1-3 harm tier coded alongside the rubric; a four-stage pipeline view of consensus collapse; a minimal audit protocol; and a single-rater pilot study (3 education prompts x 3 frontier models x 2 response styles, blinded during rating) in which the disagreement-preserving response style raised the V/T/C/B/A composite by +1.73 on a 1-5 scale (Wilcoxon p = 0.004 over 9 paired prompt-by-model cells) while raising ordinary helpfulness only +0.67 (p = 0.062), and in which 3 of 18 responses (17%) fell in a false-consensus quadrant (helpfulness >= 4 with composite <= 3), all under ordinary prompting. We describe a planned multi-rater follow-up (N>=5 trained annotators, M>=30 prompts, target Krippendorff's alpha >= 0.6) and treat the present numbers as a small-N behavioral signal rather than a benchmark. The positive cultural value proposed here is interpretive agency: users and communities should be able to see what a system has compressed, what it has bounded, and what value tradeoff supports its recommendation.