Interpretive Anchoring for Culturally Situated LLM Evaluation
Abstract
Generative AI systems increasingly produce cultural artifacts, yet their evaluation is often reduced to context-free scoring. We argue that LLM-as-a-judge systems should instead be understood as \emph{interpretive technologies}: when they evaluate generated texts, explanations, critiques, or creative outputs, they make situated judgments about genre, audience, precedent, and value. This framing reveals a failure mode we call \emph{context-induced interpretive miscalibration}, in which irrelevant or mismatched reference examples distort the judge's decision boundary. We introduce DA-RAC, a distance-aware reference anchoring method that retrieves semantically and structurally proximate labelled precedents for each judgment scenario, weights them by distance, and exposes neighbourhood difficulty as a signal for human review. On LLM-judge benchmarks, DA-RAC improves calibration relative to zero-shot, rubric-based, and static-precedent baselines; we use these results as a technical probe of reference dependence rather than as a complete cultural benchmark. The broader contribution is a design pattern for cultural AI evaluation: judgments should be grounded in relevant, inspectable, and contestable interpretive precedents. DA-RAC offers one computational mechanism for embedding contextual sensitivity upstream in AI evaluation while preserving human agency in ambiguous cultural cases.