From Error Detection to Cultural Legibility: Human-AI Cooperation for Trauma-Informed Heritage Education in Conflict Zones
Abstract
Pedagogical tools for learners living in conflict zones remain scarce, leaving volunteer educators to shoulder substantial cognitive and emotional burdens while creating trauma-sensitive, culturally grounded teaching materials under severe resource constraints. We present \textsc{Haven} (\textbf{H}eritage-\textbf{A}ugmented \textbf{V}olunteer-led \textbf{E}ducation \textbf{N}etwork), a cooperative AI system in which three agents work in concert with human volunteers who iteratively refine each agent's behaviour by embedding tacit knowledge of trauma sensitivity, cultural nuance, and pedagogical appropriateness directly into the system. Indonesian cultural heritage serves as an intercultural pedagogical bridge, grounding English lessons in culturally rich content without requiring volunteers to simulate learners' own heritage traditions. Beyond system description, this paper advances a methodological argument relevant to the workshop's central question: \emph{how should we evaluate cultural aspects of generative AI in ways that articulate success, not only catalogue failure?} We introduce a dual-prompt human evaluation protocol and apply it to AI-generated illustrations of Indonesian cultural heritage topics across four image-generation models, each produced with and without cultural grounding. Our findings show that cultural grounding shifts the \textit{type} of errors that become visible to a rater, not only their frequency, and the \textit{narrative richness} captures a dimension of cultural legibility that error taxonomies cannot. These findings underscore that human involvement in the cultural evaluation loop is not a provisional workaround pending better automation, but an epistemic necessity for non-Western heritage contexts.