Reliable for Whom? Directional Reliability in AI-Mediated Political Dialogue
Abstract
Many ML evaluation frameworks approach reliable AI through epistemic accuracy, robustness, and general harmlessness. Yet as generative AI systems increasingly mediate human-to-human communication, these model-centric notions become insufficient. This paper is an empirically grounded position paper. Drawing on a formative design probe (N=25) conducted in the aftermath of South Korea's martial law declaration (December 3, 2024), we argue that reliability in AI mediation is not a property of a model output alone, but a directional relationship between speaker, recipient, and system. The probe revealed a layered asymmetry: politically entrenched participants resisted AI intervention in both directions, while participants who held views but had been reluctant to voice them found that reception-side scaffolding lowered the threshold for participation. We offer Directional Reliability as a sensitizing concept: rather than treating harmlessness as a symmetric property of text, it asks whether an intervention preserves expressive agency on the sender side while scaffolding reception on the recipient side. This reframing surfaces a blind spot in current output-level evaluation frameworks for AI systems deployed in politically difficult dialogue.