Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation
Abstract
Lay Summary
AI systems such as chatbots are often improved using feedback from people. A common approach is to show people two possible answers from an AI system and ask which one they prefer. The system then learns from these choices so that it can produce answers that better match human expectations. However, this process assumes that the people giving feedback fairly represent the wider population. In practice, this is often not true. Some people or communities may be easier to reach, more active online, or more connected to the platforms where feedback is collected. Their opinions may therefore have more influence than those of less visible groups. This paper studies how such imbalance can affect AI training when feedback providers are connected through social, institutional, geographic, or platform-based networks. We show that standard methods may unintentionally learn preferences that reflect the most represented or most connected groups, rather than the intended population as a whole. To address this, we propose Graph-Preference Learning, a method that uses information about relationships among feedback providers and adjusts how much each person’s feedback contributes. This helps the learned preference model better reflect a chosen target population. Our experiments, using both simulated networks and language-model preference data, show that the proposed method can more accurately recover the intended population preferences and reduce performance differences across language groups. Overall, this work aims to make human feedback for AI systems more representative, transparent, and fair.