Bandit Social Leaning Dynamics with Exploration Episodes
Abstract
Lay Summary
Many online platforms rely on users learning from each other’s experiences. Examples include people repeatedly interacting with AI systems, shopping on online marketplaces, or using recommendation systems. In these settings, users naturally prefer options that already seem best, rather than trying uncertain alternatives that could help others learn. We study this phenomenon through a mathematical model based on multi-armed bandits, where each user makes a short sequence of decisions instead of only one. A natural question is whether this repeated interaction creates enough “organic” exploration for the system to learn good alternatives over time. Our main result is negative: even when users explore a little within their own sessions, the overall system still typically fails to learn efficiently. In fact, we show that these failures are not rare worst-case examples, but arise broadly across many settings and utility models. As a consequence, the collective learning process has Bayesian regret that grows linearly over time. These results suggest that externally incentivized exploration remains necessary in social learning systems, even when users already experiment on their own.