Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
Abstract
Lay Summary
Many learning systems improve by trying actions, observing feedback, and gradually learning which actions work best. But if the system starts with very uncertain beliefs, it is unclear how much this initial uncertainty should slow down learning. Prior analyses suggested that this uncertainty could keep making learning harder over the entire time horizon. We show that, for an important model of sequential decision-making with noisy feedback, this pessimism is unnecessary. For Thompson sampling, a widely used strategy that chooses actions by sampling from its current beliefs, the cost of initial uncertainty is mostly a one-time “burn-in” cost. After this early phase, performance is governed mainly by the noise in the feedback, not by how uncertain the system was at the start. We also prove that this burn-in cost is unavoidable in general: any method must spend some effort resolving large initial uncertainty. Our results give a sharper understanding of when Thompson sampling is efficient, and help explain why broad prior uncertainty need not permanently harm long-run learning.