Timezone: »

Racing Thompson: an Efficient Algorithm for Thompson Sampling with Non-conjugate Priors
Yichi Zhou · Jun Zhu · Jingwei Zhuo

Fri Jul 13 09:15 AM -- 12:00 PM (PDT) @ Hall B #155

Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors since they require to perform online Bayesian posterior inference, which is a difficult task when the prior is not conjugate. In this paper, we propose a novel algorithm for Thompson sampling which only requires to draw samples from a tractable proposal distribution. So our algorithm is efficient even when the prior is non-conjugate. To do this, we reformulate Thompson sampling as an optimization proplem via the Gumbel-Max trick. After that we construct a set of random variables and our goal is to identify the one with highest mean which is an instance of best arm identification problems. Finally, we solve it with techniques in best arm identification. Experiments show that our algorithm works well in practice.

Author Information

Yichi Zhou (Tsinghua University)
Jun Zhu (Tsinghua University)
Jingwei Zhuo (Tsinghua University)

Related Events (a corresponding poster, oral, or spotlight)

More from the Same Authors