Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
Abstract
Lay Summary
Imagine you are playing a game with multiple slot machines. Your goal is to find the balance between choosing the most promising machine to earn rewards and testing other machines to gather information. In our study, we look at a unique setup where you can separate these tasks by choosing two different machines at once. In the real world, it is often difficult to predict how rewards are generated in advance. For example, a machine's jackpots might follow a fixed pattern, or the reward process may change over time due to external manipulation or unexpected environmental shifts. To address this, we developed a robust method that guarantees optimal performance under any kind of uncertainty. Furthermore, our approach is not only much faster to compute than existing methods but also delivers better actual performance. By successfully combining robustness and efficiency, this method offers a practical solution for real-world decision-making.