Learning in Bayesian Stackelberg Games With Unknown Follower's Types
Abstract
Lay Summary
Stackelberg equilibria describe situations where one agent, called the leader, moves first, and another agent, called the follower, observes this decision and responds to it. This model captures many real-world interactions, such as a platform choosing a policy for its users, a company designing incentives for customers, or a defender allocating resources against possible attackers. In this paper, we study the case where the leader does not know the characteristics of the follower. In particular, the leader does not know how different followers may react to its decisions and must learn an optimal policy by repeatedly interacting with them. This makes learning much harder. We show that, if the leader only observes the follower’s actions, then learning a good strategy is impossible in general. However, if the leader can also observe the follower’s features after each interaction, then learning becomes possible. We provide an algorithm that learns a good leader strategy in this setting and prove that it has strong performance guarantees.