Towards Learning Representations of Policies in Two-Player Zero-Sum Games -- Wang et al.
Abstract
Towards Learning Representations of Policies in Two-Player Zero-Sum Games. -- Kevin A. Wang, Kevin Yang, Arjun Prakash, and Amy Greenwald.
Abstract: We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three types of contributions: First, we introduce methods of creating datasets of policies for a given game. Second, we propose methods to learn representations of policies. Third, we introduce downstream tasks to evaluate the effectiveness of such policy representations. For each dataset method, embedding method, and downstream task, we evaluate our methods on Kuhn and Leduc Poker. We find that while the trivial encoding methods fail to outperform a naive baseline, sophisticated methods perform significantly better. This demonstrates that useful behavioral representations are present in the learned embeddings. To our knowledge, this work is among the first to systematically compare self-supervised learning techniques for learning policy representations in this domain. Our source code will be available for others to extend in any of the three aspects.