Rationality Measurement and Theory for Reinforcement Learning Agents
Abstract
Lay Summary
Reinforcement learning agents are increasingly penetrating society, deployed in critical domains, such as robots, autonomous vehicles, and financial trading systems. Understanding their socioeconomic behaviour is thus crucial, for which establishing the rational model is fundamental. To this end, this paper develops the first mathematical framework in the literature for the rationality of reinforcement learning agents, to the best of our knowledge. We define a suite of measurements for rationality, develop rationality theory based on the measurements, and conduct experiments to showcase them. Our theory indicates that the irrational behaviours mainly come from two sources: the deficiency of the model and algorithm, as well as the shifts between training and deployment environments. Both theory and experiments suggest practical methods to improve rationality, including certain regularisation methods and diversifying training environments. This work sheds light on future development for detecting and mitigating the irrational behaviours of individual agents, understanding the new equilibria when AI goes into society, and informing mechanism design.