Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
Abstract
Lay Summary
This research studies how to make good decisions even when the necessary information is delivered with delays. Taking autonomous driving as an example, a self-driving car experiences delays between sensing its surroundings and steering the vehicle. Delays that are unavoidable no matter how quickly the agent makes decisions also arise in communication networks, robotics, healthcare, and online advertising. We develop a mathematical framework for decision-making problems under delayed observations that can be applied across a wide variety of settings. We propose a strategy for how an AI agent should act in these settings and prove that it achieves a guaranteed level of performance. Moreover, we show that this guarantee is the best possible performance that any strategy can achieve. By these results, we answer the question of how observational delays affect the difficulty of decision-making problems. While it was known that delays affect decision-making by limiting the available strategies, our results additionally reveal that they slow down collecting the relevant information.