Near-Minimax Multi-Objective RL under Predictable Adversarial Preferences and Preference-Free Exploration in Linear MDPs
Abstract
Lay Summary
Many AI systems need to balance several goals at once. For example, a robot may need to be fast, safe, and energy-efficient, while different users may care about these goals in different ways. A common shortcut is to combine past outcomes using whatever preference is chosen later, but this can make the learning process statistically unreliable. This paper studies how to learn when preferences can change from one trial to the next, or when preferences are only specified after data has already been collected. Our method keeps separate estimates for each goal and combines them only when a particular preference is requested. We also explain how to evaluate the set of trade-offs that can actually be deployed when randomized choices between decision rules are allowed. The results show that this protocol-aware approach can achieve strong learning guarantees while avoiding an expensive search over all possible preferences.