Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning
Abstract
Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-conditioned value function in long-horizon tasks remains challenging. In this paper, we identify erroneous generalization in goal-conditioned value functions as a fundamental bottleneck, and demonstrate that appropriate inductive bias in the value function is crucial for addressing the bottleneck. Building on these findings, we propose Latent-Aligned Value Learning (LAVL), an offline GCRL algorithm that integrates latent-representation-based value generalization with hierarchical planning in a unified framework. Extensive numerical experiments on OGBench demonstrate that LAVL consistently outperforms existing offline GCRL methods, achieving the highest performance on 20 out of 22 datasets. Notably, LAVL exhibits strong performance in long-horizon tasks and trajectory stitching datasets, where prior methods suffer significant performance degradation.
Lay Summary
In this work, we show that erroneous generalization that fails to reflect temporal distance is a central bottleneck in offline GCRL, and that value generalization can be substantially improved by incorporating appropriate inductive bias into the value function. Motivated by this observation, we design the Latent Alignment Network (LAN), an effective architecture for goal-conditioned value learning, and propose Latent-Aligned Value Learning (LAVL), an algorithm that integrates LAN-based value learning with a hierarchical policy framework. Through extensive experiments on OGBench, we demonstrate that LAVL improves upon existing offline GCRL algorithms by a large margin, while enabling highly stable learning, particularly in long-horizon tasks and datasets that require trajectory stitching.