Compositional Skill Chaining and Policy Blending for Hard Exploration in the BRIO Labyrinth Game
Abstract
The classic BRIO Labyrinth game presents a hard-exploration problem with sparse and deceptive rewards, requiring precise motor control and long-horizon planning that existing approaches struggle to solve without human guidance. To address this, we propose a two-phase compositional framework. Phase 1 leverages the Go-Explore paradigm to discover a high-performing trajectory. Phase 2 robustifies this trajectory via backward learning and decomposes it into specialized skills using compositional skill chaining, avoiding the instability of a monolithic policy. To ensure kinematic continuity, we introduce a policy-blending mechanism that mitigates control instability during abrupt skill transitions. Experimental results demonstrate our framework effectively solves this challenge, with policy blending actively counteracting physical momentum to achieve a 92.6% success rate, substantially outperforming standard skill chaining.