Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy
Abstract
Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actionable Parts (GAParts) understanding has attracted increasing attention from relevant researchers. However, most recent works either have insufficient design regarding the symmetry issue or require rich symmetry annotation, which severely impedes precise GAPart pose estimation in data-lacking scenarios. In this paper, we propose SAFAG, a novel Symmetry Annotation-Free framework for Generalizable and Actionable Parts Pose Estimation. Specifically, we suggest a stepwise refinement two-stage framework for candidate-to-final quaternion regression, and tackle the symmetry prediction as a probability distribution problem with self-supervised learning strategy. The experimental results demonstrate the superior performance and robustness of our SAFAG. We believe that our work has the enormous potential to be applied in many areas of embodied AI system.
Lay Summary
Robots need to understand how an object is positioned in space before they can interact with it accurately. In real-world manipulation, robots often operate on specific parts of objects rather than entire objects. However, these parts are usually inherently symmetric, so they can look nearly the same from many viewpoints even when they have different poses. This creates multiple plausible pose answers for the same observation, making it difficult for robots to understand the part’s spatial position and orientation. Many existing methods address this problem by relying on detailed symmetry information that manually annotated by human-beings. However, such annotations are difficult and time-consuming to obtain, which limits the practicality of these methods in data-lacking scenarios. In this paper, we propose a learning framework that helps estimate the pose of actionable object parts while reducing the need for detailed symmetry labels. Instead of relying on manually provided symmetry information, the method learns to handle pose ambiguity automatically.