ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
Abstract
Lay Summary
Most robots used in today’s vision-language systems are built from rigid mechanical arms, which can struggle in crowded or narrow environments. Soft robotic arms, made from flexible materials, could adapt more naturally to these situations, but they are also much harder to control because they bend and deform continuously. To study these challenges, we introduce ManiSoft, a new benchmark for training and evaluating soft robotic arms on language-guided manipulation tasks. ManiSoft includes a realistic simulation environment where soft arms can interact with objects and obstacles in physically accurate ways. We design four representative tasks that test different abilities, such as precise reaching, coordinated movement, and navigating around obstacles. We also build an automated system to generate thousands of training scenes and expert demonstrations, enabling large-scale learning and evaluation. Our experiments with several representative robot learning methods show that current approaches can perform reasonably well in simple settings, but their performance drops significantly in more complex or randomized environments. Further analysis suggests that existing models struggle to accurately infer the shape and state of soft arms from visual input and do not fully take advantage of the flexibility of soft robots during interaction. We hope ManiSoft can serve as a useful platform for advancing research on flexible, adaptable robotic systems that can better operate in real-world environments.