AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
Abstract
While Vision-Language-Action (VLA) models have achieved remarkable success in ground-based embodied intelligence, their application to Aerial Manipulation Systems (AMS) remains a largely unexplored frontier. The inherent characteristics of AMS, including floating-base dynamics, strong coupling between the UAV and the manipulator, and the multi-step, long-horizon nature of operational tasks, pose severe challenges to existing VLA paradigms designed for static or 2D mobile bases. To bridge this gap, we propose AIR-VLA, the first VLA benchmark specifically tailored for aerial manipulation. We construct a physics-based simulation environment and release a high-quality multimodal dataset comprising 3000 manually teleoperated demonstrations, covering base manipulation, object & spatial understanding, semantic reasoning, and long-horizon planning. Leveraging this platform, we systematically evaluate mainstream VLA models and state-of-the-art VLM models. Our experiments not only validate the feasibility of transferring VLA paradigms to aerial systems but also, through multi-dimensional metrics tailored to aerial tasks, reveal the capabilities and boundaries of current models regarding UAV mobility, manipulator control, and high-level planning. AIR-VLA establishes a standardized testbed and data foundation for future research in general-purpose aerial robotics. The resource of AIR-VLA will be available at https://anonymous.4open.science/r/AIR-VLA-dataset-B5CC/.
Lay Summary
Drones equipped with robotic arms have huge potential for search and rescue, logistics, and high-altitude work. However, controlling them is incredibly difficult because they must balance hovering in mid-air while performing precise movements. While artificial intelligence has recently made ground-based robots much smarter at following human instructions, flying robots have largely been left behind. To bridge this gap, we built a virtual 3D testing environment and recorded 3,000 human-guided examples of a flying robot performing various tasks. We used this data to evaluate how well the latest AI models—which translate human language into physical actions—can learn to control both the drone's flight and the robotic arm simultaneously. Our results show that it is possible to adapt these powerful AI models for the skies, though they currently still struggle with complex long-term planning and staying perfectly stable. By providing the first standardized testbed for this technology, our work lays the foundation for developing smart, fully autonomous flying assistants that can follow natural language commands in the real world.