Observing and Controlling Features in Vision-Language-Action Models
Hugo Buurmeijer ⋅ Carmen Amo Alonso ⋅ Aiden Swann ⋅ Marco Pavone
Abstract
Vision-Language-Action Models (VLAs) have shown remarkable progress towards embodied intelligence. While their architecture partially resembles that of Large Language Models (LLMs), VLAs exhibit higher complexity due to their multi-modal inputs/outputs and often hybrid nature of transformer and diffusion heads. This is part of the reason why insights from mechanistic interpretability in LLMs, which explain how the internal model representations relate to their output behavior, do not trivially transfer to VLA counterparts. In this work, we study the linear separability of action-relevant features across VLA architectures, and show that observability varies significantly by design: in the transformer-based OpenVLA, action-relevant features are poorly linearly separable, whereas in the hybrid architecture $\pi_{0.5}$, such features are accurately recovered from internal representations. Building on this, we show that lightweight linear interventions grounded in optimal control can reliably steer $\pi_{0.5}$'s behavior while preserving closed-loop capabilities, enabling alignment with user preferences and task requirements without fine-tuning.
Chat is not available.
Successful Page Load