Interpretable Signals Reveal Failure-Relevant Representations in Vision-Language-Action Models
Tony Yang ⋅ Aidan Mokalla
Abstract
Vision-language-action (VLA) models take in visual observations and language instructions to produce physical robot actions, but when a task is failed, this poses safety concerns for real-world deployment. We study whether interpretable internal signals from a VLA can predict if the model will fail on LIBERO tasks. Using $\pi_{0.5}$, we cache action-expert activations and attention weights during evaluation, and then derive two families of episode-level features: attention-derived summaries over visual tokens and activation-probe summaries for action direction and gripper state. Our analysis suggests that action-direction probe confidence is the strongest failure-relevant signal, while attention-derived features are also predictive and gripper-state probes are weaker. We further compare against success-centroid distance and random-label probe baselines, and find that meaningful action-direction labels are important for the probe signal. These results work to show that simple mechanistic features inside VLAs can reveal failure-relevant representations.
Chat is not available.
Successful Page Load