Learning to Hear Motion Before Naming It
Abstract
An agent that perceives a continuous sensory stream should first have a representation of how the stream is changing before naming or describing it. This motivates the learning of latent motion systems, considering that movement can be learned through time and space from audio, among other sensors. We instantiate this work by proposing a two-stage approach on 8-channel raw audio of vehicles approaching an occluded junction. The first stage trains a predictor on the latent representation of raw multichannel waveforms with no labels. The second attaches a temporal model with three-task classification heads (direction, count, and type of vehicles being detected). With this approach, we investigate what self-supervised learning (SSL) contributes and what supervised learning changes. We find that the SSL representation captures the movement structure of the scene and the acoustic-source identity that distinguishes the vehicles from each other. The acoustic-vehicle data underlying these experiments will be released in a forthcoming multi-modal data collection paper.