PHALAR: Phasors for Learned Musical Audio Representations
Abstract
Lay Summary
When music is mixed, every instrument must lock together in time. A drummer even slightly off-beat with the bassist makes the whole track feel wrong. Current AI models are blind to this: designed to recognize what is in a recording, they discard the timing information needed to judge whether those parts actually fit together. We introduce PHALAR, a model that fixes this by exploiting a mathematical property of sound. Shifting audio in time rotates its frequency-domain representation by a proportional angle, and rather than throwing this rotation away like standard models do, PHALAR is built to preserve it, encoding rhythmic alignment as a geometric angle in a complex-valued space. The result is a model that judges musical coherence up to 70% more accurately than the previous best approach, using half the parameters and training seven times faster, while aligning significantly closer with how human listeners perceive whether stems belong together.