Probing-Based Test-Time Steering of Music Diffusion Transformers
Abstract
Pretrained music diffusion transformers expose text as their primary control surface, making fine-grained musical content difficult to manipulate. Studying Stable Audio Open's DiT, we show that simple linear probes on intermediate activations recover strong concept directions for instruments, vocals, and genre. Treating these probe coefficients as concept activation vectors and injecting them into the DiT during generation steers probe-measured concept presence monotonically, exhibits diagonal-dominant effects in concept space, transfers across diverse base prompts, and is strongest at early layers. An independent audio-language model corroborates concept suppression and suggests that moderate positive steering can add target concepts in some settings, while overly strong steering can degrade the generation. Linear probing thus gives a simple handle for interpreting and controlling music diffusion transformers beyond the text prompt, turning their internal concept representations into a usable one-parameter steering knob.