SALSA: State Augmentation via Learned Selective Attention
Abstract
Hybrid architectures combining attention and recurrent primitives have emerged as a popular paradigm in sequence modeling, achieving favorable tradeoffs between the in-context recall of transformers and the bounded state compression of recurrences. However, existing hybrids statically interleave primitives at fixed, hand-designed ratios, invoking expensive attention operations regardless of whether the recurrent state is sufficient, raising concerns about adaptability and efficiency. We introduce SALSA, a dynamic hybrid architecture that pairs the recurrent primitive with a learned router that selectively invokes attention on a per-token basis, when projected to improve the recurrent state quality. We evaluate this approach on short and long context benchmarks at scale. At 1.3B parameters, SALSA achieves lower pretraining perplexity than the static hybrid (9.72 vs. 9.99), outperforms it on 8 of 8 LM-Eval metrics by an average of 1.5\%, and improves PG19 perplexity at every sequence length evaluated from 2k to 32k. Additionally, we investigate the learned routing patterns of SALSA to reveal sparse, structured attention patterns that suggest SALSA acquires meaningful, content-dependent strategies for when global context is necessary. Together, these results position dynamic primitive allocation as a principled and practical direction for hybrid sequence modeling.