NOSA: Native and Offloadable Sparse Attention
Yuxiang Huang ⋅ Pengjie Wang ⋅ Jicheng Han ⋅ Weilin Zhao ⋅ Zhou su ⋅ Sun Ao ⋅ Lyu Hongya ⋅ Hengyu Zhao ⋅ Yudong Wang ⋅ Chaojun Xiao ⋅ Xu Han ⋅ Zhiyuan Liu
Abstract
Decoding throughput is often limited by GPU memory dominated by the KV cache. Existing KV cache offloading reduces memory by storing context on CPU and fetching sparse KV subsets, but training-free methods suffer from long-generation quality degradation, while trainable sparse attention incurs excessive CPU--GPU transfers. We propose NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading. NOSA constrains CPU--GPU KV transfer volume to lower communication overhead and improve throughput. We further build NOSI, an offloading inference system that realizes NOSA's efficiency. Experiments on {1,3,8}B LLMs show that NOSA improves quality across general, long-input, and long-generation tasks, while boosting decoding throughput by up to $5.04\times$, $1.92\times$, and $1.83\times$ over FullAttn, InfLLMv2, and ShadowKV.
Chat is not available.
Successful Page Load