Decoupling Spatial and Semantic Token Compression for Vision-Language Model Acceleration
Seunghun Moon ⋅ JAEHYUN PYUN ⋅ Hyunwoo Yu ⋅ Suk-Ju Kang
Abstract
Vision-Language Models (VLMs) face massive inference overhead from extensive visual tokens. Existing Top-$K$ pruning methods mitigate this but suffer from severe spatial bias, information redundancy, and crucial context loss. To address this, we propose TokenNMS, a training-free two-stage framework that reframes token reduction as deterministic feature-space Non-Maximum Suppression (NMS). TokenNMS seamlessly bridges query-agnostic spatial pruning with query-aware semantic filtering, enforcing similarity constraints to penalize semantic overlap. Extensive experiments demonstrate our approach effectively preserves spatially diverse representations while accelerating inference across diverse VLMs.
Chat is not available.
Successful Page Load