Investigating Expert Semantic Specialization in Mixture-of-Expert Models
Abstract
Mixture-of-Experts (MoE) models achieve exceptional scalability through selective routing, yet our understanding of what routing actually learns remains limited. In particular, it is unclear to what extent routing induces meaningful expert semantic specialization. Sparse Autoencoders (SAEs) provide a way to extract interpretable semantic feature spaces from model representations. Building upon this capability, we introduce the Semantic Specialization Index (SSI), a quantitative metric designed to measure the degree of expert semantic specialization. Using SSI, we systematically quantify expert semantic specialization in MoE models. We further investigate the relationship between semantic specialization and language modeling performance, and find that specialization exhibits a non-monotonic trajectory during training and does not increase indefinitely with model quality, suggesting the existence of an optimal specialization regime. These findings provide a quantitative foundation for understanding expert organization in sparse models and open new opportunities for specialization-aware optimization of MoE models.