Hyperparameter Transfer with Mixture-of-Expert Layers
Abstract
Mixture-of-Experts (MoE) layers have emerged as an important tool in scaling up modern neural networks by decoupling total trainable parameters from activated parameters in the forward pass for each token. However, sparse MoEs add complexity to training due to (i) new trainable parameters (router weights) that, like all other parameter groups, require hyperparameter (HP) tuning; (ii) new architecture scale dimensions (number of and size of experts) that must be chosen and potentially taken large. To make HP selection cheap and reliable, we propose a new parameterization for transformer models with MoE layers when scaling model width, depth, number of experts, and expert (hidden) size. Our parameterization is justified by a novel dynamical mean-field theory (DMFT) analysis. When varying different model dimensions trained at a fixed token budget, we find empirically that our parameterization enables reliable HP transfer across models from 51M to 2B total parameters. We further take HPs identified from sweeping small models on a short token horizon to train larger models on longer horizons and report performant model behaviors.
Lay Summary
Training massive AI models leads to better performance, but it requires carefully tuning "hyperparameters", the underlying control settings during the training process. Because giant models require massive amounts of compute, finding the optimal settings on a large-scale model is expensive. To solve this, we study "hyperparameter transfer", a technique that involves finding the best settings on a small and cheap model and using specific rules to translate those settings to larger models. We derive these exact rules focusing on "Mixture-of-Experts" models, a modern model design that only requires fractional compute versus traditional designs, but is more sensitive to training settings. We show that our rule consistently enables good hyperparameter transfer and leads to performant models under compute budget.