AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping
Abstract
Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause of loss spikes by investigating individual factors, we observe that, in practice, such spikes are typically triggered by the confluence of heterogeneous factors. Empirically, loss spikes may arise from a combination of data outliers, hardware or transient computational faults, numerical precision issues, and hyperparameter settings. Regardless of the underlying cause, these spikes manifest as unstable optimizer updates, as abnormal gradients contaminate both first- and second-moment states. In this paper, we propose a principled gradient-centric remedy: AdaGC, an adaptive per-tensor gradient clipping scheme that mitigates such contamination by bounding gradient norms relative to a tensor-wise exponential moving average of their historical clipped values. AdaGC is optimizer-agnostic, introduces negligible memory overhead, and reduces communication costs compared to GlobalGC, particularly in hybrid-parallel distributed training. Experiments on Llama-2 7B, Mixtral 8×1B, and ERNIE 10B-A1.4B demonstrate that AdaGC robustly eliminates training instabilities, consistently reducing spike scores to zero for all models and improving downstream accuracy over GlobalGC by 1.32\%, 1.27\%, and 2.48\%, respectively. Furthermore, AdaGC seamlessly integrates with optimizers such as Muon and Lion, consistently yielding higher average accuracy and zero spike scores. The code is available at https://github.com/PaddlePaddle/PaddleFleet (see Research/AdaGC).
Lay Summary
Large language models are trained through long and expensive runs, but training can sometimes suddenly become unstable, causing loss spikes, failed runs, and wasted computing resources. These failures can be triggered by many different factors, such as unusual data, hardware faults, numerical precision issues, or sensitive training settings, making it difficult to fix every root cause directly. We propose AdaGC, a simple method that acts as a safety mechanism during training. Instead of using one fixed rule for the whole model, AdaGC monitors different parts of the model separately and limits unusually large training signals before they can disrupt future updates. This makes training more stable while adding very little extra memory cost and reducing communication overhead in large distributed systems. Across several dense and mixture-of-experts language models, AdaGC reduces training instability and improves final model performance compared with standard clipping methods.