Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex Optimization
Abstract
Lay Summary
Recent empirical studies have shown that the noise in stochastic gradients used for training language models is often heavy-tailed, meaning that unusually large gradient updates occur more frequently than standard assumptions predict. While existing theory for heavy-tailed stochastic optimization mainly relies on strong convexity and smoothness assumptions, much less is understood in the more realistic weakly convex setting. This paper studies stochastic gradient descent (SGD) under heavy-tailed noise for weakly convex optimization problems. We establish new convergence guarantees for standard SGD under different heavy-tailed noise assumptions, including both expectation-based and high-probability results. We further show that adding gradient clipping allows SGD to achieve stronger high-probability convergence guarantees even in unbounded domains under weaker assumptions on the noise distribution.