Sharper Generalization Guarantees for Asynchronous SGD: Beyond Lipschitzness, Smoothness and Data Homogeneity
Abstract
Lay Summary
Asynchronous stochastic gradient descent (ASGD) is widely used to accelerate the training of large-scale machine learning tasks. However, this acceleration introduces a bias caused by computers updating and communicating independently. We wondered if this bias affects the performance of ASGD, especially on unseen data. We mathematically quantify this performance by excess risk and use an algorithmic stability framework to analyze it, relaxing the impractical assumptions used in prior theoretical studies. Our results indicate that this bias has a provably limited effect on ASGD's performance, showing that in certain cases, it can perform as well as standard methods while enjoying a faster speed. Also, we prove that a random variant of ASGD performs well even in a more complex case where training data comes from several different distributions. Our findings deepen the theoretical understanding of ASGD under realistic assumptions.