Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Huangyu Xu ⋅ Jingqin Yang ⋅ Qianqian Xu ⋅ Jiaye Teng
Abstract
Sparse optimization is a fundamental challenge in various practical applications. A popular approach to sparse optimization is Lp regularization. However, it may encounter optimization instability due to the unbounded gradients when 0
Lay Summary
Models can actually achieve great accuracy with only a small fraction of their weights being non-zero, and removing most parameters makes models faster and cheaper to run. We propose ReWA, which writes each weight as a product of several auxiliary variables, applies a gentle pull toward zero on all weights, and gives each weight its own adaptive learning speed. In experiments, ReWA achieves better sparsity than previous methods.
Video
Chat is not available.
Successful Page Load