$\delta$-Regularized Gradient Clipping for Stable Optimization: Analysis and Empirical Evaluation
Katsiaryna Novikava ⋅ Anna Lytova ⋅ Omar Rivasplata
Abstract
This work provides an extended empirical and theoretical analysis of the proposed recently $\delta$-GClip, a variant of gradient clipping with a formal convergence guarantee. Our experiments analyze activation patterns, gradient dynamics, and dependence of $\delta$ on architectural scale across supervised benchmarks, diffusion models, and a lightweight protein‑generation task. In particular, we show that combining Adam with a brief $\delta$‑clipping warm‑up improves the stability and early‑phase optimization in diffusion model training. Using the Kurdyka–Łojasiewicz framework we further extend the convergence guarantees of $\delta$-GClip beyond the squared‑loss setting to more general smooth non‑convex objectives.
Chat is not available.
Successful Page Load