Generalizing Stochastic Smoothing for Differentiation and Gradient Estimation
Felix Petersen ⋅ Christian Borgelt ⋅ Aashwin Mishra ⋅ Stefano Ermon
Abstract
We address the problem of gradient estimation for stochastic differentiable relaxations of algorithms, operators, simulators, and other non-differentiable functions. Stochastic smoothing conventionally perturbs the input of a non-differentiable function with a differentiable density distribution with full support, smoothing it and enabling gradient estimation. Our theory starts at first principles to derive stochastic smoothing with reduced assumptions, without requiring a differentiable density nor full support, and presenting a general framework for relaxation and gradient estimation of non-differentiable black-box functions $f$. We develop variance reduction for gradient estimation from 3 orthogonal perspectives. Empirically, we benchmark 6 distributions and up to 24 variance reduction strategies for differentiable sorting and ranking, differentiable shortest-paths on graphs, differentiable rendering for pose estimation, as well as differentiable cryo-electron tomography simulations.
Lay Summary
Derivatives lie at the heart of optimization which is foundational to modern machine learning. However, many functions, algorithms, simulators are not differentiable, thus precluding their inclusion in an end-to-end machine learning workflow. To address this shortcoming, we propose and extend methods for turning non-differentiable functions differentiable through stochastic methods. This enables end-to-end optimization of workflows that include arbitrary functions.
Successful Page Load