A $p$-adic Perspective on Low-Bit Training of Neural Networks
Daniil Bershatskiy ⋅ Marina Munkhoeva ⋅ Ivan Oseledets
Abstract
We investigate low-bit neural network training from a $p$-adic perspective. In extreme low-bit regimes, the set of representable values is so small that gradient descent operates on an essentially discrete domain, making continuous analysis inadequate. This observation motivates us to model a neural network as a polynomial system over the integers modulo $p^N$. Activations and losses are replaced by piecewise polynomial approximations, and training is recast as finding roots of the resulting system. Hensel's lemma provides an iterative procedure for lifting seed roots digit-by-digit to the required precision. We formalize this approach and demonstrate its feasibility on linear regression and shallow polynomial networks.
Chat is not available.
Successful Page Load