How feature learning breaks the curse of dimensionality in isotropic data:\\ A mean field approach
Niclas Göring ⋅ Yoonsoo Nam ⋅ Chris Mingard ⋅ Jake Reid ⋅ Ard Louis
Abstract
Neural networks efficiently learn isotropic data distributions with low-dimensional target structure (canonical examples include $k$-sparse parity and sparse multi-index models), yet their neural tangent kernel limits require $d^k$ samples in the ambient dimension $d$. We trace this gap to a single mechanism: during training, networks develop strong anisotropy between weight coordinates aligned with the task and those that are not. We call this input feature selection (IFS) and show, through analysis of stochastic gradient Langevin dynamics, that it arises from coordinate-dependent effective regularisation that kernels structurally cannot exhibit. Mean-field (MF) theory is the natural interpretable framework for feature learning beyond the kernel regime, but standard MF tracks only first moments of the weight distribution and so cannot represent IFS. We introduce MF-ARD, which augments MF with coordinate-wise precisions through automatic relevance determination. With this single additional set of order parameters, MF-ARD (i) captures the sharp generalisation transitions of SGLD-trained networks on $k$-sparse parity and single-index models, and (ii) provably breaks the curse of dimensionality: its phase-transition threshold depends on the intrinsic task dimension $k$ rather than the ambient dimension $d$.
Chat is not available.
Successful Page Load