Activation Functions Control Finite-Width Concentration in Wide Neural Networks
Soumya Ganguly ⋅ Nilava Metya ⋅ Alexandre V Morozov ⋅ Anirvan Sengupta
Abstract
Wide randomly initialized neural networks are known to converge to Gaussian processes in the infinite-width limit. While this asymptotic limit is well understood, much less is known about the finite-width fluctuation behavior of empirical neural kernels and how this behavior depends on the activation function. In this work, we study finite-width concentration of random feature kernels through the lens of Orlicz and sub-Weibull tail theory. We show that activation growth directly controls the universality class of kernel fluctuations. Bounded activations such as $\tanh$ and $\operatorname{erf}$ satisfy Hoeffding-type concentration, $\operatorname{ReLU}$ activations exhibit sub-exponential Bernstein behavior, whereas polynomial activations generate sub-Weibull concentration regimes whose order depends explicitly on the polynomial degree. In particular, for $\varphi(x)=x^p$ and Gaussians $G,G_i,G_j$, the activation value $\varphi(G)$ has stretched-exponential tail order $2/p$, while the kernel summand $\varphi(G_i)\varphi(G_j)$ has sub-Weibull order $1/p$. This yields concentration inequalities with Weibull-type large-deviation behavior governed by the product tail. We derive entrywise concentration bounds and corresponding finite-dimensional operator-norm bounds for empirical neural kernels and illustrate the predicted scaling numerically in two-layer random networks. Our results suggest that activation growth provides a natural organizing principle for finite-width fluctuation regimes in wide neural networks.
Chat is not available.
Successful Page Load