Iterative Magnitude Pruning Reduces Weight-Space Coupling
Abstract
Neural networks contain many redundant and approximately equivalent parameterizations, yet it remains unclear how pruning affects this weight-space structure. We study iterative magnitude pruning through Fisher geometry and introduce a per-dimension log-determinant ratio that measures parameter coupling. Across several vision architectures and datasets, pruning consistently reduces this coupling over broad sparsity ranges. Controlled comparisons with smaller unpruned networks suggest that the effect is not explained by parameter count alone. We further show that as Fisher coupling decreases, Adam updates become more aligned with natural-gradient updates, while the same isn't true for SGD. These results suggest that winning tickets are not only sparse and trainable, but also occupy locally simpler, more coordinate-separable regions of weight space.