On the Origin of Algorithmic Progress in AI
Hans Gundlach ⋅ Alex Fogelson ⋅ Jayson Lynch ⋅ Ana Trisovic ⋅ Jonathan S Rosenfeld ⋅ Anmol Sandhu ⋅ Neil Thompson
Abstract
Algorithms have been estimated to increase AI pretraining FLOP efficiency by a factor of $22,000$ between 2012 and 2023 (Ho et al., 2024). Running ablation experiments on key innovations from this time period, we are able to account for less than $10\times$ of these gains at small compute scales. Surveying the broader literature, we estimate that additional innovations not included in our ablations also account for less than $10\times$, yielding less than $100\times$ efficiency gains total in low-compute regimes. This leads us to conduct scaling experiments, which reveal that much of this efficiency gap can be explained by algorithms with scale-dependent efficiency improvements. In particular, we conduct scaling experiments between LSTMs and Transformers, finding exponent differences in their compute-optimal scaling law while finding little scaling difference for many other innovations. These experiments demonstrate that -- contrary to standard assumptions -- an algorithm's efficiency gains are tied to compute scale. Using experimental extrapolation and literature estimates, we account for $6,930\times$ efficiency gains over the same time period, with scale-dependent innovations accounting for close to 90\% of log-efficiency gains by 2023. Our results indicate that algorithmic progress for small models has been far slower than previously assumed, and that measures of algorithmic efficiency are strongly reference-dependent.
Video
Chat is not available.
Successful Page Load