Evolution Strategies at the Hyperscale
Bidipta Sarkar ⋅ Mattie Fellows ⋅ Juan Duque ⋅ Alistair Letcher ⋅ Antonio León Villares ⋅ Anya Sims ⋅ Clarisse Wibault ⋅ Dmitry Samsonov ⋅ Dylan Cope ⋅ Jarek Liesen ⋅ Kang Li ⋅ Lukas Seier ⋅ Theo Wolf ⋅ Uljad Berdica ⋅ Valentin Mohl ⋅ Alexander D. Goldie ⋅ Aaron Courville ⋅ Karin Sevegnani ⋅ Shimon Whiteson ⋅ Jakob Foerster
Abstract
Evolution Strategies (ES) is a class of powerful black-box optimisation methods that are highly parallelisable and can handle non-differentiable and noisy objectives. However, naïve ES becomes prohibitively expensive at scale on GPUs due to the low arithmetic intensity of batched matrix multiplications with unstructured random perturbations. We introduce Evolution Guided GeneRal Optimisation via Low-rank Learning (EGGROLL), which improves arithmetic intensity by structuring individual perturbations as rank-$r$ matrices, resulting in a hundredfold increase in training speed for billion-parameter models at large population sizes, achieving up to 91\% of the throughput of pure batch inference. We provide a rigorous theoretical analysis of ES for high-dimensional parameter objectives, investigating conditions needed for ES updates to converge in high dimensions, revealing a linearising effect, and proving consistency between EGGROLL and ES as parameter dimension increases. Our experiments show that EGGROLL: (1) enables the stable pretraining of nonlinear recurrent language models that operate purely in integer datatypes, (2) is competitive with GRPO for post-training LLMs on reasoning tasks, and (3) does not compromise performance compared to ES in tabula rasa RL settings, despite being faster.
Lay Summary
Most modern AI systems are trained using methods that require every part of the model to be differentiable, i.e. smooth and easy to measure mathematically. This is powerful, but it limits the kinds of models and goals we can train. We introduce EGGROLL, an evolution-inspired method that improves models by testing many small variations and keeping the changes that work best. EGGROLL makes this process fast enough for very large AI systems, reaching up to 91% of normal inference speed while requiring machines to share only simple performance scores. We show that it can train large language models, reinforcement-learning agents, and integer-only models, opening the door to training systems that standard methods struggle to handle.
Successful Page Load