Scaling Laws of Global Weather Models
Abstract
Lay Summary
Artificial intelligence is currently revolutionizing how we predict the weather, but figuring out the most efficient way to train these massive programs is a major challenge. We wanted to understand the "scaling laws" of weather AI—meaning how a system's forecasting accuracy improves when we increase its size, the amount of data it learns from, or the computing power used to train it. By comparing several leading AI weather models, we discovered that these systems actually learn fundamentally differently than popular language models. Specifically, weather models become much more accurate when their internal mathematical networks are built "wider" rather than "deeper". We also found that if researchers have a limited computing budget, it is far better to feed a model more historical weather data than it is to simply make the model itself larger. Ultimately, these insights provide a clear blueprint for developers, showing that future weather-predicting AI should prioritize wider designs and massive training datasets to give us the most accurate and computationally efficient forecasts possible.