Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting
Abstract
Lay Summary
Time series forecasting models predict future values such as traffic, weather, or electricity demand. Current evaluations often compare models using one average error score, but this can be misleading because some examples are naturally much easier to predict than others. This paper proposes a diagnostic framework that estimates how predictable each forecasting example is before judging model performance. It also measures how well a model uses the predictable structure in the data. Across synthetic and real-world datasets, we find that this predictability estimate closely matches actual model errors and reveals that forecasting difficulty changes over time and across variables. Our results suggest that forecasting models should be evaluated with awareness of example difficulty, rather than only by average benchmark scores.