Ellipsoidal Time Series Forecasting
Abstract
Lay Summary
Forecasting models are often compared on a small set of standard benchmark datasets reused across many papers. The trouble is that these datasets may not reveal how models behave when the world changes. A model can look strong on historical tables and still break when the system shifts, when a shock hits, or when small errors snowball into large ones. We take a different approach: forecasting should be tested more like a controlled scientific experiment. We build a benchmark of 21 simulated systems where the underlying dynamics are known exactly, and where chaos, sudden shocks, and regime changes can be introduced on purpose. This lets us ask not just “which model has the lowest average error?” but “which model is still useful when the world changes?” We also develop Fern, a forecasting model that predicts not just a number but a shape: an ellipse describing where future outcomes may stretch or shrink. On these stress tests, several popular methods fail sharply, sometimes catastrophically. Fern is often much more stable, with errors more than a hundred times smaller in the most brittle cases, and its ellipses help signal when forecasts should no longer be trusted. The broader point is that forecasting models should be judged by their failure modes, not only by average accuracy on quiet historical data.